> Markdown version of [/jobs/ext/2716469-sre-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2716469-sre-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE Reliability Engineer - **Company:** FlowCode LLC - **Location:** New York, NY, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Cloud Computing, Continuous Integration, DevOps, Disaster Recovery, Distributed Systems, Github, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Key Management, Prometheus, Shell Script, Datadog, Data Logging, Delivery Pipeline, Amazon Virtual Private Cloud (VPC), Amazon Relational Database Service, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Production Code, Terraform - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-sre-engineer-flowcode-9763552 ## About the Role * 4+ years of professional experience across SRE, DevOps, or Platform Engineering domains * Technical proficiency in Kubernetes, including cluster troubleshooting and managing controllers or CRDs * Advanced Terraform or OpenTofu expertise, encompassing module architecture and production state management * Hands-on operational experience with GitOps workflows via ArgoCD and Helm-based deployments * Ability to author production-grade code in Go or Python alongside robust shell scripting * Mastery of core AWS services, specifically EKS, Networking/VPC, RDS, and IAM * Experience maintaining and scaling CI/CD automation using GitHub Actions within collaborative environments * Proven track record of leading infrastructure initiatives from initial design through to long-term operation * Background in supporting large-scale distributed systems within high-availability production environments * Adept at navigating interrupt-driven workflows, balancing strategic project delivery with day-to-day operational support, * Exposure to Crossplane or alternative Kubernetes-native solutions for infrastructure provisioning * Deep observability experience utilizing Datadog or Prometheus to engineer SLOs, high-signal dashboards, and intelligent alerting * Practical knowledge of modern secrets management frameworks and implementation * Experience optimizing cluster efficiency through autoscaling technologies such as Karpenter or Cluster Autoscaler Flowcode is not for everyone. We hire with a pinhole lens - only those with the rare combination of intellectual horsepower, execution velocity, and uncompromising drive will thrive here. If you are seeking to operate at the highest levels of performance and impact, we want to meet you. ## Description Flowcode is seeking a Senior Site Reliability Engineer (SRE) to work on reliability and infrastructure efforts across our platforms. This role will help grow and drive our infrastructure strategy, operational rigor and observability while building and supporting the systems and tooling required to support Flowcode's continued growth. As an individual contributor within our engineering organization, you will develop and operate scalable cloud infrastructure, establish best practices around deployment and reliability, and partner closely with engineering teams to ensure systems are scalable, resilient and observable., * Improve system availability, scalability, and resilience across Flowcode's platforms * Own key pieces of our EKS-based infrastructure end-to-end * Contribute to incident response and postmortems, turning findings into durable fixes * Support engineering teams with infrastructure questions, escalations, and day-to-day unblocking Cloud & Platform Engineering * Manage and scale our core AWS footprint (EKS, VPC, RDS) through Infrastructure as Code (Terraform) * Enhance disaster recovery and failover mechanisms to protect mission-critical workloads * Collaborate with product engineering to streamline and optimize internal developer experience CI/CD & Deployment Automation * Design and scale deployment pipelines using GitHub Actions * Expand GitOps practices and tooling through ArgoCD * Facilitate secure delivery with automated validation and progressive rollout strategies Observability & Monitoring * Oversee and optimize the organization's monitoring, logging, and alerting infrastructure * Develop high-signal metrics, tracing, and visualization dashboards while minimizing operational noise * Establish and monitor Service Level Objectives for managed platform components ## Related Videos - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)