> Markdown version of [/jobs/ext/616240-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/616240-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer - **Company:** ALLEN FAMILY VENTURES, LLC - **Location:** New York, United States (Remote available) - **Experience:** Expert - **Salary:** $179,000.0 - $226,000.0 - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Amazon Web Services, Big Data, Databases, Software Debugging, Programming Tools, Distributed Systems, Python (Programming Language), Software Engineering, Datadog, Build Management, Kubernetes, Build Tools, Cloudwatch, Terraform, Docker, Golang, Programming Languages - **Published:** June 24, 2026 - **Apply:** https://www.dice.com/job-detail/dc446845-ff0c-46fd-adbe-713bdf57816b ## About the Role * 10+ years of experience in infrastructure, SRE, or software engineering roles * Strong software engineering skills-you build systems, not just scripts * Experience managing production infrastructure at scale (cloud + containerized systems) * Experience with Infrastructure as Code (e.g., Terraform) * Experience running and troubleshooting distributed systems (Docker/Kubernetes) * Experience with observability and debugging tools (Datadog, CloudWatch, ELK/EFK, etc.) * Proficiency in at least one programming language (Python, Go, JavaScript, etc.) * Experience participating in on-call rotations and improving systems based on incidents * Strong communication and collaboration skills You might be a great fit if you * Default to automation over manual processes * See repetitive work and immediately want to eliminate it * Think in terms of systems, failure modes, and long-term scalability * Care about building infrastructure that other engineers can use safely and confidently * Enjoy working in a small team with high ownership and impact Nice to have * Experience running Kubernetes in production at scale * Deep familiarity with AWS * Experience building internal platforms or developer tooling * Background in distributed systems or large-scale data systems We're a lean team, so your impact will be felt immediately, and opportunities will grow as the company scales up. If this all sounds like a good fit for you, why not join us? ## Description We're looking for engineers who enjoy turning complex, fragile systems into automated, self-service platforms with strong safety guarantees. What you'll be doing Reporting to the Engineering Manager of Infrastructure, you'll: * Design and build systems to automate infrastructure management at scale (provisioning, upgrades, migrations) * Reduce operational toil by turning manual processes into reliable, repeatable workflows * Build internal tooling and platforms that enable safe self-service changes for other engineers * Improve the reliability and resilience of our infrastructure (Kubernetes, databases, services) * Implement and evolve systems for deploying and running applications in Kubernetes * Contribute to architecture decisions across infrastructure, reliability, and security * Write and review production-quality code * Participate in on-call rotations-but focus on building systems that prevent incidents, not just respond to them ## Related Videos - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)