> Markdown version of [/jobs/ext/2702207-staff-site-reliability-engineer-in-new-york](https://www.wearedevelopers.com/jobs/ext/2702207-staff-site-reliability-engineer-in-new-york). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer in New York - **Company:** Energy Jobline - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $220,000.0 - $260,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Build Automation, Continuous Integration, Distributed Systems, Github, Reliability Engineering, Software Engineering, Datadog, AWS Fargate - **Published:** September 4, 2026 - **Apply:** https://www.energyjobline.com/job/staff-site-reliability-engineer-new-york-31474500 ## About the Role * 10+ years in SRE, infrastructure, or backend engineering roles * Strong software engineering experience in one or more modern * Expertise operating distributed systems in production at scale * Deep experience with AWS, observability tooling, and CI/CD systems * Comfortable navigating ambiguity and setting direction in a fast-moving environment ## Description We're looking for a Staff Site Reliability Engineer to lead the evolution of Tabs' platform as we scale. In this role, you'll operate as a senior individual contributor, partnering closely with engineering and product teams to design, build, and operate systems that are reliable, observable, and easy to develop on. You'll own our infrastructure direction, shape how we ship software, and set the standard for operational excellence across the company. This is a high-impact role for someone who enjoys solving complex systems problems, influencing architecture, and raising the reliability bar without becoming a gatekeeper. What You'll Own * AWS infrastructure direction and platform evolution, including the migration from ECS/Fargate toward a more modern, scalable runtime * CI/CD systems with a strong emphasis on developer experience, safety, and automation (GitHub Actions today; maturing CD tomorrow) * Ephemeral environments and preview deploys to speed iteration and increase confidence in changes * Observability standards across metrics, logs, and tracing, including alert hygiene, dashboards, and SLO development * Incident response, postmortems, and the reliability culture that surrounds them What You'll Do * Define and evolve reliability standards, SLIs, SLOs, and error budgets * Improve observability, alerting, and incident processes across services * Lead high-severity incidents and drive clear, actionable follow-ups * Partner with engineering teams to design resilient, scalable systems * Build automation to reduce toil and lower operational risk * Mentor engineers and influence best practices across teams Who You Are * You've run production systems on AWS and can lead platform-level change * You think in systems: risk, rollback strategy, blast radius, and feedback loops * You treat CI/CD and environments as products that should be fast, reliable, and self-serve * You influence through trust and clarity rather than control * You balance pragmatism with long-term system health * You value learning from failure and improving processes over assigning blame * You communicate clearly and work well across teams ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Serverless landscape beyond functions](https://www.wearedevelopers.com/videos/491-serverless-landscape-beyond-functions) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)