> Markdown version of [/jobs/ext/2493164-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2493164-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Umbra Inc. - **Location:** Arlington, VA, United States - **Experience:** Expert - **Salary:** $150,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Elastic Compute Cloud, Amazon S3, Software as a Service, Cloud Computing, DevOps, Distributed Systems, Monitoring of Systems, Identity and Access Management, Scrum Methodology, Software Architecture, Reliability Engineering, Istio, Software Security, Amazon Virtual Private Cloud (VPC), Kubernetes, Infrastructure Automation Frameworks, Information Technology, Performance Monitor, Functional Programming, Terraform - **Published:** August 2, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=94111b66fd0be8c1 ## About the Role * Bachelor's degree in Computer Science or a related technical field. * 5-8+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform, with demonstrated expertise managing distributed systems. * Extensive experience with AWS services (EC2, S3, Lambda, VPC Networking) and deep knowledge of cloud infrastructure, networking, and security best practices. * Proficiency running, optimizing, and scaling Kubernetes clusters in production environments. * Experience using and writing Terraform to architect and manage production infrastructure. * Ability to create and utilize Infrastructure-as-code (IaC), GitOps practices, and automation tools to increase reliability and reduce manual tasks. * Proven success in leading teams or projects using Agile/Scrum methodologies. * Expertise in infrastructure and software architecture, capable of designing and implementing large-scale, reliable systems with minimal guidance. * Experience developing and managing comprehensive infrastructure monitoring and alerting strategies., * 10+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform, with demonstrated expertise managing distributed systems. * Advanced understanding of cloud and application security, identity management, and compliance. * Expertise in service mesh and service registration technologies, focusing on performance and reliability. * Experience in the aerospace industry. ## Description * Ensure the reliability and scalability of critical systems, meeting SLAs through proactive monitoring and effective incident response. * Develop and promote new technologies and tools, conducting research and creating proofs of concept to introduce solutions that enhance the team's capabilities. * Lead by example in fostering a culture of excellence and reliability. * Continuously evaluate and improve team processes and workflows to increase efficiency and reduce complexity. * Collaborate closely with cross-functional teams, product managers, and stakeholders to align on technical strategy and provide expert guidance. * Participate in on-call rotations, providing support and resolving complex technical issues. ## Related Videos - [Reliable scalability: How Amazon.com scales on AWS](https://www.wearedevelopers.com/videos/983-reliable-scalability-how-amazon-com-scales-on-aws) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)