> Markdown version of [/jobs/ext/3102421-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3102421-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Reliability Engineer - **Company:** Systems, Inc - **Location:** Boston, MA, United States - **Salary:** $74,090.0 - $106,496.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Linux, DevOps, Document Management Systems, Disaster Recovery, Distributed Systems, Domain Name System (DNS), Fault Tolerance, Monitoring of Systems, Python (Programming Language), Networking Basics, Oracle (Applications), Windows PowerShell, Reliability Engineering, Prometheus, Zero Trust Network Access, TCP/IP, Datadog, Data Logging, Scripting, Load Balancing, Performance Testing, Grafana, Mttr, Reliability of Systems, Cloudformation, SC Clearance, Containerization, Kubernetes, Deployment Automation, Terraform, Oracle Cloud Infrastructure, Docker, Golang, Programming Languages - **Published:** September 27, 2026 - **Apply:** https://www.careerjet.com/jobad/us7c47484779a15fc674dd2f2937792f78 ## About the Role · Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience. Additional experience may be accepted in lieu of degree. · Active Secret clearance at a minimum required to start · US citizenship required · Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services · Experience with containerized environments (Docker, Kubernetes) · Familiarity with CI/CD pipelines and deployment automation · SLOs and error budgets · Capacity modeling and performance testing · Strong understanding of: · Distributed systems and high-availability architectures · Linux/Windows system administration · Networking fundamentals (DNS, TCP/IP, load balancing) · Hands-on experience with: · Monitoring and observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor) · Infrastructure as Code (Terraform, ARM, CloudFormation) · Scripting or programming languages (Python, Bash, Go, PowerShell, or similar) · Experience supporting incident management and on-call operations Preferred Skills * Experience with USAF Cloud One or Platform 1. * Experience with Zero Trust Architecture * Cloud certifications in AWS, Azure, Google, or Oracle clouds ## Description Location: This position will be hybrid remote. Candidates will be required to work onsite as needed. Candidates preferred to be located near Hanscom AFB (Boston, MA). Requirements System Reliability & Availability * Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments * Define, measure, and report Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets * Identify reliability risks and implement mitigation strategies across the system lifecycle * Conduct capacity planning and performance modeling to ensure systems scale to meet demand Monitoring, Observability & Alerting * Implement and manage monitoring, logging, and tracing solutions to provide full system observability * Define actionable alerting thresholds that minimize noise and enable rapid incident detection * Analyze trends and metrics to proactively identify potential reliability issues Incident Response & Problem Management * Participate in on-call rotations and lead incident response activities for production systems * Coordinate troubleshooting efforts across development, infrastructure, and security teams * Conduct post-incident reviews (PIRs) and develop corrective and preventive action plans * Track recurring issues and ensure root causes are resolved Automation & Engineering Excellence * Automate operational tasks to reduce manual intervention and operational risk * Develop scripts, tools, and services that improve system reliability and reduce mean time to recovery (MTTR) * Promote "automation over toil" and standardize operational workflows Reliability-Focused Engineering * Participate in architecture and design reviews with an emphasis on reliability, resiliency, and recoverability * Validate disaster recovery (DR) and business continuity plans; test failover mechanisms * Support chaos engineering, fault injection testing, and resilience validation where appropriate Collaboration & Governance * Partner with DevOps, Platform, and Security teams to ensure reliability aligns with delivery and compliance objectives * Document system reliability standards, runbooks, and operational procedures * Support compliance and audit activities (e.g., FedRAMP, FISMA, internal operational controls), Description Leidos has an opening for a AWS Cloud Engineer supporting the U.S. Air Force Cloud One Architecture and Common Shared Services (ACSS) contract. This is an exciting op… + 3 days ago ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)