> Markdown version of [/jobs/ext/1821910-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1821910-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Reliability Engineer - **Company:** Systems, Inc - **Location:** Boston, MA, United States - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Linux, DevOps, Document Management Systems, Disaster Recovery, Distributed Systems, Domain Name System (DNS), Fault Tolerance, Monitoring of Systems, Python (Programming Language), Networking Basics, Oracle (Applications), Windows PowerShell, Reliability Engineering, Prometheus, Zero Trust Network Access, Software Engineering, TCP/IP, Datadog, Data Logging, Scripting, Load Balancing, Performance Testing, Grafana, Mttr, Reliability of Systems, Cloudformation, SC Clearance, Containerization, Kubernetes, Deployment Automation, Terraform, Oracle Cloud Infrastructure, Docker, Golang, Programming Languages - **Published:** July 2, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/17441835?backUrl=%2Fcareer%2F17441835%2FReliability-Engineer-Massachusetts-Boston ## About the Role * Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience. Additional experience may be accepted in lieu of degree. * Active Secret clearance at a minimum required to start * US citizenship required * Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services * Experience with containerized environments (Docker, Kubernetes) * Familiarity with CI/CD pipelines and deployment automation * SLOs and error budgets * Capacity modeling and performance testing * Strong understanding of: * Distributed systems and highavailability architectures * Linux/Windows system administration * Networking fundamentals (DNS, TCP/IP, load balancing) * Hands-on experience with: * Monitoring and observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor) * Infrastructure as Code (Terraform, ARM, CloudFormation) * Scripting or programming languages (Python, Bash, Go, PowerShell, or similar) * Experience supporting incident management and oncall operations Preferred Skills * Experience with USAF Cloud One or Platform 1. * Experience with Zero Trust Architecture * Cloud certifications in AWS, Azure, Google, or Oracle clouds ## Description This role supports the U.S. Air Force Cloud One Architecture and Common Shared Services contract and currently has an opening for a Reliability Engineer. The Reliability Engineer is responsible for ensuring the availability, performance, scalability, and resiliency of missioncritical systems. This role applies software engineering principles to infrastructure and operations, with a strong emphasis on automation, monitoring, incident response, and continuous reliability improvement. The reliability engineer serves as the bridge between development, operations, and platform teams to ensure production systems consistently meet defined service level objectives (SLOs) while supporting rapid, safe delivery of new capabilities., * Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments * Define, measure, and report Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets * Identify reliability risks and implement mitigation strategies across the system lifecycle * Conduct capacity planning and performance modeling to ensure systems scale to meet demand Monitoring, Observability & Alerting * Implement and manage monitoring, logging, and tracing solutions to provide full system observability * Define actionable alerting thresholds that minimize noise and enable rapid incident detection * Analyze trends and metrics to proactively identify potential reliability issues Incident Response & Problem Management * Participate in oncall rotations and lead incident response activities for production systems * Coordinate troubleshooting efforts across development, infrastructure, and security teams * Conduct postincident reviews (PIRs) and develop corrective and preventive action plans * Track recurring issues and ensure root causes are resolved Automation & Engineering Excellence * Automate operational tasks to reduce manual intervention and operational risk * Develop scripts, tools, and services that improve system reliability and reduce mean time to recovery (MTTR) * Promote "automation over toil" and standardize operational workflows ReliabilityFocused Engineering * Participate in architecture and design reviews with an emphasis on reliability, resiliency, and recoverability * Validate disaster recovery (DR) and business continuity plans; test failover mechanisms * Support chaos engineering, fault injection testing, and resilience validation where appropriate Collaboration & Governance * Partner with DevOps, Platform, and Security teams to ensure reliability aligns with delivery and compliance objectives * Document system reliability standards, runbooks, and operational procedures * Support compliance and audit activities (e.g., FedRAMP, FISMA, internal operational controls) ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [Azure-Well Architected Framework - designing mission critical workloads in practice](https://www.wearedevelopers.com/videos/1529-azure-well-architected-framework-designing-mission-critical-workloads-in-practice) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Top Must-Visit Developer Conferences in the US in 2026](https://www.wearedevelopers.com/magazine/679-top-must-visit-developer-conferences-in-the-us-in-2026) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)