Reliability Engineer

Systems, Inc
Boston, MA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Linux DevOps Document Management Systems Disaster Recovery Distributed Systems Domain Name System (DNS) Fault Tolerance Monitoring of Systems
+27 more
Python (Programming Language) Networking Basics Oracle (Applications) Windows PowerShell Reliability Engineering Prometheus Zero Trust Network Access Software Engineering TCP/IP Datadog Data Logging Scripting Load Balancing Performance Testing Grafana Mttr Reliability of Systems Cloudformation SC Clearance Containerization Kubernetes Deployment Automation Terraform Oracle Cloud Infrastructure Docker Golang Programming Languages

Job description

This role supports the U.S. Air Force Cloud One Architecture and Common Shared Services contract and currently has an opening for a Reliability Engineer. The Reliability Engineer is responsible for ensuring the availability, performance, scalability, and resiliency of missioncritical systems. This role applies software engineering principles to infrastructure and operations, with a strong emphasis on automation, monitoring, incident response, and continuous reliability improvement. The reliability engineer serves as the bridge between development, operations, and platform teams to ensure production systems consistently meet defined service level objectives (SLOs) while supporting rapid, safe delivery of new capabilities., * Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments

  • Define, measure, and report Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
  • Identify reliability risks and implement mitigation strategies across the system lifecycle
  • Conduct capacity planning and performance modeling to ensure systems scale to meet demand

Monitoring, Observability & Alerting

  • Implement and manage monitoring, logging, and tracing solutions to provide full system observability
  • Define actionable alerting thresholds that minimize noise and enable rapid incident detection
  • Analyze trends and metrics to proactively identify potential reliability issues

Incident Response & Problem Management

  • Participate in oncall rotations and lead incident response activities for production systems
  • Coordinate troubleshooting efforts across development, infrastructure, and security teams
  • Conduct postincident reviews (PIRs) and develop corrective and preventive action plans
  • Track recurring issues and ensure root causes are resolved

Automation & Engineering Excellence

  • Automate operational tasks to reduce manual intervention and operational risk
  • Develop scripts, tools, and services that improve system reliability and reduce mean time to recovery (MTTR)
  • Promote “automation over toil” and standardize operational workflows

ReliabilityFocused Engineering

  • Participate in architecture and design reviews with an emphasis on reliability, resiliency, and recoverability
  • Validate disaster recovery (DR) and business continuity plans; test failover mechanisms
  • Support chaos engineering, fault injection testing, and resilience validation where appropriate

Collaboration & Governance

  • Partner with DevOps, Platform, and Security teams to ensure reliability aligns with delivery and compliance objectives
  • Document system reliability standards, runbooks, and operational procedures
  • Support compliance and audit activities (e.g., FedRAMP, FISMA, internal operational controls)

Requirements

  • Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience. Additional experience may be accepted in lieu of degree.

  • Active Secret clearance at a minimum required to start

  • US citizenship required

  • Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services

  • Experience with containerized environments (Docker, Kubernetes)

  • Familiarity with CI/CD pipelines and deployment automation

  • SLOs and error budgets

  • Capacity modeling and performance testing

  • Strong understanding of:

  • Distributed systems and highavailability architectures

  • Linux/Windows system administration

  • Networking fundamentals (DNS, TCP/IP, load balancing)

  • Hands-on experience with:

  • Monitoring and observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor)

  • Infrastructure as Code (Terraform, ARM, CloudFormation)

  • Scripting or programming languages (Python, Bash, Go, PowerShell, or similar)

  • Experience supporting incident management and oncall operations

Preferred Skills

  • Experience with USAF Cloud One or Platform 1.
  • Experience with Zero Trust Architecture
  • Cloud certifications in AWS, Azure, Google, or Oracle clouds

Benefits & conditions

SES provides a competitive salary and the following benefits:

  • Medical
  • Dental
  • Vision
  • AD&D
  • STD
  • LTD
  • Company paid Life Insurance
  • 401k with employer contribution
  • Paid Time Off
  • Pet Insurance

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · WWC 2022

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all