Senior Cloud Site Reliability Engineer (SRE)

Peraton Inc
United States
1 day ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$104,000.0 - $166,000.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Agile Methodology Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Build Automation Automation of Tests Cloud Computing Software Quality Code Review Information Systems Continuous Integration
+23 more
DevOps Identity and Access Management Python (Programming Language) Pattern Recognition Reliability Engineering Site Reliability Engineering Practices Software Engineering Software Modules Cloud Platform System Test-Driven Development (TDD) Grafana State Machines Amazon Virtual Private Cloud (VPC) Cloudformation Servicebus Infrastructure Automation Frameworks Information Technology Functional Programming Cloudwatch Terraform Software Version Control Golang Programming Languages

Job description

Peraton is looking for a Senior Cloud Site Reliability Engineer (SRE) who will be responsible for designing and developing advanced Python-based AWS cloud solutions and engineering reliability tools for the Cloud Foundation Services (CFS) platform in the Infrastructure, Platforms & Operations organization. This person will apply software engineering practices including Infrastructure-as-Code (IaC) with Terraform to build scalable, reusable solutions and utilities that enhance platform reliability across the Federal Reserve System., * Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil, improve cloud platform reliability, and industrialize SRE practices across the system

  • Build automation scripts, APIs, and utilities in Python to reduce toil and improve platform reliability.
  • Implement observability and monitoring solutions (Grafana, AWS CloudWatch) leveraging Python for custom metrics and dashboards
  • Build and optimize Infrastructure as Code (IaC) using Terraform to manage AWS resources related to SRE solutions, incorporating cost-efficient design principles
  • Optimize Infrastructure as Code (IaC) with Terraform for AWS resources, integrating Python-based workflows.
  • Develop CI/CD pipelines and automated testing to ensure code quality, reliability, and rapid delivery of the solutions
  • Define SRE standards, best practices, and guidelines for adoption across teams; establish SRE metrics like SLI, SLOs, etc.
  • Apply software engineering best practices including version control, code reviews, test-driven development, and documentation to all development
  • Participate in incident management and on-call rotation, providing technical support for SRE tools, troubleshooting production issues, and collaborating with teams to reduce incident recurrence through proactive detection and pattern analysis
  • Stay current with emerging AWS services, SRE methodologies, and cloud-native development technologies, and drive adoption of innovative solutions
  • Collaborate within Agile and Scaled Agile frameworks with cross-functional teams to deliver integrated cloud automation solutions
  • Produce clear, blameless postmortems with actionable items and documented failure scenarios

Requirements

  • Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level Clearance
  • Bachelors Degree and 8 years of experience, or a High School diploma or equivalent and 12 years of experience
  • Must have 5+ years of advanced Python development experience, building enterprise-grade, highly available tools, APIs, and utilities for AWS
  • 7+ years of extensive experience in software development with focus on reliability and platform engineering
  • 3+ years of hands-on experience developing solutions in AWS environments with deep understanding of core services (EC2, VPC, S3, Lambda, IAM, CloudFormation, EventBridge, Step Functions etc.) and resource cost optimization
  • 3+ years of experience applying SRE principles including observability, toil automation, SLIs/SLOs and reliability engineering
  • Expert-level proficiency with Infrastructure as Code (IaC) using Terraform, including module development and state management
  • Strong experience with CI/CD pipelines, automated testing frameworks, and DevOps practices
  • Experience with observability tools and practices including Grafana, AWS CloudWatch, AWS Canary
  • Experience defining, implementing, and managing SLOs/SLIs and error budgets; familiarity with conducting RCAs and producing postmortem documentation
  • Working experience in Agile and Scaled Agile environments and familiarity with ITSM processes (incident, change, and problem management), resilience testing and chaos engineering practices

Preferred Qualifications

  • Experience with GoLang or additional programming languages is a plus
  • Bachelors Degree in Computer Science, Information Systems, or similar

Benefits & conditions

Target Salary Range: $104,000 - $166,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.

Benefits Statement: Peraton offers eligible employees a variety of benefits including medical, dental, vision, life, health savings account, short/long term disability, EAP, parental leave, 401(k), paid time off (PTO) for vacation, and company paid holidays. A full listing of available benefits can be viewed at https://www.careers.peraton.com/benefits.

About the company

Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees solve the most daunting challenges that our customers face. Visit peraton.com to learn how we’re keeping people around the world safe and secure.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar · World Congress 2025

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

Videos

See all

Related articles

See all