Site Reliability Engineer

Zachary Piper
United States
22 days ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$140,000.0 - $165,000.0
Working hours
Regular working hours

Tech stack

Kubernetes Security Amazon Web Services Microsoft Azure Cloud Computing Configuration Management Computer Programming Continuous Integration Linux DevOps Monitoring of Systems Python (Programming Language) Linux System Administration
+19 more
Octopus Deploy Object-Oriented Software Development Performance Tuning Reliability Engineering Prometheus Ruby Scripting System Availability Grafana Infrastructure as Code (IaC) Cloudformation Containerization Git Flow Kubernetes Infrastructure Automation Frameworks Deployment Automation Terraform Docker Golang

Job description

Piper Companies is seeking a Site Reliability Engineer (SRE) to support the development, maintenance, and operation of a Kubernetes-based platform within highly regulated cloud and on-premises environments. This individual will work closely with senior engineers and technical leaders to improve platform reliability, scalability, security, and operational performance while supporting compliance-driven initiatives. This is a long-tern contract opportunity with a strong focus on Kubernetes infrastructure, Linux administration, automation, and cloud technologies. This individual may sit remote in the US., · Build, maintain, and support Kubernetes clusters across on-premises and AWS environments.

· Monitor platform reliability, availability, and performance through observability, alerting, and troubleshooting activities.

· Develop and implement automation tools and processes to improve operational efficiency and reduce manual intervention.

· Collaborate with senior engineers to define and track service reliability metrics, including SLIs, SLOs, and error budgets.

· Support compliance, security, auditing, and continuous monitoring initiatives within regulated environments.

· Contribute to Infrastructure as Code (IaC) development and enhancements using tools such as Terraform, CloudFormation, or similar technologies.

· Improve CI/CD pipelines and deployment processes to support platform scalability and operational excellence.

· Partner with Security, Platform, and Application teams to resolve issues and deliver reliable infrastructure solutions.

Requirements

· 4-6 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related infrastructure-focused roles.

· Strong hands-on experience managing and supporting Kubernetes clusters in production environments.

· Experience with AWS, Azure, or similar cloud platforms; GovCloud experience is a plus.

· Strong Linux administration skills, including system troubleshooting, performance analysis, and infrastructure support.

· Experience with Infrastructure as Code technologies such as Terraform, CloudFormation, or similar tools.

· Programming or scripting experience in Python, Go, Ruby, or other object-oriented languages..

· Experience with observability and monitoring tools such as Prometheus, Grafana, and centralized logging solutions.

· Exposure to CI/CD tools, deployment automation, and GitOps practices such as ArgoCD is preferred.

· Experience supporting FedRAMP High, DoD IL5, regulated, or audited environments is strongly preferred., Keywords: Kubernetes, Site Reliability Engineering (SRE), DevOps, Platform Engineering, AWS, Azure, GovCloud, Linux, Terraform, CloudFormation, Infrastructure as Code (IaC), CI/CD, ArgoCD, Prometheus, Grafana, Monitoring, Alerting, Observability, Containerization, Docker, Networking, Automation, Python, Go, Ruby, Object-Oriented Programming, Cloud Infrastructure, Kubernetes Clusters, Production Support, Troubleshooting, Incident Response, Performance Optimization, Scalability, Reliability, High Availability, Security, Compliance, FedRAMP High, DoD IL5, Continuous Monitoring, Audit Support, GitOps, Deployment Automation, Platform Operations, Container Security, Cross-Functional Collaboration, SLI, SLO, Error Budgets, On-Call Support, Systems Administration, Infrastructure Management, Cloud-Native Technologies, AWS Infrastructure, Infrastructure Automation, Root Cause Analysis, Configuration Management, Platform Reliability, Regulated Environments.

Benefits & conditions

· Salary range: $140,000 - $165,000

· Comprehensive Benefits: Medical, Dental, Vision, 401(k), and applicable sick leave

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

50 sec

Why developer happiness matters in web frameworks

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:30 min

Falling in love with Ruby and creating Basecamp

David Heinemeier Hansson David Heinemeier Hansson +1 · Coffee With Developers

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

Videos

See all

Related articles

See all