Site Reliability Engineer (SRE), Senior (Contingent)

Wilcore Technologies Inc.
United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$120,000.0 - $142,000.0
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Amazon Web Services Automation of Tests Bash Shell Cloud Computing Cloud Engineering Continuous Integration DevOps Distributed Systems Python (Programming Language) Linux System Administration Performance Tuning
+20 more
Windows PowerShell Role-Based Access Control Reliability Engineering Site Reliability Engineering Practices Ansible Prometheus Zero Trust Network Access Software Vulnerability Management Data Logging Grafana Kubernetes Information Technology Deployment Automation Hashicorp Cloudwatch Terraform Splunk Docker Golang Programming Languages

Job description

  • Service Reliability & Ownership: Help mature SRE practices across the service lifecycle - design, deployment, and ongoing operation - defining and tracking SLIs, SLOs, and error budgets to guide engineering and reliability decisions.
  • Automation & CI/CD: Build and maintain CI/CD pipelines and Infrastructure as Code (Terraform, Ansible) to enable secure, repeatable delivery, integrating automated testing and security checks to reduce release risk.
  • Observability & Performance: Design monitoring, logging, tracing, and alerting to speed issue detection; analyze system trends to improve reliability and reduce operational toil through automation and better runbooks.
  • Cloud Engineering & Modernization: Support and modernize AWS and containerized (Kubernetes) infrastructure, contributing to deployment consistency, capacity planning, and architectural improvements.
  • Security & Compliance: Implement reliability practices aligned with Federal security requirements (secure configuration, least privilege, vulnerability remediation) in partnership with cybersecurity teams.
  • Cross-Functional Collaboration: Work across development, platform, operations, and architecture teams - and support Agile/SAFe delivery - to translate technical direction into engineering improvements and reliable release practices.
  • Incident Support & Continuous Improvement: Participate in incident response, root cause analysis, and post-incident reviews; drive corrective action through automation and process refinement, and support on-call readiness.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
  • 8+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud operations, or related roles supporting enterprise or mission-critical environments.
  • Hands-on experience with cloud platforms (AWS preferred), Linux-based environments, and distributed systems at scale.
  • Strong experience with Infrastructure as Code and automation tools such as Terraform or Ansible.
  • Experience with containers and orchestration platforms such as Kubernetes, EKS, ECS, or Docker in production.
  • Experience building and maintaining CI/CD pipelines and deployment automation.
  • Strong understanding of monitoring, observability, incident response, and performance optimization.
  • Proficiency in one or more scripting/programming languages (Python, Go, Bash, or PowerShell).
  • Ability to obtain and maintain a federal Public Trust clearance., * Experience supporting VA, Federal Government, or other regulated environments with strong security and compliance requirements.
  • Experience defining and operationalizing SLIs, SLOs, error budgets, and service health metrics.
  • Familiarity with observability tools such as Prometheus, Grafana, CloudWatch, ELK, Splunk, or OpenTelemetry.
  • Experience with FedRAMP, NIST, Zero Trust, or other Federal security frameworks.
  • Experience supporting healthcare platforms, high-availability enterprise services, or large-scale modernization initiatives.
  • Relevant certifications (e.g., AWS Certified DevOps Engineer, AWS Certified Solutions Architect, CKA, HashiCorp Terraform Associate, or SRE/DevOps certifications).

About the company

Wilcore believes that the best products are built when companies understand and value the things they are working on. We value learning and growth and the ability to make a big impact at a small company. We believe that we can make big changes happen and improve the daily lives of millions of people by bringing quality software to the federal space.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all