Site Reliability Engineer

ClearanceJobs Workforce Solutions
San Francisco, CA, United States
20 days ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Continuous Integration Extract Transform Load (ETL) Linux DevOps Distributed Systems White-Box Testing Python (Programming Language) Nagios Performance Tuning Queueing Systems Redis
+13 more
Reliability Engineering Software Engineering Datadog Data Logging Google Cloud Caching SC Clearance Infrastructure Automation Frameworks Build Tools Machine Learning Operations Amazon Simple Queue Service (SQS) Terraform Pagerduty

Job description

ClearanceJobs Worforce Solutions is actively seeking a Site Reliability Engineer (DevOps) for our client located in San Francisco, CA. ClearanceJobs is the largest career network for professionals with federal government security clearances, connecting cleared talent with the opportunities that matter most. We are committed to matching top-tier candidates with clients who depend on mission-focused professionals to support national security, defense, and intelligence community initiatives. Our team works closely with hiring managers to understand their needs and deliver candidates who are not only technically qualified but cleared and ready to contribute from day one.

Responsibilities will include, but not limited to:

  • Design, implement, and continuously optimize scalable infrastructure to support a growing production environment, including capacity planning and performance tuning.
  • Maintain and enhance whitebox and blackbox monitoring and observability (metrics, logging, tracing), using tools such as Datadog and PagerDuty to drive reliable alerting and fast incident response.
  • Support and safeguard existing ETL pipelines, partnering with data teams to ensure pipeline reliability and data quality.
  • Manage CI/CD and build systems to keep development workflows reliable and developer productivity high.
  • Maintain and iterate on existing Terraform-based infrastructure as code, focused on optimization and hardening rather than net-new builds.
  • Participate in a shared on-call rotation as the primary responder for production incidents, developing and automating playbooks to reduce downtime and manual toil., * San Francisco, CA (SoMa) -100% on-site (preferred location)
  • Remote candidates considered near Omaha, NE (proximity to Offutt AFB), Boston, MA, or the Washington, DC metro area, tied to client site needs.
  • Travel to a client site roughly every two to three months (quarterly, potentially more frequent depending on contract needs) should be expected, for both cleared and uncleared hires, across both roles.

Requirements

  • 6+ years of experience in Site Reliability Engineering or DevOps roles.
  • Proficiency with infrastructure as code tools, particularly Terraform.
  • Hands-on experience with AWS and/or Google Cloud.
  • Strong background in software engineering and CI/CD pipeline management.
  • Experience with on-call operations and incident management.
  • Familiarity with monitoring and alerting tools such as Datadog and PagerDuty.

Preferred Qualifications:

  • Candidates local to San Francisco, CA
  • Expertise in managing and scaling distributed systems.
  • Strong understanding of networking, security, and Linux systems.
  • Experience automating infrastructure and deployment processes.
  • Familiarity with message queues (SNS/SQS) and caching systems (Redis).
  • Knowledge of ML workflows and GPU resource management is a plus.
  • Comfort working in Python, the team’s primary language.
  • A collaborative, communicative style with a strong sense of initiative - this team is small and startup-paced, so self-starters who flag issues early are highly valued.

Clearance Requirements:

  • Must currently hold, or be able to obtain, an active US Secret clearance
  • Ability to work on US federal government contracts requiring an active security clearance., This position requires the ability to remain in a stationary position for extended periods of time, operate standard office equipment such as a computer, keyboard, and telephone, and occasionally move about the work environment. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

Videos

See all

Related articles

See all