> Markdown version of [/jobs/ext/2188457-site-reliability-engineer-infrastructure-operations](https://www.wearedevelopers.com/jobs/ext/2188457-site-reliability-engineer-infrastructure-operations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - Infrastructure Operations - **Company:** ClearanceJobs Workforce Solutions - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Computer Clusters, Databases, Continuous Integration, Extract Transform Load (ETL), DevOps, White-Box Testing, Python (Programming Language), Linux System Administration, Performance Tuning, Queueing Systems, Redis, Reliability Engineering, Prometheus, Datadog, Google Cloud, Load Balancing, Grafana, Mttr, Caching, Cloudformation, SC Clearance, Kubernetes, Amazon Simple Queue Service (SQS), Terraform, Splunk, Serverless Computing, Pagerduty - **Published:** August 22, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9089567/site-reliability-engineer-infrastructure-operations ## About the Role * 5+ years of hands-on experience in SRE, infrastructure operations, or DevOps roles * Python proficiency - it's our primary language; you'll write operational automation tools * Deep experience with monitoring/observability platforms (Datadog, Prometheus, Grafana, Splunk, ELK, or equivalent) * Proven on-call incident response experience * Comfort with 24/7 on-call rotations - you understand production-critical operations and can escalate appropriately * AWS or Google Cloud infrastructure experience (VPCs, load balancing, databases, serverless) * Linux systems administration at a production level * Active Secret clearance or ability to obtain one (required for this role) Preferred Skills * Infrastructure as Code (Terraform, CloudFormation) - nice to have but learnable on the job * Kubernetes operations (EKS/GKE) - operational expertise more valuable than deep design * Experience with message queues (SNS/SQS) and caching (Redis) * GPU resource optimization and ML workload observability * Experience supporting mission-critical systems for government/defense customers * Background in incident response and blameless postmortem culture ## Description ClearanceJobs Worforce Solutions is seeking a Site Reliability Engineer - Infrastructure Operations for our client based in San Francisco, CA. Our client is a cutting-edge startup delivering AI-based solutions for critical defense applications serving the U.S. Department of Defense and international defense customers. They operate 24/7 across multiple computing environments-cloud hyperscalers, on-premise GPU clusters, field-deployed systems, and supercomputing centers-with zero tolerance for downtime. The team is lean and mission-focused, with every member directly impacting the delivery of life-critical systems to their customers. This is a 24/7 on-call infrastructure operations role focused on maintaining mission-critical AI systems at 99% uptime SLA. You'll be the primary responder for infrastructure incidents, monitoring system health, and optimizing existing production pipelines for reliability and performance. This role is NOT about building new infrastructure from scratch. Our foundations are established. You'll be the operational backbone keeping them running reliably while continually optimizing for performance, cost, and customer SLA commitments., Primary Focus - On-Call Operations & Monitoring (60%) * 24/7 on-call rotation as the primary responder for infrastructure incidents and production alerts * Monitor system health and SLA metrics across all customer contracts; escalate to engineering team when needed * Triage and respond to production incidents: analyze logs/metrics, diagnose root causes, execute remediation, and document post-mortems * Build and maintain comprehensive observability: whitebox (application-level) and blackbox (system-level) monitoring * Ensure Datadog and PagerDuty alerting strategies are tuned to catch issues without alert fatigue * Develop and automate incident response playbooks to minimize MTTR (mean time to recovery) * Maintain 99% uptime SLA across defense and government contracts Secondary Focus - Infrastructure Optimization & Reliability (40%) * Optimize existing infrastructure for performance, cost, and reliability (not redesign from scratch) * Identify and address infrastructure bottlenecks through capacity planning and performance tuning * Maintain ETL pipelines and data quality for core forecasting operations * Collaborate with software engineers on deployment processes and CI/CD improvements * Document runbooks, playbooks, and operational procedures for team scalability ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)