Site Reliability Engineer - Infrastructure Operations
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+17 more
Job description
ClearanceJobs Worforce Solutions is seeking a Site Reliability Engineer - Infrastructure Operations for our client based in San Francisco, CA.
Our client is a cutting-edge startup delivering AI-based solutions for critical defense applications serving the U.S. Department of Defense and international defense customers. They operate 24/7 across multiple computing environments-cloud hyperscalers, on-premise GPU clusters, field-deployed systems, and supercomputing centers-with zero tolerance for downtime. The team is lean and mission-focused, with every member directly impacting the delivery of life-critical systems to their customers.
This is a 24/7 on-call infrastructure operations role focused on maintaining mission-critical AI systems at 99% uptime SLA. You’ll be the primary responder for infrastructure incidents, monitoring system health, and optimizing existing production pipelines for reliability and performance.
This role is NOT about building new infrastructure from scratch. Our foundations are established. You’ll be the operational backbone keeping them running reliably while continually optimizing for performance, cost, and customer SLA commitments., Primary Focus - On-Call Operations & Monitoring (60%)
- 24/7 on-call rotation as the primary responder for infrastructure incidents and production alerts
- Monitor system health and SLA metrics across all customer contracts; escalate to engineering team when needed
- Triage and respond to production incidents: analyze logs/metrics, diagnose root causes, execute remediation, and document post-mortems
- Build and maintain comprehensive observability: whitebox (application-level) and blackbox (system-level) monitoring
- Ensure Datadog and PagerDuty alerting strategies are tuned to catch issues without alert fatigue
- Develop and automate incident response playbooks to minimize MTTR (mean time to recovery)
-
Maintain 99% uptime SLA across defense and government contracts Secondary Focus - Infrastructure Optimization & Reliability (40%)
- Optimize existing infrastructure for performance, cost, and reliability (not redesign from scratch)
- Identify and address infrastructure bottlenecks through capacity planning and performance tuning
- Maintain ETL pipelines and data quality for core forecasting operations
- Collaborate with software engineers on deployment processes and CI/CD improvements
- Document runbooks, playbooks, and operational procedures for team scalability
Requirements
- 5+ years of hands-on experience in SRE, infrastructure operations, or DevOps roles
- Python proficiency - it’s our primary language; you’ll write operational automation tools
- Deep experience with monitoring/observability platforms (Datadog, Prometheus, Grafana, Splunk, ELK, or equivalent)
- Proven on-call incident response experience
- Comfort with 24/7 on-call rotations - you understand production-critical operations and can escalate appropriately
- AWS or Google Cloud infrastructure experience (VPCs, load balancing, databases, serverless)
- Linux systems administration at a production level
-
Active Secret clearance or ability to obtain one (required for this role) Preferred Skills
- Infrastructure as Code (Terraform, CloudFormation) - nice to have but learnable on the job
- Kubernetes operations (EKS/GKE) - operational expertise more valuable than deep design
- Experience with message queues (SNS/SQS) and caching (Redis)
- GPU resource optimization and ML workload observability
- Experience supporting mission-critical systems for government/defense customers
- Background in incident response and blameless postmortem culture
Benefits & conditions
- Someone to redesign infrastructure from scratch (our foundations are solid)
- Pure CI/CD specialists focused on build systems (that’s secondary here)
- Infrastructure architects without on-call operations experience
- Candidates uncomfortable with 24/7 on-call responsibilities
What You’ll Get
- Competitive salary
- Small, world-class engineering team (5 engineers + CTO) - you’ll have impact on every decision
- Mission-critical work for U.S. Air Force, Navy, and international defense customers
- Modern tech stack: Python, AWS/GCP, Datadog, PagerDuty, GitHub, Kubernetes
- Clear SRE culture: monitoring first, blameless postmortems, automation over heroics
- Bay Area location with relocation flexibility for exceptional candidates
The Reality of On-Call This is a genuinely 24/7 role:
- You’ll be on rotation with other engineers (currently 5-person team)
- Monitoring alerts notify you when systems degrade; you jump on them ASAP
- If you can’t respond, it escalates to the next person in the rotation
- Customers are U.S. Air Force, Navy, and international defense agencies - downtime matters
- 99% SLA means ~7.2 hours of allowable downtime per month across all contracts
- This is not a role for someone who wants “off-the-grid” availability The client prioritizes operational discipline over hero culture. If you’re uncomfortable with this level of on-call responsibility, this role is not a fit.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Fully Remote Software Engineer Jobs
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Dev Digest 120 - Apple and peers