Site Reliability Engineer - Infrastructure Operations

ClearanceJobs Workforce Solutions
New York, NY, United States
7 days ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Computer Clusters Databases Continuous Integration Extract Transform Load (ETL) DevOps White-Box Testing Python (Programming Language) Linux System Administration Performance Tuning Queueing Systems
+17 more
Redis Reliability Engineering Prometheus Datadog Google Cloud Load Balancing Grafana Mttr Caching Cloudformation SC Clearance Kubernetes Amazon Simple Queue Service (SQS) Terraform Splunk Serverless Computing Pagerduty

Job description

ClearanceJobs Worforce Solutions is seeking a Site Reliability Engineer - Infrastructure Operations for our client based in San Francisco, CA.

Our client is a cutting-edge startup delivering AI-based solutions for critical defense applications serving the U.S. Department of Defense and international defense customers. They operate 24/7 across multiple computing environments-cloud hyperscalers, on-premise GPU clusters, field-deployed systems, and supercomputing centers-with zero tolerance for downtime. The team is lean and mission-focused, with every member directly impacting the delivery of life-critical systems to their customers.

This is a 24/7 on-call infrastructure operations role focused on maintaining mission-critical AI systems at 99% uptime SLA. You’ll be the primary responder for infrastructure incidents, monitoring system health, and optimizing existing production pipelines for reliability and performance.

This role is NOT about building new infrastructure from scratch. Our foundations are established. You’ll be the operational backbone keeping them running reliably while continually optimizing for performance, cost, and customer SLA commitments., Primary Focus - On-Call Operations & Monitoring (60%)

  • 24/7 on-call rotation as the primary responder for infrastructure incidents and production alerts
  • Monitor system health and SLA metrics across all customer contracts; escalate to engineering team when needed
  • Triage and respond to production incidents: analyze logs/metrics, diagnose root causes, execute remediation, and document post-mortems
  • Build and maintain comprehensive observability: whitebox (application-level) and blackbox (system-level) monitoring
  • Ensure Datadog and PagerDuty alerting strategies are tuned to catch issues without alert fatigue
  • Develop and automate incident response playbooks to minimize MTTR (mean time to recovery)
  • Maintain 99% uptime SLA across defense and government contracts Secondary Focus - Infrastructure Optimization & Reliability (40%)

  • Optimize existing infrastructure for performance, cost, and reliability (not redesign from scratch)
  • Identify and address infrastructure bottlenecks through capacity planning and performance tuning
  • Maintain ETL pipelines and data quality for core forecasting operations
  • Collaborate with software engineers on deployment processes and CI/CD improvements
  • Document runbooks, playbooks, and operational procedures for team scalability

Requirements

  • 5+ years of hands-on experience in SRE, infrastructure operations, or DevOps roles
  • Python proficiency - it’s our primary language; you’ll write operational automation tools
  • Deep experience with monitoring/observability platforms (Datadog, Prometheus, Grafana, Splunk, ELK, or equivalent)
  • Proven on-call incident response experience
  • Comfort with 24/7 on-call rotations - you understand production-critical operations and can escalate appropriately
  • AWS or Google Cloud infrastructure experience (VPCs, load balancing, databases, serverless)
  • Linux systems administration at a production level
  • Active Secret clearance or ability to obtain one (required for this role) Preferred Skills

  • Infrastructure as Code (Terraform, CloudFormation) - nice to have but learnable on the job
  • Kubernetes operations (EKS/GKE) - operational expertise more valuable than deep design
  • Experience with message queues (SNS/SQS) and caching (Redis)
  • GPU resource optimization and ML workload observability
  • Experience supporting mission-critical systems for government/defense customers
  • Background in incident response and blameless postmortem culture

Benefits & conditions

  • Someone to redesign infrastructure from scratch (our foundations are solid)
  • Pure CI/CD specialists focused on build systems (that’s secondary here)
  • Infrastructure architects without on-call operations experience
  • Candidates uncomfortable with 24/7 on-call responsibilities

What You’ll Get

  • Competitive salary
  • Small, world-class engineering team (5 engineers + CTO) - you’ll have impact on every decision
  • Mission-critical work for U.S. Air Force, Navy, and international defense customers
  • Modern tech stack: Python, AWS/GCP, Datadog, PagerDuty, GitHub, Kubernetes
  • Clear SRE culture: monitoring first, blameless postmortems, automation over heroics
  • Bay Area location with relocation flexibility for exceptional candidates

The Reality of On-Call This is a genuinely 24/7 role:

  • You’ll be on rotation with other engineers (currently 5-person team)
  • Monitoring alerts notify you when systems degrade; you jump on them ASAP
  • If you can’t respond, it escalates to the next person in the rotation
  • Customers are U.S. Air Force, Navy, and international defense agencies - downtime matters
  • 99% SLA means ~7.2 hours of allowable downtime per month across all contracts
  • This is not a role for someone who wants “off-the-grid” availability The client prioritizes operational discipline over hero culture. If you’re uncomfortable with this level of on-call responsibility, this role is not a fit.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

2:19 min

Applying code assistant capabilities to infrastructure and cloud operations

Ryan J Salva · Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all