Staff Site Reliability Engineer

Stratitech Services LLC
San Francisco, CA, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$210,000.0 - $250,000.0
Working hours
Regular working hours
Job source

Tech stack

Airflow Amazon Web Services Systems Engineering Bash Shell Continuous Integration Software Debugging Linux DevOps Distributed Systems Github Python (Programming Language) Machine Learning
+12 more
Performance Tuning Reliability Engineering Prometheus Azure Machine Learning Datadog Scripting Grafana Apache Spark Apache Kafka Data Management Machine Learning Operations Docker

Job description

You will define how mission-critical machine learning and real-time analytics systems operate in production - influencing reliability strategy, deployment standards, and infrastructure architecture across engineering.

This team operates in a highly collaborative, in-person engineering environment in SOMA. Infrastructure, ML, and engineering leaders work side by side to design, build, and operate complex systems in real time. The pace is fast, the feedback loops are tight, and decisions happen quickly.

If you’ve grown from Linux systems DevOps Staff-level SRE, and you now think in terms of systemic risk, scalability, and long-term reliability strategy - this role gives you direct influence and visibility.

This role is intentionally in-person because:

  • Reliability decisions happen at architectural depth - not over Slack threads * ML, data, and infrastructure teams collaborate continuously in real time * Post-incident reviews, system design debates, and performance tuning sessions are hands-on and high impact * You will have direct access to engineering leadership and decision-makers * The infrastructure you’re operating is mission-critical and evolving quickly

If you value deep technical collaboration, tight feedback loops, and being at the center of high-scale ML systems - this environment is built for that., You’ll partner directly with engineering and data science teams to ensure ML workloads are production-ready and reliable by design.

Requirements

Do you have experience in Systems engineering?, * Deep experience operating Linux infrastructure and networking in production environments * Proven impact as a Staff SRE, Senior SRE, or senior-level DevOps/Platform Engineer supporting distributed systems * Experience supporting complex, data-intensive or ML-driven systems in production * Strong hands-on experience with Docker and Kubernetes * Infrastructure-as-Code expertise * Strong scripting ability (Bash and/or Python) * CI/CD ownership experience (GitHub Actions, ArgoCD, or similar) * Experience with modern observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry) * Ability to debug systemic failures across infrastructure, deployments, and workloads * Clear communicator who works effectively across engineering and data teams

Engineers who have evolved from infrastructure foundations into strategic reliability leaders will thrive here.

These Skills Are a Plus

  • Experience operating ML platforms at scale (training + inference) * AWS or cloud-managed services experience * Exposure to data platforms such as Spark, Airflow, or Kafka * Experience in SOC 2 or regulated environments

Benefits & conditions

Collaborative engineering culture designed for speed and depth * Competitive base compensation ($210K-$250K)

If you’re a Staff-level reliability engineer who wants real ownership and architectural influence - let’s start the conversation.

StratITech is partnering with our San Francisco client to build the next generation of high-scale ML infrastructure.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:00 min

Exploring the specific workplace responsibilities of staff software engineers

Jan Giacomelli · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

Videos

See all

Related articles

See all