Staff Site Reliability Engineer
LS&CO
Spain
2 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on levistraussandco.wd5.myworkdayjobs.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
10 years minimum
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Audit Trail
Microsoft Azure
Bash Shell
Big Data
BigQuery
Cloud Computing
Configuration Management
Code Review
Computer Networks
Information Engineering
Data Governance
+33 more
Data Security
Data Warehousing
DevOps
Data Flow Control
Identity and Access Management
Python (Programming Language)
Key Management
Reliability Engineering
Prometheus
Software Engineering
Azure Service Bus
Datadog
Data Logging
Scripting
Azure Data Factory
Cloud Monitoring
Istio
Large Language Models
Grafana
Multi-Agent Systems
Mttr
Multi-Cloud
AI Platforms
Kubernetes
Infrastructure Automation Frameworks
Information Technology
Data Lineage
Deployment Automation
Google Cloud Functions
Azure AKS
Data Management
Terraform
Dynatrace
Job description
- Define, instrument, and enforce SLOs, SLIs, and error budgets across all platform services, ensuring alignment with business and product commitments
- Drive continuous reduction in MTTD and MTTR through improved observability, automated alerting, and runbook-driven incident response
- Lead blameless post-mortems and translate findings into durable reliability improvements, ensuring systemic issues are eliminated rather than patched
Toil Reduction & Automation
- Systematically identify, measure, and eliminate operational toil; track toil percentage per sprint and enforce guardrails to keep it below 50% of engineering capacity
- Build and maintain self-serve infrastructure capabilities - enabling product and data engineering teams to provision, scale, and operate their own resources safely and consistently
- Automate deployment pipelines, configuration management, and operational workflows using Infrastructure-as-Code principles (Terraform, Helm, GitOps)
Platform Engineering & Architecture
- Serve as the primary GCP subject matter expert - architecting and optimizing workloads across GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI
- Lead multi-cloud architecture decisions across GCP and Azure, ensuring consistent security posture, cost efficiency, and operational practices across environments
- Design and implement self-healing infrastructure patterns, auto-scaling strategies, and capacity planning models to support high-availability data and AI platforms
- Champion data security and governance best practices - including encryption at rest and in transit, IAM least-privilege, secrets management, and audit logging
AI, Agentic Systems & Modern Observability
- Apply SRE principles to agentic AI workloads - defining reliability expectations for LLM-based and multi-agent systems, including latency SLOs, fallback patterns, and model observability
- Partner with AI Platform teams to productionize agentic pipelines with robust monitoring, drift detection, and rollback capabilities
- Drive adoption of AI-assisted operations tooling to enhance observability, anomaly detection, and predictive incident management
Leadership & Culture
- Guide and mentor junior and mid-level SREs - conducting code reviews, running reliability reviews, and elevating the team’s engineering craft
- Collaborate cross-functionally with Data Engineering, Software Engineering, Security, and Product teams to embed reliability as a shared value from design through deployment
- Champion a culture of psychological safety, continuous learning, and reliability excellence
- Communicate platform health, risk posture, and reliability roadmaps clearly to both technical and executive audiences
Requirements
- Master’s degree in Computer Science, Engineering, or related field (or equivalent practical experience)
- 10+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering, with a strong track record in large-scale production environments
- Deep, hands-on expertise in GCP - including GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI
- Proficiency with Infrastructure-as-Code tools (Terraform, Helm) and GitOps workflows (ArgoCD, Flux)
- Strong command of observability tooling - distributed tracing, structured logging, metrics pipelines, and alerting platforms (e.g., Cloud Monitoring, Datadog, Prometheus/Grafana)
- Proven experience defining and operating against SLOs, SLIs, and error budgets in production environments
- Solid understanding of data security principles: IAM, encryption, secrets management, network policies, and compliance frameworks
- Experience with multi-cloud environments (GCP + Azure), including cross-cloud networking, identity federation, and cost governance
- Demonstrated ability to lead without authority - influencing engineers across teams and driving reliability improvements at the organizational level
- Excellent written and verbal communication skills; ability to translate complex reliability concepts for non-technical stakeholders
Technical Depth
- Fluency in at least one systems or scripting language (Python, Go, or Bash) for automation and tooling
- Experience with container orchestration (Kubernetes/GKE), service mesh, and traffic management patterns
- Familiarity with data engineering patterns: batch and streaming pipelines, data warehouses, and the operational challenges of large-scale data platforms
- Understanding of agentic AI architectures and the unique reliability challenges of LLM-based, event-driven, and multi-agent systems
- Working knowledge of data governance frameworks, data lineage tooling, and platform-level data quality enforcement
Desirable Experience
- Experience operating data platforms in retail or e-commerce environments
- Familiarity with SRE principles in practice - error budget policies, CRE engagements, production readiness reviews
- Exposure to FinOps practices - cloud cost attribution, commitment optimization, and unit economics for data workloads
- Experience with Azure-native services (AKS, Azure Data Factory, Event Hubs) and cross-cloud identity management
- Prior experience in a Staff or Principal-level SRE role with organization-wide scope
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on levistraussandco.wd5.myworkdayjobs.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
IK
Igor Khokhriakov
about 1 month ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
over 2 years ago
AH
Aishi Huang
Trustworthy AI Starts at Deployment: 5 Checks Before You Ship
about 1 month ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
over 2 years ago
EM
Eli McGarvie
Software Engineer Salary London
over 3 years ago