Staff Software Engineer focused on Reliability
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+17 more
Job description
A Staff Software Engineer focused on Reliability is needed to own reliability across the entire platform and drive the practices that ensure system availability, resilience, and observability for mission-critical infrastructure.
You will build reliability from first principles: architecting failover systems, implementing chaos engineering, and improving the observability foundation to maintain 99.9%+ uptime as the company scales into new markets.
As the technical owner of the reliability posture, you will tackle challenges like external service failover, dependency mirroring, and database replication - working alongside highly technical teams across the organization to influence architecture decisions and establish company-wide reliability standards.
This role sits on the Product Foundations team, building the foundational infrastructure that powers a large-scale mobility and commerce platform.
Tech challenge
- Maintain 99.9%+ uptime as the platform scales to new markets
- External service failover, dependency mirroring, and database replication at production scale, * Own the overall reliability posture for the platform - practices, metrics, and systems that ensure 99.9%+ uptime across all services
- Design and implement automatic failover for critical external dependencies (e.g. SMS/voice and payments providers) with circuit breakers, retry policies, and degraded-mode operations
- Architect and build active-passive or active-active regional deployment strategies with database replication, automated failover, and DNS-based traffic routing - including disaster recovery planning and testing
- Establish comprehensive monitoring using Datadog (or equivalent) for APM, logs, and metrics correlation
- Implement synthetic monitoring, SLO-based alerting, on-call rotation, and escalation policies; build service health dashboards that show customer impact
- Own the incident management process - workflows, tooling, post-mortem culture, runbook automation, and MTTR reduction from detection to resolution
- Drive adoption of resilience patterns across services: health checks, graceful degradation, feature flags, rate limiting, backpressure, and chaos engineering
- Build and maintain local mirrors for critical dependencies - artifact caching, dependency pinning, and vulnerability scanning to prevent build failures from upstream outages
Requirements
- 10+ years of engineering experience in software engineering, reliability engineering, SRE practices, or production operations at scale
- Expert-level reliability engineering: multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery
- Production observability at scale - deep experience with monitoring, alerting, tracing, and logging; Datadog or similar APM in high-load environments
- Strong systems thinking - design resilient distributed systems that handle failures, network partitions, and external dependency outages
- Database and data systems knowledge: replication strategies, backup/restore, connection pooling, query optimization; relational and NoSQL experience
- AWS production experience: multi-region deployments, load balancing, DNS-based failover
- Experience with AI-powered development tools (e.g. GitHub Copilot or similar agentic coding tools)
- Expert-level Java and/or Scala - JVM performance, concurrency, and operational characteristics
- Strong technical communication; ability to influence architecture across teams, document complex systems, run post-mortems, and establish org-wide reliability standards
Preferred
- Scala experience
- SRE or Reliability Engineering experience at companies known for operational excellence (e.g. large-scale tech companies or high-growth startups where you built reliability practices from the ground up)
- Incident response leadership: incident management processes, blameless post-mortems, MTTR reduction in production
- Chaos engineering with tools like Chaos Monkey, Gremlin, or similar - including game days and failure injection testing
- Performance optimization: profiling, benchmarking, capacity planning, and system tuning at hyperscale
- Open source contributions or technical writing demonstrating depth in reliability engineering, distributed systems, or production operations
Ideal candidate
- Builds reliability from first principles
- Works alongside highly technical teams to influence architecture and establish company-wide reliability standards
- Excellent technical communication - documents complex systems, conducts post-mortems, and drives reliability standards organization-wide
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Résumé-Driven Development: How IT trends affect the job market for software developers
Dev Digest 120 - Apple and peers
Highest Paying Tech Companies for Developers