Staff Software Engineer focused on Reliability

VANHACK TECHNOLOGIES INC.
New York, NY, United States
3 months ago
Apply on vanhack.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Artificial Intelligence Amazon Web Services Profiling Databases Data Systems Programming Tools Disaster Recovery Distributed Systems Domain Name System (DNS) Fault Tolerance Java Virtual Machine (JVM)
+17 more
NoSQL Open Source Technology Performance Tuning Query Optimization Reliability Engineering Site Reliability Engineering Practices Software Engineering Datadog Data Logging Load Balancing GitHub Copilot System Availability Concurrency Mttr Deployment Automation Database Replication Vulnerability Analysis

Job description

A Staff Software Engineer focused on Reliability is needed to own reliability across the entire platform and drive the practices that ensure system availability, resilience, and observability for mission-critical infrastructure.

You will build reliability from first principles: architecting failover systems, implementing chaos engineering, and improving the observability foundation to maintain 99.9%+ uptime as the company scales into new markets.

As the technical owner of the reliability posture, you will tackle challenges like external service failover, dependency mirroring, and database replication - working alongside highly technical teams across the organization to influence architecture decisions and establish company-wide reliability standards.

This role sits on the Product Foundations team, building the foundational infrastructure that powers a large-scale mobility and commerce platform.

Tech challenge

  • Maintain 99.9%+ uptime as the platform scales to new markets
  • External service failover, dependency mirroring, and database replication at production scale, * Own the overall reliability posture for the platform - practices, metrics, and systems that ensure 99.9%+ uptime across all services
  • Design and implement automatic failover for critical external dependencies (e.g. SMS/voice and payments providers) with circuit breakers, retry policies, and degraded-mode operations
  • Architect and build active-passive or active-active regional deployment strategies with database replication, automated failover, and DNS-based traffic routing - including disaster recovery planning and testing
  • Establish comprehensive monitoring using Datadog (or equivalent) for APM, logs, and metrics correlation
  • Implement synthetic monitoring, SLO-based alerting, on-call rotation, and escalation policies; build service health dashboards that show customer impact
  • Own the incident management process - workflows, tooling, post-mortem culture, runbook automation, and MTTR reduction from detection to resolution
  • Drive adoption of resilience patterns across services: health checks, graceful degradation, feature flags, rate limiting, backpressure, and chaos engineering
  • Build and maintain local mirrors for critical dependencies - artifact caching, dependency pinning, and vulnerability scanning to prevent build failures from upstream outages

Requirements

  • 10+ years of engineering experience in software engineering, reliability engineering, SRE practices, or production operations at scale
  • Expert-level reliability engineering: multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery
  • Production observability at scale - deep experience with monitoring, alerting, tracing, and logging; Datadog or similar APM in high-load environments
  • Strong systems thinking - design resilient distributed systems that handle failures, network partitions, and external dependency outages
  • Database and data systems knowledge: replication strategies, backup/restore, connection pooling, query optimization; relational and NoSQL experience
  • AWS production experience: multi-region deployments, load balancing, DNS-based failover
  • Experience with AI-powered development tools (e.g. GitHub Copilot or similar agentic coding tools)
  • Expert-level Java and/or Scala - JVM performance, concurrency, and operational characteristics
  • Strong technical communication; ability to influence architecture across teams, document complex systems, run post-mortems, and establish org-wide reliability standards

Preferred

  • Scala experience
  • SRE or Reliability Engineering experience at companies known for operational excellence (e.g. large-scale tech companies or high-growth startups where you built reliability practices from the ground up)
  • Incident response leadership: incident management processes, blameless post-mortems, MTTR reduction in production
  • Chaos engineering with tools like Chaos Monkey, Gremlin, or similar - including game days and failure injection testing
  • Performance optimization: profiling, benchmarking, capacity planning, and system tuning at hyperscale
  • Open source contributions or technical writing demonstrating depth in reliability engineering, distributed systems, or production operations

Ideal candidate

  • Builds reliability from first principles
  • Works alongside highly technical teams to influence architecture and establish company-wide reliability standards
  • Excellent technical communication - documents complex systems, conducts post-mortems, and drives reliability standards organization-wide

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on vanhack.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · World Congress 2025

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

Videos

See all

Related articles

See all