> Markdown version of [/jobs/ext/542979-staff-software-engineer-focused-on-reliability](https://www.wearedevelopers.com/jobs/ext/542979-staff-software-engineer-focused-on-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer focused on Reliability - **Company:** VANHACK TECHNOLOGIES INC. - **Location:** New York, NY, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Profiling, Databases, Data Systems, Programming Tools, Disaster Recovery, Distributed Systems, Domain Name System (DNS), Fault Tolerance, Java Virtual Machine (JVM), NoSQL, Open Source Technology, Performance Tuning, Query Optimization, Reliability Engineering, Site Reliability Engineering Practices, Software Engineering, Datadog, Data Logging, Load Balancing, GitHub Copilot, System Availability, Concurrency, Mttr, Deployment Automation, Database Replication, Vulnerability Analysis - **Published:** June 10, 2026 - **Apply:** https://vanhack.com/job/11935 ## About the Role * 10+ years of engineering experience in software engineering, reliability engineering, SRE practices, or production operations at scale * Expert-level reliability engineering: multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery * Production observability at scale - deep experience with monitoring, alerting, tracing, and logging; Datadog or similar APM in high-load environments * Strong systems thinking - design resilient distributed systems that handle failures, network partitions, and external dependency outages * Database and data systems knowledge: replication strategies, backup/restore, connection pooling, query optimization; relational and NoSQL experience * AWS production experience: multi-region deployments, load balancing, DNS-based failover * Experience with AI-powered development tools (e.g. GitHub Copilot or similar agentic coding tools) * Expert-level Java and/or Scala - JVM performance, concurrency, and operational characteristics * Strong technical communication; ability to influence architecture across teams, document complex systems, run post-mortems, and establish org-wide reliability standards Preferred * Scala experience * SRE or Reliability Engineering experience at companies known for operational excellence (e.g. large-scale tech companies or high-growth startups where you built reliability practices from the ground up) * Incident response leadership: incident management processes, blameless post-mortems, MTTR reduction in production * Chaos engineering with tools like Chaos Monkey, Gremlin, or similar - including game days and failure injection testing * Performance optimization: profiling, benchmarking, capacity planning, and system tuning at hyperscale * Open source contributions or technical writing demonstrating depth in reliability engineering, distributed systems, or production operations Ideal candidate * Builds reliability from first principles * Works alongside highly technical teams to influence architecture and establish company-wide reliability standards * Excellent technical communication - documents complex systems, conducts post-mortems, and drives reliability standards organization-wide ## Description A Staff Software Engineer focused on Reliability is needed to own reliability across the entire platform and drive the practices that ensure system availability, resilience, and observability for mission-critical infrastructure. You will build reliability from first principles: architecting failover systems, implementing chaos engineering, and improving the observability foundation to maintain 99.9%+ uptime as the company scales into new markets. As the technical owner of the reliability posture, you will tackle challenges like external service failover, dependency mirroring, and database replication - working alongside highly technical teams across the organization to influence architecture decisions and establish company-wide reliability standards. This role sits on the Product Foundations team, building the foundational infrastructure that powers a large-scale mobility and commerce platform. Tech challenge * Maintain 99.9%+ uptime as the platform scales to new markets * External service failover, dependency mirroring, and database replication at production scale, * Own the overall reliability posture for the platform - practices, metrics, and systems that ensure 99.9%+ uptime across all services * Design and implement automatic failover for critical external dependencies (e.g. SMS/voice and payments providers) with circuit breakers, retry policies, and degraded-mode operations * Architect and build active-passive or active-active regional deployment strategies with database replication, automated failover, and DNS-based traffic routing - including disaster recovery planning and testing * Establish comprehensive monitoring using Datadog (or equivalent) for APM, logs, and metrics correlation * Implement synthetic monitoring, SLO-based alerting, on-call rotation, and escalation policies; build service health dashboards that show customer impact * Own the incident management process - workflows, tooling, post-mortem culture, runbook automation, and MTTR reduction from detection to resolution * Drive adoption of resilience patterns across services: health checks, graceful degradation, feature flags, rate limiting, backpressure, and chaos engineering * Build and maintain local mirrors for critical dependencies - artifact caching, dependency pinning, and vulnerability scanning to prevent build failures from upstream outages ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Psychological Safety in Software Engineering - Jenny-Margrethe Vej & Alexandra Hou Aldershaab](https://www.wearedevelopers.com/videos/2142-psychological-safety-in-software-engineering-jenny-margrethe-vej-alexandra-hou-aldershaab) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)