> Markdown version of [/jobs/ext/2181988-software-engineer-load-and-fault-environments](https://www.wearedevelopers.com/jobs/ext/2181988-software-engineer-load-and-fault-environments). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Load and Fault Environments - **Company:** RPEV INFRASTRUCTURE DEVELOPMENT, LLC - **Location:** Menlo Park, CA, United States - **Experience:** Expert - **Salary:** $196,000.0 - $230,000.0 - **Contract:** Permanent contract - **Skills:** Distributed Systems, Python (Programming Language), Load Testing, Microsoft Office, Reliability Engineering, Software Engineering, Build Management, Kubernetes, Production Code, Software Coding - **Published:** August 22, 2026 - **Apply:** https://www.dice.com/job-detail/55ebc1b4-ba58-4348-bbb1-3c9dc7d0277c ## About the Role * 5+ years of software engineering experience with strong coding skills in Python, Go, or a similar language - this role requires building production systems, not just operating them. * Experience designing or contributing to load testing, fault injection, chaos engineering, or developer productivity infrastructure in a distributed systems environment. * Solid understanding of distributed systems fundamentals - how services fail, how failures propagate, and how to design experiments that surface meaningful signal without causing unintended outages. * Familiarity with Kubernetes and container-based infrastructure; experience integrating resilience tooling into CI/CD pipelines is a plus. * A builder's approach to developer tooling - you care about ergonomics, adoption, and making it easy for other engineers to do the right thing. ## Description As a Senior Software Engineer on the Load and Fault Environments team, you will design and build the infrastructure that lets Robinhood's engineers simulate load, inject faults, and validate system behavior under stress - at the scale of a fast-growing financial platform. You'll own meaningful components of the load testing and fault injection platform, write production-quality code, and collaborate with engineers across infrastructure and product teams to ensure the tooling you build gets adopted and drives real reliability improvements. This is a high-impact individual contributor role where your work directly shapes how Robinhood builds resilient systems! This role is based in our Menlo Park, CA office, with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. What you'll do * Design, build, and maintain load testing and fault injection infrastructure that enables engineering teams across Robinhood to validate system resilience as part of their standard development workflow. * Own platform components end-to-end - from architecture and implementation through testing, deployment, and iteration - ensuring the tools you build are reliable, scalable, and developer-friendly. * Collaborate with infrastructure and product engineering teams to understand their resilience testing needs, translate those needs into platform capabilities, and drive adoption of shared tooling. * Contribute to defining steady-state hypotheses, blast radius controls, and fault injection primitives that make chaos experiments safe, repeatable, and actionable at scale. * Write clean, well-tested code and participate in technical design discussions that raise the engineering bar across the team. ## Related Videos - [Answering the Million Dollar Question: Why did I Break Production?](https://www.wearedevelopers.com/videos/1171-answering-the-million-dollar-question-why-did-i-break-production) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Load Testing AI: Aiming at a Moving Target](https://www.wearedevelopers.com/videos/1914-load-testing-ai-aiming-at-a-moving-target) - [Designing the Future of Human<>Agent Collaboration](https://www.wearedevelopers.com/videos/1447-designing-the-future-of-human-agent-collaboration) - [Platform Engineering vs. DevOps Why not both?](https://www.wearedevelopers.com/videos/885-platform-engineering-vs-devops-why-not-both) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Hard?](https://www.wearedevelopers.com/magazine/448-is-software-engineering-hard)