> Markdown version of [/jobs/ext/2814960-staff-systems-software-engineer](https://www.wearedevelopers.com/jobs/ext/2814960-staff-systems-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Systems Software Engineer - **Company:** REMOTE USA, LLC - **Location:** Los Angeles, CA, United States (Remote available) - **Salary:** $162,500.0 - $219,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Test Suite, Artificial Intelligence, Automation of Tests, C++ (Programming Language), Software as a Service, Cloud Computing, Computer Clusters, CMake, Code Review, Linux, Programming Tools, Disaster Recovery, Distributed Systems, Game Engine, Groovy, Apache JMeter, Python (Programming Language), Load Testing, Regression Analysis, Platform as a Service (PAAS), Perforce, Regression Testing, Reliability Engineering, Ansible, Shell Script, TCP/IP, Strategies of Testing, Toolchain, TypeScript, Mttr, Gatling, Multi-Cloud, Caching, Backend, Gitlab, Git, Kotlin, Gitlab-ci, Kubernetes, Bare Metal, Build Tools, Slurm, Terraform, Device Compatibility, Docker, Jenkins, Golang - **Published:** September 10, 2026 - **Apply:** https://startup.jobs/staff-systems-software-engineer-thatgamecompany-9975809 ## About the Role * Languages: Python, Golang, C/C++, Jenkins Groovy; Linux shell scripting; fluent English * Testing & Quality Systems: Test framework design; load/stress tooling (k6, Locust, Gatling, JMeter); fault injection/chaos engineering; coverage instrumentation; flake detection * Performance & Analysis: CPU/GPU/memory/I/O profiling; statistical analysis of noisy benchmarks; regression detection and commit attribution; observability * Build & CI at Scale: Jenkins/GitLab CI/CD; Make/CMake; GCC/Clang; distributed build systems/caching; Git/GitLab and Perforce (large-binary workflows) * Infrastructure: Linux/macOS administration; bare-metal/on-prem cluster ops; Docker/Podman; Kubernetes; Terraform/OpenTofu; hardware capacity planning * Systems Fundamentals: Distributed systems failure modes; TCP/IP and UDP; resilience patterns; disaster recovery planning and execution Preferred Skills Mobile device farm operation, Android/iOS build and profiling toolchains, GPU cluster scheduling (Slurm, K8s device plugins, MIG), game engine internals/engine-level test automation, deterministic simulation/replay testing, live service game experience (or SaaS/PaaS/XaaS equivalent), Windows administration, Ansible, TypeScript/Java/Kotlin/Swift, hybrid/multi-cloud management ## Description Most engineering roles at Thatgamecompany ship features to players. This one ships velocity and quality to engineers. Cloud, Gameplay Backend, and Cross-Platform Backend build the services and features the game runs on; you build the systems that prove that work is correct, fast, and ready - before it reaches players. You will build the systems that prove the gameplay feature works - at scale, on every device, on every commit - and the infrastructure that lets a hundred engineers find that out in minutes instead of days. If it runs in production, another team owns it. If it exists to test, measure, or accelerate what those teams build - test harnesses, load/chaos infrastructure, device labs, build/GPU clusters - you own it. Your output is the studio's velocity and the game's quality floor. On any given day, you might: Prove correctness and quality before players find the bugs * Build and own the studio's functional test automation framework - runner, fixtures, game-state harness, reporting - plus coverage instrumentation that makes untested systems visible * Build regression analysis tooling that separates real regressions from noise and attributes them to a commit automatically * Improve test suite speed/trust via flake detection, quarantine, parallel sharding, deterministic replay Find the breaking point on purpose * Design load/stress testing systems modeling realistic player behavior, capable of driving launch-scale synthetic traffic against pre-production * Build chaos testing capability - fault injection, latency/partition simulation, dependency failure - run safely against real infrastructure * Run disaster recovery exercises, measure actual RTO/RPO vs. stated targets, and turn findings into permanent regression tests Prove quality on real devices * Build app quality testing automation - frame time, memory, thermal, battery, cold-start captured automatically and gated in CI * Own the device compatibility lab, orchestration, and certification matrix so device support is verified every commit, not the week before ship Own the infrastructure engineers build on * Build/operate on-prem build clusters (distributed compilation, caching, artifacts) and on-prem GPU clusters (scheduling, multi-tenancy, utilization), backing buy-vs-build with real data * Design cross-platform developer tools - including AI-assisted tooling (agents, skills, MCP servers) - that reduce toil across Cloud, Backend, and Client Engineering Make it land * Define, instrument, and publish developer velocity/quality metrics (build time, test cycle time, flake rate, escaped defect rate, MTTR) and use them to prioritize your roadmap * Drive adoption across teams that don't report to you * Mentor engineers in testing strategy, performance analysis, systems thinking * Raise the bar through code review, documentation, reducing bus factor We expect you to * Be a passionate gamer who puts players first * Treat developer experience as a product - users, adoption curves, a quality bar - not internal plumbing * Operate at Staff level: set technical direction studio-wide, influence without authority * Be relentless about systems that prevent problems, not just report them * Embrace experimentation for ambiguous, complex problems * Communicate trade-offs clearly across distributed teams ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The 8 Best Code Testing Tools](https://www.wearedevelopers.com/magazine/402-the-8-best-code-testing-tools) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)