> Markdown version of [/jobs/ext/3276543-distributed-systems-engineer](https://www.wearedevelopers.com/jobs/ext/3276543-distributed-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Distributed Systems Engineer - **Company:** Fluidstack Ltd - **Location:** San Francisco, CA, United States - **Salary:** $208,000.0 - $269,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Dynamic Host Configuration Protocol, Distributed Systems, Domain Name System (DNS), Infrastructure as a Service (IaaS), Kubernetes, Data Pipelines - **Published:** September 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=ca993e4bef9cba3a ## About the Role The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would. ## Description * Make tens of thousands of GPUs legible in real time: build the observability platform that turns raw telemetry into signal, from site-level health down to individual device and link. At this scale, you cannot operate what you cannot see. * Build the control plane every team at Fluidstack depends on: replace one-off tooling with a stable, versioned API surface that covers unified machine management, actual state inspection, and distributed command execution. One interface for the whole company, not a hundred scripts. * Make the system's view of itself always match reality: integrate fleet state as a machine-readable source of truth across provisioning, operations, and customer-facing platforms, so every new site and GPU generation lands cleanly from day zero., * Own the observability platform. Build and operate the data pipelines, decoration and correlation engine, and healthcheck framework that make the fleet legible - from site down to device and link. No other team should need to scrape production directly to answer a question. * Define and build the API surface for infrastructure. Design the contracts between production infrastructure and every tool that touches it. All other teams at Fluidstack use your tooling to manage and operate our hyperscale fleet. * Build the production control plane. Unified machine management, actual state inspection, distributed command execution - and the Kubernetes-based infrastructure that underpins it all. * Own fleet state as source of truth. SLOs, site lifecycle state, and integration with internal infrastructure management and customer-facing operations platforms. What the system says about itself should match reality, and you're accountable when it doesn't. * Land new hardware into the platform cleanly. ZTP, DHCP, DNS, artifacts - every new XPU generation and site integration goes through IaaS before production. ## Related Videos - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [Best Practices for AI-Assisted Development of Distributed Systems](https://www.wearedevelopers.com/videos/100200-best-practices-for-ai-assisted-development-of-distributed-systems) - [The Future of Cloud is Abstraction - Why Kubernetes is not the Endgame for STACKIT ](https://www.wearedevelopers.com/videos/413-the-future-of-cloud-is-abstraction-why-kubernetes-is-not-the-endgame-for-stackit) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Throwing off the burdens of scale in engineering](https://www.wearedevelopers.com/videos/608-throwing-off-the-burdens-of-scale-in-engineering) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers)