> Markdown version of [/jobs/ext/3077375-staff-software-engineer](https://www.wearedevelopers.com/jobs/ext/3077375-staff-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer - **Company:** Hedra, Inc - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Continuous Integration, Software Debugging, Distributed Systems, Python (Programming Language), Performance Tuning, Software Engineering, System Testing, Strategies of Testing, Backend, Kubernetes, Low Latency, Bare Metal, Build Tools, Machine Learning Operations - **Published:** September 25, 2026 - **Apply:** https://startup.jobs/senior-staff-software-engineer-distributed-systems-hedra-8120883 ## About the Role * 5+ years of software engineering experience, with significant experience building distributed backend or infrastructure systems. * A track record of designing and operating high-availability production services at meaningful scale. * Strong distributed systems fundamentals, including experience reasoning about concurrency, queues, retries, failure recovery, consistency, backpressure, capacity, and system behavior under load. * Experience debugging complex production systems across multiple layers of the stack. * Strong judgment around architectural tradeoffs, particularly reliability, performance, complexity, and operational cost. * Experience with CI/CD, comprehensive testing and validation strategies, monitoring, alerting, and production operations. * Strong programming fundamentals and experience working in systems-oriented backend languages. Python experience is helpful given our stack. * Experience using agentic coding tools and workflows beyond basic prompting. * Comfort operating in an environment where problems are often underspecified and engineers are expected to independently determine the right approach. * Ability to communicate technical decisions clearly and collaborate with engineers across infrastructure, research, and product. Nice to Have * Experience with inference or model-serving infrastructure. * GPU or accelerator infrastructure. * Compute scheduling, orchestration, or resource management. * High-throughput or low-latency systems. * Kubernetes or other cluster orchestration systems. * Performance optimization and profiling. * Developer infrastructure, APIs, SDKs, or platform engineering. * Experience operating infrastructure across cloud and/or bare-metal environments. ## Description We're looking for a Senior or Staff Software Engineer with deep experience building and operating distributed production systems. You'll work on the infrastructure underlying Hedra's inference platform: systems that schedule and route compute, serve models efficiently, handle high-throughput workloads, and remain reliable as both traffic and the number of models we support grow. The problems are often ambiguous and don't have obvious answers. We're looking for someone who can reason from first principles, identify bottlenecks and failure modes before they become problems, and make thoughtful tradeoffs across performance, reliability, complexity, and cost. You do not need to come from an AI company or already be an expert in model inference. We care much more about depth in distributed systems and your ability to apply that experience to a new problem space. What You'll Do * Design, build, and operate distributed systems that power Hedra's inference infrastructure. * Build systems for scheduling, routing, and managing compute-intensive workloads across heterogeneous resources. * Improve throughput, latency, reliability, and resource utilization across our serving infrastructure. * Design systems that remain predictable and resilient under load, partial failures, changing capacity, and unpredictable workloads. * Own production systems end to end, including deployment, CI/CD, testing and validation, observability, monitoring, alerting, debugging, and incident response. * Identify architectural bottlenecks and failure modes and drive solutions rather than waiting for problems to be fully specified. * Work across infrastructure, model serving, APIs, and developer-facing systems as the platform evolves. * Use modern agentic coding workflows as part of how you design, build, debug, and ship software. * Help shape the technical direction of a small engineering organization where individual engineers have substantial ownership. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Single Server, Global Reach: Running a Worldwide Marketplace on Bare Metal in a Cloud-Dominated World](https://www.wearedevelopers.com/videos/1206-single-server-global-reach-running-a-worldwide-marketplace-on-bare-metal-in-a-cloud-dominated-world) - [Scaling: from 0 to 20 million users](https://www.wearedevelopers.com/videos/676-scaling-from-0-to-20-million-users) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)