> Markdown version of [/jobs/ext/2716732-staff-software-engineer-ai-runtime](https://www.wearedevelopers.com/jobs/ext/2716732-staff-software-engineer-ai-runtime). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, AI Runtime - **Company:** Databricks - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $190,000.0 - $265,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Extract Transform Load (ETL), Data Structures, Software Debugging, Distributed Computing Environment, Distributed Systems, InfiniBand, Dynamic Routing, High Performance Computing, Pytorch, Information Technology, Machine Learning Operations, Databricks - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/staff-software-engineer-ai-runtime-databricks-8016184 ## About the Role * 10+ years of experience building and operating large-scale distributed systems, with significant depth in GPU training infrastructure, high-performance computing, or ML systems. * Hands-on experience with distributed training frameworks (such as PyTorch, FSDP, DeepSpeed, or Megatron) and the parallelism strategies (data, tensor, pipeline, and sequence parallelism) used to train large models. * Strong understanding of training resilience patterns, including checkpointing, failure detection, and automatic recovery for long-running, multi-node jobs. * Solid grasp of GPU performance fundamentals, including accelerator architecture, high-speed interconnects (such as NVLink and InfiniBand or RoCE), collective communication, and the bottlenecks that govern training throughput and utilization. * Experience building and operating managed, multi-tenant platform products in the cloud, with clear SLAs and SLOs for availability, performance, and reliability. * Strong foundation in algorithms, data structures, and system design as applied to performance-sensitive, large-scale distributed systems. * Proven ability to deliver technically complex, high-impact initiatives that create clear customer or business value. * Strong communication skills and the ability to collaborate across product, research, and infrastructure teams in a fast-moving environment. * Strategic, product-oriented mindset with the ability to align technical execution to a long-term vision, and a passion for mentoring engineers and fostering technical excellence. * BS in Computer Science or a related field (MS or PhD preferred). Pay Range Transparency ## Description * Solve the hardest problems in large-scale training, including multi-node orchestration, distributed parallelism strategies, GPU scheduling and dynamic routing, high-throughput data loading, and checkpoint and restore for very long-running jobs. * Push GPU efficiency and training performance, raising utilization (such as model FLOPs utilization and end-to-end throughput) and lowering cost per training run across diverse model architectures and hardware generations. * Build the resilience and observability foundations that keep multi-node jobs healthy, detecting and recovering from hardware and software failures with minimal disruption to customers. * Partner with product, research, and platform teams to shape the APIs, CLI, and developer experience that make it easy to launch, monitor, and debug production training jobs. * Lead end-to-end engineering efforts, from design through production rollout, holding a high bar for performance, correctness, and reliability. * Make direct, high-impact contributions to the core systems behind AIR, and help bring up support for the latest accelerators and new regions as the fleet grows. * Champion engineering excellence, mentor other engineers through design reviews and technical discussions, and help shape Databricks' long-term technical direction in AI training infrastructure. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [OLTP in the Lakehouse: Redefining Data for AI Workloads](https://www.wearedevelopers.com/videos/2038-oltp-in-the-lakehouse-redefining-data-for-ai-workloads) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies in Europe](https://www.wearedevelopers.com/magazine/162-highest-paying-tech-companies-in-europe) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Are Software Engineer Wages Being Pushed Down? A Report on Tech Salaries](https://www.wearedevelopers.com/magazine/417-are-software-engineer-wages-being-pushed-down-a-report-on-tech-salaries)