> Markdown version of [/jobs/ext/503857-senior-software-engineer-ai-runtime](https://www.wearedevelopers.com/jobs/ext/503857-senior-software-engineer-ai-runtime). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, AI Runtime - **Company:** Databricks - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Salary:** $160,000.0 - $225,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Extract Transform Load (ETL), Data Structures, Software Debugging, Distributed Computing Environment, Distributed Systems, InfiniBand, Dynamic Routing, High Performance Computing, Pytorch, Information Technology, Machine Learning Operations, Databricks - **Published:** June 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f8bf14ac7a28f7f3 ## About the Role Do you have experience in Scalable systems?, Do you have a Bachelor's degree?, * 5+ years of experience building and operating large-scale distributed systems, with experience in GPU training infrastructure, high-performance computing, or ML systems. * Experience with distributed training frameworks (such as PyTorch, FSDP, DeepSpeed, or Megatron) and the parallelism strategies (data, tensor, pipeline, and sequence parallelism) used to train large models. * Strong understanding of training resilience patterns, including checkpointing, failure detection, and automatic recovery for long-running, multi-node jobs. * Solid grasp of GPU performance fundamentals, including accelerator architecture, high-speed interconnects (such as NVLink and InfiniBand or RoCE), collective communication, and the bottlenecks that govern training throughput and utilization. * Experience building and operating managed, multi-tenant platform products in the cloud, with clear SLAs and SLOs for availability, performance, and reliability. * Strong foundation in algorithms, data structures, and system design as applied to performance-sensitive, large-scale distributed systems. * Proven ability to deliver technically complex, high-impact initiatives that create clear customer or business value. * Strong communication skills and the ability to collaborate across product, research, and infrastructure teams in a fast-moving environment. * Customer-focused mindset with the ability to align implementation details with product goals, and a passion for mentoring engineers and fostering technical excellence. * BS in Computer Science or a related field (MS or PhD preferred). Pay Range Transparency ## Description * Solve the hardest problems in large-scale training, including multi-node orchestration, distributed parallelism strategies, GPU scheduling and dynamic routing, high-throughput data loading, and checkpoint and restore for very long-running jobs. * Push GPU efficiency and training performance, raising utilization (such as model FLOPs utilization and end-to-end throughput) and lowering cost per training run across diverse model architectures and hardware generations. * Build the resilience and observability foundations that keep multi-node jobs healthy, detecting and recovering from hardware and software failures with minimal disruption to customers. * Partner with product, research, and platform teams to shape the APIs, CLI, and developer experience that make it easy to launch, monitor, and debug production training jobs. * Lead end-to-end engineering efforts, from design through production rollout, holding a high bar for performance, correctness, and reliability. * Make direct, high-impact contributions to the core systems behind AIR, and help bring up support for the latest accelerators and new regions as the fleet grows. * Champion engineering excellence, mentor other engineers through design reviews and technical discussions, and contribute to Databricks' technical direction in AI training infrastructure. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [OLTP in the Lakehouse: Redefining Data for AI Workloads](https://www.wearedevelopers.com/videos/2038-oltp-in-the-lakehouse-redefining-data-for-ai-workloads) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)