> Markdown version of [/jobs/ext/1247259-sr-staff-production-engineer-data-platform](https://www.wearedevelopers.com/jobs/ext/1247259-sr-staff-production-engineer-data-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Staff Production Engineer - Data Platform - **Company:** Databricks - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Salary:** $228,600.0 - $314,250.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Data Analysis, Data Infrastructure, Distributed Data Store, Distributed Systems, Python (Programming Language), Reliability Engineering, Site Reliability Engineering Practices, Sysadmin, Delivery Pipeline, Large Language Models, Multi-Cloud, Information Technology, Databricks - **Published:** July 12, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=dbde16bf5fc21979 ## About the Role * BS/MS/PhD in Computer Science, or a related field * Technical Depth: 10+ years of production-level experience as a Software Engineer or SRE in highly distributed, multi-cloud environments. * Engineering Persona: You write code to solve operational problems. You are not a traditional sys-admin; you build frameworks, automation, and tooling (Scala, Java, Go, or Python) to eliminate toil. * Platform & AI Mindset: Deep understanding of distributed data platforms and a passion for leveraging AI/ML to revolutionize infrastructure management. Familiarity with LLM infrastructure, training/inference pipelines, or agentic frameworks is a significant plus. * Operational Grit: Proven ability to remain calm and decisive under pressure. You have navigated large-scale distributed systems through hyper-growth and have a track record of driving incident-to-roadmap loops. * Strategic Influence: Experience building long-range technical roadmaps and driving cross-functional alignment. You are comfortable challenging senior leadership with data-driven insights. Pay Range Transparency ## Description As a Sr. Staff Production Engineer, you will lead the strategic vision for the operational stability of our internal "Databricks-on-Databricks" environment. You will transition our infrastructure from traditional SRE models toward an agent-driven, self-healing architecture, ensuring that our platform-and the agents operating within it-remain rock-solid for mission-critical customer workloads., * Architecting Agentic Reliability: Define and drive the design of future "self-healing" infrastructure at scale where AI agents proactively detect, diagnose, and remediate production incidents before they impact customers. * Data Platform Optimization: Own the operational integrity of the Data Platform that powers our internal AI models, ensuring 99.99% availability for the compute, storage, and control plane services used by thousands of Databricks engineers. * High-Scale Operational Excellence: Establish the next generation of "Change Safety" protocols, utilizing automation and agentic guardrails to manage complex deployments across 100+ global regions. * Leadership in Chaos & Scale: Serve as a technical bar-raiser for the team, evangelizing modern SRE practices (including Chaos Engineering) to navigate the structural transformation of the industry toward agentic, autonomous systems. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Hacking MSSQL on Cloud. All of them. How I became sysadmin on Azure, AWS, GCP and Alibaba.](https://www.wearedevelopers.com/videos/100339-hacking-mssql-on-cloud-all-of-them-how-i-became-sysadmin-on-azure-aws-gcp-and-alibaba) - [The Open-source Java SDK for Multi-Cloud Development - Sandeep Pal](https://www.wearedevelopers.com/videos/2113-the-open-source-java-sdk-for-multi-cloud-development-sandeep-pal) - [OLTP in the Lakehouse: Redefining Data for AI Workloads](https://www.wearedevelopers.com/videos/2038-oltp-in-the-lakehouse-redefining-data-for-ai-workloads) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies in Europe](https://www.wearedevelopers.com/magazine/162-highest-paying-tech-companies-in-europe)