Sr. Staff Production Engineer - Data Platform
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
As a Sr. Staff Production Engineer, you will lead the strategic vision for the operational stability of our internal āDatabricks-on-Databricksā environment. You will transition our infrastructure from traditional SRE models toward an agent-driven, self-healing architecture, ensuring that our platform-and the agents operating within it-remain rock-solid for mission-critical customer workloads., * Architecting Agentic Reliability: Define and drive the design of future āself-healingā infrastructure at scale where AI agents proactively detect, diagnose, and remediate production incidents before they impact customers.
- Data Platform Optimization: Own the operational integrity of the Data Platform that powers our internal AI models, ensuring 99.99% availability for the compute, storage, and control plane services used by thousands of Databricks engineers.
- High-Scale Operational Excellence: Establish the next generation of āChange Safetyā protocols, utilizing automation and agentic guardrails to manage complex deployments across 100+ global regions.
- Leadership in Chaos & Scale: Serve as a technical bar-raiser for the team, evangelizing modern SRE practices (including Chaos Engineering) to navigate the structural transformation of the industry toward agentic, autonomous systems.
Requirements
- BS/MS/PhD in Computer Science, or a related field
- Technical Depth: 10+ years of production-level experience as a Software Engineer or SRE in highly distributed, multi-cloud environments.
- Engineering Persona: You write code to solve operational problems. You are not a traditional sys-admin; you build frameworks, automation, and tooling (Scala, Java, Go, or Python) to eliminate toil.
- Platform & AI Mindset: Deep understanding of distributed data platforms and a passion for leveraging AI/ML to revolutionize infrastructure management. Familiarity with LLM infrastructure, training/inference pipelines, or agentic frameworks is a significant plus.
- Operational Grit: Proven ability to remain calm and decisive under pressure. You have navigated large-scale distributed systems through hyper-growth and have a track record of driving incident-to-roadmap loops.
- Strategic Influence: Experience building long-range technical roadmaps and driving cross-functional alignment. You are comfortable challenging senior leadership with data-driven insights.
Pay Range Transparency
Benefits & conditions
Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above.
Zone 1 Pay Range $228,600-$314,250 USD, At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees.
About the company
At Databricks, we are passionate about enabling data teams to solve the worldās toughest problems - from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the worldās best data and AI infrastructure platform so our customers can use deep data insights to improve their business. Founded by engineers - and customer obsessed - we leap at every opportunity to tackle technical challenges, from designing next-gen UI/UX for interfacing with data to scaling our services and infrastructure across millions of virtual machines. And weāre only getting started.
Our Data Platform organization is the backbone of the worldās leading Data and AI platform. We operate a massive, multi-cloud (AWS, Azure, GCP), multi-AI, multi-region stack that powers thousands of the worldās most demanding workloads. As we move into the era of Agentic AI, we are not just scaling infrastructure; we are building a new generation of Agentic Observability and Reliability systems., Databricks is the data and AI company. More than 10,000 organizations worldwide - including Comcast, CondĆ© Nast, Grammarly, and over 50% of the Fortune 500 - rely on the Databricks Data Intelligence Platform to unify and democratize data, analytics and AI. Databricks is headquartered in San Francisco, with offices around the globe and was founded by the original creators of Lakehouse, Apache Spark , Delta Lake and MLflow. To learn more, follow Databricks on Twitter, LinkedIn and Facebook.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role ā technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Data Engineer Salary UK
Top-Paying Tech Jobs (with Salaries)
How to Become an AI Engineer
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production