> Markdown version of [/jobs/ext/3079928-remote-lead-mlops-engineer-databricks](https://www.wearedevelopers.com/jobs/ext/3079928-remote-lead-mlops-engineer-databricks). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # REMOTE - Lead MLOps Engineer (Databricks) - **Company:** The Information Technology - **Location:** Bloomington, IL, United States (Remote available) - **Experience:** Expert - **Salary:** $100,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Application Layers, Application Performance Management, Command-Line Interface, Code Review, Continuous Integration, Data Validation, Data Infrastructure, Software Debugging, Python (Programming Language), Open Source Technology, Systems Development Life Cycle, SQL Databases, Strategies of Testing, Enterprise Software Applications, Apache Spark, Technical Debt, Data Lakes, Pyspark, Data Lineage, Deployment Automation, Machine Learning Operations, Databricks - **Published:** September 25, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=83f44037af353907 ## About the Role * Databricks (Non-Negotiable) - At least a year running Databricks in production at the application layer. You've designed and operated medallion architecture, and you know Delta Lake, Unity Catalog, and PySpark cold. * Production ML (Non-Negotiable) - You've kept models healthy in production, not just trained them and handed them off. You've been paged, written the RCA, and caught drift and silent degradation before they became someone else's problem. * Agentic-Native (Non-Negotiable) - GenAI and agentic tools are already part of how you work every day, not an experiment. You're at home on the command line and always looking for the next thing to automate. ## Description At the State Farm Applied Innovation Lab, we're building an Agentic SDLC control plane, where GenAI helps with every part of how we build, test, deploy, and run software. We're looking for someone who has run real workloads on Databricks in production and knows Delta Lake, Unity Catalog, and MLflow inside and out. Just as important, they already work the way we do, reaching for agentic tools naturally whether they're building something new or tracking down a bug. If you're happiest at the command line, you'll feel right at home here. In this role, you'll own the performance, quality, and production-readiness of select innovation projects running on Databricks - the ones where we dig deep to learn what excellence really looks like, then turn those lessons into patterns we can reuse on future projects. Other teams take care of the platform and infrastructure; you'll own the layer above it: the data foundations, governance, FinOps, lineage, runbooks, and the handoff to Enterprise Technology. When a project leaves our hands, it's fully governed, tagged at the project level, and optimized for cost, with a low L3 incident rate, complete runbooks, and self-healing wherever possible. This is a lead individual contributor role: you'll set technical direction and raise the bar across teams without managing people. You'll partner with teams across the enterprise, take on hard problems on Databricks, and help shape how we put the platform's newest capabilities to work on the goals that matter most to State Farm. It's a great fit if you want to grow your skills at the leading edge of Databricks and agentic AI, work on high-visibility initiatives, and do your best work in a fast-moving, experimental setting. What You'll Own Data Foundations - Design and run medallion architecture (Bronze Silver * Gold) with solid data quality, schema management, and Delta Lake tuning. Everything downstream depends on getting this right. * Application Performance - Hunt down bottlenecks in Spark queries, partitioning, caching, and Databricks cluster sizing, then fix them and keep them from coming back. * Unity Catalog & Lineage - Own schemas, access controls, and project-level tagging, with lineage-driven ticketing so upstream breaks surface right away. * MLOps / MLflow 3.0 - Take models through their full lifecycle, from experiment to production to retirement, including retraining and drift detection. * FinOps - Make Databricks cost visible by default: track, alert on, and optimize spend for every project. * ET Handoff - Define and hold the line on "done": low L3 incident rate, complete runbooks, self-healing where possible, and governance and FinOps buttoned up. When a project crosses the handoff plane, it's ready. * Patterns & Practices - Turn what you learn into templates and standards that make every future project better. * Agentic Integration - Bring agentic AI into everything you do: data validation, pipeline generation, testing, monitoring, documentation. If it can be augmented, it should be. * Technical Leadership & Collaboration - Set technical direction, mentor engineers, and pair with pod teams on architecture, code reviews, and debugging. You teach by doing., * MLflow 3.0 - Hands-on with the full lifecycle: registry, experiment tracking, deployment automation, retraining pipelines, and drift detection. * FinOps - You've tracked and optimized Databricks spend at the project level, and you treat compute budgets with the same rigor as SLAs. * Data Lineage & Observability - You've built or run systems that track lineage end to end and open tickets automatically on upstream breaks, schema changes, or SLA misses. * Production Handoff - You've prepared systems for operations teams with runbooks, self-healing patterns, incident targets, and governance checklists. * Systems Thinking - You see the whole system, not just the model or the pipeline. You've read Hidden Technical Debt in Machine Learning Systems, and you've lived it. * Engineering - Production-grade Python, PySpark, SQL, and AWS; CI/CD for data and ML pipelines; and testing strategies for non-deterministic systems. * Omnigent (Nice to Have) - Bonus points if you've worked with Omnigent, Databricks' open-source agent meta-harness, * Gold) with solid data quality, schema management, and Delta Lake tuning. Everything downstream depends on getting this right. * Application Performance - Hunt down bottlenecks in Spark queries, partitioning, caching, and Databricks cluster sizing, then fix them and keep them from coming back. * Unity Catalog & Lineage - Own schemas, access controls, and project-level tagging, with lineage-driven ticketing so upstream breaks surface right away. * MLOps / MLflow 3.0 - Take models through their full lifecycle, from experiment to production to retirement, including retraining and drift detection. * FinOps - Make Databricks cost visible by default: track, alert on, and optimize spend for every project. * ET Handoff - Define and hold the line on "done": low L3 incident rate, complete runbooks, self-healing where possible, and governance and FinOps buttoned up. When a project crosses the handoff plane, it's ready. * Patterns & Practices - Turn what you learn into templates and standards that make every future project better. * Agentic Integration - Bring agentic AI into everything you do: data validation, pipeline generation, testing, monitoring, documentation. If it can be augmented, it should be. * Technical Leadership & Collaboration - Set technical direction, mentor engineers, and pair with pod teams on architecture, code reviews, and debugging. You teach by doing. * Databricks (Non-Negotiable) - At least a year running Databricks in production at the application layer. You've designed and operated medallion architecture, and you know Delta Lake, Unity Catalog, and PySpark cold. * Production ML (Non-Negotiable) - You've kept models healthy in production, not just trained them and handed them off. You've been paged, written the RCA, and caught drift and silent degradation before they became someone else's problem. * Agentic-Native (Non-Negotiable) - GenAI and agentic tools are already part of how you work every day, not an experiment. You're at home on the command line and always looking for the next thing to automate. * MLflow 3.0 - Hands-on with the full lifecycle: registry, experiment tracking, deployment automation, retraining pipelines, and drift detection. * FinOps - You've tracked and optimized Databricks spend at the project level, and you treat compute budgets with the same rigor as SLAs. * Data Lineage & Observability - You've built or run systems that track lineage end to end and open tickets automatically on upstream breaks, schema changes, or SLA misses. * Production Handoff - You've prepared systems for operations teams with runbooks, self-healing patterns, incident targets, and governance checklists. * Systems Thinking - You see the whole system, not just the model or the pipeline. You've read Hidden Technical Debt in Machine Learning Systems, and you've lived it. * Engineering - Production-grade Python, PySpark, SQL, and AWS; CI/CD for data and ML pipelines; and testing strategies for non-deterministic systems. * Omnigent (Nice to Have) - Bonus points if you've worked with Omnigent, Databricks' open-source agent meta-harness ## Related Videos - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [MLOps - What’s the deal behind it?](https://www.wearedevelopers.com/videos/392-mlops-what-s-the-deal-behind-it) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [OLTP in the Lakehouse: Redefining Data for AI Workloads](https://www.wearedevelopers.com/videos/2038-oltp-in-the-lakehouse-redefining-data-for-ai-workloads) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)