> Markdown version of [/jobs/ext/3049766-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/3049766-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - **Company:** Press Ganey Associates, LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $130,000.0 - $190,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Automation of Tests, Microsoft Azure, Cloud Database, Software Quality, Code Review, Continuous Integration, Database Queries, Distributed Systems, Fault Tolerance, Python (Programming Language), Load Testing, Search Technologies, Software Engineering, Unstructured Data, Management of Software Versions, Speech Recognition, Data Processing, Retrieval-Augmented Generation, Large Language Models, Model Validation, AI Platforms, Information Technology, Low Latency, Machine Learning Operations, Text Analysis, Data Pipelines, Serverless Computing, Databricks - **Published:** September 24, 2026 - **Apply:** https://www.builtincolorado.com/auth/login?destination=/job/staff-ml-engineer/11346753 ## About the Role * Bachelor's degree in Computer Science, Electrical Engineering, or a related technical discipline, or equivalent practical experience. * 8+ years of professional software engineering experience, including at least 3 years owning ML or LLM systems in production and their operational support. * Proven track record of independently leading complex technical initiatives from requirements through production, making architectural decisions and coordinating delivery across stakeholders. * Advanced proficiency in Python for production services and data processing, with strong SQL skills. * Experience designing and operating high-throughput distributed systems, with a strong understanding of failure recovery, multi-tenancy, and capacity planning. * Hands-on experience deploying and operating LLM-based applications, including evaluation, output validation, observability, and cost management. * Strong production engineering practices across automated testing, CI/CD, monitoring, incident response, and root-cause analysis. * Demonstrated technical leadership through system design, hands-on implementation, code review, and mentorship. * Ability to communicate technical decisions and tradeoffs clearly to engineering, research, product, and governance stakeholders., * Experience with Databricks or comparable cloud-based data and AI platforms for workflow orchestration, scalable processing, model deployment, and evaluation. * Experience with NLP, text analytics, or large-scale processing of unstructured data. * Experience building shared infrastructure for inference, evaluation, and model lifecycle management. * Familiarity with retrieval-augmented generation, semantic search, and LLM orchestration frameworks. * Experience with speech-to-text, speaker diarization, or processing conversational audio data. * Experience deploying and operating cloud-native services on AWS or Azure. * Experience with healthcare or other regulated environments, including sensitive data handling, auditability, and model governance. ## Description As a technical leader, you will guide engineers through design and implementation, contribute directly to the codebase, and lead architectural decisions across teams. Working primarily with Python and cloud-based data and AI platforms, with a growing emphasis on Databricks, you will shape the shared infrastructure supporting both established products and new AI applications. What You'll Do * Own technical delivery from prototype through deployment and ongoing production support, partnering with AI Scientists and Product to define requirements, plan implementation, and resolve cross-team dependencies. * Evaluate proposed AI solutions for production suitability, identify technical risks, and choose architectures that meet quality, reliability, and cost requirements without unnecessary complexity. * Architect and build high-volume AI services and processing pipelines, including fault tolerance, backpressure, retries, idempotency, and recovery from partial failures. * Lead the evolution of our Python services and Databricks-based platform for distributed processing, model serving, and integration of traditional ML and LLM-based components. * Establish production engineering standards for automated testing, CI/CD, model and prompt versioning, load testing, controlled rollouts, and rollback. * Build evaluation and monitoring capabilities to detect AI quality regressions and track service reliability, throughput, latency, and inference cost. * Partner with Product and Responsible AI teams to define release criteria and implement requirements for model validation, data privacy, security, and governance. * Optimize processing and inference workloads, balancing model quality, throughput, latency, capacity, and cost. * Mentor engineers and lead architecture and code reviews, maintaining consistent standards for software quality and maintainability. ## Related Videos - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [OLTP in the Lakehouse: Redefining Data for AI Workloads](https://www.wearedevelopers.com/videos/2038-oltp-in-the-lakehouse-redefining-data-for-ai-workloads) - [How to Avoid LLM Pitfalls - Mete Atamel and Guillaume Laforge](https://www.wearedevelopers.com/videos/1328-how-to-avoid-llm-pitfalls-mete-atamel-and-guillaume-laforge) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)