> Markdown version of [/jobs/ext/3530612-lead-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/3530612-lead-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Machine Learning Engineer - **Company:** JPMorgan Chase & Co. - **Location:** Palo Alto, CA, United States - **Experience:** Expert - **Salary:** $171,000.0 - $260,000.0 - **Contract:** Permanent contract - **Skills:** Airflow, Big Data, Cloud Computing, Nvidia CUDA, Monitoring of Systems, Python (Programming Language), Machine Learning, Tensorflow, Azure Machine Learning, Data Storage Technologies, Apache Spark, Kubernetes, Information Technology, SGLang, Machine Learning Operations, Docker - **Published:** October 2, 2026 - **Apply:** https://dejobs.org/x/x/DAB8BFC1C1DE43C38AA3BE47CB508BB4/job/ ## About the Role * BS in Computer Science or related Engineering field with 6+ years of experience Or MS degree in Computer Science or related Engineering field with 4+ years experience. * Solid knowledge and extensive experience in Python and in cloud computing, along with ML frameworks (i.e. pytorch, tensorflow) * Deep knowledge and passion for data science fundamentals, training and deploying models * Experience in monitoring and observability tools to monitor model input/output and features stats * Operational experience in big data/ML tools such as Ray, Spark and in training/inference systems such as Ray, vllm/SGLang * Solid grounding in engineering fundamentals and enterprise system design Preferred qualifications, capabilities, and skills * Experience with recommendation and personalization systems is a plus. * CUDA experience is a big plus * Solid fundamentals and experience in containers (docker ecosystem), container orchestration systems [Kubernetes, ECS], DAG orchestration [Airflow, Kubeflow etc] * Good knowledge of data storage solutions and strategies (online and offline) ## Description * Build, deploy, and maintain robust pipelines for distributed training on GPU-enabled clusters to support scalable machine learning workflows. * Develop and manage pipelines for model promotion and other capabilities related to MDLC. * Optimize training throughput for large data sources * Establish and maintain integrations to platforms and tools related to model monitoring and observability * Collaborate with cross-functional teams to integrate new technologies and improve the capabilities of our ML Platform. * Partner with product, architecture, modeling, and engineering to design robust solutions that power our Digital channels