> Markdown version of [/jobs/ext/2025219-lead-software-engineer-machine-learning-platform](https://www.wearedevelopers.com/jobs/ext/2025219-lead-software-engineer-machine-learning-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Software Engineer - Machine Learning Platform - **Company:** JPMorgan Chase & Co. - **Location:** Palo Alto, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Microsoft Access, Artificial Intelligence, Airflow, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Automation of Tests, Big Data, Cloud Engineering, Software Quality, Code Review, Continuous Integration, Information Engineering, Extract Transform Load (ETL), Software Debugging, Distributed Computing Environment, Identity and Access Management, Python (Programming Language), Machine Learning, Modular Design, Software Tools, Tensorflow, Secure Coding, Software Engineering, Strategies of Testing, Toolchain, Parquet, Cloud Platform System, Pytorch, Large Language Models, Apache Spark, Deep Learning, Caching, Amazon Virtual Private Cloud (VPC), Kubernetes, Dask, Cloudwatch, Code Restructuring - **Published:** August 11, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/27919766/Lead-Software-Engineer-Machine-Learning-Platform-California-Palo-Alto-7463 ## About the Role * Formal training or certification on software engineering concepts and 5+ years applied experience. * Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code. * Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management). * Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows). * Hands on experience with deep learning training workflows and at least one major framework (eg., PyTorch or TensorFlow). * Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology. * Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics, scaling and bottleneck analysis). * Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling). * Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (eg., EKS/ECR, S3, IAM, VPC/networking, Cloudwatch, EC2) * Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security. * Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices., * Experience running training workloads across multiple cloud platforms and managing portability, performance, and governance across environments. * Familiarity with cloud-native networking/storage patterns for high-throughput training and artifact management. * Experience optimizing training input pipelines (sharding, prefetching, caching, format choices such as Parquet/WebDataset) and working with large datasets. * Familiarity with distributed compute frameworks (Spark, Ray, Dask) for feature/dataset generation. * Familiarity with workflow orchestration tools (Airflow-like systems, Argo Workflows-like patterns) and model registry concepts. * Experience optimizing training cost/performance (right-sizing, scheduling policies, interruptible capacity strategies where applicable budge guardrails, quota planning). * Strong observability practice for training systems: metrics/logs/traces, GPU telemetry, dashboards, and alert tuning. FEDERAL DEPOSIT INSURANCE ACT: This position is subject to Section 19 of the Federal Deposit Insurance Act. As such, an employment offer for this position is contingent on JPMorganChase's review of criminal conviction history, including pretrial diversions or program entries. ## Description * Design, build, and maintain end-to-end ML training platform. * Run and optimize GPU training workloads (single-node and distributed), improving throughput, utilization and reproducibility. * Build and operate training infrastructure on Kubernetes (e.g., EKS and other manage Kubernetes platforms), including resource management and workload troubleshooting. * Enable Gen AI/LLM training and fine-tuning workflows (e.g., supervised fine-tuning), including evaluation harnesses, artifact/version governance, and scalable GPU execution patterns aligned to enterprise controls. * Implement observability for training systems: metrics, logs, dashboards, alerting, and operational runbooks. * Partner with data engineering and platform teams to define interfaces, standards, and guardrails (security, access, cost controls) * Improve developer experience for training: standardized containers, CI/CD, templates, documentation, and self-service workflow * Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team. * Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation. ## Related Videos - [Agentic employees in world's most downloaded FinTech app](https://www.wearedevelopers.com/videos/100123-agentic-employees-in-world-s-most-downloaded-fintech-app) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Event based cache invalidation in GraphQL](https://www.wearedevelopers.com/videos/433-event-based-cache-invalidation-in-graphql) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [The Prompt Engineer ✍️](https://www.wearedevelopers.com/magazine/216-the-prompt-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)