> Markdown version of [/jobs/ext/3530318-sr-staff-software-engineer-product-ml-infrastructure](https://www.wearedevelopers.com/jobs/ext/3530318-sr-staff-software-engineer-product-ml-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Staff Software Engineer, Product ML Infrastructure - **Company:** Pinterest - **Location:** Palo Alto, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $245,402.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, C++ (Programming Language), Software Code Optimization, Profiling, Extract Transform Load (ETL), Python (Programming Language), Machine Learning, Recommender Systems, System Programming, Discretization, Information Technology, Machine Learning Operations - **Published:** October 2, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/18507812?backUrl=%2Fcareer%2F18507812%2FSr-Staff-Software-Engineer-Product-Ml-Infrastructure-California-Palo-Alto ## About the Role * A track record of setting technical strategy and delivering company-wide infrastructure initiatives in ambiguous environments. * Deep expertise in distributed ML systems, including production experience with both large-scale training and online inference. * Strong GPU performance knowledge, such as profiling, distributed execution, kernel and memory optimization, compilation, or quantization. * Experience with AI/ML modeling, recommender systems, Ads ranking, retrieval, feature platforms, or similarly demanding ML workloads. * Strong systems programming and design skills in C++, Java, or Python. * High ownership and sound judgment in reliability, security, cost, and operational excellence. * Demonstrated ability to use AI to improve speed and critically evaluate AI-assisted work, with accountability for correctness, quality, and sensitive data. * Bachelor's degree in Computer Science, Engineering, a related field, or equivalent experience. ## Description Pinterest's Product ML Infrastructure (PMLI) team enables fast, safe, and efficient delivery of AI/ML solutions across Ads and Core critical products. We build unified data, training, feature, and inference infrastructure; this role will set technical direction across model training and serving, with a focus on GPU efficiency and large-scale ranking systems. What you'll do: * Set the technical vision and roadmap for model training and serving across PMLI, with reusable interfaces to data and feature infrastructure. * Lead architectures for distributed training, fine-tuning, distillation, evaluation, and high-scale CPU/GPU inference. * Improve efficiency across data loading, distributed execution, GPU kernels and memory, compilation, quantization, scheduling, and capacity. * Build reliable, observable platforms with strong quality guarantees and training/serving consistency. * Partner with Ads and Core AI/ML teams to productionize features and models safely at Pinterest scale. * Drive cross-organizational architecture decisions, migrations, and operational standards; mentor senior engineers and raise the engineering bar. * Use AI-assisted development and analysis to accelerate prototyping, performance diagnosis, and validation while maintaining rigorous correctness and data safeguards., * We recognize that the ideal environment for work is situational and may differ across departments. What this looks like day-to-day can vary based on the needs of each organization or role. * This role will need to be in the office for in-person collaboration 1-2 times per quarter and therefore needs to be within a commutable distance of our Palo Alto, CA or San Francisco, CA office. ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Implementing an Event Sourcing strategy on Azure](https://www.wearedevelopers.com/videos/688-implementing-an-event-sourcing-strategy-on-azure) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Enhancing Workload Security in Kubernetes](https://www.wearedevelopers.com/videos/356-enhancing-workload-security-in-kubernetes) - [How We Built a Machine Learning-Based Recommendation System (And Survived to Tell the Tale)](https://www.wearedevelopers.com/videos/752-how-we-built-a-machine-learning-based-recommendation-system-and-survived-to-tell-the-tale) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)