> Markdown version of [/jobs/ext/1334537-member-of-technical-staff-machine-learning-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1334537-member-of-technical-staff-machine-learning-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Member of Technical Staff - Machine Learning Infrastructure Engineer - **Company:** Preference Model, Inc. - **Location:** San Francisco, CA, United States - **Salary:** $180,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Automation of Tests, Data Infrastructure, Software Debugging, Distributed Computing Environment, Distributed Systems, Machine Learning, Software Tools, Tensorflow, Data Logging, Pytorch, Large Language Models, Kubernetes, Low Latency, Data Pipelines - **Published:** July 18, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=282a869d58e4487a ## About the Role * Have strong software engineering fundamentals, experience building production-grade infrastructure (ideally for ML or data-intensive systems), and proficiency in core ML frameworks such as PyTorch or JAX * Understand distributed systems principles, and have hands-on experience with cloud platforms (AWS, GCP) and container orchestration (Kubernetes), building systems for high-throughput, low-latency workloads * Have experience with data engineering tools and building robust, scalable data pipelines * Have some familiarity with LLM training/inference internals (transformers, distributed training, inference libraries like vLLM or SGLang) - deep expertise is a plus, not a requirement * Can balance production rigor with the pace of fast-moving research, and communicate infrastructure tradeoffs clearly to researchers who aren't infra specialists ## Description * Design, build, and scale the compute, scheduling, and data infrastructure that powers post-training research on our in-house RL environments * Develop and maintain core ML framework primitives and internal tooling that researchers rely on daily, accelerating reproducible experimentation and reducing time from idea to result * Build evaluation and benchmarking infrastructure, monitoring, logging, and debugging tooling, and automated testing and deployment systems, so failures are caught early and infrastructure stays reliable as it scales * Partner directly with Research Engineers to translate research needs into infrastructure requirements, and ship fast in response to their feedback ## Related Videos - [The state of MLOps - machine learning in production at enterprise scale](https://www.wearedevelopers.com/videos/369-the-state-of-mlops-machine-learning-in-production-at-enterprise-scale) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [How We Built a Machine Learning-Based Recommendation System (And Survived to Tell the Tale)](https://www.wearedevelopers.com/videos/752-how-we-built-a-machine-learning-based-recommendation-system-and-survived-to-tell-the-tale) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)