> Markdown version of [/jobs/ext/2847599-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/2847599-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - **Company:** Toyota Research Institute - **Location:** Los Altos, CA, United States - **Experience:** Expert - **Salary:** $287,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon S3, C++ (Programming Language), Code Review, Nvidia CUDA, Continuous Integration, Information Engineering, Distributed Computing Environment, Monitoring of Systems, Python (Programming Language), Machine Learning, Open Source Technology, Tensorflow, Software Systems, Pytorch, Large Language Models, Prompt Engineering, Containerization, Kubernetes, ONNX (Open Neural Network Exchange) Format, Machine Learning Operations, GPT, Software Version Control, Data Pipelines, Docker - **Published:** September 11, 2026 - **Apply:** https://www.thejobnetwork.com/job/add154ad-acdd-4798-96dd-1c01b04cf4be/senior-machine-learning-engineer ## About the Role The ideal candidate is a strong generalist who can move fluently across the ML stack, from cloud training infrastructure to LLM integrations to data pipelines. We also value candidates who bring deep expertise in two or more specific areas - for example, someone who combines strong MLOps fundamentals with hands-on edge/embedded ML experience, or who pairs LLM systems expertise with robust data engineering skills., * Candidates should have one of the following: a BS with 6-10 years, an MS with 5-9 years, a PhD with 3-7 years, or no degree with 9-13 years of equivalent experience; specific degree fields are flexible, with demonstrated experience prioritized over pedigree. * Strong proficiency in PyTorch and/or TensorFlow, with hands-on experience building, fine-tuning (including adapter-based methods such as LoRA and QLoRA), evaluating, and deploying large language models. * Experience working with multimodal data-including text, images, sensor/telemetry data, and speech-and understanding the associated data characteristics and pipeline requirements. * Proven track record of deploying and maintaining machine learning systems in production environments. * Experience with AWS services for machine learning workloads (e.g., Bedrock, SageMaker, ECS/Batch, S3), strong Python fundamentals, and comfort working within polyglot codebases. * Ability to consult effectively with researchers, translate ambiguous technical requirements into actionable solutions, operate autonomously on cross-team problems, and communicate clearly in both written and verbal contexts., * Experience optimizing models for resource-constrained hardware through quantization, pruning, and compilation frameworks (e.g., TFLite, LiteRT, ONNX), along with proficiency in C/C++ and/or CUDA for performance-critical inference. * Familiarity with MLOps practices such as experiment tracking (MLflow, Weights & Biases), CI/CD for ML, and model versioning (e.g., DVC), as well as containerization (Docker required; ECS/Batch preferred; Kubernetes a plus), distributed training across multi-GPU and multi-node setups, and experience with Vertex AI in addition to AWS. * Background in robotics, autonomous systems, materials science, or energy domains, with experience translating published research into production systems ("paper-to-production"), working in academic or industry R&D environments, and developing agentic AI systems with tool use and multi-step reasoning. * AWS certifications (e.g., Solutions Architect, ML Specialty) and contributions to open-source machine learning projects. ## Description * Build and maintain machine learning infrastructure, including training pipelines, distributed compute systems, model serving platforms, and monitoring tools that research teams rely on daily. * Integrate and evaluate large language models (LLMs) and foundation models by developing retrieval-augmented generation (RAG) systems, performing full and adapter-based fine-tuning, applying prompt engineering techniques, and benchmarking performance across providers such as AWS Bedrock, Gemini, Claude, GPT, and open-source models. * Design scalable data pipelines to support multimodal data, including text, images, sensor data, speech, video, and structured scientific datasets. * Consult with research teams to understand machine learning requirements, evaluate potential approaches, and propose solutions aligned with TRI's technology stack and engineering standards. * Support edge and embedded machine learning by optimizing, quantizing, and deploying models to onboard hardware platforms such as robotics systems and vehicles. * Bridge the gap between research and production by translating experimental notebooks and prototypes into maintainable, scalable, and deployable systems while preserving research innovation. * Stay current with advancements in machine learning and artificial intelligence by evaluating emerging techniques and assessing their potential adoption within TRI. * Drive technical quality by participating in code reviews, producing clear documentation, and fostering knowledge sharing across the Research Software Engineering (RSE) team. ## Related Videos - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)