> Markdown version of [/jobs/ext/2044465-machine-learning-engineer-ai-inference-solutions-university-grad](https://www.wearedevelopers.com/jobs/ext/2044465-machine-learning-engineer-ai-inference-solutions-university-grad). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer, AI Inference Solutions (University Grad) - **Company:** General Motors - **Location:** Sunnyvale, CA, United States - **Experience:** Starter - **Salary:** $119,250.0 - $150,850.0 - **Contract:** Internship / Graduate position - **Skills:** Artificial Intelligence, Airflow, Computer Vision, C++ (Programming Language), Compilers, Program Optimization, Profiling, Code Review, Nvidia CUDA, Data Structures, Cursor (Graphical User Interface Elements), Linux, Distributed Computing Environment, Distributed Systems, Python (Programming Language), Machine Learning, Natural Language Processing, Open Source Technology, Azure Machine Learning, Secure Coding, Software Deployment, GitHub Copilot, Pytorch, Delivery Pipeline, Large Language Models, Deep Learning, Gpu Programming, Kubernetes, Information Technology, ONNX (Open Neural Network Exchange) Format, Production Code, Free and Open-Source Software, Machine Learning Operations, TensorRT, Software Coding - **Published:** August 13, 2026 - **Apply:** https://generalmotors.wd5.myworkdayjobs.com/Careers_GM/job/Sunnyvale-California-United-States-of-America/Machine-Learning-Engineer--AI-Inference-Solutions--University-Grad-_JR-202610103 ## About the Role * Recently completed or completing a Bachelor's or Master's degree by Spring 2026 in Computer Science, ECE, or a related technical field. (Degree must be completed before your start date.) * Strong computer science fundamentals (e.g., data structures, algorithms, operating systems, computer architecture) and solid coding skills in Python and/or C++, demonstrated through coursework, internships, or substantial projects. * Hands-on experience in AI/ML (e.g., machine learning, deep learning, computer vision, NLP, or ML systems) via classes, research, internships, or personal projects. * Depth in at least one of: computer architecture, operating systems, distributed systems, or compilers. * Demonstrated software-engineering experience (internships, coursework, open-source, research code, or competitions) showing good judgment around reliability, correctness, and clean abstractions. * Experience with-or strong interest in-using coding assistants/agents (e.g., Cursor, Claude Code, GitHub Copilot) as part of your workflow. * Ability to work effectively in collaborative, cross-functional teams and communicate clearly-both in writing and verbally-including explaining technical work partners What Will Give You a Competitive Edge (Preferred Qualifications) * Internship, research, or advanced coursework in ML systems, ML compilers, GPU programming (CUDA, OpenAI Triton), inference optimization, or distributed training/serving infrastructure. * Familiarity with PyTorch and modern ML compiler/runtime stacks (e.g., torch.compile, TensorRT, ONNX, Triton Inference Server, vLLM, or equivalent). * Exposure to model optimization (quantization, pruning, distillation) or GPU profiling tools (Nsight Systems, Nsight Compute, PyTorch Profiler). * Familiarity with workflow/ML platforms such as Airflow, Temporal, Flyte, Ray, or Kubeflow. * Experience building agentic or LLM-powered tools or workflows. * Open-source contributions related to PyTorch, TensorRT, vLLM, OpenAI Triton, or similar projects. * Coursework, projects, or publications touching ML systems (e.g., MLSys, OSDI, ASPLOS, HPCA, NeurIPS systems track). * Familiarity with a systems language (e.g., C++) and development in a Linux environment. ## Description As an early career Engineer on the Model Deployment & Inference Solutions team, you'll contribute across both sides of our mission: building the ML deployment platform and optimizing models for on-vehicle inference. You'll work with and learn from senior engineers on real production deployments, platform features, and model-optimization workflows that ship to GM's Super Cruise fleet at large scale, with structured mentorship and a clear onboarding plan. You'll also collaborate closely with our sister teams (kernels, compiler, reduced precision, and parity) on the end-to-end path that takes trained models from research frameworks to ultra-efficient, safety-critical inference on the car. This is an early-career / new graduate role designed for candidates who have recently or will be completing their degree by June 2026. What You'll Do (Responsibilities) * Contribute production code across the ML deployment platform, model-optimization workflows, and inference benchmarking/profiling infrastructure. * Pair with senior engineers on deployment workflows, performance investigations, model-optimization experiments (e.g., quantization, pruning, distillation), and platform tooling. * Build, test, and maintain platform tools (e.g., validators, performance probes, parity and sensitivity analyzers, agentic specialists) with technical guidance and code review support. * Investigate and help root-cause production deployment or performance issues; learn and apply the diagnostic playbook for compiler, kernel, runtime, and parity bugs. * Collaborate with cross-functional teams across the AV organization; including kernels, compiler, reduced-precision, parity, and model-development groups-to plan and execute model deployments to the AV stack, working under the guidance of senior engineers * Participate in code reviews, design discussions, and technical documentation to ensure reliability, correctness, and clear abstractions in a large-scale codebase. * Learn and follow secure coding, safety, and compliance practices required for on-vehicle autonomous driving software., * This role is categorized as hybrid. This means the selected candidate is expected to report to a specific location at least 3 times a week. ## Related Videos - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)