> Markdown version of [/jobs/ext/519895-ml-systems-engineer](https://www.wearedevelopers.com/jobs/ext/519895-ml-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Systems Engineer - **Company:** CHIP'S GENERAL CARPENTRY, LLC - **Location:** Santa Barbara, CA, United States - **Salary:** $150,000.0 - $350,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Business Logic, C++ (Programming Language), Profiling, Nvidia CUDA, Software Debugging, Distributed Computing Environment, General-Purpose Computing on Graphics Processing Units, Python (Programming Language), Graphics Processing Unit (GPU), Pytorch, Large Language Models, Parallel Computation, Information Technology, Low Latency, Machine Learning Operations - **Published:** June 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=43a551dbf6a35041 ## About the Role Do you have experience in Research?, * B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience). * Experience with large-scale ML systems, GPU computing, or high-performance inference optimization. * Strong proficiency in Python and C++/CUDA; hands-on experience with SGLang, vLLM, PyTorch, or similar inference frameworks. * Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms. * Experience deploying and optimizing LLMs in production: model serving, batching strategies, distributed inference, or quantization. * Strong systems-level debugging and profiling skills; comfort working at multiple layers of the stack from CUDA kernels to application logic. * Familiarity with distributed computing frameworks (Ray, multi-node training/inference) is a plus. * Self-directed problem solver who is interested in working on ambitious optimization challenges. ## Description We are seeking an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering our agentic AI platform. This is a technical role focused on low-level systems optimization. You will implement performance optimizations, build evaluation harnesses, and architect multi-node clusters for training and inference that push the limits of LLM throughput and latency. Your work will directly impact the responsiveness and cost-efficiency of AI agents used by leading semiconductor companies to design chips., * Design, deploy, and optimize LLM inference systems across multi-node clusters, maximizing throughput and minimizing latency for production workloads. * Implement and benchmark concrete inference optimizations. * Profile and analyze inference bottlenecks at the systems level-from GPU kernel execution to memory bandwidth constraints. * Build robust evaluation harnesses and benchmarking frameworks that measure accuracy, throughput, latency, and resource utilization across various parallelism strategies. * Collaborate with research scientists to integrate new model architectures and optimizations into production inference infrastructure. * Investigate and apply emerging techniques from research papers and open-source projects to continuously improve inference performance. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)