> Markdown version of [/jobs/ext/3080524-machine-learning-systems-engineer](https://www.wearedevelopers.com/jobs/ext/3080524-machine-learning-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Systems Engineer - **Company:** Atlas Data Systems, L. L.C. - **Location:** Seattle, WA, United States (Remote available) - **Experience:** Expert - **Salary:** $206,100.0 - $269,075.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Continuous Delivery, Continuous Integration, Open Source Technology, Performance Tuning, Software Engineering, System Programming, Load Balancing, Autoscaling, Large Language Models, Caching, Parallel Computation, Low Latency, Machine Learning Operations, TensorRT - **Published:** September 25, 2026 - **Apply:** https://www.themuse.com/jobs/atlassian/senior-machine-learning-systems-engineer-4ddf46 ## About the Role * 5+ years of software engineering experience, 2+ years of system performance optimization experience * Deep low-level systems programming (C/C++ or Rust) * Experience with large-scale, high-concurrent production serving. * Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.). * Strong background in system optimizations: batching, caching, load balancing, parallelism. It would be great, but not required if you have * Low-level inference optimizations: GPU kernels * Algorithmic inference optimizations: quantization, speculative decoding, distillation * Experience with testing, benchmarking, and reliability of inference services. * Experience designing and implementing CI/CD infrastructure for inference. ## Description * Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache). * Optimize latency and throughput of model inference under real production workloads. * Build reliable, high-concurrency serving systems that serve billions of requests reliably * Benchmark, fine-tune, and accelerate inference engines. * Create robust CI/CD infrastructure for seamless model deployment and inference engine updates. * Partner with senior ML engineers to finetune open-source LLMs and deploy ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Fifty Shades of Kubernetes Autoscaling](https://www.wearedevelopers.com/videos/813-fifty-shades-of-kubernetes-autoscaling) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Event based cache invalidation in GraphQL](https://www.wearedevelopers.com/videos/433-event-based-cache-invalidation-in-graphql) - [Serverless-Native Java with Quarkus](https://www.wearedevelopers.com/videos/243-serverless-native-java-with-quarkus) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)