> Markdown version of [/jobs/ext/1490266-ml-systems-engineer](https://www.wearedevelopers.com/jobs/ext/1490266-ml-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Systems Engineer - **Company:** Bright Vision Technologies - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $145,000.0 - $165,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, C++ (Programming Language), Distributed Systems, Python (Programming Language), Azure Machine Learning, Data Logging, Graphics Processing Unit (GPU), Autoscaling, Large Language Models, Kubernetes, Information Technology, Free and Open-Source Software, Machine Learning Operations, TensorRT - **Published:** July 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=13c27dd88fb9ce5f ## About the Role Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position., * Bachelor's or Master's degree in Computer Science or a related field. * Six or more years of experience in distributed systems, infrastructure, or ML platform engineering. * Strong proficiency in Python and a systems language such as Go, Rust, or C++. * Deep experience operating high-throughput, low-latency services in production. * Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM. * Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization. * Familiarity with Kubernetes, autoscaling, and modern cloud platforms. * Experience with observability stacks including metrics, tracing, and structured logging. * Solid grounding in performance engineering and capacity planning. * Strong communication and incident response skills., * Open-source contributions to model serving infrastructure. * Experience with multi-region or globally distributed AI serving. * Familiarity with model quantization, distillation, and compression techniques. * Exposure to FinOps for AI workloads and cost-efficient serving design. * Experience supporting external-facing AI APIs at scale. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Fifty Shades of Kubernetes Autoscaling](https://www.wearedevelopers.com/videos/813-fifty-shades-of-kubernetes-autoscaling) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)