> Markdown version of [/jobs/ext/1255896-machine-learning-systems-engineer-networking](https://www.wearedevelopers.com/jobs/ext/1255896-machine-learning-systems-engineer-networking). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Systems Engineer, Networking - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Experienced - **Salary:** $152,000.0 - $241,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Computer Programming, Databases, Data Centers, Python (Programming Language), Machine Learning, Network Architecture, Data Streaming, Feature Engineering, Data Ingestion, Information Technology, Low Latency, Apache Kafka, Machine Learning Operations - **Published:** July 13, 2026 - **Apply:** https://www.juju.com/job/00000000gfyyam ## About the Role + A BS (or equivalent experience) and 5+ years of experience, MS and 3+ years, or PhD with 1+ years in Computer Science, Statistics, or a related field + Strong mathematical foundation: statistics, probability, linear algebra, and algorithm analysis + Proven experience implementing and optimizing ML algorithms in production - this is a coding-first role; strong implementation skills are required + Strong programming skills in one or more of Go, C/C++, Rust, or Scala; Python working knowledge is a plus + Familiarity with time-series databases and streaming data architectures + Ability to work independently and navigate ambiguity in a fast-paced engineering environment Ways to stand out from the crowd: + Data Science background with hands-on experience building and validating ML models - bridging research and production implementation + Experience implementing ML algorithms directly in systems languages for latency-sensitive or resource-constrained environments + Research experience: knowing the latest ML literature and translating advances into practical improvements + Experience with Kafka-based streaming pipelines and real-time feature engineering at scale ## Description Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. As an ML Engineer on this team, you'll design and implement ML algorithms that run in real-time streaming pipelines, detecting anomalies and surfacing insights across massive-scale infrastructure before they impact AI training and inference. The core challenge of this role is building ML algorithms that are simultaneously accurate and efficient -processing millions of telemetry streams in real time within tight CPU and memory budgets. You'll need both the data science depth to design and validate algorithms and the engineering discipline to implement them in production at scale. What you'll be doing: + Implement production ML algorithms in Go - optimized for real-time streaming pipelines operating at massive scale under strict resource constraints + Design and develop new ML algorithms where needed: anomaly detection, health scoring, and predictive analytics on high-volume time-series telemetry from GPU and network infrastructure + Improve and extend existing algorithms and experiment with new approaches suited to real-time streaming constraints + Build and maintain end-to-end ML pipelines - from data ingestion and schema design through model inference - optimized for on-premises, latency-sensitive deployments ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)