> Markdown version of [/jobs/ext/1301092-ai-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1301092-ai-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Infrastructure Engineer - **Company:** NVIDIA Ltd. - **Location:** Ann Arbor, MI, United States (Remote available) - **Experience:** Expert - **Salary:** $170,000.0 - $210,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Nvidia CUDA, Data Centers, Distributed Systems, Fault Tolerance, Python (Programming Language), Machine Learning, Open Source Technology, Prometheus, Software Engineering, AI Infrastructure, Rust (Programming Language), Datadog, Cloud Platform System, Grafana, Backend, Kubernetes, Low Latency, Optimization Algorithms, Machine Learning Operations, TensorRT, Hardware Infrastructure, Terraform, Docker, Golang - **Published:** July 16, 2026 - **Apply:** https://www.workingnomads.com/job/go/1732117/ ## About the Role * 5+ years of software engineering experience with a strong focus on AI infrastructure, backend systems, or distributed systems * Hands-on experience with AI model serving frameworks (e.g., vLLM, SGLang, Triton, TensorRT, TorchServe, or similar) * Understanding of container orchestration and cluster management (Kubernetes, Docker) * Experience deploying and operating infrastructure across both datacenter and on-prem environments * Strong knowledge of GPU workloads and the tradeoffs that come with them - you understand how inference differs from training, and why it matters * Proficiency in Python; C++, CUDA, Go, Rust a plus * Excellent communication skills and comfort working cross-functionally in a lean, fast-moving environment * Willingness to travel up to 10% of time Enhanced Qualifications (Nice to Have) * Dynamo experience a plus * Experience with edge AI deployments or constrained compute environments * Familiarity with infrastructure as code (Terraform, Helm) * Experience with observability platforms (Datadog, Prometheus, Grafana) * Background in energy, utilities, or industrial IoT * Contributions to open-source ML infrastructure projects ## Description Utilidata is a fast-growing NVIDIA-backed AI company enabling AI data centers to dynamically orchestrate power and unlock more compute capacity from existing energy infrastructure. For over a decade, we have applied AI to the electric grid - bringing real-time visibility and power-flow control to complex energy infrastructure. Our Karman platform, built on a custom NVIDIA module, brings that same capability to AI data centers, giving operators a way to better use the power already available to them. The AI Infrastructure Engineer is responsible for designing, building, and owning the end-to-end infrastructure that serves Utilidata's AI and ML models across edge deployments, cloud environments, and data center integrations. They are also responsible for designing, building, and owning the integration of power data with AI inference software. This is Utilidata's first dedicated role of this kind, and will serve as the foundational function for how the company deploys and operates AI capabilities in production. The role requires deep technical expertise in ML model serving, distributed systems, and GPU infrastructure, with a strong emphasis on reliability, performance, and scalability. This position works cross-functionally with product, engineering, and data science teams and is open to fully remote candidates, with periodic travel expected for company retreats and key on-site engagements. Responsibilities * Lead the design and build of Utilidata's AI inference platform - establishing architecture patterns, deployment standards, and operational practices that will scale with the company * Own end-to-end model serving infrastructure for Utilidata's AI infrastructure (on-prem and datacenter) * Build and maintain fault-tolerant, high-performance systems for serving AI models at scale, with a focus on low latency, reliability, and cost efficiency * Collaborate closely with algorithms engineers to integrate AI inference data and configuration with power optimization algorithms * Optimize GPU utilization and inference performance across our hardware fleet, including NVIDIA accelerators central to Utilidata's edge AI platform * Establish MLOps best practices including CI/CD pipelines for model deployment, monitoring, and rollback across environments * Contribute to infrastructure roadmap decisions, including build vs. buy tradeoffs, tooling selection, and platform evolution as the team grows, Utilidata is a fast-growing NVIDIA-backed AI company enabling AI data centers to dynamically orchestrate power and unlock more compute capacity from existing energy infrastructure. For over a decade, we have applied AI to the electric grid - bringing real-time visibility and power-flow control to complex energy infrastructure. Our Karman platform, built on a custom NVIDIA module, brings that same capability to AI data centers, giving operators a way to better use the power already available to them. The AI Infrastructure Engineer is responsible for designing, building, and owning the end-to-end infrastructure that serves Utilidata's AI and ML models across edge deployments, cloud environments, and data center integrations. They are also responsible for designing, building, and owning the integration of power data with AI inference software. This is Utilidata's first dedicated role of this kind, and will serve as the foundational function for how the company deploys and operates AI capabilities in production. The role requires deep technical expertise in ML model serving, distributed systems, and GPU infrastructure, with a strong emphasis on reliability, performance, and scalability. This position works cross-functionally with product, engineering, and data science teams and is open to fully remote candidates, with periodic travel expected for company retreats and key on-site engagements. Responsibilities * Lead the design and build of Utilidata's AI inference platform - establishing architecture patterns, deployment standards, and operational practices that will scale with the company * Own end-to-end model serving infrastructure for Utilidata's AI infrastructure (on-prem and datacenter) * Build and maintain fault-tolerant, high-performance systems for serving AI models at scale, with a focus on low latency, reliability, and cost efficiency * Collaborate closely with algorithms engineers to integrate AI inference data and configuration with power optimization algorithms * Optimize GPU utilization and inference performance across our hardware fleet, including NVIDIA accelerators central to Utilidata's edge AI platform * Establish MLOps best practices including CI/CD pipelines for model deployment, monitoring, and rollback across environments * Contribute to infrastructure roadmap decisions, including build vs. buy tradeoffs, tooling selection, and platform evolution as the team grows ## Related Videos - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Super scaling for the Super Bowl: How to survive 30 million users hitting your backend in 30 minutes](https://www.wearedevelopers.com/videos/100356-super-scaling-for-the-super-bowl-how-to-survive-30-million-users-hitting-your-backend-in-30-minutes) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)