> Markdown version of [/jobs/ext/2715147-inference-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2715147-inference-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Inference Infrastructure Engineer - **Company:** Rhoda ai - **Location:** United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Computer Clusters, Data Transport Utility, Distributed Systems, Tensorflow, Management of Software Versions, Pytorch, Delivery Pipeline, HybridCloud, Kubernetes, Low Latency, Apache Kafka, Slurm, Machine Learning Operations, Hardware Infrastructure, Stream Processing, Grpc - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/inference-infrastructure-engineer-rhoda-ai-8292495 ## About the Role * 3+ years of experience in ML infrastructure, MLOps, or distributed systems * Strong proficiency with Kubernetes and containerized deployment pipelines * Experience with GPU orchestration and resource scheduling across large distributed jobs * Experience with cloud providers (e.g., AWS, GCP) and hybrid cloud/on-prem infrastructure * Familiarity with ML frameworks (e.g., PyTorch, JAX) and model serving tools (e.g., Triton, Ray Serve, TorchServe) * Strong debugging instincts and ownership mentality - comfortable driving issues to resolution across the stack Nice to Have (But Not Required) * Experience with streaming systems or high-throughput data transport (e.g., Kafka, gRPC, NATS) * Background in networking, low-latency systems, or network-aware scheduling * Experience with edge/cloud hybrid deployment patterns and the latency constraints that come with them * Familiarity with on-robot or embedded inference environments * Experience with large-scale cluster topology and scheduling systems (e.g., SLURM, Ray, Volcano) ## Description We're looking for an Inference Infrastructure Engineer to help build and operate the systems that power our model deployment stack. You'll be responsible for running large foundation models efficiently and reliably across cloud and on-prem environments, with a focus on resource management, scheduling, and infrastructure scalability. What You'll Do * Design and operate large-scale infrastructure to run model workloads across cloud and on-prem environments * Build and maintain Kubernetes-based deployment pipelines for managing distributed ML workloads * Own resource scheduling and orchestration across GPU clusters - optimizing utilization, workload balancing, and cost-performance tradeoffs * Integrate and manage ML frameworks and model serving systems (e.g., Triton, Ray Serve, TorchServe) across research and production use cases * Build tooling for model deployment, versioning, and observability to support fast iteration cycles * Contribute to the reliability and scalability of the infrastructure stack as model complexity and deployment footprint grow ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Exploring the Power of gRPC-Gateway for Writing RESTful Services](https://www.wearedevelopers.com/videos/2072-exploring-the-power-of-grpc-gateway-for-writing-restful-services) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)