> Markdown version of [/jobs/ext/3416610-ml-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/3416610-ml-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Infrastructure Engineer - **Company:** REBAR - **Location:** New York, NY, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Software as a Service, Cloud Computing, Software Debugging, Python (Programming Language), Azure Machine Learning, Management of Software Versions, Datadog, Graphics Processing Unit (GPU), Pytorch, Large Language Models, ONNX (Open Neural Network Exchange) Format, Machine Learning Operations, TensorRT - **Published:** September 20, 2026 - **Apply:** https://startup.jobs/software-engineer-ml-infrastructure-rebar-8356558 ## About the Role You should feel confident designing developer-facing APIs and SDKs, integrating disparate cloud and SaaS services into coherent systems, and obsessing over the experience of the engineers who use what you build. We're seeking someone with strong platform-engineering instincts who enjoys turning fragmented workflows into products teams actually want to use. This role is a great fit if you have taste in abstractions, opinions about developer experience, and a track record of making ML or data teams meaningfully more productive., * 6+ years of experience building production backend systems, with significant time on internal developer platforms, ML platforms, or integration-heavy infrastructure work. * Strong Python + PyTorch, including profiling and debugging below the model code * 3+ years of experience with cloud infrastructure (AWS preferred) * Hands on GPU performance and inference work. Ideally you have ptimizations in production: batching, mixed precision, quantization, ONNX Runtime/TensorRT or torch.compile * A proven track record operating inference at large scale across a range of model types - detection, segmentation, recognition, and LLM/VLM workloads. * Experience managing a model zoo / model registry - versioning, promotion, and governance of models from experiment to production. Nice to Have * Experience integrating common ML tooling - experiment trackers (W&B, MLflow), feature stores, model serving frameworks - into broader platforms. * Experience with DAG / workflow orchestration frameworks such as Temporal, Prefect, or Apache Airflow. * Built a Backstage-style internal developer portal or comparable internal platform. * Familiarity with GPU compute providers (AWS, Lambda Labs, CoreWeave, RunPod). * Some ML practitioner background - you've trained or deployed models yourself and understand the workflow from the user's side. * Experience with deployment and monitoring pipelines for ML systems. ## Description We're hiring a Software Engineer to own and expand the ML infrastructure platform our ML engineers depend on to rapidly iterate, experiment, and ship models. This work will span feature pipelines, training infra, evaluation, deployment, and monitoring. You should be well versed with GPU architecture and performance focused - understanding what is blocking the engineers as well as our MFU. You'll be joining a small group of talented engineers focused on delivering practical, production-ready ML systems in a fast-moving startup context. The role is ideal for someone who is obsessive over the developer experience of the engineers they support, is interested in training, and meticulous over performance. Our team is still (somewhat) lean and engineers wear many hats. Your purview will span the whole ML lifecycle. Responsibilities * Platform and Developer Experience + You'll help build services that act as the single front door to our ML platform * Infrastructure Consolidation and Integration + Smooth the seams for GPU inference and our cloud stack + Right now, we utilize AWS + Temporal + Modal all in tandem, but some of this has some friction and sticking points. You will optimize and refine our architecture * Observability and Operations + We utilize DataDog for our observability. You should be comfortable extending DD and Modal support to have clear GPU util metrics as well as clean insights into our entire inference flow * Collaboration and Roadmap + You will work closely with ML engineers to understand their workflows, turn one-off scripts into self-serve platform features, and participate in architecture and roadmap decisions. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [The Memory Leak That Ate Our Cluster: A Postmortem](https://www.wearedevelopers.com/videos/2057-the-memory-leak-that-ate-our-cluster-a-postmortem) - [Fully Orchestrating Databricks from Airflow](https://www.wearedevelopers.com/videos/336-fully-orchestrating-databricks-from-airflow) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)