> Markdown version of [/jobs/ext/2717213-ml-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2717213-ml-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Infrastructure Engineer - **Company:** REBAR - **Location:** New York, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Software as a Service, Cloud Computing, Identity and Access Management, Python (Programming Language), Azure Machine Learning, Management of Software Versions, AWS Cdk, Pulumi, Large Language Models, Backend, Machine Learning Operations, Terraform - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-ml-infrastructure-engineer-rebar-8356558 ## About the Role You should feel confident designing developer-facing APIs and SDKs, integrating disparate cloud and SaaS services into coherent systems, and obsessing over the experience of the engineers who use what you build. We're seeking someone with strong platform-engineering instincts who enjoys turning fragmented workflows into products teams actually want to use. This role is a great fit if you have taste in abstractions, opinions about developer experience, and a track record of making ML or data teams meaningfully more productive., * Bachelor's degree or higher in Computer Science, Electrical Engineering, or other relevant field - or equivalent industry experience. * 3+ years of experience building production backend systems, with significant time on internal developer platforms, ML platforms, or integration-heavy infrastructure work. * Expert-level Python; comfortable picking up other languages as the tooling demands. * 2+ years of experience with cloud infrastructure (AWS preferred), including IAM, networking, and cost management. * Proficiency with infrastructure-as-code (IaC) tooling such as Terraform, AWS CDK, or Pulumi for managing reproducible, version-controlled cloud environments. * Hands-on experience with managed ML inference and serving platforms such as AWS SageMaker and GCP Vertex AI. * A proven track record operating inference at large scale across a range of model types - detection, segmentation, recognition, and LLM/VLM workloads. * Experience managing a model zoo / model registry - versioning, promotion, and governance of models from experiment to production. * Proven ability to design clean, composable APIs and SDKs that internal users adopt willingly. * Deep understanding of ML related workflows and requirements. Nice to Have * Experience integrating common ML tooling - experiment trackers (W&B, MLflow), feature stores, model serving frameworks - into broader platforms. * Experience with DAG / workflow orchestration frameworks such as Temporal, Prefect, or Apache Airflow. * Built a Backstage-style internal developer portal or comparable internal platform. * Familiarity with GPU compute providers (AWS, Lambda Labs, CoreWeave, RunPod). * Some ML practitioner background - you've trained or deployed models yourself and understand the workflow from the user's side. * Experience with deployment and monitoring pipelines for ML systems. ## Description Platform & Developer Experience: Design and build the CLI, SDK, and services that serve as the single front door to our ML platform. Make launching a training job, tracking an experiment, or shipping a model feel like one coherent product. Infrastructure Integration: Wire together our cloud and SaaS stack - compute providers, storage, experiment tracking, model serving - into a unified system, codified as version-controlled infrastructure-as-code. Own the abstractions for compute orchestration, feature store, model registry, and model deployment. Observability & Operations: Build cost attribution, usage dashboards, and monitoring across the platform. Surface what's running where, catch problems early, and keep production model serving - across detection, segmentation, recognition, and LLM/VLM workloads - reliable and cost-efficient at scale. Collaboration and Roadmap: Work closely with ML engineers to understand their workflows, turn one-off scripts into self-serve platform features, and participate in architecture and roadmap decisions. ## Related Videos - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Building Reliable Serverless Applications with AWS CDK and Testing](https://www.wearedevelopers.com/videos/812-building-reliable-serverless-applications-with-aws-cdk-and-testing) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)