> Markdown version of [/jobs/ext/2716641-software-engineer-ml-infrastructure](https://www.wearedevelopers.com/jobs/ext/2716641-software-engineer-ml-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, ML Infrastructure - **Company:** Voxel, Inc - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Computer Vision, Continuous Integration, DevOps, Python (Programming Language), Software Deployment, Software Systems, Pytorch, ONNX (Open Neural Network Exchange) Format, Build Tools, Machine Learning Operations, TensorRT - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-ml-infrastructure-voxelai-com-7973398 ## About the Role * 4+ years of experience building and shipping large scale software solutions. * Hands-on experience building ML training pipelines in PyTorch. * Hands-on experience with ML experiment tracking and lifecycle tools (Weights & Biases, MLflow, ClearML, or similar). * Experience with AWS (S3, EC2, EKS, or similar) for ML workloads. * Strong Python. Write performant code that scales well in production environments. * Track record of owning infrastructure end-to-end: scoping, building, shipping, and improving systems that internal teams depend on. * Bias toward shipping. You'd rather ship something good this week than something perfect next quarter. * Strong communication skills., * Experience with modern ML orchestration tools (Ray, Sematic, Flyte, Metaflow, Prefect, or similar) * Familiarity with GPU performance profiling and optimization (Nsight, PyTorch profiler, or similar) * Background in computer vision model training ## Description Voxel's perception system is the technical core of everything we ship. Our models detect human activity, equipment interactions, environmental hazards, and operational state in real time across thousands of cameras in manufacturing, logistics, retail, and pharmaceutical environments. Safety was our wedge; it proved our platform works. Now customers are pulling us into operations: equipment utilization, workflow compliance, process efficiency. Every new use case runs through the perception team. We're hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models. You'll build systems that let our applied ML team train multiple models concurrently, manage experiments and ship optimized models to production. You'll set technical direction, write code, make architecture calls, and partner closely with applied CV, ML Data and Platform engineers. What You'll Do * Build and maintain training infrastructure that lets the applied ML team train multiple models concurrently, manage experiments, and iterate quickly on new architectures. * Own the train-to-deploy handoff - export trained models to optimized inference formats (TensorRT, ONNX), quantify accuracy and latency impact, and partner with Platform on production deployment. * Establish ML experiment tracking and lifecycle management - pick the right tools (Weights & Biases, MLflow, ClearML, or similar) so researchers can run, compare, and reproduce experiments efficiently. * Establish DevOps-for-ML best practices on AWS (IaC, CI/CD, observability, cost monitoring) so researchers can iterate quickly and safely. * Understand the infra needs of applied ML/CV engineers and design scalable solutions that support model development. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers)