> Markdown version of [/jobs/ext/2718474-staff-software-engineer-ml-infrastructure](https://www.wearedevelopers.com/jobs/ext/2718474-staff-software-engineer-ml-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, ML Infrastructure - **Company:** Voxel, Inc - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Computer Vision, Big Data, Continuous Integration, DevOps, Python (Programming Language), Machine Learning, Software Deployment, Software Systems, Pytorch, Delivery Pipeline, ONNX (Open Neural Network Exchange) Format, Machine Learning Operations, TensorRT - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/staff-software-engineer-ml-infrastructure-voxelai-com-7973399 ## About the Role * 7+ years building and shipping large-scale software systems, with at least 3 years focused on ML infrastructure or large-scale data infrastructure * A track record of being the person who decides the architecture, not just the person who implements it. You've owned tool selection, framework choices, and build-vs-buy calls for systems other engineers depend on * Deep fluency in PyTorch and the modern ML training stack. You know what good experiment tracking looks like, what makes a training pipeline reliable at scale, and where the failure modes live * Strong Python. Performant, maintainable code that holds up in production * A pragmatic shipping orientation. You can tell the difference between architectural decisions that need to be right and ones that can be revisited later, and you don't over-engineer the latter * Strong communication skills. You can explain complex tradeoffs clearly to ML researchers, infra peers, and leadership, * Production experience on AWS (S3, EC2, EKS, or similar) for ML workloads * Hands-on experience with model export and inference optimization (TensorRT, ONNX, or similar), including measuring accuracy and latency tradeoffs against training-time baselines * Experience with modern ML orchestration tools (Ray, Sematic, Flyte, Metaflow, Prefect, or similar) * Familiarity with GPU performance profiling and optimization (Nsight, PyTorch profiler, or similar) * Background in computer vision model training ## Description Voxel's perception system is the technical core of everything we ship. Our models detect human activity, equipment interactions, environmental hazards, and operational state in real time across thousands of cameras in manufacturing, logistics, retail, and pharmaceutical environments. Safety was our wedge; it proved our platform works. Now customers are pulling us into operations: equipment utilization, workflow compliance, process efficiency. Every new use case runs through the perception team., We're hiring a Staff Software Engineer to own ML Infrastructure at Voxel. Our applied ML team is shipping vision models into production every week, across thousands of cameras at Fortune 500 customers, and the infrastructure underneath determines how fast we can move. You'll set the technical direction for how we train, track, and ship vision models, build the foundational systems that the applied ML team relies on, and shape the architectural decisions that will define our ML stack for the next several years. This is a hands-on role. You'll write code, make architecture calls, and own outcomes end to end. You'll partner closely with applied CV engineers, the ML Data team, and the Platform team, and you'll be the technical voice in the room when ML infrastructure tradeoffs come up. What You'll Do * Set the technical direction for ML infrastructure at Voxel: what we build, what we buy, and how the pieces fit together as the team and model portfolio scale * Architect and build the training infrastructure that lets the applied ML team run multiple experiments concurrently and iterate quickly on new architectures (PyTorch, AWS) * Own the train-to-deploy handoff: export trained models to optimized inference formats (TensorRT, ONNX), quantify accuracy and latency impact, and partner with Platform on production deployment * Pick and roll out the experiment tracking and lifecycle stack (Weights & Biases, MLflow, ClearML, or similar) so researchers can run, compare, and reproduce experiments efficiently * Establish DevOps-for-ML best practices (IaC, CI/CD, observability, cost monitoring) so researchers can iterate quickly and safely * Mentor engineers across Vision & AI on ML infrastructure best practices, raising the bar for how the org thinks about training, evaluation, and deployment * Anticipate where the infrastructure needs to be in 12 to 18 months, including the upcoming move to on-device inference, and architect for that future ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)