> Markdown version of [/jobs/ext/3037847-software-engineer-ml-systems](https://www.wearedevelopers.com/jobs/ext/3037847-software-engineer-ml-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, ML Systems - **Company:** Voxel, Inc - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Computer Vision, C++ (Programming Language), Configuration Management, Code Review, Encodings, Distributed Computing Environment, Python (Programming Language), BIG-IP Global Traffic Manager (GTM), DataOps, Data Streaming, State Machines, Backend, Build Management, Kubernetes, Infrastructure Automation Frameworks, Apache Flink, Apache Kafka, Machine Learning Operations, TensorRT, Terraform, Docker - **Published:** September 23, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-ml-systems-voxelai-com-10158033 ## About the Role * 5+ years building and shipping production ML or computer vision systems, with a track record of moving techniques from research to reliable product * Strong Python; working knowledge of C++ or Go is a plus * Experience with model serving and inference optimization (Triton, TensorRT, or similar) * Experience with streaming or distributed data processing (Flink, Kafka, or similar) * Solid grounding in evaluation: building datasets, defining metrics, and detecting regressions before customers do * Familiarity with Kubernetes, Docker, and infrastructure as code (Terraform, ArgoCD); AWS experience Nice to have * Experience with the full CV data lifecycle: collection, sampling, labeling, and curation * Experience with auto-labeling or embedding-based data search * Exposure to customer deployments or field-facing ML systems The soft skills that matter most * Technical leadership without authority: you drive alignment across engineering, platform, GTM, and data ops through clear thinking and communication * Ownership: you follow a feature past "the model works" to "customers trust it" * Pragmatism: you know when a research result is ready for production and when it needs more work * Mentorship: you raise the bar for teammates through design reviews, code review, and sharing context * Comfort with ambiguity: you help shape an architecture in flux and can validate it with a fast proof of concept Why this role You'll have real influence over the architecture of a system that's being rebuilt, with a direct line from your work to customer outcomes. ## Description Our perception team turns raw video into reliable, customer-facing incidents. We're re-architecting our perception system around stream-based processing, and we need an engineer who can take research-grade computer vision techniques and make them work in production: scalable, observable, and integrated cleanly into the wider product. You'll be a technical leader on the team, setting direction through design and execution rather than people management. What you'll do * Turn research ideas (transformers, ensemble models, new detection and classification approaches) into production features, from prototype through evaluation, rollout, and iteration * Design and build perception pipeline components on our stream-processing architecture (Flink/PyFlink), working with the platform team on model serving, configuration management, and orchestration * Own and improve production CV models, including training, evaluation, and inference optimization (Triton, TensorRT, multi-backend serving) * Design the logic that converts model outputs into trustworthy incidents: tracking, filtering, probabilistic reasoning, and state machines * Build observability into everything you ship: metrics, dashboards, and evaluation loops that show whether a model or pipeline is working in the field * Reduce end-to-end latency and cost while keeping detection quality high * Partner with GTM, customer success, and data operations to plan customer rollouts, tune deployments, and improve labeling and review quality * Evaluate and integrate third-party ML tooling (orchestration, labeling, dataset curation) and drive adoption across the team * Write clear design docs and requirements, and de-risk major changes with proofs of concept before full buildouts ## Related Videos - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [MLOps - What’s the deal behind it?](https://www.wearedevelopers.com/videos/392-mlops-what-s-the-deal-behind-it) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)