Software Engineer, ML Infrastructure

Voxel, Inc
San Francisco, CA, United States
3 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Computer Vision Continuous Integration DevOps Python (Programming Language) Software Deployment Software Systems Pytorch ONNX (Open Neural Network Exchange) Format Build Tools
+2 more
Machine Learning Operations TensorRT

Job description

Voxel’s perception system is the technical core of everything we ship. Our models detect human activity, equipment interactions, environmental hazards, and operational state in real time across thousands of cameras in manufacturing, logistics, retail, and pharmaceutical environments. Safety was our wedge; it proved our platform works. Now customers are pulling us into operations: equipment utilization, workflow compliance, process efficiency. Every new use case runs through the perception team.

We’re hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models. You’ll build systems that let our applied ML team train multiple models concurrently, manage experiments and ship optimized models to production. You’ll set technical direction, write code, make architecture calls, and partner closely with applied CV, ML Data and Platform engineers.

What You’ll Do

  • Build and maintain training infrastructure that lets the applied ML team train multiple models concurrently, manage experiments, and iterate quickly on new architectures.
  • Own the train-to-deploy handoff - export trained models to optimized inference formats (TensorRT, ONNX), quantify accuracy and latency impact, and partner with Platform on production deployment.
  • Establish ML experiment tracking and lifecycle management - pick the right tools (Weights & Biases, MLflow, ClearML, or similar) so researchers can run, compare, and reproduce experiments efficiently.
  • Establish DevOps-for-ML best practices on AWS (IaC, CI/CD, observability, cost monitoring) so researchers can iterate quickly and safely.
  • Understand the infra needs of applied ML/CV engineers and design scalable solutions that support model development.

Requirements

  • 4+ years of experience building and shipping large scale software solutions.
  • Hands-on experience building ML training pipelines in PyTorch.
  • Hands-on experience with ML experiment tracking and lifecycle tools (Weights & Biases, MLflow, ClearML, or similar).
  • Experience with AWS (S3, EC2, EKS, or similar) for ML workloads.
  • Strong Python. Write performant code that scales well in production environments.
  • Track record of owning infrastructure end-to-end: scoping, building, shipping, and improving systems that internal teams depend on.
  • Bias toward shipping. You’d rather ship something good this week than something perfect next quarter.
  • Strong communication skills., * Experience with modern ML orchestration tools (Ray, Sematic, Flyte, Metaflow, Prefect, or similar)
  • Familiarity with GPU performance profiling and optimization (Nsight, PyTorch profiler, or similar)
  • Background in computer vision model training

Benefits & conditions

  • Total compensation includes base salary, annual bonus, and equity
  • Comprehensive health, dental, and vision insurance
  • Competitive paid parental leave
  • Unlimited PTO and flexible work arrangements
  • Daily meals in-office, team events, annual company onsite

About the company

Voxel is building the future of Computer Vision and Machine Learning for operations, risk, and safety. We use computer vision and AI to enable existing security cameras to automatically detect hazards and high-risk activities, keep people safe and drive operational efficiencies. Our technology addresses the key cost drivers for workers’ compensation, general liability, and property damage, which cost US employers over $500 billion annually. Our customers include Fortune 500 companies across grocery, retail, manufacturing, food and beverage, logistics, and pharmaceutical distribution. We’ve passed $10M ARR with strong expansion revenue. Based in SF, backed by industry-leading VCs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Open-source community and machine learning frameworks

Gian Marco Iodice Gian Marco Iodice · World Congress 2025

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

Videos

See all

Related articles

See all