Machine Learning Inference Engineer

Oscar Technology
San Francisco, CA, United States
17 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$250,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Python (Programming Language) Machine Learning Pytorch TensorRT Hardware Infrastructure Stable Diffusion Microservices

Requirements

An AI Unicorn startup is hiring a Senior Machine Learning Inference Engineer for a full-time role.

You will be responsible for improving efficiency for AI-native infrastructure powered by generative and multimodal models. The ideal candidate has over 3 years of professional experience and a strong understanding of GPU infrastructure, Python, and PyTorch. This is a highly autonomous role with significant ownership across inference systems and model performance in production.

This role is hybrid in San Francisco Bay Area and offers full benefits and equity.

Experience:

  • Building AI applications at scale from the ground up
  • Strong understanding of GPU infrastructure including Triton, TensorRT, or vLLM frameworks Hands-on experience with Python and PyTorch

  • Building model-serving Microservices
  • Diffusion and Multimodal model experience is a plus

Benefits & conditions

  • Competitive base salary
  • Equity
  • $401k matching
  • Medical coverage

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Open-source community and machine learning frameworks

Gian Marco Iodice Gian Marco Iodice · World Congress 2025

3:07 min

Transitioning architecture to microservices at Netflix

Steve Upton Steve Upton · World Congress 2022

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

3:38 min

The convergence of mobile engineering and machine learning

Sasha Denisov Sasha Denisov · World Congress 2026 Europe

Videos

See all

Related articles

See all