Machine Learning Inference Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Requirements
An AI Unicorn startup is hiring a Senior Machine Learning Inference Engineer for a full-time role.
You will be responsible for improving efficiency for AI-native infrastructure powered by generative and multimodal models. The ideal candidate has over 3 years of professional experience and a strong understanding of GPU infrastructure, Python, and PyTorch. This is a highly autonomous role with significant ownership across inference systems and model performance in production.
This role is hybrid in San Francisco Bay Area and offers full benefits and equity.
Experience:
- Building AI applications at scale from the ground up
-
Strong understanding of GPU infrastructure including Triton, TensorRT, or vLLM frameworks Hands-on experience with Python and PyTorch
- Building model-serving Microservices
- Diffusion and Multimodal model experience is a plus
Benefits & conditions
- Competitive base salary
- Equity
- $401k matching
- Medical coverage
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
MLOps And AI Driven Development
What Are Large Language Models?
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production