Machine Learning Engineer, Ops

Cantina Labs
United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$125,000.0 - $165,000.0
Working hours
Regular working hours
Job source

Tech stack

Cloud Engineering Continuous Integration Python (Programming Language) Machine Learning Software Engineering Speech Recognition Delivery Pipeline Backend Kubernetes Machine Learning Operations Speech Synthesis

Job description

We are looking for an MLOps Engineer to build and scale the inference infrastructure for our generative audio models, including Text-to-Speech (TTS), voice conversion, and Automatic Speech Recognition (ASR). You will be responsible for designing and deploying high-performance systems that ensure low-latency, reliable, and scalable model serving for both streaming and batch inference. This role is central to bridging the gap between research and production, ensuring our audio models are optimized for performance and cost-efficiency as we scale.

What You’ll Do:

  • Design and maintain inference infrastructure for generative audio model architectures.
  • Implement and manage high-performance inference engines.
  • Orchestrate service deployments using Kubernetes (K8S), implementing advanced autoscaling paradigms to handle varying traffic loads efficiently.
  • Develop and automate robust CI/CD pipelines to streamline the testing and deployment of model artifacts and inference configurations.
  • Monitor production systems, establishing observability practices to track latency, resource utilization, and overall model performance.
  • Collaborate closely with research teams to optimize model serving paths and evaluate various inference strategies.
  • Optimize inference performance for both streaming and batch applications.

Requirements

  • Strong hands-on experience with Kubernetes (K8S), container orchestration, and implementing autoscaling strategies for production workloads.
  • Solid background in MLOps, including CI/CD automation and managing scalable cloud infrastructure.
  • Proficiency in software engineering principles and experience with Python or Go for infrastructure tooling and backend services.
  • Experience with GPU-accelerated inference and performance profiling techniques.
  • Familiarity with high-performance inference engines (e.g., Triton Inference Server, vLLM-Omni) is a plus.

Benefits & conditions

Pulled from the full job description

  • Parental leave
  • Health insurance
  • Paid time off
  • Vision insurance
  • Dental insurance, The anticipated annual base salary range for this role is between $125,000-$165,000 (€110,000-€145,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits for U.S.-based roles:

  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance - 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including:
  • 15 PTO days
  • 10 sick days
  • 15 company holidays
  • 2 floating holidays
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account - $500/month to use however you’d like
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership, and more!

About the company

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:35 min

Serving production machine learning workloads with KubeFlow and KServe

Aarno Aukia · LIVE

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

58 sec

Real-world application of MLOps architectural patterns at Wayfair

Bas Geerdink · LIVE

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · World Congress 2024

Videos

See all

Related articles

See all