technical ML infrastructure engineer

Quince Therapeutics, Inc.
Palo Alto, United States
4 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$218,000.0 - $285,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Big Data Cloud Computing Code Review Continuous Integration Tensorflow Pulumi Cloud Platform System Data Ingestion Pytorch Apache Spark Kubernetes
+10 more
Infrastructure Automation Frameworks Apache Flink Deployment Automation Performance Monitor Apache Kafka Machine Learning Operations Terraform Software Version Control Data Pipelines Docker

Job description

  • Architect the ML Infrastructure Foundation: Own the end-to-end technical design of Quince’s ML platform - including model training, serving, feature pipelines, and monitoring - ensuring it is modular, scalable, and built for long-term extensibility.
  • Build the “Paved Road” for Production: Design and implement the core developer experience for Quince’s Data Scientists and AI Researchers, enabling them to move from “idea to production” with minimal friction and maximum reliability.
  • Drive Technical Excellence Across the Stack: Set and uphold engineering standards in CI/CD for ML, Infrastructure as Code (IaC), model versioning, experiment tracking, and deployment strategies (blue-green, canary) - and build the tooling that makes those standards the path of least resistance.
  • Own High-Impact System Design Decisions: Lead the technical evaluation and selection of core platform components - from inference runtimes and feature stores to orchestration frameworks - with a clear-eyed view of build vs. buy tradeoffs.
  • Optimize Compute Performance & Cost: Design and implement GPU utilization optimizations, model batching strategies, and cloud cost controls to maximize performance per dollar across training and inference workloads.
  • Ensure Production Scalability & Reliability: Architect ML serving infrastructure that gracefully handles traffic surges, seasonal spikes, and model version transitions, with robust monitoring, alerting, and automated recovery.
  • Mentor and Elevate the Engineering Team: Provide deep technical mentorship to junior and mid-level engineers through design reviews, code reviews, and pairing sessions - raising the collective technical bar without adding process overhead.
  • Champion Operational Excellence: Lead root-cause analyses (RCAs) for production failures and drive systemic, permanent fixes over reactive patches. Model a culture of rigorous on-call discipline and accountability., Quince is committed to providing reasonable accommodations to qualified individuals with disabilities. If you need a reasonable accommodation to complete your application or to perform the essential functions of a role at Quince, please let us know by completing this accommodation form. We review all requests individually and will work with you to determine appropriate accommodations on a case-by-case basis.

Employment is contingent upon successful completion of a background check. Quince will conduct background checks in compliance with applicable federal, state, and local laws.

Requirements

The ideal candidate is a deeply technical ML infrastructure engineer who combines hands-on mastery with system-level thinking. You have built and operated production-grade ML systems at scale - from distributed training pipelines and feature stores to high-throughput inference serving - and you take pride in engineering platforms that other engineers love to use. You don’t just build for today’s requirements; you design for extensibility, observability, and resilience.

You are the kind of engineer who gravitates toward the hardest problems - whether that’s optimizing GPU utilization at the tail of the cost curve, designing a zero-downtime model deployment system, or defining the architectural patterns that will define how Quince industrializes AI at scale. You operate with high autonomy, hold yourself to exceptional standards, and elevate the engineers around you through code reviews, technical mentorship, and by setting a bar for what great looks like., * 8+ years of industry experience, with at least 4+ years of focused, hands-on work in ML Infrastructure, MLOps, or large-scale Data Platform engineering.

  • Proven track record of designing and building MLOps platforms that support the full model lifecycle - from data ingestion and distributed training to real-time inference and model governance.
  • Deep expertise in cloud-native infrastructure (preferably AWS), Kubernetes (EKS), Docker, and Infrastructure as Code tools (Terraform/Pulumi).
  • Hands-on mastery of ML frameworks such as PyTorch, TensorFlow, Kubeflow, or SageMaker, with strong opinions on building a cohesive, high-leverage developer experience.
  • Expertise in building Feature Stores and high-throughput data pipelines (Spark, Flink, Kafka), with a strong understanding of training/serving skew and data consistency.
  • Expert-level knowledge of CI/CD for ML, including model versioning, experiment tracking, and deployment strategies such as blue-green and canary rollouts.
  • Demonstrated ability to optimize GPU utilization, implement model batching, and systematically reduce cloud infrastructure costs.
  • Strong operational instincts, with a history of improving reliability through rigorous on-call practices, proactive monitoring, and root-cause analysis.
  • You understand the hustle of a startup and are good at handling ambiguity. You are a curious, quick learner who loves to experiment and thrives at a rapid pace.

Benefits & conditions

Pay Range: $218,000-$285.000 (base) + bonus and stock

All posted ranges are reflective of base salary and may vary depending upon experience level and location. Bonus and equity may also be provided for eligible roles. Pay Range $218,000-$285,000 USD

About the company

Joining Quince means being part of a mission-driven team reshaping retail. You will work alongside talented colleagues, tackle meaningful challenges, and contribute to building a more sustainable, accessible future for customers and partners alike.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all