Software Engineer, Platform & Inference Organization

Chai Discovery, Inc.
San Francisco, CA, United States
3 days ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Cloud Computing Distributed Systems Software Engineering Graphics Processing Unit (GPU) Autoscaling Caching Machine Learning Operations

Job description

Platform engineers make Chai’s models fast, cheap, and reliable at scale, and enable the outer loop that accelerates research: the infrastructure and software abstractions used to train, eval, and understand models.

You’ll own the serving stack that turns our frontier models into a product scientists depend on: latency, throughput, GPU efficiency, batching, and autoscaling across a large multi-cloud GPU fleet. You’ll also contribute to the work that enables turning raw models into product-ready pipelines, and the experiment and observability tooling that lets a researcher ship faster.

You’ve built high-performance services that developers love, moved ML systems into production at scale, and can see around corners before they become outages.

You’ll work closely with the researchers who train the models, the product engineers who build on them, and the commercial team deploying them to the world’s largest pharma companies.

Requirements

We index on systems judgment, ownership, and the scars that come from having run production infrastructure before. We’re looking for engineers who get obsessed with hard problems and don’t give up easily. We look for:

  • 4+ years building production systems, with real depth in performance, distributed systems, or ML serving
  • Experience optimizing model inference: GPU utilization, batching, quantization, caching, or kernel-level work
  • A platform mindset: you like building the tools and abstractions that make other engineers and researchers faster
  • End-to-end ownership of 24/7 systems, including observability, alerting, and incident response
  • Experience across both 0-to-1 buildouts and 1-to-n scale-ups, with an always-evolving playbook you bring wherever you go
  • The instinct to treat cost and efficiency as first-class constraints, not afterthoughts

A background in biology is not required. What makes the difference is technical excellence, curiosity about the domain, and grit. We offer

The opportunity to work at the vanguard of AI research and frontier biology, with world-class people, on a mission that matters. We protect & promote a culture of high velocity and ownership. We compensate our team accordingly.

Skills: Artificial Intelligence (AI), Autoscaling, Biochemistry, Biology, Biotech and Pharmaceutical, Cloud Computing, Distributed Computing, Drug Discovery, GPU (Graphics Processing Unit), Incident Response, Infrastructure Software, Machine Tool, Product Engineering, Production Systems, Research Skills, Software Engineering, Training/Teaching, Vehicle Fleets

About the company

Chai builds the design suite for molecules. We train frontier models that learn the underlying foundations of biochemical structure and interaction, so scientists can move faster and pursue targets that other methods cannot reach.

AI is reinventing life sciences the same way it reinvented software engineering, and Chai is at the forefront of this shift. Leading pharmaceutical companies like Eli Lilly, Pfizer, and Novartis are adopting our platform to power their drug discovery programs.

We value diverse perspectives and are ready to find greatness in unexpected places.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:42 min

Dissecting artificial intelligence layers from compute to applications

Christian Nagel Christian Nagel +3 · World Congress 2026 Europe

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

3:33 min

Managing node capacity with default cluster autoscaling

Mario-Leander Reimer · World Congress 2023

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo · World Congress 2024

4:13 min

Empowering developers with comprehensive AI software stacks

Markus Hacker Markus Hacker +1 · World Congress 2025

2:33 min

Maintaining prompt structures for prefix caching

Douglas Reiser Douglas Reiser · Europe 2026 Virtual

Videos

See all

Related articles

See all