Infrastructure Engineer

Tamarind Bio Inc.
San Francisco, CA, United States
3 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Microsoft Azure Bash Shell Computational Biology Nvidia CUDA DevOps Amazon DynamoDB Monitoring of Systems Python (Programming Language)
+18 more
Machine Learning Open Source Technology Tensorflow Value Engineering Web Applications Pulumi Graphics Processing Unit (GPU) Cloud Platform System Pytorch ReactJS System Availability Kubernetes Infrastructure Automation Frameworks Build Tools Slurm Software Coding Terraform Docker

Job description

We’re looking for two Infrastructure Engineers to lead the scaling of our machine learning inference system. You’ll be responsible for architecting and maintaining infrastructure that serves 150+ biological ML models, scaling our platform several orders of magnitude to meet rapidly growing demand.

You’ll work closely with the founders to design to the constraints of customer needs, unpredictable workloads, and unique Bio-ML models. You’ll work with Kubernetes and other tools to orchestrate containerized workloads, optimize resource allocation, and ensure high availability across our model serving infrastructure.

Most importantly, you should thrive in a fast-paced startup environment where you’ll wear multiple hats, learn new technologies quickly, and help solve novel technical challenges. We value engineering judgment, problem-solving ability, and the capacity to build systems that can evolve with our growing needs.

Techstack:

  • Python, React, AWS (EC2, S3, DynamoDB), Docker, CUDA, Conda, TensorFlow/PyTorch; notebooks; bash/Slurm; APIs & web apps., Our technology sits at the intersection of DevOps, MLOps, and Computational Biology. We deal with problems ranging from scaling ML inference on AWS for hundreds of GPUs to dissecting pdb files with Biopython. We deploy a wide range of open source ML models for customers, navigating between Docker containers, Colab notebooks, bash scripts, slurm jobs, and more.

Requirements

  • Solid programming and automation skills
  • Experience with containerization and orchestration concepts
  • Cloud platform knowledge (AWS/GCP/Azure)
  • Located in the SF Bay Area or able to relocate to the Bay Area
  • Onsite expectation: Team currently onsite in SF ~5 days/week.

Preferred

  • Experience scaling production systems
  • Kubernetes experience
  • Infrastructure as code tools (Terraform, Pulumi)
  • Monitoring and observability tools
  • Experience with GPU workloads

About the company

About Tamarind Bio

We enable any scientist to access AI-powered drug discovery. Thousands of scientists from large pharma companies, top biotechs, and academic institutions use Tamarind to design protein drugs, improve industrial enzymes, and create cutting edge molecules that weren’t feasible until now.

New AI models are quickly eclipsing physics-based tools in computational drug discovery. Scientists often struggle to fine-tune, deploy, and scale these models, leaving breakthroughs on the table. Tamarind provides a simple interface to the vast array of tools being released daily.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

3:20 min

Overview of infrastructure as code tools

Alexander Bubeck · World Congress 2023

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all