Lead Platform Engineer

DKMRBH Inc.
South San Francisco, CA, United States
2 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Big Data Cloud Computing Cloud Engineering Extract Transform Load (ETL) Software Debugging Linux Programming Tools Microprocessors Distributed Systems
+24 more
Django Web Framework Fault Tolerance Python (Programming Language) Operational Data Store Queueing Systems RabbitMQ Service Development Studio Software Engineering Workflow Management Systems Graphics Processing Unit (GPU) Reliability of Systems Backend Fastapi Event Driven Architecture Containerization Infrastructure Automation Frameworks Apache Kafka Celery Hardware Infrastructure Restful APIs Amazon Simple Queue Service (SQS) Serverless Computing Docker Microservices

Job description

We are seeking two hands-on Lead Platform Engineers to help build and scale the backend and cloud infrastructure supporting a global DNA-sequencing operation., The role will initially be highly hands-on, with the engineer learning the existing systems and taking ownership of major technical areas. Over time, the position will transition into approximately:

  • 70% hands-on architecture and software engineering
  • 30% technical leadership, mentoring, and team leadership, The platform supports a globally distributed sequencing operation involving:
  • Laboratory robots and DNA sequencers
  • On-premises Linux infrastructure
  • AWS cloud environments
  • Terabytes of sequencing data movement
  • Hundreds of thousands to millions of bioinformatics jobs
  • Distributed queues and asynchronous workloads
  • Workflow orchestration and scheduling
  • Retry and failure-handling mechanisms
  • Distributed state management
  • Fault-tolerant production systems
  • Large-scale compute infrastructure
  • AI agents that interact with internal tools, APIs, operational data, and physical workflows

The infrastructure currently supports approximately 5,000 CPUs, 12.5 TB of RAM, and 100+ GPUs, with significant production AWS usage., * Design, build, and operate high-scale distributed backend and platform systems.

  • Develop production backend services primarily using Python.
  • Build and maintain cloud-native services and infrastructure on AWS.
  • Design systems for asynchronous processing, distributed queues, workflow orchestration, scheduling, retries, state management, and fault tolerance.
  • Build reliable systems for moving large volumes of data between laboratory equipment, on-premises infrastructure, and AWS.
  • Develop services that orchestrate large numbers of compute and bioinformatics workloads.
  • Design and implement scalable APIs, microservices, workers, and event-driven services.
  • Work across application code, cloud infrastructure, Linux systems, data movement, and operational tooling.
  • Architect and implement production systems from concept through deployment and ongoing operation.
  • Troubleshoot complex production issues and improve system reliability, performance, and scalability.
  • Lead major technical projects from design through production.
  • Establish engineering patterns, standards, and best practices for a growing platform team.
  • Mentor and provide technical guidance to engineers while remaining deeply hands-on.
  • Collaborate with engineering, infrastructure, security, and scientific teams.
  • Explore and implement practical applications of AI agents and modern AI development tools within production systems., * Python
  • FastAPI, Django, or comparable Python backend frameworks
  • REST APIs
  • Microservices
  • Event-driven architecture
  • Asynchronous processing
  • Backend service development

AWS & Cloud

  • Amazon Web Services (AWS)
  • ECS
  • AWS Batch
  • AWS Step Functions
  • SQS
  • Lambda
  • Cloud-native architecture
  • Infrastructure as Code
  • Production cloud environments

Distributed Systems & Orchestration

  • Distributed systems
  • Distributed computing
  • Task queues
  • Job queues
  • Message queues
  • Workflow orchestration
  • Job orchestration
  • Scheduling
  • Retry mechanisms
  • State management
  • Fault tolerance
  • Failure recovery
  • High-volume workload processing
  • Event-driven services

Requirements

The ideal candidate is a Senior or Staff-level engineer with strong Python, AWS, distributed-systems, and cloud-orchestration experience, who has worked in an early-stage startup environment and personally built systems from the ground up., * 6+ years of professional experience in backend engineering, platform engineering, infrastructure engineering, distributed systems, or closely related software engineering roles.

  • Strong professional experience with Python backend development.
  • Strong hands-on experience with AWS cloud infrastructure and services.
  • Significant experience designing and building distributed production systems.
  • Hands-on experience with task queues, asynchronous processing, event-driven systems, workflow orchestration, scheduling, retries, state management, and fault tolerance.
  • Experience building systems from scratch or from early-stage prototypes through production.
  • Meaningful experience working in an early-stage startup or rapidly scaling engineering environment.
  • Demonstrated ownership of major technical systems or projects.
  • Experience scaling systems, infrastructure, workloads, or engineering platforms as a company grows.
  • Experience providing technical leadership and mentoring engineers.
  • Strong Linux and cloud-infrastructure fundamentals.
  • Ability and willingness to remain approximately 70% hands-on with architecture, coding, debugging, and production engineering., * Celery
  • Amazon SQS
  • Kafka
  • RabbitMQ
  • gRPC
  • Similar distributed messaging or task-processing technologies

Infrastructure

  • Linux
  • Docker / containerized environments
  • Cloud infrastructure
  • Infrastructure automation
  • Production monitoring and troubleshooting

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:37 min

Core concepts of Celery and message broker integration

Jan Giacomelli · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all