Lead Platform Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+24 more
Job description
We are seeking two hands-on Lead Platform Engineers to help build and scale the backend and cloud infrastructure supporting a global DNA-sequencing operation., The role will initially be highly hands-on, with the engineer learning the existing systems and taking ownership of major technical areas. Over time, the position will transition into approximately:
- 70% hands-on architecture and software engineering
- 30% technical leadership, mentoring, and team leadership, The platform supports a globally distributed sequencing operation involving:
- Laboratory robots and DNA sequencers
- On-premises Linux infrastructure
- AWS cloud environments
- Terabytes of sequencing data movement
- Hundreds of thousands to millions of bioinformatics jobs
- Distributed queues and asynchronous workloads
- Workflow orchestration and scheduling
- Retry and failure-handling mechanisms
- Distributed state management
- Fault-tolerant production systems
- Large-scale compute infrastructure
- AI agents that interact with internal tools, APIs, operational data, and physical workflows
The infrastructure currently supports approximately 5,000 CPUs, 12.5 TB of RAM, and 100+ GPUs, with significant production AWS usage., * Design, build, and operate high-scale distributed backend and platform systems.
- Develop production backend services primarily using Python.
- Build and maintain cloud-native services and infrastructure on AWS.
- Design systems for asynchronous processing, distributed queues, workflow orchestration, scheduling, retries, state management, and fault tolerance.
- Build reliable systems for moving large volumes of data between laboratory equipment, on-premises infrastructure, and AWS.
- Develop services that orchestrate large numbers of compute and bioinformatics workloads.
- Design and implement scalable APIs, microservices, workers, and event-driven services.
- Work across application code, cloud infrastructure, Linux systems, data movement, and operational tooling.
- Architect and implement production systems from concept through deployment and ongoing operation.
- Troubleshoot complex production issues and improve system reliability, performance, and scalability.
- Lead major technical projects from design through production.
- Establish engineering patterns, standards, and best practices for a growing platform team.
- Mentor and provide technical guidance to engineers while remaining deeply hands-on.
- Collaborate with engineering, infrastructure, security, and scientific teams.
- Explore and implement practical applications of AI agents and modern AI development tools within production systems., * Python
- FastAPI, Django, or comparable Python backend frameworks
- REST APIs
- Microservices
- Event-driven architecture
- Asynchronous processing
- Backend service development
AWS & Cloud
- Amazon Web Services (AWS)
- ECS
- AWS Batch
- AWS Step Functions
- SQS
- Lambda
- Cloud-native architecture
- Infrastructure as Code
- Production cloud environments
Distributed Systems & Orchestration
- Distributed systems
- Distributed computing
- Task queues
- Job queues
- Message queues
- Workflow orchestration
- Job orchestration
- Scheduling
- Retry mechanisms
- State management
- Fault tolerance
- Failure recovery
- High-volume workload processing
- Event-driven services
Requirements
The ideal candidate is a Senior or Staff-level engineer with strong Python, AWS, distributed-systems, and cloud-orchestration experience, who has worked in an early-stage startup environment and personally built systems from the ground up., * 6+ years of professional experience in backend engineering, platform engineering, infrastructure engineering, distributed systems, or closely related software engineering roles.
- Strong professional experience with Python backend development.
- Strong hands-on experience with AWS cloud infrastructure and services.
- Significant experience designing and building distributed production systems.
- Hands-on experience with task queues, asynchronous processing, event-driven systems, workflow orchestration, scheduling, retries, state management, and fault tolerance.
- Experience building systems from scratch or from early-stage prototypes through production.
- Meaningful experience working in an early-stage startup or rapidly scaling engineering environment.
- Demonstrated ownership of major technical systems or projects.
- Experience scaling systems, infrastructure, workloads, or engineering platforms as a company grows.
- Experience providing technical leadership and mentoring engineers.
- Strong Linux and cloud-infrastructure fundamentals.
- Ability and willingness to remain approximately 70% hands-on with architecture, coding, debugging, and production engineering., * Celery
- Amazon SQS
- Kafka
- RabbitMQ
- gRPC
- Similar distributed messaging or task-processing technologies
Infrastructure
- Linux
- Docker / containerized environments
- Cloud infrastructure
- Infrastructure automation
- Production monitoring and troubleshooting
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Find a Developer Job: 12 Best Job Sites For Developers
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Fully Remote Software Engineer Jobs
The 7 Most Popular Backend Frameworks for Developers