Sr. SWE Datacenter Automation

Ziplines, Inc.
South San Francisco, CA, United States
6 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Link Aggregation (Ethernet) Application Programming Interfaces (APIs) Intelligent Platform Management Interface Border Gateway Protocol BIOS Continuous Integration VMware ESX Servers Firmware Python (Programming Language) Kernel-Based Virtual Machine Network Troubleshooting Network Configuration and Change Management
+14 more
Routing Quick EMUlator (QEMU) Ansible Virtual Local Area Networks Ceph (Software) Computer Networking Systems Mttr Reliability of Systems Generative AI Kubernetes Bare Metal Software Coding Terraform Nvme

Job description

You will be a Senior Software Engineer on Zipline’s Infrastructure team, owning the infrastructure that runs our global autonomous delivery platform. Zipline operates safety-critical logistics at scale: we run private datacenters and edge compute to support flight operations, simulation, CI/CD, and production services that deliver millions of flights and time-sensitive medical deliveries. This role sits at the intersection of hardware, virtualization, orchestration, and automation. Your work directly impacts deployment velocity, compute cost, and the reliability of systems supporting Zipline’s operations., * Own end-to-end lifecycle for datacenter compute and storage: bare-metal provisioning, hypervisor management, SAN/NVMe storage clusters, network configuration, and Kubernetes cluster lifecycle.

  • Design, build, and operate automation that reduces manual setup time and increases deployment velocity: PXE/firmware workflows, dynamic inventory, image generation, fleet-wide configuration drift detection, and automated recovery playbooks.
  • Deliver measurable reliability and scale improvements: set SLIs/SLOs for provisioning time, node commissioning success rate, cluster upgrade success rate, and mean time to recover (MTTR); own meeting those targets.
  • Lead cross-functional runbook and incident ownership for infra incidents affecting flight operations or telemetry: on-call rotation, incident commander for datacenter platform incidents, postmortems and action items.
  • Instrument and maintain monitoring, alerting, and dashboards for hardware health, hypervisor performance, storage latency, Kubernetes control plane health, and cluster autoscaling behavior.
  • Implement cost, capacity, and lifecycle management: capacity planning for compute/storage, automated reclamation, firmware/BIOS/hypervisor patch pipelines, and cold-standby / failover procedures for critical systems.
  • Execute hands-on tasks when required: racking and cabling in datacenters, troubleshooting hardware failures, capture forensic logs, and coordinate physical repairs with vendors and field ops.

Requirements

  • 5+ years of engineering experience with at least 4 years owning production datacenter, virtualization, or infrastructure automation systems.
  • Deep, hands-on expertise with bare-metal provisioning and imaging (PXE/iPXE, IPMI, Redfish), hypervisors (KVM/qemu, ESXi or equivalent), and storage systems (Ceph, NVMeoF, SAN) at scale.
  • Proven Kubernetes operations experience: cluster provisioning, upgrades, control-plane HA, kubeadm/cluster API or equivalent, CNI and CSI troubleshooting, and workload scheduling at multi-cluster scale.
  • Production-grade automation and coding skills in one or more languages (Python, Go, or Rust) and experience with CI/CD pipelines, Terraform/Ansible/Helm, and GitOps practices.
  • Strong networking fundamentals: VLANs, BGP/EVPN at leaf/spine, LACP, routing, and network troubleshooting for cluster networking and storage fabrics.
  • On-call and incident experience: you have owned postmortems, SLIs/SLOs, and driven reliability improvements under operational pressure.
  • Physical datacenter readiness: able to work on-site in South San Francisco HQ with regular in-office cadence, plus occasional travel to partner datacenters or field sites and hands-on rack/cable/repair work when required.
  • Security and safety mindset: experience operating in regulated or safety-sensitive environments, following change control and audit processes.
  • Clear communication and cross-team ownership: you will partner with flight software, field ops, hardware, and SRE teams and must translate operational needs into automated, testable systems.

About the company

Zipline is the world’s largest and most experienced drone delivery service. We are on a mission to serve all humans equally by ensuring access to food, medicine and essential goods anytime, anywhere. We design, build, and operate the world’s largest autonomous logistics system, delivering critical supplies quickly and reliably. Today, Zipline operates on four continents, makes a delivery somewhere in the world every 30 seconds, and has completed millions of deliveries to date, including blood, vaccines, medical supplies, food, and retail products.

Our customers include the world’s largest and most prominent healthcare systems, governments, retailers, restaurants and global businesses who rely on us to save lives, reduce emissions, increase economic opportunity, and provide delivery from point A to point B as fast as possible. The drone is only 15% of what we’ve built to enable seamless, reliable, global operations.

Our system strengthens supply chains, reduces congestion, and gives people time back. With more than 140 million commercial autonomous miles safely flown, Zipline is redefining access to healthcare, consumer products, and food across the globe.

We operate at a global scale and are looking for practical problem solvers who thrive on real-world challenges and rapid growth. Our team is motivated by building systems that have a direct, meaningful impact on people’s lives and by scaling the future of logistics. We are seeking people who sculpt from first principles, enjoy facing adversity, and can do the impossible at record breaking speeds.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

41 sec

Massive client data loss and bio-digital storage

Chris Heilmann +1 · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

4:01 min

Implementing the barbell strategy and focusing on recovery time

Jan de Vries Jan de Vries · World Congress 2026 Europe

Videos

See all

Related articles

See all