> Markdown version of [/jobs/ext/2272476-swe-datacenter-automation-organization](https://www.wearedevelopers.com/jobs/ext/2272476-swe-datacenter-automation-organization). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SWE Datacenter Automation Organization - **Company:** Ziplines, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Link Aggregation (Ethernet), Application Programming Interfaces (APIs), Intelligent Platform Management Interface, Border Gateway Protocol, BIOS, Continuous Delivery, Continuous Integration, VMware ESX Servers, Failover, Firmware, Hardware Virtualization, Hypervisor, Python (Programming Language), Kernel-Based Virtual Machine, Network Troubleshooting, Network Configuration and Change Management, Routing, Network Administration, Quick EMUlator (QEMU), Reliability Engineering, Cadence Virtuoso, Ansible, Service Pack, Software Engineering, Virtual Local Area Networks, Virtual Machines, Virtualization Technology, Ceph (Software), Network Routing, Computer Networking Systems, Computer Network Operations, Autoscaling, Mttr, Reliability of Systems, Generative AI, Kubernetes, Bare Metal, Software Coding, Terraform, Nvme - **Published:** August 27, 2026 - **Apply:** https://www.careerbuilder.com/job-details/sr-swe-datacenter-automation-san-francisco-ca--e4b9aabd-789f-4fe9-b852-0ffc3870cc68 ## About the Role * 5+ years of engineering experience with at least 4 years owning production datacenter, virtualization, or infrastructure automation systems. * Deep, hands-on expertise with bare-metal provisioning and imaging (PXE/iPXE, IPMI, Redfish), hypervisors (KVM/qemu, ESXi or equivalent), and storage systems (Ceph, NVMeoF, SAN) at scale. * Proven Kubernetes operations experience: cluster provisioning, upgrades, control-plane HA, kubeadm/cluster API or equivalent, CNI and CSI troubleshooting, and workload scheduling at multi-cluster scale. * Production-grade automation and coding skills in one or more languages (Python, Go, or Rust) and experience with CI/CD pipelines, Terraform/Ansible/Helm, and GitOps practices. * Strong networking fundamentals: VLANs, BGP/EVPN at leaf/spine, LACP, routing, and network troubleshooting for cluster networking and storage fabrics. * On-call and incident experience: you have owned postmortems, SLIs/SLOs, and driven reliability improvements under operational pressure. * Physical datacenter readiness: able to work on-site in South San Francisco HQ with regular in-office cadence, plus occasional travel to partner datacenters or field sites and hands-on rack/cable/repair work when required. * Security and safety mindset: experience operating in regulated or safety-sensitive environments, following change control and audit processes. * Clear communication and cross-team ownership: you will partner with flight software, field ops, hardware, and SRE teams and must translate operational needs into automated, testable systems., Application Programming Interface (API), Automation, Automation Systems, Autoscaling, BGP, Building Systems, Cadence, Capacity Management, Change Control, Commissioning, Communication Skills, Computer Firmware, Continuous Deployment/Delivery, Continuous Integration, Engineering, Failover, Genetics, Government Reporting, Hardware Virtualization, Healthcare, Hypervisors, IPMI (Intelligent Platform Management Interface), Identify Issues, International Operations, K Virtual Machine (KVM), Link Aggregation Control Protocol (LACP), Logistics, Medical Products, Medicine, Network Administration/Management, Network Configuration Management, Network Operations Center, Network Routing, Operational Improvement, Operational Support, Problem Solving Skills, Process Improvement, Reliability Engineering, Reporting Dashboards, Restaurant, Retail, Sculpture, Simulation, Software Engineering, Software Patches, State Laws and Regulations, Storage Area Network (SAN), Systems Administration/Management, Systems Reliability, Telemetry, Time Management, Unmanned Aircraft Systems (UAS), VLAN (Virtual Local Area Network), VMWare ESX/ESXi, Vaccination, Vehicle Fleets, Virtualization, Willing to Travel ## Description You will be a Senior Software Engineer on Zipline's Infrastructure team, owning the infrastructure that runs our global autonomous delivery platform. Zipline operates safety-critical logistics at scale: we run private datacenters and edge compute to support flight operations, simulation, CI/CD, and production services that deliver millions of flights and time-sensitive medical deliveries. This role sits at the intersection of hardware, virtualization, orchestration, and automation. Your work directly impacts deployment velocity, compute cost, and the reliability of systems supporting Zipline's operations. What You'll Do * Own end-to-end lifecycle for datacenter compute and storage: bare-metal provisioning, hypervisor management, SAN/NVMe storage clusters, network configuration, and Kubernetes cluster lifecycle. * Design, build, and operate automation that reduces manual setup time and increases deployment velocity: PXE/firmware workflows, dynamic inventory, image generation, fleet-wide configuration drift detection, and automated recovery playbooks. * Deliver measurable reliability and scale improvements: set SLIs/SLOs for provisioning time, node commissioning success rate, cluster upgrade success rate, and mean time to recover (MTTR); own meeting those targets. * Lead cross-functional runbook and incident ownership for infra incidents affecting flight operations or telemetry: on-call rotation, incident commander for datacenter platform incidents, postmortems and action items. * Instrument and maintain monitoring, alerting, and dashboards for hardware health, hypervisor performance, storage latency, Kubernetes control plane health, and cluster autoscaling behavior. * Implement cost, capacity, and lifecycle management: capacity planning for compute/storage, automated reclamation, firmware/BIOS/hypervisor patch pipelines, and cold-standby / failover procedures for critical systems. * Execute hands-on tasks when required: racking and cabling in datacenters, troubleshooting hardware failures, capture forensic logs, and coordinate physical repairs with vendors and field ops. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [How to Submit CFPs and Get into Public Speaking - Moran Weber](https://www.wearedevelopers.com/videos/2112-how-to-submit-cfps-and-get-into-public-speaking-moran-weber) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)