Sr. SWE Datacenter Automation

Posted 3 Hours Ago
Be an Early Applicant
South San Francisco, CA, USA
In-Office
Senior level
Aerospace • Hardware • Logistics • Robotics • Software • Transportation
Zipline democratizes access to critical medical supplies through instant drone delivery.
The Role
Own and automate datacenter compute and storage lifecycle (bare‑metal, hypervisors, storage, networking, Kubernetes). Build provisioning and recovery automation, set SLIs/SLOs, run on‑call incident response, monitor hardware and clusters, manage capacity/patch pipelines, and perform hands‑on datacenter tasks and vendor coordination.
Summary Generated by Built In
About Zipline

Zipline is the world’s largest and most experienced drone delivery service. We are on a mission to serve all humans equally by ensuring access to food, medicine and essential goods anytime, anywhere. We design, build, and operate the world’s largest autonomous logistics system, delivering critical supplies quickly and reliably. Today, Zipline operates on four continents, makes a delivery somewhere in the world every 30 seconds, and has completed millions of deliveries to date, including blood, vaccines, medical supplies, food, and retail products. 

Our customers include the world’s largest and most prominent healthcare systems, governments, retailers, restaurants and global businesses who rely on us to save lives, reduce emissions, increase economic opportunity, and provide delivery from point A to point B as fast as possible. The drone is only 15% of what we’ve built to enable seamless, reliable, global operations.

Our system strengthens supply chains, reduces congestion, and gives people time back. With more than 140 million commercial autonomous miles safely flown, Zipline is redefining access to healthcare, consumer products, and food across the globe.

We operate at a global scale and are looking for practical problem solvers who thrive on real-world challenges and rapid growth. Our team is motivated by building systems that have a direct, meaningful impact on people’s lives and by scaling the future of logistics. We are seeking people who sculpt from first principles, enjoy facing adversity, and can do the impossible at record breaking speeds.

About You and The Role 

You will be a Senior Software Engineer on Zipline’s Infrastructure team, owning the infrastructure that runs our global autonomous delivery platform. Zipline operates safety‑critical logistics at scale: we run private datacenters and edge compute to support flight operations, simulation, CI/CD, and production services that deliver millions of flights and time‑sensitive medical deliveries. This role sits at the intersection of hardware, virtualization, orchestration, and automation. Your work directly impacts deployment velocity, compute cost, and the reliability of systems supporting Zipline’s operations.

What You'll Do
  • Own end‑to‑end lifecycle for datacenter compute and storage: bare‑metal provisioning, hypervisor management, SAN/NVMe storage clusters, network configuration, and Kubernetes cluster lifecycle.
  • Design, build, and operate automation that reduces manual setup time and increases deployment velocity: PXE/firmware workflows, dynamic inventory, image generation, fleet-wide configuration drift detection, and automated recovery playbooks.
  • Deliver measurable reliability and scale improvements: set SLIs/SLOs for provisioning time, node commissioning success rate, cluster upgrade success rate, and mean time to recover (MTTR); own meeting those targets.
  • Lead cross‑functional runbook and incident ownership for infra incidents affecting flight operations or telemetry: on‑call rotation, incident commander for datacenter platform incidents, postmortems and action items.
  • Instrument and maintain monitoring, alerting, and dashboards for hardware health, hypervisor performance, storage latency, Kubernetes control plane health, and cluster autoscaling behavior.
  • Implement cost, capacity, and lifecycle management: capacity planning for compute/storage, automated reclamation, firmware/BIOS/hypervisor patch pipelines, and cold‑standby / failover procedures for critical systems.
  • Execute hands‑on tasks when required: racking and cabling in datacenters, troubleshooting hardware failures, capture forensic logs, and coordinate physical repairs with vendors and field ops.
What You'll Bring
  • 5+ years of engineering experience with at least 4 years owning production datacenter, virtualization, or infrastructure automation systems.
  • Deep, hands‑on expertise with bare‑metal provisioning and imaging (PXE/iPXE, IPMI, Redfish), hypervisors (KVM/qemu, ESXi or equivalent), and storage systems (Ceph, NVMeoF, SAN) at scale.
  • Proven Kubernetes operations experience: cluster provisioning, upgrades, control‑plane HA, kubeadm/cluster API or equivalent, CNI and CSI troubleshooting, and workload scheduling at multi‑cluster scale.
  • Production‑grade automation and coding skills in one or more languages (Python, Go, or Rust) and experience with CI/CD pipelines, Terraform/Ansible/Helm, and GitOps practices.
  • Strong networking fundamentals: VLANs, BGP/EVPN at leaf/spine, LACP, routing, and network troubleshooting for cluster networking and storage fabrics.
  • On‑call and incident experience: you have owned postmortems, SLIs/SLOs, and driven reliability improvements under operational pressure.
  • Physical datacenter readiness: able to work on‑site in South San Francisco HQ with regular in‑office cadence, plus occasional travel to partner datacenters or field sites and hands‑on rack/cable/repair work when required.
  • Security and safety mindset: experience operating in regulated or safety‑sensitive environments, following change control and audit processes.
  • Clear communication and cross‑team ownership: you will partner with flight software, field ops, hardware, and SRE teams and must translate operational needs into automated, testable systems.
What Else You Need To Know

This role is based in South San Francisco with an expectation of regular on‑site presence and participation in on‑call rotations and occasional travel to datacenter or field locations.

Zipline is an equal opportunity employer and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws or our own sensibilities.

We value diversity at Zipline and welcome applications from those who are traditionally underrepresented in tech. If you like the sound of this position but are not sure if you are the perfect fit, please apply!

Voluntary Self-Identification

For government reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file.

As set forth in Zipline ’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Skills Required

  • 5+ years of engineering experience with at least 4 years owning production datacenter, virtualization, or infrastructure automation systems.
  • Hands‑on bare‑metal provisioning and imaging experience (PXE/iPXE, IPMI, Redfish).
  • Experience with hypervisors such as KVM/qemu and ESXi (or equivalent) at scale.
  • Experience with storage systems at scale (Ceph, NVMeoF, SAN).
  • Proven Kubernetes operations experience: cluster provisioning, upgrades, control‑plane HA, kubeadm/Cluster API, CNI and CSI troubleshooting, multi‑cluster workload scheduling.
  • Production‑grade automation and coding skills in one or more languages (Python, Go, or Rust); experience with CI/CD pipelines.
  • Infrastructure as code and automation tooling experience (Terraform, Ansible, Helm) and GitOps practices.
  • Strong networking fundamentals including VLANs, BGP/EVPN (leaf/spine), LACP, routing, and network troubleshooting for cluster and storage fabrics.
  • On‑call and incident response experience: ownership of postmortems, SLIs/SLOs, and driving reliability improvements.
  • Ability to work on‑site in South San Francisco regularly, occasional travel to partner datacenters/field sites, and hands‑on racking/cabling/repair work.
  • Experience operating in regulated or safety‑sensitive environments, following change control and audit processes.
  • Clear communication and cross‑team ownership across flight software, field ops, hardware, and SRE teams.

Zipline Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Zipline and has not been reviewed or approved by Zipline.

  • Healthcare Strength Healthcare coverage is consistently described as comprehensive and high quality, with multiple plan choices, low out‑of‑pocket costs, and additions like HRA, One Medical, and fertility support.
  • Parental & Family Support Parental and family leave is highlighted as generous and meaningfully used, reinforcing support for major life events.
  • Leave & Time Off Breadth Paid time off, sick days, and holidays are portrayed as solid and broadly available, contributing to overall benefits depth.

Zipline Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: South San Francisco, CA
375 Employees
Year Founded: 2014

What We Do

Zipline is the world's largest autonomous delivery network and is powered entirely by fixed-wing drones. Our fleet circles the equivalent distance of the equator every 2.5 days, and we have shipped hundreds of thousands of critical medical products across Rwanda, Ghana, and now beginning in the United States.

Why Work With Us

Zipline is the perfect intersection of super cutting-edge tech, deep social mission, and extremely compelling business case. Our small, scrappy, customer-obsessed, humble, and mission-driven team has set the bar for what is possible in the drone logistics industry globally, and has designed some incredibly elegant technology in the process.

Gallery

Gallery

Similar Jobs

Order.co Logo Order.co

Quality Assurance Manager

eCommerce • Fintech • Payments • Software
Remote or Hybrid
United States
152 Employees
135K-150K Annually

Allen Control Systems Logo Allen Control Systems

Software Engineer

Defense • Manufacturing
In-Office
2 Locations
12 Employees
In-Office
2 Locations
12 Employees
165K-225K Annually
In-Office
2 Locations
12 Employees
74K-134K Annually

Similar Companies Hiring

Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account