Head of Engineering, Compute

Posted 3 Days Ago
Be an Early Applicant
Hiring Remotely in United States
Remote
214K-352K Annually
Expert/Leader
Software
The Role
Leads the strategy, architecture, development, and operations of Temporal’s large-scale, multi-tenant compute platform. Builds and manages a high-ownership engineering team, drives roadmap execution, reliability, incident response, capacity planning, fleet utilization, cost efficiency, and customer alignment. Guides distributed-systems decisions involving workload isolation, security, scheduling, resource management, performance, virtualization, serverless infrastructure, and future accelerated GPU compute.
Summary Generated by Built In
Head of Engineering, Compute

Companies at the frontier of the AI revolution run on Temporal. OpenAI runs on Temporal, handling millions of requests. Cursor runs its cloud coding agents on Temporal at over 50 million actions a day across 7M+ workflows, and more than a third of the pull requests its users merge now come from those agents. Lovable, Abridge, and Hebbia build their agents on it too. In the last year alone, AI-native companies executed 1.86 trillion actions on Temporal Cloud, and the curve is still bending upwards. Backed by a recent $300M Series D at a $5B valuation, we are building the durable execution layer the agentic era depends on.

The Compute team owns the layer all of that runs on. We are looking for a Head of Engineering to lead the effort to make any aspects of Temporal's compute invisible to our customers, allowing them to focus on application layer innovation, while we handle the compute muck. This is a rare, build-the-foundation mandate: the compute substrate that the world's most demanding AI workloads will run on. We want a leader who has operated compute at planet scale, thinks in fleets, goodput, and cost-per-unit-of-compute, and pairs that with the operational rigor to run a service that frontier-AI companies bet production on.

What You'll Do
  • Strategic direction for Compute: Own the strategy and standards of excellence for the compute layer that the world's agents run on, across design, delivery, and operations. Build a culture of ownership, quality, and customer-first decision-making.

  • Technical leadership: Lead, hire, and grow a high-ownership team; roll up sleeves, ready to do deep into the trenches, by staying close to design docs and code, rather than managing from a distance. Coach engineers, level them up, and clear the friction that slows them down.

  • Roadmap & trajectory: Drive the arc from today's compute toward the next-generation of compute platforms. Ground prioritization in customer and design-partner feedback, and turn ambiguous, fast-moving requirements into predictable, iterative delivery.

  • Operational excellence: When you run frontier AI in production, reliability is the product. Own operations, run on-call and incident response, and drive blameless postmortems and the systemic fixes that prevent recurrence.

  • Technical depth: Guide the hard architectural decisions for large-scale, multi-tenant compute, where technical concerns cut across workload isolation and security, scheduling, fleet efficiency / utilization / goodput, and performance, while ensuring the platform is reliable and efficient for the workloads that depend on it.

  • Capacity, supply & economics: Own utilization, capacity and supply planning, and the cost-per-unit-of-compute and margin profile of the fleet, across CPU compute today and accelerated compute ahead.

  • Cross-team & customer execution: Partner with leadership, Product, SDK, UX/DX, Security, and design-partner customers to align priorities and unblock delivery. Communicate progress, tradeoffs, and risk clearly to technical and non-technical audiences alike.

What You'll Bring
  • Proven experience leading software engineering teams that build and operate large-scale compute platforms or fleets, with strong operational practices.

  • 15+ years in software and/or infrastructure engineering, including 10+ years of people management and demonstrated ownership of delivery and live-site outcomes.

  • Deep distributed-systems and compute infrastructure depth, with the hands-on judgment to guide architecture and execution rather than from a distance.

  • Experience operating multi-tenant compute that other people's production workloads depend on.

  • Bachelor's degree in Computer Science or related field, or equivalent practical experience; advanced degree a plus.

  • Excellent communication skills, with the ability to partner across engineering, product, and leadership and fold customer feedback into the roadmap.

  • Strong leadership, coaching, and performance management; ability to grow engineers and build a healthy, accountable, high-ownership team.

  • Excellence in execution: planning, prioritization, and delivering iterative milestones in an ambiguous, fast-moving environment while managing unplanned work.

  • Fleet thinking: utilization, goodput, capacity and supply planning, and cost discipline as first-class engineering concerns.

  • Live-site reliability craft: on-call, incident management & response, and postmortem-driven continuous improvement.

  • Strong command of the building blocks of a compute platform: multi-tenant isolation and security, scheduling, and resource management.

  • Ability to review and raise the bar on technical artifacts (design docs, code reviews) across a distributed-systems codebase.

Nice to Have
  • MicroVMs and virtualization (Firecracker, gVisor, Edera) or managed-compute primitives (AWS Fargate, GCP Cloud Run, AWS Lambda), and/or Kubernetes internals.

  • Building serverless or hosted-compute products from 0 to 1, including the rapid-delivery-vs-durable-platform tradeoffs that come with it.

  • Multi-cloud delivery across AWS and GCP.

  • Cold-start, warm-pool, and scheduling/latency optimization for on-demand compute.

  • Agent sandboxes, secure execution of untrusted code, or other AI-agent infrastructure.

  • GPU / accelerated compute: fractional GPUs (MIG, MPS, time-slicing), GPU scheduling, training vs. inference fleets, and multi-tenant GPU isolation.

---

The ideal candidate is a strategic thinker with a hands-on approach, energized by building the compute foundation the AI era runs on. They are comfortable shaping a space that doesn't fully exist yet, obsessed with reliability when customers bet production on it, default to working backwards from customers, and balance the speed to ship 0-to-1 against the durable design a planet-scale platform demands.

Temporal Technologies is an Equal Opportunity Employer. Temporal Technologies does not discriminate on the basis of race, religion, color, sex, gender identity, sexual orientation, age, non-disqualifying physical or mental disability, national origin, veteran status, or any other basis covered by appropriate law. All employment is decided on the basis of qualifications, merit, and business need. We embrace and celebrate differences and diversity.

Temporal is committed to providing access, equal opportunity, and reasonable accommodation for individuals with disabilities in employment, its services, programs, and activities. If you need to request a reasonable accommodation, please let your Recruiter know so we can assist.

Skills Required

  • 15+ years of experience in software and/or infrastructure engineering
  • 10+ years of people management experience
  • Experience leading software engineering teams that build and operate large-scale compute platforms or fleets
  • Strong operational practices and ownership of delivery and live-site outcomes
  • Deep distributed-systems and compute infrastructure expertise
  • Experience operating multi-tenant compute supporting production workloads
  • Bachelor’s degree in Computer Science or a related field, or equivalent practical experience
  • Excellent communication and cross-functional partnership skills
  • Strong leadership, coaching, and performance management abilities
  • Ability to plan, prioritize, and deliver iterative milestones in ambiguous environments
  • Experience with utilization, goodput, capacity and supply planning, and cost discipline
  • Experience with on-call, incident management, incident response, and postmortem-driven improvement
  • Strong understanding of multi-tenant isolation and security, scheduling, and resource management
  • Ability to review design documents and code across a distributed-systems codebase
  • Advanced degree
  • Experience with microVMs, virtualization, or managed compute primitives
  • Experience with Kubernetes internals
  • Experience building serverless or hosted-compute products from zero to one
  • Multi-cloud delivery across AWS and GCP
  • Experience optimizing cold starts, warm pools, scheduling, or latency
  • Experience with agent sandboxes or secure execution of untrusted code
  • GPU or accelerated compute experience, including GPU scheduling and multi-tenant GPU isolation

Temporal Technologies Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Temporal Technologies and has not been reviewed or approved by Temporal Technologies.

  • Healthcare Strength Healthcare coverage is described as 100% employer-paid for medical, dental, and vision, with AD&D, short- and long-term disability, and life insurance included. Feedback suggests this breadth and cost coverage is a strong differentiator for a remote-first employer.
  • Leave & Time Off Breadth Time off includes unlimited PTO alongside 12 standard holidays and 2 floating holidays. Feedback suggests this structure supports rest and recharge across teams.
  • Wellbeing & Lifestyle Benefits Wellbeing and remote-work support include a home office stipend, internet reimbursement, WFH meals, Calm app access, a lifestyle spending account, and learning/professional membership budgets. Feedback suggests these perks enhance overall total rewards beyond base pay.

Temporal Technologies Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Bellevue, Washington
501 Employees
Year Founded: 2019

What We Do

Temporal develops and distributes the world's leading open source durable execution system. We make code fault tolerant, durable and simple. Innovative companies like Datadog, Glovo, Indeed, Netflix, Qualtrics, Remitly, Snap and Yum! Brands build their services and applications with Temporal to make them reliable to run, productive to enhance and easy to troubleshoot and repair. More than a decade in the making, Temporal is powered by veterans behind some of the industry's most loved systems technologies, programming frameworks and open source communities as well as investors like Amplify Partners, Sequoia Capital and Index Ventures.

Similar Jobs

Boeing Logo Boeing

Senior Quality Engineer

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Remote
California, USA
170000 Employees
131K-177K Annually

Boeing Logo Boeing

UAS Operator 3 (Remote/Deployable)

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office or Remote
Bingen, WA, USA
170000 Employees
30-42 Hourly

MetLife Logo MetLife

Consultant

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
55K-70K Annually
Easy Apply
Remote or Hybrid
3 Locations
4405 Employees
111K-168K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account