Staff Software Engineer, Replication Foundations

Posted 2 Days Ago
Hiring Remotely in Seattle, WA, USA
In-Office or Remote
170K-278K Annually
Entry level
Software
The Role
Lead the architecture, design, implementation, and operation of Temporal’s distributed replication stack. Drive replication protocols for high availability, cross-cluster and cross-region failover, and workload migration. Define consistency, ordering, recovery, and performance guarantees; improve scalability and reliability; debug production issues; influence technical direction across teams; mentor engineers; and contribute to Temporal OSS and cloud infrastructure.
Summary Generated by Built In
Role Summary

We’re hiring a Staff Software Engineer to join the Replication Foundations team within Temporal’s Cloud Global Services (CGS) organization.

Replication Foundations owns and evolves Temporal’s core replication stack in Temporal OSS—the distributed systems backbone behind key Temporal Cloud capabilities such as High Availability namespaces, cross-cluster and cross-region failover, and migration products that enable customers to move workloads between self-hosted Temporal and Temporal Cloud. The team also builds foundational scalability and reliability mechanisms that support Temporal at scale.

In this role, you’ll help set the technical direction for Temporal’s distributed replication systems. You’ll lead complex, correctness-critical initiatives spanning architecture, design, implementation, rollout, and operations. You’ll work across teams to evolve reliable and scalable replication capabilities that support both the open source project and Temporal Cloud.

 
What You'll Do
  • Set the technical direction and evolve the architecture of Temporal’s OSS replication stack, from problem definition through rollout and operational support.

  • Lead the design and implementation of replication protocols that power:

  • High Availability namespaces

  • Cross-cluster and cross-region replication

  • Migration between Temporal clusters, including cloud-to-self-hosted and cloud-to-cloud scenarios

  • Drive scalability and reliability initiatives such as:

  • Multi-cell namespaces

  • Enabling a namespace to span multiple clusters

  • Improving load distribution and handling hot spots

  • Define and communicate system-level guarantees, including consistency models, ordering, idempotency, failure recovery, performance, and operational behavior.

  • Identify architectural risks and opportunities, and shape the technical roadmap for replication capabilities that support current and future cloud products.

  • Partner with Cloud Enablement, CGS, Product, and other engineering teams to align OSS replication foundations with customer and product needs.

  • Lead design reviews, raise the quality of implementation and testing practices, mentor engineers, and provide technical guidance across the organization.

  • Lead or contribute to debugging complex production issues, incident response, and follow-up improvements related to replication and core system behavior.

What You'll Bring
  • A track record of designing and delivering complex production distributed systems, including systems where correctness, availability, and performance are critical.

  • Deep understanding of distributed systems fundamentals such as replication, partitioning, consistency, fault tolerance, durability, concurrency, and failure recovery.

  • Experience defining system architecture, protocol behavior, invariants, and trade-offs in ambiguous or evolving problem spaces.

  • Experience debugging complex production issues, including concurrency bugs, data inconsistencies, partial failures, and performance bottlenecks.

  • Proficiency writing production-quality concurrent code in Go; experience with Java, C++, or similar systems languages is also welcome.

  • Strong written and verbal communication skills, including the ability to explain complex designs and trade-offs to both technical and cross-functional audiences.

  • Demonstrated ability to influence technical direction across teams, build alignment without direct authority, and mentor engineers.

  • A thoughtful and curious approach to understanding how systems behave under load, failure, and changing workload conditions.

Nice to Have
  • Experience designing or maintaining replication protocols or data-plane infrastructure.

  • Experience with multi-cluster or multi-region architectures, including active-active or active-passive systems.

  • Familiarity with database internals, log-based replication, or event-sourced systems.

  • Prior contributions to large open source projects or distributed systems infrastructure.

Temporal Technologies is an Equal Opportunity Employer. Temporal Technologies does not discriminate on the basis of race, religion, color, sex, gender identity, sexual orientation, age, non-disqualifying physical or mental disability, national origin, veteran status, or any other basis covered by appropriate law. All employment is decided on the basis of qualifications, merit, and business need. We embrace and celebrate differences and diversity.

Temporal is committed to providing access, equal opportunity, and reasonable accommodation for individuals with disabilities in employment, its services, programs, and activities. If you need to request a reasonable accommodation, please let your Recruiter know so we can assist.

Skills Required

  • Track record designing and delivering complex production distributed systems where correctness, availability, and performance are critical
  • Deep understanding of replication, partitioning, consistency, fault tolerance, durability, concurrency, and failure recovery
  • Experience defining system architecture, protocol behavior, invariants, and trade-offs in ambiguous or evolving problem spaces
  • Experience debugging complex production issues, including concurrency bugs, data inconsistencies, partial failures, and performance bottlenecks
  • Proficiency writing production-quality concurrent code in Go
  • Strong written and verbal communication skills for explaining complex designs and trade-offs
  • Ability to influence technical direction across teams and build alignment without direct authority
  • Experience mentoring engineers and providing technical guidance
  • Thoughtful and curious approach to understanding system behavior under load, failure, and changing workload conditions
  • Experience with Java, C++, or similar systems languages
  • Experience designing or maintaining replication protocols or data-plane infrastructure
  • Experience with multi-cluster or multi-region architectures, including active-active or active-passive systems
  • Familiarity with database internals, log-based replication, or event-sourced systems
  • Prior contributions to large open source projects or distributed systems infrastructure

Temporal Technologies Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Temporal Technologies and has not been reviewed or approved by Temporal Technologies.

  • Healthcare Strength — Healthcare coverage is described as 100% employer-paid for medical, dental, and vision, with AD&D, short- and long-term disability, and life insurance included. Feedback suggests this breadth and cost coverage is a strong differentiator for a remote-first employer.
  • Leave & Time Off Breadth — Time off includes unlimited PTO alongside 12 standard holidays and 2 floating holidays. Feedback suggests this structure supports rest and recharge across teams.
  • Wellbeing & Lifestyle Benefits — Wellbeing and remote-work support include a home office stipend, internet reimbursement, WFH meals, Calm app access, a lifestyle spending account, and learning/professional membership budgets. Feedback suggests these perks enhance overall total rewards beyond base pay.

Temporal Technologies Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Bellevue, Washington
501 Employees
Year Founded: 2019

What We Do

Temporal develops and distributes the world's leading open source durable execution system. We make code fault tolerant, durable and simple. Innovative companies like Datadog, Glovo, Indeed, Netflix, Qualtrics, Remitly, Snap and Yum! Brands build their services and applications with Temporal to make them reliable to run, productive to enhance and easy to troubleshoot and repair. More than a decade in the making, Temporal is powered by veterans behind some of the industry's most loved systems technologies, programming frameworks and open source communities as well as investors like Amplify Partners, Sequoia Capital and Index Ventures.

Similar Jobs

Toast Logo Toast

Senior Product Designer

Cloud • Fintech • Food • Information Technology • Software • Hospitality
Remote
USA
5000 Employees
159K-254K Annually

Pluralsight Logo Pluralsight

Director, Growth & Digital Marketing

Edtech • Information Technology • Software
Remote or Hybrid
USA
1000 Employees
148K-195K Annually

Boomi Logo Boomi

Enterprise Account Executive

Cloud • Information Technology • Productivity • Software • Automation
Remote
United States of America
2200 Employees
315K-394K Annually

Imprivata Logo Imprivata

Regional Sales Manager

Healthtech • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
United States
1372 Employees
210K-320K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account