We’re hiring a Staff Software Engineer to join the Replication Foundations team within Temporal’s Cloud Global Services (CGS) organization.
Replication Foundations owns and evolves Temporal’s core replication stack in Temporal OSS—the distributed systems backbone behind key Temporal Cloud capabilities such as High Availability namespaces, cross-cluster and cross-region failover, and migration products that enable customers to move workloads between self-hosted Temporal and Temporal Cloud. The team also builds foundational scalability and reliability mechanisms that support Temporal at scale.
In this role, you’ll help set the technical direction for Temporal’s distributed replication systems. You’ll lead complex, correctness-critical initiatives spanning architecture, design, implementation, rollout, and operations. You’ll work across teams to evolve reliable and scalable replication capabilities that support both the open source project and Temporal Cloud.
Set the technical direction and evolve the architecture of Temporal’s OSS replication stack, from problem definition through rollout and operational support.
Lead the design and implementation of replication protocols that power:
High Availability namespaces
Cross-cluster and cross-region replication
Migration between Temporal clusters, including cloud-to-self-hosted and cloud-to-cloud scenarios
Drive scalability and reliability initiatives such as:
Multi-cell namespaces
Enabling a namespace to span multiple clusters
Improving load distribution and handling hot spots
Define and communicate system-level guarantees, including consistency models, ordering, idempotency, failure recovery, performance, and operational behavior.
Identify architectural risks and opportunities, and shape the technical roadmap for replication capabilities that support current and future cloud products.
Partner with Cloud Enablement, CGS, Product, and other engineering teams to align OSS replication foundations with customer and product needs.
Lead design reviews, raise the quality of implementation and testing practices, mentor engineers, and provide technical guidance across the organization.
Lead or contribute to debugging complex production issues, incident response, and follow-up improvements related to replication and core system behavior.
A track record of designing and delivering complex production distributed systems, including systems where correctness, availability, and performance are critical.
Deep understanding of distributed systems fundamentals such as replication, partitioning, consistency, fault tolerance, durability, concurrency, and failure recovery.
Experience defining system architecture, protocol behavior, invariants, and trade-offs in ambiguous or evolving problem spaces.
Experience debugging complex production issues, including concurrency bugs, data inconsistencies, partial failures, and performance bottlenecks.
Proficiency writing production-quality concurrent code in Go; experience with Java, C++, or similar systems languages is also welcome.
Strong written and verbal communication skills, including the ability to explain complex designs and trade-offs to both technical and cross-functional audiences.
Demonstrated ability to influence technical direction across teams, build alignment without direct authority, and mentor engineers.
A thoughtful and curious approach to understanding how systems behave under load, failure, and changing workload conditions.
Experience designing or maintaining replication protocols or data-plane infrastructure.
Experience with multi-cluster or multi-region architectures, including active-active or active-passive systems.
Familiarity with database internals, log-based replication, or event-sourced systems.
Prior contributions to large open source projects or distributed systems infrastructure.
Temporal Technologies is an Equal Opportunity Employer. Temporal Technologies does not discriminate on the basis of race, religion, color, sex, gender identity, sexual orientation, age, non-disqualifying physical or mental disability, national origin, veteran status, or any other basis covered by appropriate law. All employment is decided on the basis of qualifications, merit, and business need. We embrace and celebrate differences and diversity.
Temporal is committed to providing access, equal opportunity, and reasonable accommodation for individuals with disabilities in employment, its services, programs, and activities. If you need to request a reasonable accommodation, please let your Recruiter know so we can assist.
Skills Required
- Track record designing and delivering complex production distributed systems where correctness, availability, and performance are critical
- Deep understanding of replication, partitioning, consistency, fault tolerance, durability, concurrency, and failure recovery
- Experience defining system architecture, protocol behavior, invariants, and trade-offs in ambiguous or evolving problem spaces
- Experience debugging complex production issues, including concurrency bugs, data inconsistencies, partial failures, and performance bottlenecks
- Proficiency writing production-quality concurrent code in Go
- Strong written and verbal communication skills for explaining complex designs and trade-offs
- Ability to influence technical direction across teams and build alignment without direct authority
- Experience mentoring engineers and providing technical guidance
- Thoughtful and curious approach to understanding system behavior under load, failure, and changing workload conditions
- Experience with Java, C++, or similar systems languages
- Experience designing or maintaining replication protocols or data-plane infrastructure
- Experience with multi-cluster or multi-region architectures, including active-active or active-passive systems
- Familiarity with database internals, log-based replication, or event-sourced systems
- Prior contributions to large open source projects or distributed systems infrastructure
Temporal Technologies Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Temporal Technologies and has not been reviewed or approved by Temporal Technologies.
-
Healthcare Strength — Healthcare coverage is described as 100% employer-paid for medical, dental, and vision, with AD&D, short- and long-term disability, and life insurance included. Feedback suggests this breadth and cost coverage is a strong differentiator for a remote-first employer.
-
Leave & Time Off Breadth — Time off includes unlimited PTO alongside 12 standard holidays and 2 floating holidays. Feedback suggests this structure supports rest and recharge across teams.
-
Wellbeing & Lifestyle Benefits — Wellbeing and remote-work support include a home office stipend, internet reimbursement, WFH meals, Calm app access, a lifestyle spending account, and learning/professional membership budgets. Feedback suggests these perks enhance overall total rewards beyond base pay.
Temporal Technologies Insights
What We Do
Temporal develops and distributes the world's leading open source durable execution system. We make code fault tolerant, durable and simple. Innovative companies like Datadog, Glovo, Indeed, Netflix, Qualtrics, Remitly, Snap and Yum! Brands build their services and applications with Temporal to make them reliable to run, productive to enhance and easy to troubleshoot and repair. More than a decade in the making, Temporal is powered by veterans behind some of the industry's most loved systems technologies, programming frameworks and open source communities as well as investors like Amplify Partners, Sequoia Capital and Index Ventures.









