Principal Distributed Systems Engineer - Observability

Posted 3 Days Ago
Be an Early Applicant
Pleasanton, CA, USA
In-Office
187K-334K Annually
Expert/Leader
Cloud • Fintech • HR Tech
The Role
Own the technical vision and architecture for Workday’s distributed tracing observability platform. Design and build multi-petabyte ingestion, storage, and query infrastructure using ClickHouse, Tempo, Kafka, Spark or Flink, Iceberg, S3, and AWS. Lead scalability, performance, high availability, disaster recovery, security, and operational excellence. Partner with AI stakeholders on anomaly detection and root-cause analysis, while mentoring engineers, setting cross-team technical direction, and serving as the tracing platform’s primary architect and escalation point.
Summary Generated by Built In

Your work days are brighter here.

We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re shaping the future of work so teams can reach their potential and focus on what matters most. The minute you join, you’ll feel it. Not just in the products we build, but in how we show up for each other. Our culture is rooted in integrity, empathy, and shared enthusiasm. We’re in this together, tackling big challenges with bold ideas and genuine care. We look for curious minds and courageous collaborators who bring sun-drenched optimism and drive. Whether you're building smarter solutions, supporting customers, or creating a space where everyone belongs, you’ll do meaningful work with Workmates who’ve got your back. In return, we’ll give you the trust to take risks, the tools to grow, the skills to develop and the support of a company invested in you for the long haul. So, if you want to inspire a brighter work day for everyone, including yourself, you’ve found a match in Workday, and we hope to be a match for you too.

About the Team

The Data Platform and Observability Engineering (DPOE) team is building Workday's next-generation, multi-petabyte scale Observability Platform. We own the libraries, distributed services, and infrastructure that power ingestion, storage, and query across the observability stack — Iceberg, ClickHouse, Tempo, Grafana, S3, Kafka, and Elasticsearch — serving traces, metrics, and logs for every workload at Workday. Our roadmap directly shapes how the company detects, diagnoses, and eventually predicts operational issues at scale.

About the Role

To own the technical vision and architecture for distributed tracing as a first-class pillar of Workday's Observability Platform, built on ClickHouse and/or Grafana Tempo, backed by a big-data pipeline (Kafka, Spark/Flink, Iceberg, Clickhouse, Tempo,S3) running on AWS. This is a hands-on, high-autonomy role for an engineer who can design and build multi-petabyte, low-latency tracing infrastructure end-to-end — and who is equally excited to help define where Observability AI goes next: using traces, logs, and metrics as the substrate for automated root-cause analysis, anomaly detection, and AI-driven incident triage.

You'll set technical direction across multiple teams, mentor senior and staff engineers, and act as the primary architect and escalation point for the tracing subsystem — from ingestion and storage design through query performance and platform reliability.

  • Architect and build Workday's distributed tracing platform on ClickHouse/Tempo, designed for multi-petabyte scale ingestion and sub-second interactive query performance.
  • Own the big-data pipeline feeding tracing data — Kafka-based ingestion, Spark/Flink stream and batch processing, and Iceberg-on-S3 storage — including schema design, partitioning, compaction, and lifecycle management.
  • Drive performance and scaling across ingestion and query paths: storage format optimization (Parquet/Iceberg), compression strategy, partitioning/indexing, and query engine tuning under real production load.
  • Lead HA/DR design for tracing services — multi-region/multi-AZ resilience, failover, backup/restore, and recovery time/point objectives appropriate to a tier-1 platform.
  • Design security architecture for the platform, including authentication/authorization (authn/authz) for multi-tenant data access across ingestion and query layers.
  • Own operational excellence for distributed tracing: monitoring, logging, alerting, capacity planning, and participation in an on-call rotation for the platform.
  • Evaluate and introduce new technologies — open source and cloud-native — that materially improve the platform's scalability, cost efficiency, or capability.
  • Shape the future of Observability AI: partner with ML/AI stakeholders to define how tracing data feeds automated anomaly detection, root-cause analysis, and AI-assisted incident management.
  • Evangelize the platform: publish best practices, mentor engineers across DPOE and partner teams, and act as a technical thought leader for the modern observability/data stack internally.
  • Operate with high autonomy in a fast-moving, ambiguous environment — setting technical direction with minimal oversight while aligning with broader platform strategy.

About You

Basic Qualification 

14+ years experience in software development engineering.
6+ years experience specifically focused on designing, building, and operating complex distributed system architectures, evidenced by successful deployment of systems with high availability (e.g., 99.9% uptime) and fault tolerance.
8+ years experience with at least two of the following programming languages (e.g., Java, Python, Go), including experience in writing production-level code for distributed systems.
Bachelor’s degree in a relevant field such as Computer Science, Engineering, or a related discipline; a Master's degree (e.g., MS in Computer Science, Distributed Systems, or related field) is strongly preferred or equivalent practical experience.

Other Qualification

Expert-level ability in Algorithmic Thinking, including [insert specific advanced algorithms or data structures relevant to distributed systems], to architect highly efficient and scalable solutions for complex
Deep expertise in API Development, including understanding of advanced API protocols or architectural patterns
Deep understanding of Distributed Systems Software principles, like distributed consensus or fault tolerance mechanisms
Proven ability to design and implement High Availability strategies for critical distributed systems
Extensive experience with Large Scale Data Processing technologies and frameworks
Deep understanding of Large Scale Systems design principles like distributed data management or scalability strategies
Strong understanding of System Security principles and best practices relevant to securing complex distributed environments
Proven ability to lead Team Collaboration within and across distributed software development teams and drive architectural direction
Strong skills in creating Technical Writing Documentation and Presentation


Workday Pay Transparency Statement

The annualized base salary ranges for the primary location and any additional locations are listed below.  Workday pay ranges vary based on work location. As a part of the total compensation package, this role may be eligible for the Workday Bonus Plan or a role-specific commission/bonus, as well as annual refresh stock grants. Recruiters can share more detail during the hiring process. Each candidate’s compensation offer will be based on multiple factors including, but not limited to, geography, experience, skills, job duties, and business need, among other things. For more information regarding Workday’s comprehensive benefits, please click here.

Primary Location: USA.CA.Pleasanton


 

Primary Location Base Pay Range: $222,900 USD - $334,300 USD


 

Additional US Location(s) Base Pay Range: $187,100 USD - $334,300 USD


Our Approach to Flexible Work
 

With Flex Work, we’re combining the best of both worlds: in-person time and remote. Our approach enables our teams to deepen connections, maintain a strong community, and do their best work. We know that flexibility can take shape in many ways, so rather than a number of required days in-office each week, we simply spend at least half (50%) of our time each quarter in the office or in the field with our customers, prospects, and partners (depending on role). This means you'll have the freedom to create a flexible schedule that caters to your business, team, and personal needs, while being intentional to make the most of time spent together. Those in our remote "home office" roles also have the opportunity to come together in our offices for important moments that matter.

Pursuant to applicable Fair Chance law, Workday will consider for employment qualified applicants with arrest and conviction records.

Workday is an Equal Opportunity Employer including individuals with disabilities and protected veterans.


Workday is committed to providing reasonable accommodations for qualified individuals during our application process, in order to perform one or more essential functions of their job, as well as regarding the use of AI tools for employment decision-making to any degree. Please see below for more details including how to request an accommodation as a qualified veteran, due to a disability or for religious reasons, or as otherwise provided under applicable law.


Workday prohibits taking adverse action against any candidate or employee for reporting a possible violation of this policy, requesting one or more work accommodations, exercising a privacy right, or cooperating in an investigation in accordance with applicable law. Any employee who retaliates against a candidate or employee for doing so may be subject to disciplinary action, up to and including termination of employment, to the fullest extent allowable under applicable law.


If you require a reasonable accommodation, you may email [email protected], as far in advance as possible.


Are you being referred to one of our roles? If so, ask your connection at Workday about our Employee Referral process!

At Workday, we value our candidates’ privacy and data security.  Workday will never ask candidates to apply to jobs through websites that are not Workday Careers. 

  

Please be aware of sites that may ask for you to input your data in connection with a job posting that appears to be from Workday but is not.

  

In addition, Workday will never ask candidates to pay a recruiting fee, or pay for consulting or coaching services, in order to apply for a job at Workday.

Skills Required

  • 14+ years of software development engineering experience
  • 6+ years designing, building, and operating complex distributed system architectures with high availability and fault tolerance
  • 8+ years of experience with at least two programming languages such as Java, Python, or Go, including production distributed-systems code
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience
  • Master’s degree in Computer Science, Distributed Systems, or a related field
  • Expert-level algorithmic thinking involving advanced algorithms or data structures relevant to distributed systems
  • Deep expertise in API development and advanced API protocols or architectural patterns
  • Deep understanding of distributed-systems principles, including distributed consensus and fault tolerance
  • Proven ability to design and implement high-availability strategies for critical distributed systems
  • Extensive experience with large-scale data-processing technologies and frameworks
  • Deep understanding of large-scale systems design, distributed data management, and scalability strategies
  • Strong understanding of system-security principles and best practices for complex distributed environments
  • Proven ability to lead collaboration across distributed software-development teams and drive architectural direction
  • Strong technical writing, documentation, and presentation skills

Workday Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Workday and has not been reviewed or approved by Workday.

  • Healthcare Strength Health coverage is positioned as broad and well-supported, with multiple medical carrier options, virtual care access, and some locations offering onsite clinic/pharmacy services. Mental health support is described as notably strong, including therapy sessions and confidential support availability for household members.
  • Parental & Family Support Family-related benefits are portrayed as extensive, including paid bonding and caregiver leave alongside fertility, adoption, and surrogacy reimbursement. Added support like parenting resources, milk-shipping/lactation assistance during travel, and backup child/elder care is explicitly outlined.
  • Strong & Reliable Incentives Equity participation and savings-oriented programs are presented as meaningful components of total rewards, including an ESPP discount with a lookback feature. Additional programs like a student-loan pathway to earn the 401(k) match are included as financial-support enhancements.

Workday Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Pleasanton, CA
14,894 Employees
Year Founded: 2005

What We Do

Workday is a leading provider of enterprise cloud applications for finance, HR, and planning. Founded in 2005, Workday delivers financial management, human capital management, and analytics applications designed for the world’s largest companies, educational institutions, and government agencies. Organizations ranging from medium-sized businesses to Fortune 50 enterprises have selected Workday.

Similar Jobs

Capital One Logo Capital One

Artificial Intelligence Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
4 Locations
55000 Employees
230K-286K Annually

Capital One Logo Capital One

Artificial Intelligence Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
4 Locations
55000 Employees
230K-286K Annually

Optum Logo Optum

Service Desk Analyst

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Irvine, CA, USA
160000 Employees
24-43 Hourly

Optum Logo Optum

Principal Data Scientist

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Irvine, CA, USA
160000 Employees
227K-315K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account