Director - Application Site-Reliability Engineering

Posted 13 Days Ago
Be an Early Applicant
75039, Irving, TX, USA
In-Office
Expert/Leader
Artificial Intelligence • Healthtech • Biotech
Where Molecular Science Meets Artificial Intelligence – Revolutionizing Cancer Care.
The Role
Directs application reliability and production operations for regulated clinical software. Responsibilities include leading incident response, managing on-call operations, defining SLOs and error budgets, governing production access, overseeing deployments and rollbacks, building runbooks and automation, improving observability, generating audit evidence through CI/CD, and hiring and developing the App-SRE team. The role also represents reliability engineering in leadership and compliance forums and participates in senior escalation support.
Summary Generated by Built In

At Caris, we understand that cancer is an ugly word—a word no one wants to hear, but one that connects us all. That’s why we’re not just transforming cancer care—we’re changing lives.

 

We introduced precision medicine to the world and built an industry around the idea that every patient deserves answers as unique as their DNA. Backed by cutting-edge molecular science and AI, we ask ourselves every day: “What would I do if this patient were my mom?” That question drives everything we do.

 

But our mission doesn’t stop with cancer. We're pushing the frontiers of medicine and leading a revolution in healthcare—driven by innovation, compassion, and purpose.

 

Join us in our mission to improve the human condition across multiple diseases. If you're passionate about meaningful work and want to be part of something bigger than yourself, Caris is where your impact begins.

Position Summary

Caris Life Sciences is one of the largest precision-oncology platforms in the world, serving hundreds of thousands of molecular cases a year and growing at double-digit rates. Behind every case is a matched molecular, imaging, and clinical-outcomes data estate few organizations anywhere can rival, and the clinical software that drives the lab's instruments and processes, captures results, and delivers each patient's report. When that software degrades, patient care waits; keeping it reliable is this role's charter.

Reporting to the Corporate Vice President for Clinical Software Products, the Director owns application reliability and production operations for the clinical software portfolio: production support, incident response, on-call operations, and SLO management for applications under SOX financial controls and FDA regulatory requirements. The Director hires and develops the team, sets standards and selects tooling, establishes production-access governance and segregation-of-duties controls with the information-security, quality, and infrastructure organizations, and participates directly in incident response. The role carries wide latitude to shape how reliability engineering is done here.

Frontier AI coding assistants are standard-issue tooling, with agentic workflows spanning incident diagnostics, runbook authoring, and operational automation. Characterization tests and golden-master replay validation serve as executable evidence for regulated change. Delivery runs on CI/CD with application-level observability and risk-based release governance aligned with FDA Computer Software Assurance guidance. Modernization of established systems is active engineering work, not deferred maintenance.

The infrastructure organization owns the platform and observability runtime; this role owns application-layer reliability on top of it. The operating model is automation-first: recurring manual work is engineered away rather than staffed, and operational and compliance evidence is produced by pipelines rather than assembled by hand.

Location: Irving, Texas (Dallas–Fort Worth), on-site/hybrid, minimum three days per week on campus; co-located with Caris's laboratory and clinical operations.

Job Responsibilities

  • Own the production-support model for the clinical application portfolio: the on-call rotation, escalation procedures, and incident-response playbooks.

  • Establish and audit production-access and segregation-of-duties controls with engineering, information-security, quality, and infrastructure partners; keep the evidence audit-ready for SOX ITGCs and applicable FDA requirements, including access grants, role changes, and privileged-action logs.

  • Define clinical SLOs, error budgets, and availability targets with product and engineering leadership; track attainment and operational-health metrics such as mean time to detect, mean time to recover, and on-call burden; intervene while an error budget is burning, not after it is spent.

  • Lead incident response for high-severity production events as incident commander or senior technical responder; coordinate cross-functional teams, run post-incident reviews, and drive systemic remediation to closure.

  • Develop the runbook library and approved operational automation; ensure engineers can execute standard interventions safely within documented procedures.

  • Drive the automation-first operating model: convert recurring manual interventions into reviewed automation, measure and reduce toil, and favor self-service tooling over ticket-driven request work.

  • Coordinate production deployments with engineering teams, verify deployment health, and own rollback decisions.

  • Advance automated generation of change and deployment evidence in CI/CD pipelines: deployment records, approval trails, and change documentation as an audit-ready by-product of release.

  • Hire and develop the App-SRE team; set performance expectations, on-call responsibilities, and career growth frameworks.

  • Own application-layer observability alongside the infrastructure and observability platform teams; keep production dashboards, alerting thresholds, and SLO monitors accurate and actionable.

  • Represent App-SRE in engineering leadership forums, operational reviews, and compliance audits; translate operational health and risk into clear executive communication.

  • Run the function AI-first: make AI-assisted practice the team's daily norm, from runbook automation to incident analysis and operational tooling.

Required Qualifications

  • Bachelor's degree in Computer Science, Software Engineering, Information Systems, or a closely related technical field, or equivalent practical experience.

  • 10+ years of professional experience in SRE, DevOps, platform engineering, or production operations.

  • 4+ years of direct experience in a people-management or team-lead role within an SRE or production-operations function.

  • Hands-on experience leading incident response for Tier 1 or business-critical production systems, including serving as incident commander or senior technical responder.

  • Experience defining and implementing SLOs, error budgets, and associated alerting and on-call workflows in a production environment.

  • Experience building or significantly maturing a production-support, on-call, or SRE function, including runbook development and on-call-rotation design.

  • Experience operating production systems under a formal regulatory or financial-controls framework, such as CAP/CLIA, FDA regulations, SOX ITGCs, HIPAA, or equivalent, including producing documentation that holds up in audit.

  • A record of applying AI-assisted practice to operations or engineering work, personally or through a team.

Preferred Qualifications

  • Domain experience in clinical diagnostics, laboratory information systems, molecular pathology, or digital health software.

  • Direct experience supporting SOX ITGC audit cycles or CAP/CLIA laboratory inspections, including evidence gathering for access-control, change-management, and monitoring controls.

  • Working knowledge of modern cloud-native observability at the application-instrumentation layer, including open standards for telemetry and tracing and application-performance-monitoring platforms.

  • Experience with deployment pipelines, release-management workflows, and rollback procedures in a continuous-delivery environment.

  • Ability to operate as a player-coach, contributing directly to technical work while building and leading a team.

  • Track record of reducing operational toil through automation programs in an SRE or production-operations organization.

  • Experience presenting operational strategy and risk posture to senior or executive audiences.

Physical Demands

  • Ability to sit, stand, and work at a computer for extended periods.

Training

  • All job-specific, safety, and compliance training is assigned based on the job functions associated with this employee.

Other

  • This role serves in the senior escalation tier of the production on-call rotation, with after-hours response to high-severity incidents as incident commander or senior escalation point. Periodic travel may be required to support business needs, team on-sites, and leadership reviews.

Conditions of Employment:  Individual must successfully complete pre-employment process, which includes criminal background check, drug screening, credit check ( applicable for certain positions) and reference verification.

This job description reflects management’s assignment of essential functions. Nothing in this job description restricts management’s right to assign or reassign duties and responsibilities to this job at any time.

 

Caris Life Sciences is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, religion, color, national origin, gender, gender identity, sexual orientation, age, status as a protected veteran, among other things, or status as a qualified individual with disability.

Skills Required

  • Bachelor's degree in Computer Science, Software Engineering, Information Systems, a related technical field, or equivalent practical experience
  • 10+ years of professional experience in SRE, DevOps, platform engineering, or production operations
  • 4+ years of direct people-management or team-lead experience within an SRE or production-operations function
  • Experience leading incident response for Tier 1 or business-critical production systems as incident commander or senior technical responder
  • Experience defining and implementing SLOs, error budgets, alerting, and on-call workflows in production
  • Experience building or significantly maturing a production-support, on-call, or SRE function, including runbook development and on-call rotation design
  • Experience operating production systems under a formal regulatory or financial-controls framework, such as CAP/CLIA, FDA regulations, SOX ITGCs, HIPAA, or equivalent
  • Record of applying AI-assisted practices to operations or engineering work
  • Experience in clinical diagnostics, laboratory information systems, molecular pathology, or digital health software
  • Direct experience supporting SOX ITGC audits or CAP/CLIA laboratory inspections
  • Working knowledge of application-layer cloud-native observability, open telemetry and tracing standards, and application-performance-monitoring platforms
  • Experience with deployment pipelines, release-management workflows, and rollback procedures in a continuous-delivery environment
  • Ability to operate as a player-coach while contributing technically and leading a team
  • Track record of reducing operational toil through automation programs
  • Experience presenting operational strategy and risk posture to senior or executive audiences

Caris Life Sciences Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Caris Life Sciences and has not been reviewed or approved by Caris Life Sciences.

  • Fair & Transparent Compensation Pay is considered competitive or fair across many roles and locations. Shift differentials and overtime opportunities in certain lab roles can further boost take‑home pay.
  • Healthcare Strength Medical coverage is described as strong, with the employer covering the majority of premiums and health insurance frequently cited positively. Day‑one eligibility and company‑paid short‑ and long‑term disability reinforce core health protections.
  • Retirement Support A 401(k) with immediate vesting and a defined employer match supports long‑term savings. Plan details are presented clearly in benefits materials.

Caris Life Sciences Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Irving, TX
1,700 Employees
Year Founded: 2008

What We Do

Caris Life Sciences was founded in 2008 with a simple but powerful purpose – to help improve the lives of as many people as possible. With transformative technologies informed by massive amounts of big data, we are revolutionizing healthcare to provide physicians and patients with the highest quality information about their disease – from detecting it early and determining how best to treat it, to developing the next wave of novel therapies.

Similar Jobs

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
8 Locations
40000 Employees
45K-85K Annually

PwC Logo PwC

Strategy& Deals Tech Strategy AI & Tech Value Creation Director

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
12 Locations
370000 Employees
155K-410K Annually

General Motors Logo General Motors

Senior Software Engineer

Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Hybrid
2 Locations
165000 Employees

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Data Steward: Personal Lines Product, Pricing and Underwriting data domain

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
7 Locations
40000 Employees
120K-225K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account