Site-Reliability Engineer, Application Operations

Posted Yesterday
Be an Early Applicant
75039, Irving, TX, USA
In-Office
Senior level
Artificial Intelligence • Healthtech • Biotech
Where Molecular Science Meets Artificial Intelligence – Revolutionizing Cancer Care.
The Role
Operates production clinical applications in a regulated environment as part of an App-SRE team. Responsibilities include on-call incident response, runbook execution, application monitoring and SLO tracking, deployment verification, rollbacks, post-incident reviews, audit documentation, and automating recurring operational work. The role collaborates with application engineering teams, improves observability and alert quality, reduces operational toil, and supports SOX- and FDA-compliant production operations.
Summary Generated by Built In

At Caris, we understand that cancer is an ugly word—a word no one wants to hear, but one that connects us all. That’s why we’re not just transforming cancer care—we’re changing lives.

 

We introduced precision medicine to the world and built an industry around the idea that every patient deserves answers as unique as their DNA. Backed by cutting-edge molecular science and AI, we ask ourselves every day: “What would I do if this patient were my mom?” That question drives everything we do.

 

But our mission doesn’t stop with cancer. We're pushing the frontiers of medicine and leading a revolution in healthcare—driven by innovation, compassion, and purpose.

 

Join us in our mission to improve the human condition across multiple diseases. If you're passionate about meaningful work and want to be part of something bigger than yourself, Caris is where your impact begins.

Position Summary

Caris Life Sciences is one of the largest precision-oncology platforms in the world, serving hundreds of thousands of molecular cases a year and growing at double-digit rates. Behind every case is a matched molecular, imaging, and clinical-outcomes data estate few organizations anywhere can rival, and the clinical software that drives the lab's instruments and processes, captures results, and delivers each patient's report. When that software degrades, patient care waits; keeping it healthy is this role's mission.

The Site-Reliability Engineer, Application Operations, works on the dedicated Application Site-Reliability Engineering (App-SRE) team. The work is site-reliability engineering across the clinical application portfolio, with real production telemetry and a direct hand in shaping the run-operate discipline.

This is a production-operations craft role, distinct from application feature development, operating under governed privileged access and segregation-of-duties discipline. The engineer serves as a primary responder in the on-call rotation; executes runbooks and approved maintenance scripts with audit-grade discipline; implements application instrumentation, SLO monitors, and error-budget tracking; supports releases, deployment health verification, and rollback execution; and contributes to post-incident reviews. The team runs automation-first: recurring manual work is engineering backlog, and the engineer progressively automates away the toil they encounter rather than absorbing it.

Frontier AI coding assistants are standard-issue tooling, with agentic workflows spanning runbook authoring, alert-quality analysis, and operational automation. Characterization tests and golden-master replay validation serve as executable evidence for regulated change. Delivery runs on CI/CD with application-level observability, spanning application-performance monitoring and production telemetry, and risk-based release governance aligned with FDA Computer Software Assurance guidance. Modernization of established systems is active engineering work, not deferred maintenance.

Reporting to the Director, Application Site-Reliability Engineering, this practitioner-level individual contributor operates production clinical applications subject to SOX financial controls and FDA regulatory requirements, working safely and accurately under established, documented procedures.

Location: Irving, Texas (Dallas–Fort Worth), on-site/hybrid, minimum three days per week on campus; co-located with Caris's laboratory and clinical operations.

Job Responsibilities

  • Participate in a scheduled on-call rotation as a primary responder for production incidents across the SOX- and FDA-regulated clinical application portfolio; execute triage, escalation, and initial remediation steps in accordance with documented runbooks and incident-response playbooks.

  • Execute approved runbooks and maintenance scripts in the production environment; document all privileged actions in compliance with SOX ITGCs and FDA audit-trail requirements.

  • Implement and maintain application-layer instrumentation, including dashboards, alert thresholds, and SLO monitors, in coordination with the centrally operated observability platform.

  • Support production releases and deployments: coordinate deployment health verification, monitor application behavior after deployment, and execute rollback procedures when required.

  • Participate in post-incident reviews: contribute timeline reconstructions, identify contributing factors, and track remediation action items through to closure.

  • Maintain audit-ready production-access logs, privileged-action records, and role-change documentation to support SOX ITGC and FDA regulatory audits.

  • Automate recurring operational work: convert repeated manual interventions, diagnostics, and maintenance procedures into scripted, reviewed, pipeline-executed automation under the team's change controls.

  • Track the manual-intervention rate for assigned services and drive it down over time by retiring runbook steps into automation.

  • Author and update runbook entries and known-issue documentation as operational knowledge is gained; contribute to continuous improvement of the operate discipline.

  • Collaborate with application engineering teams to gather context during incidents and to validate fixes deployed to the production environment.

  • Monitor application SLO attainment and error-budget consumption; escalate proactively when budgets are at risk.

  • Contribute to the ongoing development of on-call tooling, alert quality, and operational dashboards within the application-layer observability framework.

  • Work AI-first: use AI coding assistants and agentic workflows as daily practice in monitoring, triage, runbook work, and automation scripting, with review as the quality gate.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Systems, Software Engineering, or a closely related technical field, or equivalent practical experience.

  • 5+ years of professional experience in SRE, production operations, DevOps, or platform engineering in a production-support capacity.

  • 2+ years of direct, hands-on experience participating in an on-call rotation as a primary production responder for Tier 1 or business-critical systems.

  • Experience executing production runbooks, maintenance scripts, or change procedures in a production environment with documented privileged-action controls.

  • Experience implementing or maintaining application monitoring, alerting thresholds, or dashboards in a production observability platform.

  • Experience participating in post-incident review or blameless retrospective processes, including timeline reconstruction and corrective-action tracking.

  • Hands-on use of AI coding assistants for automation, scripting, or operational tooling.

Preferred Qualifications

  • Familiarity with SLO frameworks, error budgets, and associated alerting design patterns.

  • Experience reducing operational toil through scripting or automation in an SRE or production-operations setting.

  • Experience working in a SOX-controlled IT environment or a CLIA/CAP-regulated laboratory setting, including change-control ticketing, access-review processes, or audit-evidence collection.

  • Working knowledge of modern cloud-native observability at the application-instrumentation layer, including open standards for telemetry and tracing and application-performance-monitoring platforms.

  • Domain experience in clinical diagnostics, laboratory information systems, or digital health software.

  • Familiarity with SAST/DAST tooling and secure CI/CD pipelines, including pipeline-embedded security scanning.

  • Familiarity with deployment pipelines, container orchestration, and release-automation tooling from an operate and release-support perspective.

  • Certifications in cloud platforms or information-security disciplines relevant to production operations.

Physical Demands

  • Ability to sit, stand, and work at a computer for extended periods.

Training

  • All job-specific, safety, and compliance training is assigned based on the job functions associated with this employee.

Other

  • This role includes participation in a scheduled on-call rotation with required after-hours response to production incidents and critical service events, including evenings and weekends. Periodic travel may be required to support business needs and team on-sites.

Conditions of Employment:  Individual must successfully complete pre-employment process, which includes criminal background check, drug screening, credit check ( applicable for certain positions) and reference verification.

This job description reflects management’s assignment of essential functions. Nothing in this job description restricts management’s right to assign or reassign duties and responsibilities to this job at any time.

 

Caris Life Sciences is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, religion, color, national origin, gender, gender identity, sexual orientation, age, status as a protected veteran, among other things, or status as a qualified individual with disability.

Skills Required

  • Bachelor’s degree in Computer Science, Information Systems, Software Engineering, or a closely related technical field, or equivalent practical experience.
  • 5+ years of professional experience in SRE, production operations, DevOps, or platform engineering in a production-support capacity.
  • 2+ years of direct hands-on experience participating in an on-call rotation as a primary production responder for Tier 1 or business-critical systems.
  • Experience executing production runbooks, maintenance scripts, or change procedures with documented privileged-action controls.
  • Experience implementing or maintaining application monitoring, alerting thresholds, or dashboards in a production observability platform.
  • Experience participating in post-incident reviews or blameless retrospectives, including timeline reconstruction and corrective-action tracking.
  • Hands-on use of AI coding assistants for automation, scripting, or operational tooling.
  • Familiarity with SLO frameworks, error budgets, and alerting design patterns.
  • Experience reducing operational toil through scripting or automation in an SRE or production-operations setting.
  • Experience working in a SOX-controlled IT environment or a CLIA/CAP-regulated laboratory setting.
  • Working knowledge of cloud-native observability at the application-instrumentation layer, including telemetry, tracing, and application-performance-monitoring platforms.
  • Experience in clinical diagnostics, laboratory information systems, or digital health software.
  • Familiarity with SAST/DAST tooling and secure CI/CD pipelines with embedded security scanning.
  • Familiarity with deployment pipelines, container orchestration, and release-automation tooling.
  • Certifications in cloud platforms or information-security disciplines relevant to production operations.

Caris Life Sciences Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Caris Life Sciences and has not been reviewed or approved by Caris Life Sciences.

  • Fair & Transparent Compensation Pay is considered competitive or fair across many roles and locations. Shift differentials and overtime opportunities in certain lab roles can further boost take‑home pay.
  • Healthcare Strength Medical coverage is described as strong, with the employer covering the majority of premiums and health insurance frequently cited positively. Day‑one eligibility and company‑paid short‑ and long‑term disability reinforce core health protections.
  • Retirement Support A 401(k) with immediate vesting and a defined employer match supports long‑term savings. Plan details are presented clearly in benefits materials.

Caris Life Sciences Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Irving, TX
1,700 Employees
Year Founded: 2008

What We Do

Caris Life Sciences was founded in 2008 with a simple but powerful purpose – to help improve the lives of as many people as possible. With transformative technologies informed by massive amounts of big data, we are revolutionizing healthcare to provide physicians and patients with the highest quality information about their disease – from detecting it early and determining how best to treat it, to developing the next wave of novel therapies.

Similar Jobs

PwC Logo PwC

Front Office Strategy Consulting - Pharma Life Sciences Customer Analytics - Senior Associate

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
9 Locations
370000 Employees
77K-202K Annually

PwC Logo PwC

Specialized Tax Services - Research & Development Tax - Senior Manager

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
29 Locations
370000 Employees
124K-335K Annually

PwC Logo PwC

Procurement Senior Manager - Data

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
68 Locations
370000 Employees
91K-322K Annually

PwC Logo PwC

Consultant

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
9 Locations
370000 Employees
77K-202K Annually

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account