Senior Site Reliability Engineer II

Posted 8 Days Ago
San Jose, CA, USA
In-Office
105K-175K Annually
Senior level
Information Technology • Legal Tech • Analytics
The Role
Lead site reliability initiatives for highly available production systems across cloud and on-premises environments. Responsibilities include incident response, postmortems, root-cause analysis, remediation tracking, monitoring and observability, infrastructure automation, Kubernetes and systems support, change management, disaster recovery, and operational readiness. The role partners with engineering, security, operations, vendors, and business stakeholders while guiding less-experienced team members and improving system reliability, performance, availability, and operational quality.
Summary Generated by Built In

About the Role:


The SRE role is responsible for improving the reliability, availability, performance, and operational quality of production systems. This role provides technical input into project plans, schedules, methodologies, and operational strategies across multiple system environments.


Job Functions

  • Lead and participate in incident response, postmortems, root-cause analysis, and gap assessments.
  • Identify reliability, availability, performance, security, and operational risks across production environments.
  • Develop, prioritize, and track corrective and preventive actions through completion.
  • Follow up with engineering, development, security, support, and business stakeholders to ensure timely resolution of incidents and identified gaps.
  • Respond to system-management alerts and operational exceptions within assigned enterprise systems and product offerings.
  • Provide technical input into project plans, schedules, implementation methodologies, and operational readiness activities.
  • Support the triage, planning, execution, documentation, and closure of changes, service requests, and operational tasks.
  • Lead or contribute to Operations Team projects involving cloud, on-premises infrastructure, security, Kubernetes, automation, monitoring, and system modernization.
  • Improve production quality and availability by creating new operational capabilities and remediating weaknesses in existing systems and processes.

Qualifications

  • 5+ years of experience in Site Reliability Engineering, Systems Engineering, DevOps, Infrastructure Engineering, or a related field.
  • Bachelor’s degree in Engineering, Computer Science, Information Technology, or equivalent professional experience.
  • Demonstrated experience supporting highly available production systems.
  • Experience leading incident reviews, postmortems, root-cause analysis, and remediation planning.
  • Experience working across infrastructure, application, security, and operations teams.
  • Strong problem-solving, analytical, organizational, and communication skills.
  • Ability to manage multiple priorities and drive work to completion in a fast-paced operational environment.

Technical Skills


  • Strong experience in Site Reliability Engineering (SRE), production operations, and IT service management processes including incident, problem, change, and service request management.
  • Hands-on expertise with cloud and on-premises infrastructure, Kubernetes, containerized workloads, virtualization, and distributed systems.
  • Advanced knowledge of Linux/UNIX and Windows environments, storage and file systems, including installation, configuration, troubleshooting, lifecycle management, backup, disaster recovery, and business continuity.
  • Experience with monitoring, alerting, logging, observability, and performance analysis, including the ability to analyze system diagnostics, logs, traces, resource utilization, and operational metrics.
  • Strong automation and infrastructure engineering skills, including Infrastructure as Code (IaC), configuration management, scripting (Python, Shell, PowerShell), system provisioning, deployments, remediation, and security risk mitigation.

Accountabilities

  • Monitor assigned environments, respond to alerts and incidents, diagnose system and performance issues, and coordinate escalation and recovery.
  • Track remediation activities and stakeholder commitments through completion to improve production quality, reliability, and availability.
  • Design and maintain automation, scripts, integrations, runbooks, and workflows for provisioning, health checks, deployments, remediation, and routine operations.Install, configure, troubleshoot, and support hardware, software, storage, network, cloud, Kubernetes, and other infrastructure services.
  • Establish logging, monitoring, alerting, metrics, and tracing standards; improve alert quality by reducing noise and ensuring alerts are actionable.Build dashboards and visualizations that communicate system health, availability, performance, capacity, service-level objectives, and incident trends.
  • Develop and maintain recovery procedures and participate in disaster-recovery, resilience, and business-continuity exercises.Partner with development, operations, security, support teams, vendors, and stakeholders to coordinate work, resolve issues, and meet delivery commitments.
  • Lead or contribute to Operations Team projects from planning and implementation through documentation, transition to support, and closure.
  • Plan, risk-assess, obtain approval for, implement, document, and close changes, service requests, and operational tasks.
  • Review and improve technical procedures, scripts, automation, and operational documentation while providing guidance to less-experienced team members.

Working for you:


We know that your wellbeing and happiness are key to a long and successful career. These are some of the benefits we are delighted to offer:


  • Health Benefits: Comprehensive, multi-carrier program for medical, dental and vision benefits
  • Retirement Benefits: 401(k) with match and an Employee Share Purchase Plan
  • Wellbeing: Wellness platform with incentives, Headspace app subscription, Employee Assistance and Time-off Programs
  • Short-and-Long Term Disability, Life and Accidental Death Insurance, Critical Illness, and Hospital Indemnity
  • Family Benefits, including bonding and family care leaves, adoption and surrogacy benefits
  • Health Savings, Health Care, Dependent Care and Commuter Spending Accounts
  • In addition to annual Paid Time Off, we offer up to two days of paid leave each to participate in Employee Resource Groups and to volunteer with your charity of choice



U.S. National Base Pay Range: $104,900 - $174,700. Geographic differentials may apply in some locations to better reflect local market rates. This job is eligible for an annual incentive bonus.

We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits. Click here to access benefits specific to your location.

We are committed to providing a fair and accessible hiring process. If you have a disability or other need that requires accommodation or adjustment, please let us know by completing our Applicant Request Support Form or please contact 1-855-833-5120.

Criminals may pose as recruiters asking for money or personal information. We never request money or banking details from job applicants. Learn more about spotting and avoiding scams here.

Please read our Candidate Privacy Policy.

We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.

USA Job Seekers:

EEO Know Your Rights.

Skills Required

  • 5+ years of experience in Site Reliability Engineering, Systems Engineering, DevOps, Infrastructure Engineering, or a related field
  • Bachelor's degree in Engineering, Computer Science, Information Technology, or equivalent professional experience
  • Experience supporting highly available production systems
  • Experience leading incident reviews, postmortems, root-cause analysis, and remediation planning
  • Experience working across infrastructure, application, security, and operations teams
  • Strong problem-solving, analytical, organizational, and communication skills
  • Ability to manage multiple priorities and drive work to completion in a fast-paced operational environment
  • Strong experience in Site Reliability Engineering, production operations, and IT service management processes
  • Hands-on expertise with cloud and on-premises infrastructure, Kubernetes, containerized workloads, virtualization, and distributed systems
  • Advanced knowledge of Linux/UNIX and Windows environments, storage and file systems, installation, configuration, troubleshooting, lifecycle management, backup, disaster recovery, and business continuity
  • Experience with monitoring, alerting, logging, observability, performance analysis, diagnostics, traces, resource utilization, and operational metrics
  • Strong automation and infrastructure engineering skills, including Infrastructure as Code, configuration management, Python, Shell, PowerShell, provisioning, deployments, remediation, and security risk mitigation

RELX Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about RELX and has not been reviewed or approved by RELX.

  • Retirement Support — Retirement support is positioned as a meaningful part of total rewards through a 401(k) plan with matching contributions, alongside other financial protections such as life and disability coverage. Tuition reimbursement and share purchase access further broaden the financial value of the package beyond base salary.
  • Leave & Time Off Breadth — Leave and time off breadth appears strong, with generous vacation allowances, mental health days, and options like sabbaticals and tiered PTO by tenure. Parental and caregiving leaves are described in detail, reinforcing time-away benefits as a standout component of the overall package.
  • Wellbeing & Lifestyle Benefits — Wellbeing and lifestyle benefits are supported by offerings such as mental health support (e.g., app access), EAP resources, gym-related perks, and wellness incentives. Flexible working hours and related work-life supports add to the perceived day-to-day value of benefits.

RELX Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
10,001 Employees
Year Founded: 1880

What We Do

RELX is a global provider of information-based analytics for professional and business customers across industries. We help scientists make new discoveries, doctors and nurses improve the lives of patients and lawyers win cases. We prevent online fraud and money laundering, and help insurance companies evaluate and predict risk. Our events enable customers to learn about markets, source products and complete transactions. In short, we enable our customers to make better decisions, get better results and be more productive. We do this by leveraging a deep understanding of our customers to create innovative solutions which combine content and data with analytics and technology in global platforms. RELX serves customers in more than 180 countries and has offices in about 40 countries. It employs approximately 30,000 people of whom almost half are in North America. We operate in four major market segments: Scientific, Technical & Medical; Risk & Business Analytics; Legal; and Exhibitions.

Similar Jobs

Akamai Technologies Logo Akamai Technologies

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees
146K-264K Annually

Akamai Technologies Logo Akamai Technologies

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees
146K-264K Annually

MongoDB Logo MongoDB

Site Reliability Engineer

Big Data • Cloud • Software • Database
Easy Apply
Remote or Hybrid
10 Locations
5550 Employees
127K-249K Annually

Alembic Logo Alembic

Senior Site Reliability Engineer

Artificial Intelligence • Marketing Tech • Software • Big Data Analytics
In-Office
San Francisco, CA, USA
43 Employees
210K-240K Annually

Similar Companies Hiring

NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
900 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account