Site Reliability Engineering Lead

Posted Yesterday
Be an Early Applicant
Hiring Remotely in Alpharetta, GA, USA
In-Office or Remote
112K-264K Annually
Senior level
Information Technology • Legal Tech • Analytics
The Role
Leads Site Reliability Engineering strategy and a team of SREs supporting mission-critical platforms. Responsibilities include improving reliability, scalability, observability, security, and performance; driving cloud modernization across Azure and AWS; implementing infrastructure automation, SLOs, incident management, governance, and disaster recovery; and mentoring engineers while coordinating cross-functional initiatives across product, platform, security, and business teams.
Summary Generated by Built In

Ready to lead the reliability, scalability, and operational excellence of mission-critical platforms while shaping the future of Site Reliability Engineering?


Would you like to mentor high-performing engineers, drive cloud modernization, and influence enterprise-wide engineering practices in a highly collaborative environment?


About the Business

LexisNexis Risk Solutions is the essential partner in the assessment of risk. Within our Insurance vertical, we provide customers with solutions and decision tools that combine public and industry specific content with advanced technology and analytics to assist them in evaluating and predicting risk and enhancing operational efficiency. Our insurance risk solutions help drive better data-driven decisions across the insurance policy lifecycle, all while reducing risk. You can learn more about LexisNexis Risk at 

https://risk.lexisnexis.com/insurance


About our Team

The ICS (Insurance Core Services) team is responsible for establishing and driving reliability, observability, automation, and operational excellence standards across Insurance technology platforms. The team partners closely with application, infrastructure, database, and cloud engineering teams to improve platform availability, scalability, performance, and resilience.


ICS leads strategic initiatives including SLO/SLI implementation, observability platform adoption, cloud modernization, operational readiness reviews, performance engineering, and reliability automation. The team also develops reusable engineering frameworks, standards, and best practices that enable product teams to build and operate highly reliable cloud-native services at scale.


The SRE Lead will play a key role in shaping reliability strategy, mentoring engineers, driving cross-functional initiatives, and partnering with business and technology stakeholders to improve service reliability and operational maturity across the organization.


About the Role

As a Site Reliability Engineering Lead, you will provide technical leadership and strategic direction for Site Reliability Engineering initiatives across multiple product portfolios. You will lead a team of SREs responsible for ensuring the reliability, scalability, security, performance, and operational excellence of mission-critical applications and platforms.

The SRE Lead will partner closely with engineering, architecture, security, operations, and business stakeholders to drive cloud modernization, operational maturity, observability excellence, automation, and continuous improvement. This role combines hands-on technical expertise with people leadership, mentoring, strategic planning, and cross-functional collaboration.


Responsibilities

  • Lead and mentor a team of Site Reliability Engineers, fostering a culture of ownership, operational excellence, collaboration, and continuous learning.
  • Define and drive SRE strategy, standards, best practices, and operational frameworks across engineering organizations.
  • Partner with product and platform teams to improve application reliability, scalability, security, performance, and resilience.
  • Establish and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets.
  • Lead major incident management, root cause analysis, problem management, and post-incident review processes.
  • Drive cloud modernization initiatives and support application migrations to Azure, AWS, and containerized environments.
  • Champion automation and Infrastructure as Code (IaC) practices using tools such as Terraform, GitHub, GitLab, Jenkins, and Ansible.
  • Develop and implement observability strategies utilizing metrics, logs, traces, alerting, and dashboards.
  • Collaborate with security and compliance teams to ensure platform adherence to enterprise security and regulatory requirements.
  • Lead architecture reviews and provide guidance on cloud-native and highly resilient application designs.
  • Drive capacity planning, performance optimization, cost management, and operational efficiency initiatives.
  • Establish engineering guardrails, governance controls, and deployment standards for production environments.
  • Support organizational transformation toward DevOps and SRE practices.
  • Manage operational risk and ensure business continuity and disaster recovery preparedness.
  • Collaborate with stakeholders to prioritize reliability improvements and platform investments.
  • Build and maintain strong relationships with product owners, engineering leaders, vendors, and business partners.

Leadership Responsibilities

  • Lead, coach, mentor, and develop a high-performing team of Site Reliability Engineers.
  • Conduct resource planning and support hiring, onboarding, and career development activities.
  • Establish team objectives aligned with business and technology strategies.
  • Promote accountability, innovation, and operational excellence within the team.
  • Act as a trusted advisor and subject matter expert for reliability engineering across the organization.
  • Drive cross-team collaboration and alignment on strategic initiatives.

Essential Skills and Attributes

  • Strong leadership experience managing technical engineering teams.
  • Deep expertise in Site Reliability Engineering, DevOps, Cloud Engineering, or Platform Engineering disciplines.
  • Extensive experience with Azure and/or AWS cloud platforms.
  • Strong understanding of Kubernetes, AKS, EKS, containerization, Docker, and cloud-native architectures.
  • Expertise with Infrastructure as Code tools such as Terraform and Ansible.
  • Strong background in observability platforms such as Grafana, Prometheus, OpenTelemetry, Splunk, Dynatrace, Datadog, or similar technologies.
  • Experience managing large-scale production environments with stringent availability requirements.
  • Strong understanding of security, compliance, networking, and cloud governance principles.
  • Experience designing highly available, fault-tolerant, and resilient systems.
  • Strong proficiency in at least one scripting or programming language such as Python, Go, PowerShell, Bash, or C#.
  • Experience with CI/CD pipelines and software delivery automation.
  • Exceptional troubleshooting and problem-solving capabilities.
  • Excellent communication and stakeholder management skills.
  • Strong documentation and presentation skills.
  • Ability to influence technical direction across multiple engineering organizations.

Desired Skills

  • Experience building and managing enterprise-scale observability platforms.
  • Knowledge of FinOps, cloud cost optimization, and operational efficiency practices.
  • Experience with secret management platforms such as HashiCorp Vault, Akeyless or cloud-native alternatives.
  • Familiarity with Chaos Engineering and resilience testing.
  • Experience supporting regulated environments and compliance frameworks.
  • Experience leading cloud transformation and modernization programs.
  • Agile and Lean delivery experience.
  • Experience supporting global, distributed engineering teams.

Qualifications

  • 8+ years of experience in Cloud Engineering, DevOps, Platform Engineering, Infrastructure Engineering, or Site Reliability Engineering.
  • 2+ years of leadership or people management experience leading engineering teams.
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
  • Azure, AWS, Kubernetes, Terraform, or related certifications preferred.
  • Proven track record leading reliability and operational excellence initiatives in large-scale enterprise environments.

 

Risk benefit statement

Learn more about the LexisNexis Risk team and how we work https://relx.wd3.myworkdayjobs.com/RiskSolutions/page/21c296c982531000b79663f3194b0000


U.S. National Base Pay Range: $118,300 - $219,800. Geographic differentials may apply in some locations to better reflect local market rates. Base Pay Range for CO is $118,300 - $219,800. Base Pay Range for IL is $124,200 - $230,800. Base Pay Range for Chicago, IL is $130,200 - $241,800. Base Pay Range for MD is $124,200 - $230,800. Base Pay Range for NY is $130,200 - $241,800. Base Pay Range for New York City is $142,000 - $263,800. Base Pay Range for Rochester, NY is $118,300 - $219,800. Base Pay Range for OH is $112,400 - $208,800. Base Pay Range for NJ is $149,765- $239,235. This job is eligible for an annual incentive bonus. Application deadline is 12/23/2026.

We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits. Click here to access benefits specific to your location.

We are committed to providing a fair and accessible hiring process. If you have a disability or other need that requires accommodation or adjustment, please let us know by completing our Applicant Request Support Form or please contact 1-855-833-5120.

Criminals may pose as recruiters asking for money or personal information. We never request money or banking details from job applicants. Learn more about spotting and avoiding scams here.

Please read our Candidate Privacy Policy.

We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.

USA Job Seekers:

EEO Know Your Rights.

Skills Required

  • 8+ years of experience in Cloud Engineering, DevOps, Platform Engineering, Infrastructure Engineering, or Site Reliability Engineering
  • 2+ years of leadership or people management experience leading engineering teams
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent practical experience
  • Strong leadership experience managing technical engineering teams
  • Deep expertise in Site Reliability Engineering, DevOps, Cloud Engineering, or Platform Engineering
  • Extensive experience with Azure and/or AWS
  • Strong understanding of Kubernetes, AKS, EKS, Docker, containerization, and cloud-native architectures
  • Expertise with Terraform and Ansible
  • Strong background with observability platforms such as Grafana, Prometheus, OpenTelemetry, Splunk, Dynatrace, or Datadog
  • Experience managing large-scale production environments with stringent availability requirements
  • Strong understanding of security, compliance, networking, and cloud governance principles
  • Experience designing highly available, fault-tolerant, and resilient systems
  • Proficiency in at least one of Python, Go, PowerShell, Bash, or C#
  • Experience with CI/CD pipelines and software delivery automation
  • Proven track record leading reliability and operational excellence initiatives in large-scale enterprise environments
  • Azure, AWS, Kubernetes, Terraform, or related certifications
  • Experience building and managing enterprise-scale observability platforms
  • Knowledge of FinOps, cloud cost optimization, and operational efficiency practices
  • Experience with HashiCorp Vault, Akeyless, or cloud-native secret management alternatives
  • Familiarity with Chaos Engineering and resilience testing
  • Experience supporting regulated environments and compliance frameworks
  • Experience leading cloud transformation and modernization programs
  • Agile and Lean delivery experience
  • Experience supporting global, distributed engineering teams

RELX Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about RELX and has not been reviewed or approved by RELX.

  • Retirement Support Retirement support is positioned as a meaningful part of total rewards through a 401(k) plan with matching contributions, alongside other financial protections such as life and disability coverage. Tuition reimbursement and share purchase access further broaden the financial value of the package beyond base salary.
  • Leave & Time Off Breadth Leave and time off breadth appears strong, with generous vacation allowances, mental health days, and options like sabbaticals and tiered PTO by tenure. Parental and caregiving leaves are described in detail, reinforcing time-away benefits as a standout component of the overall package.
  • Wellbeing & Lifestyle Benefits Wellbeing and lifestyle benefits are supported by offerings such as mental health support (e.g., app access), EAP resources, gym-related perks, and wellness incentives. Flexible working hours and related work-life supports add to the perceived day-to-day value of benefits.

RELX Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
10,001 Employees
Year Founded: 1880

What We Do

RELX is a global provider of information-based analytics for professional and business customers across industries. We help scientists make new discoveries, doctors and nurses improve the lives of patients and lawyers win cases. We prevent online fraud and money laundering, and help insurance companies evaluate and predict risk. Our events enable customers to learn about markets, source products and complete transactions. In short, we enable our customers to make better decisions, get better results and be more productive. We do this by leveraging a deep understanding of our customers to create innovative solutions which combine content and data with analytics and technology in global platforms. RELX serves customers in more than 180 countries and has offices in about 40 countries. It employs approximately 30,000 people of whom almost half are in North America. We operate in four major market segments: Scientific, Technical & Medical; Risk & Business Analytics; Legal; and Exhibitions.

Similar Jobs

Remote
United States
350 Employees
250K-280K Annually

AuthZed Logo AuthZed

Senior Site Reliability Engineer

Artificial Intelligence • Information Technology • Software • Database
Remote
2 Locations
30 Employees
150K-195K Annually

MongoDB Logo MongoDB

Site Reliability Engineer

Big Data • Cloud • Software • Database
Easy Apply
Remote or Hybrid
10 Locations
5550 Employees
127K-249K Annually

Similar Companies Hiring

NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
900 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account