Senior Site Reliability Engineer I

Posted Yesterday
2 Locations
In-Office or Remote
95K-159K Annually
Senior level
Artificial Intelligence • Healthtech • Information Technology • Other • Analytics
The Role
Build and maintain reliable, scalable platforms and services. Responsibilities include monitoring, incident response, post-mortems, disaster recovery testing, production automation, AI service deployment, infrastructure documentation, and reliability testing. The role requires advanced Terraform, AWS operations, GitHub Actions CI/CD, containerized ECS Fargate environments, networking and security expertise, Linux and scripting skills, observability, and developer enablement across engineering teams.
Summary Generated by Built In

Senior Site Reliability Engineer

 Are you passionate about building resilient, scalable systems that power mission-critical applications?
Do you thrive on automating operations, improving reliability, and ensuring exceptional system performance?

 

About the team:

Embedded Innovation Teams are cross-functional squads embedded within our segments to rapidly turn internal AI experimentation into validated, reusable solutions, building the capabilities we need to deliver customer value and growth. We work problem-first rather than tool-first, directly inside segment and function teams, improving the internal workflows that help our people deliver better outcomes for customers, faster.


About the role: 

As a Senior Site Reliability Engineer (SRE), you will play a key role in ensuring the reliability, scalability, and performance of our critical platforms and services. You will lead complex reliability initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences.


You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve service availability, streamline operations, and enhance system recovery capabilities.


You will hold a high bar on code quality, flag risks and blockers early, and work alongside host-function stakeholders to make sure what you build fits real workflows, not assumed ones. You will also support handover and capability-building so the solution is owned and operable after the squad moves on.


 Key Responsibilities:


  • Creating monitoring queries and establishes service level baselines.
  • Supporting senior engineers during incidents.
  • Making contributions during post-mortems and RCAs.
  • Participating in disaster recovery tests.
  • Implementing automation and executes code in production environments.
  • Contributing to SRE knowledge documentation.
  • Supporting  the deployment, monitoring, and reliability of services integrating AI tools.  
  • Supporting architecture and senior engineers in the creation of infrastructure topology drawings and deployment workflows.
  • Carrying out the testing of availability, reliability, and recoverability in non-production environments.

    

Requirements:


  • Advanced Terraform: Expertise in modules, providers, state management, lifecycle controls, drift detection, safe refactoring, and remote state (S3, locking, cross-stack dependencies).
  • AWS Operations: Hands-on experience managing production, multi-account, multi-region AWS environments across ECS, RDS, ALB, VPC, IAM, Route53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch.
  • GitHub Actions CI/CD: Experience building and troubleshooting reusable workflows, OIDC authentication, approval gates, runners, Terraform deployments, application deployments, and migration pipelines.
  • ECS Fargate & Containers: Knowledge of Docker, ECR, ECS task definitions/services, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks.
  • AWS Networking & Security: Proficiency in VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security best practices.
  • Incident Response & Observability: Skilled in troubleshooting using logs, metrics, alarms, deployment history, root cause analysis, rollback decisions, and operational runbooks.
  • Linux & Automation: Strong Linux and Git fundamentals with Bash/Python scripting for AWS CLI automation, CI/CD, and operational tooling.
  • AI Tooling Deployment: Hands-on experience integrating and operating AI services and APIs in production, including monitoring, reliability, and security practices for AI-powered features. 
  • Developer Enablement: Ability to support multiple engineering teams, troubleshoot across infrastructure and application layers, document solutions, and enable secure self-service practices.

Elsevier is a renowned global information analytics company that primarily focuses on providing scientific, technical, and medical (STM) research content, tools, and services. It is one of the largest publishers of academic journals and scholarly literature in the world.

Elsevier operates in various domains, including science, technology, medicine, social sciences, and more. They publish a vast number of peer-reviewed journals covering a wide range of disciplines. These journals act as platforms for researchers and academics to share their findings and contribute to the advancement of knowledge in their respective fields.

In addition to publishing, Elsevier offers a suite of digital solutions and services to support researchers, scientists, and professionals in their work. They provide online platforms like ScienceDirect, Scopus, and Mendeley, which offer access to a vast repository of scholarly articles, research papers, and other scientific content. These platforms often serve as essential resources for software developers seeking to stay updated with the latest scientific advancements.

U.S. National Base Pay Range: $95,300 - $158,800. Geographic differentials may apply in some locations to better reflect local market rates. This job is eligible for an annual incentive bonus.

We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits. Click here to access benefits specific to your location.

We are committed to providing a fair and accessible hiring process. If you have a disability or other need that requires accommodation or adjustment, please let us know by completing our Applicant Request Support Form or please contact 1-855-833-5120.

Criminals may pose as recruiters asking for money or personal information. We never request money or banking details from job applicants. Learn more about spotting and avoiding scams here.

Please read our Candidate Privacy Policy.

We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.

USA Job Seekers:

EEO Know Your Rights.

Skills Required

  • Expertise with advanced Terraform, including modules, providers, state management, lifecycle controls, drift detection, safe refactoring, remote state, locking, and cross-stack dependencies
  • Hands-on experience managing production, multi-account, multi-region AWS environments
  • Experience with AWS services including ECS, RDS, ALB, VPC, IAM, Route53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch
  • Experience building and troubleshooting reusable GitHub Actions workflows, OIDC authentication, approval gates, runners, Terraform deployments, application deployments, and migration pipelines
  • Knowledge of Docker, ECR, ECS task definitions and services, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks
  • Proficiency in AWS networking and security, including VPCs, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security best practices
  • Skill in incident response, observability, troubleshooting logs and metrics, root cause analysis, rollback decisions, and operational runbooks
  • Strong Linux and Git fundamentals
  • Experience with Bash or Python scripting for AWS CLI automation, CI/CD, and operational tooling
  • Hands-on experience integrating and operating AI services and APIs in production, including monitoring, reliability, and security
  • Ability to support multiple engineering teams, troubleshoot across infrastructure and application layers, document solutions, and enable secure self-service practices

Elsevier Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Elsevier and has not been reviewed or approved by Elsevier.

  • Leave & Time Off Breadth — Feedback suggests paid time off spans vacation, holidays, sick days, bereavement, military leave, and volunteer time. Family-related leave options are also emphasized as part of the package.
  • Healthcare Strength — Feedback suggests medical, dental, vision, life insurance, wellness initiatives, and an EAP form a comprehensive health offering. Gym support and wellbeing hubs reinforce an ongoing health and wellness focus.
  • Retirement Support — Feedback suggests retirement programs include a 401(k)/retirement plan and pension options, with long-term savings vehicles such as an employee stock purchase plan also available. These elements contribute to a sense of financial security beyond base pay.

Elsevier Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Amsterdam
Year Founded: 1880

What We Do

Elsevier is a world-leading provider of information solutions that enhance the performance of science, health, and technology professionals, empowering them to make better decisions, and deliver better care. Because informed decisions lead to better outcomes, Elsevier is a leader in information and analytics for customers across the global research and health ecosystems. Elsevier helps researchers and healthcare professionals advance science and improve health outcomes for the benefit of society. We do this by facilitating insights and critical decision-making for customers across the global research and health ecosystems.

Similar Jobs

RELX Logo RELX

Senior Site Reliability Engineer

Information Technology • Legal Tech • Analytics
In-Office or Remote
2 Locations
10001 Employees
95K-159K Annually

AuthZed Logo AuthZed

Senior Site Reliability Engineer

Artificial Intelligence • Information Technology • Software • Database
Remote
2 Locations
30 Employees
150K-195K Annually

Cox Enterprises Logo Cox Enterprises

Manager, Key Account Marketing

Artificial Intelligence • Automotive • Greentech • Information Technology • Machine Learning • Software • Cybersecurity
Remote or Hybrid
United States
30000 Employees
89K-134K Annually
In-Office or Remote
2 Locations
175633 Employees
133K-251K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account