Incident Management Reliability Engineer

Posted 2 Days Ago
Be an Early Applicant
Barcelona, Cataluña, ESP
Hybrid
50K-74K Annually
Senior level
Healthtech
The Role
Lead major incident management and coordinate technical teams during critical outages. Drive root-cause analysis, post-incident improvements, reliability engineering, observability, monitoring, automated recovery, capacity planning, and SLO/SLI/SLA management. Partner with service owners and platform teams to improve fault tolerance and eliminate recurring issues through automation, runbooks, AIOps, and structured service-management practices.
Summary Generated by Built In

  • Location: Barcelona, Spain

About the Job


As an Incident Management Reliability Engineer within our Service Quality team, you'll play a pivotal role in safeguarding the stability and resilience of critical IT services that support a global biopharmaceutical company dedicated to improving patients' lives.

Sitting at the heart of Sanofi's Digital organization, you'll blend deep incident management expertise with reliability engineering principles to minimize disruptions, accelerate recovery, and continuously raise the bar on system performance. This is more than keeping the lights on — it's about transforming every challenge into an opportunity for growth, and ensuring that the technology powering life-changing medicines never misses a beat

 

About Sanofi


Sanofi is dedicated to supporting people through their health challenges. We are a global biopharmaceutical company focused on human health, committed to chasing the miracles of science to improve people's lives. Our team, across some 100 countries, is dedicated to transforming the practice of medicine by working to turn the impossible into the possible. We provide potentially life-changing treatment options and life-saving vaccine protection to millions of people globally, while putting sustainability and social responsibility at the center of our ambitions

At Sanofi, we believe our people are our greatest strength, and we are committed to creating an inclusive environment where diverse perspectives drive innovation. Our values — Aim Higher, Act for Patients, Be Bold, and Lead Together — guide everything we do

 

Main responsibilities:


Incident Management


  • Partner Lead end-to-end management of Major Incidents (P1/P2), ensuring timely resolution and clear stakeholder communication
  • Serve as Serve as command centre lead during critical outages, coordinating seamlessly across technical and business teams
  • Maintain accurate and detailed incident documentation — root cause, timeline, and resolution steps
  • Drive post-incident reviews and ensure action items are implemented to prevent recurrence
  • Uphold consistent communication and escalation processes aligned with ITSM best practices (ITIL)
  • Provide technical leadership and guidance during major incidents, challenging and validating diagnosis, troubleshooting approaches and proposed resolution actions to improve the quality and speed of technical decision-making
  • Bring the right technical expertise together, facilitate cross-functional troubleshooting, and help teams identify and resolve complex technical issues, particularly during high-severity and business-critical outages
  • Drive a structured, AIOps and data-driven approach to diagnosis and resolution, ensuring that technical teams remain focused on restoring service while maintaining appropriate risk controls

Reliability Engineering


  • Partner with service owners and platform teams to enhance service reliability, observability, and fault tolerance
  • Implement proactive monitoring, alerting, and automated recovery mechanisms
  • Analyse incident trends and develop targeted reliability improvement plans
  • Contribute to capacity planning, change reviews, and failure mode analysis to anticipate and mitigate risks
  • Define and track SLOs/SLIs/SLAs to measure and communicate service health

Continuous Improvement


  • Collaborate with problem management to identify recurring issues and lead root cause elimination initiatives
  • Automate operational tasks and enhance service recovery using scripts, runbooks, and AIOps tools
  • Shape the evolution of the Major Incident Process, embedding best practices across the organization

About you:


  • Experience: 8+ years in incident management, reliability engineering, or a related IT operations discipline
  • Technical skills: Strong foundation in networking and infrastructure as code; additional experience in cloud technologies (AWS, Azure, or GCP), virtualization, containerization, automation, databases, and middleware/scheduling is a plus
  • Certifications (preferred): ITIL v4 or Service Operations, SRE Foundation/Practitioner, cloud certifications (AWS, Azure, or GCP), or Incident Command System (ICS) / equivalent crisis leadership training
  • Soft skills: Exceptional communicator — calm under pressure, clear with stakeholders at all levels, and a natural collaborator who thrives in fast-moving, high-stakes environments
  • Languages: Fluent English (written and verbal)

Why choose us?


  • Global mission, real impact — your work directly supports the technology that brings life-changing medicines to patients worldwide
  • Be the guardian of resilience — own the reliability of critical systems that thousands of colleagues and patients depend on every day
  • Drive continuous innovation — leverage AIOps, automation, and modern SRE practices in a large-scale enterprise environment
  • Collaborative, service-first culture — join a team where quality is a shared purpose, not just a metric
  • Grow with us — access world-class learning opportunities, certifications, and a clear path to deepen your expertise in reliability and service excellence

Pursue Progress. Discover Extraordinary.


Join Sanofi and step into a new era of science - where your growth can be just as transformative as the work we do. We invest in you to reach further, think faster, and do what’s never-been-done-before. You’ll help push boundaries, challenge convention, and build smarter solutions that reach the communities we serve. Ready to chase the miracles of science and improve people’s lives? Let’s Pursue Progress and Discover Extraordinary – together.


At Sanofi, we provide equal opportunities to all regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, protected veteran status or other characteristics protected by law.


#LI-Hybrid #BarcelonaHub #SanofiHubs

Pursue progress, discover extraordinary

Better is out there. Better medications, better outcomes, better science. But progress doesn’t happen without people – people from different backgrounds, in different locations, doing different roles, all united by one thing: a desire to make miracles happen. So, let’s be those people.

At Sanofi, we provide equal opportunities to all regardless of race, colour, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, ability or gender identity.

Watch our ALL IN video and check out our Diversity Equity and Inclusion actions at sanofi.com!

The salary range for this position is :€49.600,00 - €74.400,00

Final compensation will be determined based on demonstrated experience, skills, location, and other relevant factors. Employees may be eligible to participate in Company employee benefit programs.

Skills Required

  • 8+ years of experience in incident management, reliability engineering, or related IT operations
  • Strong foundation in networking
  • Experience with infrastructure as code
  • Fluent English, written and verbal
  • Experience with cloud technologies such as AWS, Azure, or GCP
  • Experience with virtualization
  • Experience with containerization
  • Experience with automation
  • Experience with databases
  • Experience with middleware and scheduling systems
  • ITIL v4 or Service Operations certification
  • SRE Foundation or Practitioner certification
  • AWS, Azure, or GCP cloud certification
  • Incident Command System or equivalent crisis leadership training

Sanofi Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Sanofi and has not been reviewed or approved by Sanofi.

  • Retirement Support — Retirement support stands out through a notably strong 401K matching structure (e.g., 150% match up to a 6% contribution), which materially boosts total rewards for long-tenured employees.
  • Parental & Family Support — Parental and family support is positioned as unusually robust, including a global gender-neutral paid parental leave standard (14 weeks) and added supports such as childcare assistance and adoption/surrogacy/infertility help.
  • Equity Value & Accessibility — Equity participation is made more accessible via an Employee Stock Purchase Plan that includes a meaningful purchase discount and matching/free-share mechanics, increasing perceived total compensation beyond base pay.

Sanofi Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Paris
85,000 Employees
Year Founded: 1973

What We Do

We are Sanofi, an innovative global healthcare company. We chase the miracles of science to improve people’s lives. Our team, across some 100 countries, is dedicated to transforming the practice of medicine by working to turn the impossible into the possible. We provide potentially life-changing treatment options and life-saving vaccine protection to millions of people globally, while putting sustainability and social responsibility at the center of our ambitions. Interactions with this account must comply with the Terms: https://bit.ly/sanofi-terms

Similar Jobs

Benchling Logo Benchling

Account Executive

Cloud • Healthtech • Social Impact • Software • Biotech
Remote or Hybrid
27 Locations
605 Employees

MongoDB Logo MongoDB

Senior Solutions Architect

Big Data • Cloud • Software • Database
Easy Apply
Hybrid
3 Locations
5550 Employees

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Architect

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office or Remote
3 Locations
85422 Employees

Academia.edu Logo Academia.edu

Peer Review Assistant

Consumer Web • Digital Media • Edtech • Information Technology • Social Impact • Software
Remote or Hybrid
26 Locations
110 Employees

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees
Vitalize Thumbnail
Artificial Intelligence • Healthtech • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account