Senior Site Reliability Engineer (AWS Cloud)

Posted Yesterday
Be an Early Applicant
Hiring Remotely in RO
Remote
Senior level
Artificial Intelligence • Information Technology • Machine Learning • Software • Virtual Reality • Analytics
The Role
Design, implement, and drive reliability, scalability, and performance of a multi-cloud provisioning platform. Build end-to-end automation and CI/CD pipelines, define and monitor SLIs/SLOs, lead incident response and RCA, collaborate with product and engineering teams to embed reliability and security, and simplify systems to remove single points of failure and reduce operational toil.
Summary Generated by Built In
Company Description

👋🏼 We're Nagarro.

We are a digital product engineering company that is scaling in a big way! We build products, services, and experiences that inspire, excite, and delight. We work at scale — across all devices and digital mediums, and our people exist everywhere in the world (18 000+ experts across 39 countries, to be exact). Our work culture is dynamic and non-hierarchical. We're looking for great new colleagues. That's where you come in!

By this point in your career, it is not just about the tech you know or how well you can code. It is about what more you want to do with that knowledge. Can you help your teammates proceed in the right direction? Can you tackle the challenges our clients face while always looking to take our solutions one step further to succeed at an even higher level? Yes? You may be ready to join us.

Job Description

  • Architect & Drive the reliability, scalability, and performance of our multi-cloud provisioning platform across all production stacks.
  • Architect and implement end-to-end automation pipelines to eliminate manual intervention, actively identifying and reducing technical toil.
  • Define, monitor, and improve critical system health indicators (SLIs/SLOs), including latency, throughput, error rates, and capacity usage, making data-driven architectural recommendations.
  • Lead Collaboration with product and cross-functional engineering teams to embed reliability and security considerations early into the software development lifecycle (SDLC).
  • Own Incident Response Management: Design robust detection mechanisms, triage critical incidents, lead deep Root Cause Analysis (RCA), and implement long-term preventative engineering solutions.
  • Simplify Complex Systems: Continually audit platform operations to identify bottlenecks, eliminate single points of failure, and reduce structural complexity.

Qualifications

  • 8+ years of experience working within the cloud environment, in roles such as SRE (Site Reliability Engineer) or Cloud Platform/Reliability Engineer
  • Strong experience in cloud development and multi-cloud environments, preferably with a strong exposure to AWS cloud
  • Knowledge of cloud architecture, scalability, and high-availability design
  • Hands-on experience with Kubernetes and container orchestration
  • Experience with Terraform and Infrastructure as Code (IaC)
  • Experience designing automation and CI/CD pipelines to reduce operational toil
  • Strong understanding of SRE principles, SLIs/SLOs, monitoring, and observability
  • Proven experience with Incident Management, Root Cause Analysis (RCA), and reliability engineering
  • Ability to identify performance bottlenecks, single points of failure, and architectural risks

Skills Required

  • 8+ years of experience in cloud environments in SRE or Cloud Platform/Reliability roles
  • Strong experience with cloud development and multi-cloud environments, with strong exposure to AWS
  • Knowledge of cloud architecture, scalability, and high-availability design
  • Hands-on experience with Kubernetes and container orchestration
  • Experience with Terraform and Infrastructure as Code (IaC)
  • Experience designing automation and CI/CD pipelines to reduce operational toil
  • Strong understanding of SRE principles, SLIs/SLOs, monitoring, and observability
  • Proven experience with Incident Management, Root Cause Analysis (RCA), and reliability engineering
  • Ability to identify performance bottlenecks, single points of failure, and architectural risks

Nagarro Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Nagarro and has not been reviewed or approved by Nagarro.

  • Pay Growth & Progression Compensation is at times described as competitive, with salary hikes and perks occurring on certain occasions. Better growth opportunities and compensation are also positioned as an advantage versus other service-based companies.
  • Flexible Benefits Work arrangements are framed around a “work-from-anywhere” mindset with flexitime and family-friendly working models. This flexibility appears to add meaningful value to the overall rewards package for many roles.
  • Healthcare Strength Medical, dental, and vision coverage are described as available for employees and dependents, alongside life insurance. Mental-health support is also included via an Employee Assistance Program (EAP).

Nagarro Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Munich
19,994 Employees
Year Founded: 1996

What We Do

Nagarro helps future-proof your business through a forward-thinking, fluidic, and CARING mindset. We excel at digital engineering and help our clients become human-centric, digital-first organizations, augmenting their ability to be responsive, efficient, intimate, creative, and sustainable. Today, we are 19,000 experts across 36 countries, forming a Nation of Nagarrians, ready to help our customers succeed.

Similar Jobs

Wizeline Logo Wizeline

Senior Site Reliability Engineer

Information Technology • Consulting
Remote
România
1444 Employees

Tulip Logo Tulip

Technical Account Manager

Enterprise Web • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
29 Locations
310 Employees

Mondelēz International Logo Mondelēz International

Change Manager o9 MEU, Demand Planning

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
9 Locations
90000 Employees

Smartling Logo Smartling

Business Development Representative

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Natural Language Processing • Software
Easy Apply
Remote
27 Locations
117 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account