Specialist Cloud Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Pune, Mahārāshtra, IND
In-Office
Senior level
Cloud • Software • Analytics
The Role
Lead reliability efforts for cloud-native systems across multi/hybrid-cloud (primarily AWS). Design scalable infrastructure, build IaC and CI/CD automation, own observability (metrics, logs, traces), define SLOs, run incident response and RCA, ensure platform security/compliance, and mentor SREs while partnering with product and platform teams to improve uptime and performance.
Summary Generated by Built In

At NiCE, we don’t limit our challenges. We challenge our limits. Always. We’re ambitious. We’re game changers. And we play to win. We set the highest standards and execute beyond them. And if you’re like us, we can offer you the ultimate career opportunity that will light a fire within you.

So, what's the role all about? 

NICE is looking for a Senior Site Reliability Engineer to join our core Reliability Engineering team, responsible for ensuring the scalability, reliability, and performance of mission-critical systems and observability platforms across multiple environments and regions. This role is ideal for someone who thrives in fast-paced environments, enjoys automation, and has a strong background in cloud-native operations, observability stacks, and incident management. You'll collaborate closely with product, platform, and development teams to drive reliability-first design, proactive observability, and operational excellence. 

How will you make an impact? 

Reliability & Performance 
· Design and implement scalable, reliable, and resilient systems across hybrid or multi-cloud environments (primarily AWS/EKS/ECS/Lambda) 
· Drive improvements in system uptime, latency, and overall service health metrics (SLOs, SLIs, SLAs) 

Automation & Infrastructure as Code 
· Build and manage infrastructure automation using Terraform, Helm, and Kubernetes 
· Improve CI/CD pipelines using Jenkins, GitHub Actions, ensuring safe and automated rollouts, monitoring, and rollbacks 

Observability & Monitoring 
· Own and enhance the observability stack (Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Mimir, etc.) 
· Define and implement SLOs and error budgets; enable development teams to monitor and act on reliability metrics 

Incident Management 
· Lead major incident response, root cause analysis (RCA), and blameless postmortems 
· Partner with product teams to define and enforce operational readiness standards before production releases 

Security & Compliance 
· Ensure platform-level security and compliance with organizational and regulatory standards 
· Collaborate with InfoSec and compliance teams to maintain a secure and auditable infrastructure 

Technical Leadership 
· Mentor junior SREs and developers on reliability practices, automation, and observability 
· Contribute to technical roadmaps and reliability-focused design reviews 

Have you got what it takes? 

  •  8+ Years Strong experience with Kubernetes, EKS, ECS and containerized workloads in production
    · Expertise in AWS services (EC2, Lambda, IAM, RDS, S3, ALB/NLB, VPC, PrivateLink, etc.)
    · Proficiency with Terraform, Helm, Jenkins, GitHub Actions, and GitOps (ArgoCD or Flux) 
    · Deep understanding of observability frameworks – metrics, logs, traces, and distributed monitoring 
    · Hands-on experience with Prometheus, Grafana, Loki, Tempo, Alloy, OpenTelemetry, or equivalent tools 
    · Strong knowledge of Linux, networking fundamentals, and system performance tuning 
    · Familiarity with Python, Go, or Shell scripting for automation and custom tooling 
    · Practical experience in incident response, RCA, and on-call operations 

What's in it for you? 

Join an ever-growing, market disrupting, global company where the teams – comprised of the best of the best – work in a fast-paced, collaborative, and creative environment! As the market leader, every day at NICE is a chance to learn and grow, and there are endless internal career opportunities across multiple roles, disciplines, domains, and locations. 

Enjoy NICE-FLEX! 

At NICE, we work according to the NICE-FLEX hybrid model, which enables maximum flexibility: 2 days working from the office and 3 days of remote work, each week. 

Requisition ID - 11465

Reporting into: Tech Manager 

 Role Type: Individual Contributor 

About NiCE

NICE Ltd. (NASDAQ: NICE) software products are used by 25,000+ global businesses, including 85 of the Fortune 100 corporations, to deliver extraordinary customer experiences, fight financial crime and ensure public safety. Every day, NiCE software manages more than 120 million customer interactions and monitors 3+ billion financial transactions.

Known as an innovation powerhouse that excels in AI, cloud and digital, NiCE is consistently recognized as the market leader in its domains, with over 8,500 employees across 30+ countries.

NiCE is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, age, sex, marital status, ancestry, neurotype, physical or mental disability, veteran status, gender identity, sexual orientation or any other category protected by law.


Skills Required

  • 8+ years experience with Kubernetes, EKS, ECS and containerized workloads in production
  • Expertise with AWS services (EC2, Lambda, IAM, RDS, S3, ALB/NLB, VPC, PrivateLink)
  • Proficiency with Terraform, Helm, and Kubernetes for infrastructure automation
  • Experience improving CI/CD pipelines using Jenkins and GitHub Actions and implementing GitOps (ArgoCD or Flux)
  • Deep understanding of observability frameworks and defining SLOs/SLIs/SLAs
  • Hands-on experience with Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Mimir or equivalent tools
  • Strong knowledge of Linux, networking fundamentals, and system performance tuning
  • Familiarity with Python, Go, or Shell scripting for automation and custom tooling
  • Practical experience in incident response, root cause analysis (RCA), and on-call operations
  • Ability to design and implement scalable, resilient systems across hybrid or multi-cloud environments

NICE Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NICE and has not been reviewed or approved by NICE.

  • Healthcare Strength Benefits are described as broad and comprehensive, spanning medical, dental, vision, life, disability, and mental-health support. Added programs like FSA options and fitness stipends contribute to a well-rounded health and wellness offering.
  • Retirement Support A 401(k) is part of the package, sometimes paired with match details that are described as typical to stronger depending on role and time period. Employee stock participation is also positioned as an additional long-term wealth-building component for eligible roles.
  • Flexible Benefits Flexible work arrangements are emphasized, including hybrid setups and remote options for some roles. Flex scheduling, paid holidays, and paid sick time add to the perceived flexibility of the overall rewards package.

NICE Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Hoboken, NJ
10,130 Employees
Year Founded: 1986

What We Do

NICE (Nasdaq: NICE) is the worldwide leading provider of both cloud and on-premises enterprise software solutions that empower organizations to make smarter decisions based on advanced analytics of structured and unstructured data. NICE helps organizations of all sizes deliver better customer service, ensure compliance, combat fraud and safeguard citizens. Over 25,000 organizations in more than 150 countries, including over 85 of the Fortune 100 companies, are using NICE solutions. www.nice.com.

Similar Jobs

TransUnion Logo TransUnion

Senior Engineer

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Hybrid
Pune, Mahārāshtra, IND
13000 Employees

Optum Logo Optum

Product Manager

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Pune, Mahārāshtra, IND
160000 Employees

Capco Logo Capco

Product Manager

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

CSC Logo CSC

Security Engineer

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
8500 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account