Site Reliability Engineer

Posted 6 Hours Ago
Be an Early Applicant
2 Locations
In-Office
Senior level
Information Technology • Business Intelligence • Consulting
The Role
Ensures the reliability, availability, performance, and scalability of systems and infrastructure. Responsibilities include incident response, troubleshooting, root cause analysis, monitoring, infrastructure automation, infrastructure as code, capacity planning, security compliance, CI/CD, deployment optimization, and post-incident improvement. The role collaborates with development, operations, and security teams to automate processes and maintain resilient, highly available environments.
Summary Generated by Built In

Make an impact with NTT DATA
Join a company that is pushing the boundaries of what is possible. We are renowned for our technical excellence and leading innovations, and for making a difference to our clients and society. Our workplace embraces diversity and inclusion – it’s a place where you can grow, belong and thrive.

Your day at NTT DATA
The Site Reliability Engineer (SRE) is a seasoned subject matter expert, responsible for ensuring the reliability, availability, and performance of company systems and infrastructure.
This Site Reliability Engineer (SRE) works closely with development teams, operations teams, and other stakeholders to enhance system resiliency, automate processes, and improve overall system reliability.

Key responsibilities:
  • Monitors system health, performance metrics, and alerts to identify and respond to incidents promptly and diagnoses issues, troubleshoots problems, and restores services in a timely manner.
  • Implements incident response processes to minimize downtime and improve system availability.
  • Designs, develops, and maintains automation tools, scripts, and processes to streamline system management tasks, deployments, and configuration changes.
  • Implements infrastructure-as-code principles to ensure consistency and repeatability.
  • Optimizes system resources, configurations, and processes to enhance performance, scalability, and efficiency.
  • Uses monitoring tools and performance testing to identify bottlenecks and implement optimizations.
  • Collaborates with teams to forecast system resource needs, plans for capacity growth, and ensures adequate scalability.
  • Leads incident response efforts, coordinates with cross-functional teams, and drives the resolution of system issues.
  • Performs thorough post-incident analysis to identify root causes and implements preventive measures to minimize future incidents.
  • Identifies opportunities for automation and drives the implementation of self-healing, monitoring, and deployment of automation tools and frameworks.
  • Continuously improves operational efficiency, system reliability, and availability through process enhancements and automation.
  • Ensures consistency across environments, tracks changes, and enforces configuration standards.
  • Works closely with development teams, operations teams, and other stakeholders to ensure effective collaboration, knowledge sharing, and alignment on reliability goals.
  • Implements security best practices, works with security teams to assess and address vulnerabilities, and ensures compliance with security standards and regulations.
  • Performs any other related task as required.

To thrive in this role, you need to have:
  • Seasoned technical expertise in Linux/Unix systems, networking, and system administration.
  • Seasoned proficiency in scripting or programming languages, such as Python, Go, Java, or Ruby.
  • Seasoned knowledge of cloud platforms (such as AWS, Azure, or Google Cloud) and associated services.
  • Seasoned proven expertise in performance monitoring, optimization, and troubleshooting using tools such as Prometheus, Grafana, or New Relic.
  • Seasoned expertise in incident management, root cause analysis, and post-incident reviews
  • Excellent problem-solving and analytical skills, with a keen attention to detail.
  • Excellent communication, collaboration, and leadership skills.
  • Seasoned ability to optimize system performance, scalability, and reliability. experience with performance monitoring and tuning tools (for example, Prometheus, Grafana, or New Relic) to identify bottlenecks, analyze performance data, and implement optimization strategies.
  • Seasoned understanding of security principles, best practices, and compliance requirements. experience in designing and implementing security controls, performing security assessments, and ensuring compliance with industry standards.

Academic qualifications and certifications:
  • Bachelor's degree or equivalent in Computer Science, Information Technology, or a related field.
  • Relevant certifications, such as AWS Certified DevOps Engineer - Professional, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator (CKA) preferred.

Required experience:
  • Seasoned hands-on experience in a Site Reliability Engineering role or related roles, including experience in designing and maintaining highly available and scalable systems.
  • Seasoned hands-on experience with Linux/Unix systems, networking, and system administration is crucial. In-depth knowledge of cloud platforms (such as AWS, Azure, or Google Cloud) and associated services is essential.
  • Seasoned proficiency in multiple programming languages like Python, Java, Go, or Ruby is important for developing and maintaining automation tools, frameworks, and complex system integrations. Expertise in scripting languages like Bash or PowerShell is beneficial.
  • Seasoned understanding of complex infrastructure architectures, including scalable and fault-tolerant designs. experience with infrastructure-as-code tools (such as Terraform or CloudFormation) and containerization technologies (such as Docker or Kubernetes) is essential.
  • Seasoned experience in designing and implementing robust automation frameworks, CI/CD pipelines, and deployment strategies. Proficiency in tools like Jenkins, GitLab CI/CD, or CircleCI to build, test, and deploy applications with a focus on reliability and scalability.
  • Seasoned experience in incident management, troubleshooting complex system issues, and conducting post-incident analysis. Advanced ability to lead incident response efforts, drive root cause analysis, and implement preventive measures.
  • Seasoned understanding of DevOps principles, Agile methodologies, and a strong commitment to continuous improvement and learning. experience in promoting a DevOps culture and driving the adoption of best practices

Workplace type:

Hybrid Working

About NTT DATA
NTT DATA is a $30+ billion business and technology services leader, serving 75% of the Fortune

Global 100. We are committed to accelerating client success and positively impacting society through

responsible innovation. We are one of the world’s leading AI and digital infrastructure providers, with

unmatched capabilities in enterprise-scale AI, cloud, security, connectivity, data centers and

application services. Our consulting and industry solutions help organizations and society move

confidently and sustainably into the digital future. As a Global Top Employer, we have experts in more

than 70 countries. We also offer clients access to a robust ecosystem of innovation centers as well as

established and start-up partners. NTT DATA is part of NTT Group, which invests over $3 billion each

year in R&D.


Equal Opportunity Employer
NTT DATA is proud to be an Equal Opportunity Employer with a global culture that embraces diversity. We are committed to providing an environment free of unfair discrimination and harassment. We do not discriminate based on age, race, colour, gender, sexual orientation, religion, nationality, disability, pregnancy, marital status, veteran status, or any other protected category. Join our growing global team and accelerate your career with us. Apply today.


Third parties fraudulently posing as NTT DATA recruiters 

NTT DATA recruiters will never ask job seekers or candidates for payment or banking information during the recruitment process, for any reason. Please remain vigilant of third parties who may attempt to impersonate NTT DATA recruiters whether in writing or by phone in order to deceptively obtain personal data or money from you. All email communications from an NTT DATA recruiter will come from an @nttdata.com email address. If you suspect any fraudulent activity, please contact us.

Skills Required

  • Bachelor's degree or equivalent in Computer Science, Information Technology, or a related field
  • Hands-on experience in Site Reliability Engineering or related roles
  • Experience designing and maintaining highly available and scalable systems
  • Hands-on experience with Linux or Unix systems, networking, and system administration
  • Knowledge of cloud platforms such as AWS, Azure, or Google Cloud and associated services
  • Proficiency in multiple programming languages, including Python, Java, Go, or Ruby
  • Experience with scripting languages such as Bash or PowerShell
  • Understanding of scalable and fault-tolerant infrastructure architectures
  • Experience with infrastructure-as-code tools such as Terraform or CloudFormation
  • Experience with containerization technologies such as Docker or Kubernetes
  • Experience designing and implementing automation frameworks, CI/CD pipelines, and deployment strategies
  • Proficiency with Jenkins, GitLab CI/CD, or CircleCI
  • Experience in incident management, complex troubleshooting, root cause analysis, and post-incident reviews
  • Understanding of DevOps principles and Agile methodologies
  • Knowledge of performance monitoring, optimization, and troubleshooting tools such as Prometheus, Grafana, or New Relic
  • Understanding of security principles, security controls, security assessments, and compliance requirements
  • AWS Certified DevOps Engineer Professional, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator certification
  • Excellent problem-solving, analytical, communication, collaboration, and leadership skills

NTT DATA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NTT DATA and has not been reviewed or approved by NTT DATA.

  • Fair & Transparent Compensation — Feedback suggests salary bands and grades are clearly defined, making ranges and promotion criteria easier to understand. Standardized HR processes provide visibility into levels across common delivery roles.
  • Healthcare Strength — Feedback suggests the package includes comprehensive medical, dental, and vision options with HSA/FSA eligibility. Global materials emphasize comprehensive insurance and wellbeing as baseline offerings across regions.
  • Wellbeing & Lifestyle Benefits — Flexible work options, including remote/hybrid arrangements, are highlighted as core benefits and can support work–life balance. Some delivery teams report more manageable hours than strategy consultancies, improving perceived value for time.

NTT DATA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Tokyo
55,092 Employees
Year Founded: 1988

What We Do

NTT DATA, Inc. is a trusted global innovator of business and technology services. We're committed to helping clients innovate, optimize and transform for long-term success. Our R&D investments help organizations and society move confidently and sustainably into the digital future. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure, and connectivity

Similar Jobs

MetLife Logo MetLife

Site Reliability Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Hyderabad, Telangana, IND
43000 Employees
Hybrid
Hyderabad, Telangana, IND
3062 Employees

Regeneron Logo Regeneron

Site Reliability Engineer

Biotech • Pharmaceutical
In-Office
Hyderabad, Telangana, IND
15000 Employees

Regeneron Logo Regeneron

Site Reliability Engineer

Biotech • Pharmaceutical
In-Office
Hyderabad, Telangana, IND
15000 Employees

Similar Companies Hiring

Compa Thumbnail
Artificial Intelligence • HR Tech • Software • Business Intelligence
Irvine, California
75 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account