Site Reliability Engineer II

Posted An Hour Ago
Be an Early Applicant
Hyderabad, Telangana, IND
Hybrid
Mid level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
At MetLife, we’re a purpose-driven company that helps our customers build a more confident future.
The Role
Ensure reliability and performance of production services by monitoring, responding to incidents, maintaining runbooks, automating operational tasks, supporting observability and SRE practices (SLIs/SLOs/error budgets), conducting postmortems, and collaborating with engineering, cloud, and infrastructure teams.
Summary Generated by Built In
Description and Requirements
Role Reasonability
MetLife is seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, and performance of critical applications and platforms.
The SRE Engineer will monitor production systems, respond to incidents, improve observability, maintain runbooks, and automate operational tasks. Working closely with engineering, cloud, and infrastructure teams, the role supports SRE practices, operational readiness, and service reliability through SLIs, SLOs, and Error Budget management.
Core Responsibilities
  • Monitoring: Monitor service health, dashboards, alerts, and key reliability indicators for assigned applications and platforms.
  • Incident Response: Respond to alerts, support bridge calls, gather evidence, execute runbooks, communicate status, and escalate when required.
  • Observability Support: Create and maintain dashboards, log queries, telemetry checks, alert validation, and actionable monitoring signals.
  • Runbook Management: Document operational procedures, update recovery steps, validate readiness with service owners, and support knowledge sharing.
  • Automation: Create scripts for repetitive checks, data collection, remediation, operational reporting, and toil reduction.
  • Problem Follow-up: Support root cause analysis, postmortem documentation, and closure of assigned corrective/preventive action items.
  • Continuous Improvement: Identify alert noise, toil, monitoring gaps, and preventive improvements for senior SRE review.
  • SRE Alignment: Support adoption of SLOs, SLIs, SLAs, error budgets, operational readiness reviews, and production support standards.
  • AI Readiness: Use or help improve AI-assisted tools for anomaly detection, incident correlation, root cause hints, and operational knowledge retrieval.
  • Collaboration: Work with engineering, infrastructure, cloud, and application teams to align service performance with business goals.

Skills and Experience
  • Foundations: Linux, networking fundamentals, application support, cloud fundamentals, production operations, and ITIL-style incident/change processes.
  • Scripting: Python, PowerShell, Bash, or equivalent scripting for automation, diagnostics, evidence collection, and reporting.
  • Tools: Git, CI/CD basics, ServiceNow or equivalent ticketing; exposure to Elastic/ELK, Grafana, Prometheus, Splunk, APM, and Azure Monitor preferred.
  • Cloud & Containers: Azure services, Docker, Kubernetes, and hybrid cloud operations exposure; Terraform or infrastructure-as-code awareness preferred.
  • Reliability: Basic understanding of SLIs, SLOs, SLAs, error budgets, alerting, incident response, postmortems, and operational runbooks.
  • AI / AIOps Readiness: Ability to use AI-assisted investigation, anomaly detection, and correlation tools responsibly, with strong validation of evidence.
  • Database: Hands-on SQL skills for operational diagnostics, data validation, and service health checks.
  • Execution: Disciplined follow-through, evidence capture, documentation, escalation hygiene, and collaboration during incidents.
  • Learning Mindset: Willingness to deepen skills in cloud, Kubernetes, observability, automation, resilience engineering, and secure operations.

Minimum Qualifications
  • 3-7 years in production support, DevOps, infrastructure, cloud operations, or software engineering.
  • Experience supporting business-critical systems and working in incident, problem, and change management processes.
  • Ability to script and automate standard operational tasks using Python, PowerShell, Bash, or equivalent.
  • Bachelor's degree in computer science, engineering, or equivalent practical experience.
  • Exposure to regulated enterprise, insurance, banking, or financial services environments preferred.
  • Business proficiency in English; Japanese language skills are a plus.

Preferred Exposure
  • Hybrid cloud platforms including on-premises and Azure-hosted services.
  • Observability platforms such as ELK/Elastic, Grafana, Prometheus, Splunk, Azure Monitor, and Azure Application Insights.
  • GitHub, Azure DevOps, pipelines, repositories, and operational change controls.
  • Kubernetes-based production services and containerized application support.
  • SRE practices including operational readiness reviews, service health reviews, and toil reduction initiatives.

About MetLife
Recognized on Fortune magazine's list of the "World's Most Admired Companies" and Fortune World's 25 Best Workplaces™, MetLife, through its subsidiaries and affiliates, is one of the world's leading financial services companies; providing insurance, annuities, employee benefits and asset management to individual and institutional customers. With operations in more than 40 markets, we hold leading positions in the United States, Latin America, Asia, Europe, and the Middle East.
Our purpose is simple - to help our colleagues, customers, communities, and the world at large create a more confident future. United by purpose and guided by our core values - Win Together, Do the Right Thing, Deliver Impact Over Activity, and Think Ahead - we're inspired to transform the next century in financial services. At MetLife, it's #AllTogetherPossible . Join us!
#BI-Hybrid

Skills Required

  • 3-7 years in production support, DevOps, infrastructure, cloud operations, or software engineering
  • Experience supporting business-critical systems and incident, problem, and change management processes
  • Ability to script and automate operational tasks using Python, PowerShell, or Bash
  • Foundational knowledge: Linux, networking fundamentals, application support, production operations, ITIL-style incident/change processes
  • Hands-on SQL skills for operational diagnostics and data validation
  • Bachelor's degree in computer science, engineering, or equivalent practical experience
  • Experience with cloud and containers: Azure services, Docker, Kubernetes (exposure required)
  • Familiarity with observability and monitoring tools (Elastic/ELK, Grafana, Prometheus, Splunk, APM, Azure Monitor)
  • Experience with Git, CI/CD basics, GitHub or Azure DevOps, and ServiceNow or equivalent ticketing
  • Terraform or infrastructure-as-code awareness
  • Exposure to regulated enterprise, insurance, banking, or financial services environments
  • Business proficiency in English (Japanese language skills a plus)

What the Team is Saying

Chelsea
Nick
Naren
Laura
Bill
Jing Huang
Sara Strauch
Dan Xiao
Cara Mootz
Cara Mootz
Naren Peri
Patricia Hixson
Jing Huang
Samantha Heron
Pete Clarke
Erica Provido
Rene Rivera

MetLife Compensation & Benefits Highlights

  • Retirement Support Feedback suggests retirement programs pair a 401(k) with company match and additional retirement plan structures, indicating a mature long‑term savings setup. This depth signals stability for employees prioritizing retirement security.
  • Healthcare Strength Comprehensive medical, dental, vision, disability, and life insurance are paired with wellness initiatives and 24/7 EAP access, indicating broad core coverage. HSAs/FSAs and second‑opinion resources further reinforce the healthcare offering.
  • Parental & Family Support Fertility coverage, personalized child/elder‑care guidance, adoption assistance, parental leave, and breast‑milk shipping for traveling nursing parents are highlighted. These supports extend benefits meaningfully beyond the basics.

MetLife Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
43,000 Employees
Year Founded: 1868

What We Do

We're honored to be No. 10 on Great Place to Work's World's Best Workplaces and recognized in the Fortune 100 Best Companies to Work For® list in 2025. At MetLife, we're leading the global transformation of an industry we’ve defined for over 157 years. At MetLife, every innovation and line of code is a lifeline for our customers and their families—from victims of natural disasters to people living with disabilities and beyond. With operations in more than 40 markets and leading positions across the globe, MetLife fosters an inclusive culture where our people are energized and inspired to deliver for our customers and communities. Join our remarkable journey—one in which you help write the next century of innovation in financial services—because with MetLife, making the world a better place is All Together Possible.

Why Work With Us

At MetLife, you’ll be working for a company whose purpose is to help customers throughout their life’s journey, and often in their most critical time of need. You’ll be a part of developing leading-edge platforms that will have a lasting impact on the lives and well-being of tens of millions of customers.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

MetLife Teams

Team
Product + Tech
About our Teams

MetLife Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

MetLife's current workplace policies classify roles as Office, Hybrid or Virtual based on the nature of work, encouraging new ways of working together

Typical time on-site: Flexible
Company Office Image
HQNew York City, NY
Company Office Image
Pune, IN
Company Office Image
Mexico City, MX
Company Office Image
Bridgewater, NJ
Company Office Image
Cary, NC
Company Office Image
Clark Summit, PA
Company Office Image
Greenville, SC
Company Office Image
Hyderabad, IN
Company Office Image
Tampa, FL
Company Office Image
Whippany, NJ
Learn more

Similar Jobs

MetLife Logo MetLife

Software Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Hyderabad, Telangana, IND
43000 Employees

MetLife Logo MetLife

MSEPL- Team Leader - FP & A

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Hyderabad, Telangana, IND
43000 Employees

MetLife Logo MetLife

Senior Data Governance Analyst II

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Hyderabad, Telangana, IND
43000 Employees

MetLife Logo MetLife

Senior Data Goverance Analyst II

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Hyderabad, Telangana, IND
43000 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account