Principal Infrastructure Engineer - Major Incident Manager

Posted 14 Days Ago
Be an Early Applicant
Bangalore, Bengaluru Urban, Karnataka, IND
In-Office
Expert/Leader
Financial Services
The Role
Lead incident command and technical response for high-severity production incidents across applications, infrastructure, cloud, and network. Coordinate cross-functional teams, drive triage, recover services, perform RCA/PIRs, and deliver executive reporting. Partner with SRE and engineering to improve reliability through SLOs/SLIs, observability, automation, and process improvements while ensuring ITSM and regulatory compliance.
Summary Generated by Built In

FC Global Services India LLP (First Citizens India), a part of First Citizens BancShares, Inc., a top 20 U.S. financial institution, is a global capability center (GCC) based in Bengaluru. Our India-based teams benefit from the company’s over 125-year legacy of strength and stability. First Citizens India is responsible for delivering value and managing risks for our lines of business. We are particularly proud of our strong, relationship-driven culture and our long-term approach, which are deeply ingrained in our talented workforce. This is evident across all key areas of our operations, including Technology, Enterprise Operations, Finance, Cybersecurity, Risk Management, and Credit Administration. We are seeking talented individuals to join us in our mission of providing solutions fit for our clients’ greatest ambitions.

Job Description:

Value Proposition

Responsible for enhancing the reliability, resilience, and stability of enterprise IT services through effective Major Incident Management and rapid service restoration.

Leading the command, coordination, and technical response for critical production incidents impacting business-critical applications, infrastructure, cloud platforms, and customer-facing services.

Acting as the central Incident Commander during high-severity incidents, driving technical triage, root cause identification, resolution, and minimizing business impact and operational risk.

Partnering with Engineering, SRE, Infrastructure, Cybersecurity, and Application teams to strengthen platform reliability, operational resilience, monitoring, automation, and service maturity.

Job Details

Position Title:  Principal infrastructure Engineer – Major Incident Manager

Career Level:  P4

Job Category: Assistant Vice President

Role Type: Hybrid

Job Location:  Bangalore

About the Team:

The Major Incident Management Team serves as the central command function for critical incidents across global operations, providing 24x7 coverage through teams in India and the US.

Focused on rapid recovery and business continuity, the team leads coordinated incident response, promotes industry best practices, drives continuous improvement, and strengthens operational resilience to support a world-class financial institution.

Key Deliverables (Duties and Responsibilities)

  • Major Incident Command & Technical Leadership

  • Serve as the Incident Commander for major incidents across enterprise applications, cloud platforms, infrastructure, databases, and network services.

  • Lead end-to-end major incident management activities including incident assessment, prioritization, escalation, responder engagement, technical coordination, recovery execution, and service restoration.

  • Direct technical bridge calls and facilitate collaboration among Infrastructure, Cloud, Network, Database, Middleware, Application Development, Cybersecurity, Vendor, and SRE teams.

  • Drive structured technical triage and ensure investigation efforts remain focused on business recovery and root cause isolation.

  • Challenge incomplete technical updates, validate remediation approaches, and remove operational blockers during critical situations.

  • Make informed decisions under pressure while maintaining clear ownership, accountability, and resolution momentum.

  • Ensure incidents are managed in accordance with established ITSM, operational risk, and regulatory requirements.

  • Service Reliability & SRE Partnership

  • Partner with Site Reliability Engineering (SRE), Production Support, and Engineering teams to improve service reliability and operational resilience.

  • Promote adoption of reliability engineering practices including:

    • Service Level Indicators (SLIs)

    • Service Level Objectives (SLOs)

    • Service Level Agreements (SLAs)

    • Error Budgets

    • Observability

    • Proactive Monitoring

    • Event Correlation

    • Capacity Planning

    • Operational Readiness Reviews

  • Identify recurring incidents, systemic weaknesses, and reliability risks across technology platforms.

  • Contribute to the continuous improvement of operational stability through automation, monitoring enhancements, and process optimization.

  • Support operational readiness for major releases, infrastructure transformations, and cloud migrations.

  • Incident Analysis & Continuous Improvement

  • Lead Post-Incident Reviews (PIRs) and Root Cause Analysis (RCA) activities.

  • Partner with Problem Management teams to identify corrective and preventive actions.

  • Ensure action items are tracked through closure and measurable service improvements are achieved.

  • Drive improvements to runbooks, knowledge articles, operational procedures, escalation paths, and response frameworks.

  • Identify opportunities to reduce Mean Time to Detect (MTTD), Mean Time to Restore (MTTR), and recurring service disruptions.

  • Reporting & Operational Insights

  • Produce executive-level incident communications and reporting.

  • Develop and publish operational dashboards and metrics including:

    • MTTD

    • MTTA

    • MTTR

    • Incident Volumes

    • Availability Metrics

    • Escalation Trends

    • SLA Compliance

    • Recurring Incident Analysis

  • Provide actionable insights to technology leadership regarding operational health and service reliability.

  • Stakeholder Management & Communication

  • Provide timely, accurate, and concise communication to executive leadership, business stakeholders, and technical teams during major incidents.

  • Translate complex technical issues into business impact language suitable for senior leadership.

  • Lead stakeholder communications throughout the incident lifecycle.

  • Ensure high-quality incident documentation, timelines, and executive summaries.

  • Governance, Risk & Compliance

  • Ensure adherence to ITIL, Incident Management, Problem Management, Change Management, and Operational Risk frameworks.

  • Support audit, compliance, regulatory, and risk management requirements.

  • Maintain incident governance standards, documentation, playbooks, and escalation procedures.

  • Participate in regulatory and operational resilience initiatives within the organization.

Skills and Qualification:

Qualifications Required

Bachelor’s degree in Computer Science, Information Technology, Engineering, or related discipline and 12+ years of experience in Technology Operations, Production Support, Incident Management, Site Reliability Engineering, Infrastructure & application support Operations or Enterprise Technology Services.

Required Experience

  • 12-15 years of experience supporting enterprise-scale technology environments.

  • Hands-on Major Incident Management, Incident Command, Production Operations, SRE or Operations Engineering experience.

  • Proven experience leading Severity 1 / Critical production incidents within large-scale banking enterprise environments.

  • Demonstrated experience coordinating technical teams across infrastructure, cloud, applications, databases, networking, cybersecurity and vendor organizations.

  • Experience within Banking, Insurance or other regulated Financial Services environments is mandatory.

  • Strong experience working within ITIL-aligned Incident, Problem, and Change Management frameworks.

  • Experience driving operational improvements, service reliability initiatives, and incident reduction programs.

  • Proven ability to influence technical teams and senior stakeholders during high-pressure situations.

Technical Proficiency

The successful candidate should possess a good technical foundation and the ability to effectively engage engineering teams during incident diagnosis and recovery efforts.

Infrastructure & Platform Knowledge

Hands-on expertise and understanding of:

Network Infrastructure

Windows and Virtualization Platforms

Linux & AIX Server Platforms

Backup & Storage Technologies

Data Centers

Databases

Middleware Technologies

API Integrations

Cloud & Modern Platforms

Good working knowledge of:

Cloud Technologies (AWS/Azure/GCP)

Container Platforms (Docker/Kubernetes)

Microservices Architectures

Distributed Systems

Monitoring & Observability

Hands-on familiarity with enterprise monitoring and observability platforms such as, but not limited to:

Splunk

Dynatrace

SolarWinds

Prometheus

Zabbix

ITSM & Incident Tooling

Experience with incident management and collaboration platforms including:

ServiceNow

Everbridge

Jira

Confluence

Microsoft Teams

Good to Have skills:

Automation & Scripting

Experience with automation and operational tooling is preferred.

Working knowledge of:

Ansible

Python

Shell Scripting

PowerShell

Artificial Intelligence & Productivity Tools

Preferred Certifications

ITIL Foundation V3/V4 or higher

Site Reliability Engineering (SRE) Certification

AWS Certified Solutions Architect

Microsoft Azure Certification

Google Cloud Certification

DevOps Certifications

Six Sigma Green Belt or higher

Relationships & Collaboration

Reports to: Director - Technology

Partners: Senior leaders and cross-functional teams across geographies ( India and US team)

Accessibility Needs

We are committed to providing an inclusive and accessible hiring process. If you require accommodations at any stage (e.g. application, interviews, onboarding) please let us know, and we will work with you to ensure a seamless experience.

Equal Employment Opportunity

FC Global Services India LLP (First Citizens India) is an Equal Employment Opportunity Employer. We are committed to fostering an inclusive and accessible environment and prohibit all forms of discrimination on the basis of gender, religion, caste, disability, sexual orientation, economic status or any other characteristics protected by the law. We strive to foster a safe and respectful environment in which all individuals are treated with respect and dignity. Our EEO policy ensures fairness throughout the employee life cycle.

Skills Required

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related discipline
  • 12+ years experience in Technology Operations, Production Support, Incident Management, SRE, Infrastructure or Enterprise Technology Services
  • Hands-on Major Incident Management, Incident Command, Production Operations, SRE or Operations Engineering experience
  • Proven experience leading Severity 1 / Critical production incidents in large-scale banking enterprise environments
  • Experience coordinating technical teams across infrastructure, cloud, applications, databases, networking, cybersecurity and vendor organizations
  • Experience working within Banking, Insurance or other regulated Financial Services environments
  • Strong experience with ITIL-aligned Incident, Problem, and Change Management frameworks
  • Hands-on expertise with Network Infrastructure, Windows, Virtualization, Linux & AIX, Backup & Storage, Data Centers, Databases, Middleware, and API integrations
  • Working knowledge of cloud technologies (AWS/Azure/GCP), container platforms (Docker/Kubernetes), microservices and distributed systems
  • Hands-on familiarity with monitoring/observability tools (Splunk, Dynatrace, SolarWinds, Prometheus, Zabbix)
  • Experience with incident management and collaboration platforms (ServiceNow, Everbridge, Jira, Confluence, Microsoft Teams)
  • Experience with automation and scripting (Ansible, Python, Shell, PowerShell)
  • Preferred certifications: ITIL Foundation V3/V4, SRE Certification, AWS/Azure/GCP certifications, DevOps certifications, Six Sigma Green Belt
  • Experience driving operational improvements, service reliability initiatives, and incident reduction programs
  • Ability to produce executive-level incident communications, operational dashboards, and metrics (MTTD, MTTA, MTTR, availability, SLA compliance)

Silicon Valley Bank Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Silicon Valley Bank and has not been reviewed or approved by Silicon Valley Bank.

  • Healthcare Strength Medical, dental, and vision coverage are broad, complemented by mental health resources, healthcare advocacy, and wellness programs. Multiple plan options and added programs (e.g., digital health and well-being services) indicate depth in healthcare support.
  • Retirement Support Retirement programs include a matched 401(k) and additional savings vehicles such as ESOP and ESPP. Match levels and potential discretionary contributions signal strong employer support for long-term savings.
  • Parental & Family Support Parental leave provides up to 12 weeks at full base pay for birth or adoption. Family support extends to back-up care, special needs assistance, college coaching, and inclusive family-building benefits like fertility, adoption, and surrogacy support.

Silicon Valley Bank Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
3,970 Employees
Year Founded: 1983

What We Do

Silicon Valley Bank (SVB), a division of First Citizens Bank, is the bank of the world’s most innovative companies and investors. SVB provides commercial and private banking to individuals and companies in the technology, life science and healthcare, private equity, venture capital and premium wine industries. SVB operates in centers of innovation throughout the United States, serving the unique needs of its dynamic clients with deep sector expertise, insights and connections. SVB's parent company, First Citizens Bancshares, Inc. (NASDAQ: FCNCA) is a top 20 U.S. financial institution with more than $200 billion in assets. First Citizens Bank, Member FDIC. Learn more at svb.com.

Similar Jobs

SciPlay Logo SciPlay

VIP Service Representative | Gaming Exp Mandate

Gaming • Marketing Tech • Mobile • Software • App development
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
1000 Employees
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account