We are seeking a highly accomplished Principal Incident Commander / Director – Incident Management to lead enterprise-wide response to critical incidents across complex, large-scale, and globally distributed infrastructure environments.
This role operates at the intersection of technology leadership, crisis management, and business continuity, requiring the ability to make high-stakes decisions, influence senior stakeholders, and drive rapid resolution during mission-critical outages. The individual will serve as the ultimate authority during major incidents, ensuring minimal business disruption and long-term resilience.
RequirementsStrategic Responsibilities
- Own and lead enterprise-level incident management strategy across global operations.
- Act as the executive Incident Commander for P0/P1 incidents impacting business-critical systems.
- Establish and drive incident governance frameworks, SLAs, and response protocols
- Lead cross-functional crisis response involving Network, Cloud, Infrastructure, Security, and Field Operations
- Influence and align with C-suite and senior leadership during high-impact incidents
- Drive business continuity and service resilience initiatives
- Command and orchestrate war rooms and global bridge calls with multiple stakeholders
- Serve as the highest escalation point for critical outages and service disruptions.
- Ensure rapid triage, containment, and resolution of incidents with minimal downtime
- Drive real-time decision-making under ambiguity and pressure
- Oversee post-incident reviews and enforce accountability across teams
- Deep expertise in enterprise networking and distributed systems:
- BGP, OSPF, EIGRP, TCP/IP, QoS
- WAN, SD-WAN, Data Center architectures (Spine-Leaf)
- BGP, OSPF, EIGRP, TCP/IP, QoS
- Strong understanding of:
- Load balancing, DNS, DHCP, Network Security
- Latency, packet loss, and performance optimization
- Load balancing, DNS, DHCP, Network Security
- Familiarity with cloud platforms and hybrid infrastructure environments
- Ability to engage in hands-on technical triage when required
- Lead Root Cause Analysis (RCA) at an organizational level
- Drive preventive engineering, automation, and process maturity
- Establish a culture of proactive monitoring and early detection
- Enhance incident response playbooks, runbooks, and training programs
- ITIL Expert / Advanced Incident Management certifications
- Exposure to Disaster Recovery (DR) & Business Continuity Planning (BCP)
- Experience with automation, observability platforms, and AI-driven monitoring
- Track record of driving transformation in incident management practices
- 5 to 7 years of experience in Network Engineering, SRE, NOC, or Cloud Operations
- Proven experience handling enterprise-scale, high-impact incidents globally
- Prior experience in large enterprises / telecom / hyperscalers / global tech organizations
- Strong leadership presence with the ability to influence without authority
- Experience working in 24x7, mission-critical environments
Benefits
Skills Required
- Enterprise-level incident management strategy and governance experience
- Experience serving as Incident Commander for P0/P1 or other enterprise-scale, high-impact incidents
- Ability to lead cross-functional crisis response across Network, Cloud, Infrastructure, Security, and Field Operations teams
- Strong executive communication and ability to influence senior stakeholders without authority
- Deep expertise in enterprise networking and distributed systems, including BGP, OSPF, EIGRP, TCP/IP, QoS, WAN, SD-WAN, and Spine-Leaf data center architectures
- Strong understanding of load balancing, DNS, DHCP, network security, latency, packet loss, and performance optimization
- Familiarity with cloud platforms and hybrid infrastructure environments
- Ability to perform hands-on technical triage when required
- Experience leading organizational-level root-cause analysis and preventive engineering initiatives
- Five to seven years of experience in Network Engineering, SRE, NOC, or Cloud Operations
- Proven experience handling enterprise-scale, high-impact incidents globally
- Experience working in 24x7, mission-critical environments
- ITIL Expert or advanced Incident Management certification
- Exposure to Disaster Recovery and Business Continuity Planning
- Experience with automation, observability platforms, and AI-driven monitoring
- Track record of transforming incident management practices
- Prior experience in large enterprises, telecom, hyperscalers, or global technology organizations
What We Do
PowerBridge is an India-based audio visual and IT infrastructure solutions and services provider. It helps customers modernize technology through project delivery, infrastructure upgrades, and telecom initiatives, working with industry-leading technology partners under time-bound service-level agreements. With headquarters in Bangalore and an office in Hyderabad, the company supports customers across India through cross-functional teams focused on reliable, hassle-free technology implementation.









