Incident & Escalation Manager

Posted 2 Days Ago
Be an Early Applicant
Santa Clara, CA, USA
In-Office
165K-200K Annually
Expert/Leader
Artificial Intelligence • Information Technology • Software
The Role
Build and lead Selector AI’s Incident & Escalation function by defining operating models, severity standards, workflows, escalation paths, customer communications, and incident metrics. Run critical incident response, develop Incident Commanders, drive post-incident corrective actions, and establish accountability across NOC, SRE, Engineering, Product, Support, and Account teams. Use automation and AI to improve detection, triage, investigation, and response while converting recurring issues into measurable product and engineering improvements.
Summary Generated by Built In

About Us

Selector is building an operational intelligence platform for digital infrastructure. Using an AI/ML-based analytics approach, the platform provides actionable, multi-dimensional insights to network, cloud, and application operators. It helps operations teams meet their KPIs through seamless collaboration, a search-driven conversational experience, and automated data engineering pipelines.

Our solutions are used by leading Telecom, Media, Health Care, Finance, Retail, Professional Sports, and Fortune 500 enterprise organizations around the world. Our novel approach and rapidly expanding footprint position us for continued growth as a category leader.

Title: Incident & Escalation Manager

Location: Santa Clara HQ preferred 

About the Role

Selector AI is building a dedicated Incident & Escalation function to strengthen how we manage critical incidents and customer escalations as we scale.

We are looking for an experienced leader who can define the vision, build the operating model, and drive execution. You will bring together existing practices across NOC, SRE, Solution Engineering, Engineering, Product, Support, and Account teams and establish a consistent, scalable approach to incident and escalation management.

This is not simply an incident coordination role. You will build the function—from vision and process through execution, training, metrics, and continuous improvement

What You Will Own

  • Define the model: Establish what is an incident vs. escalation, severity levels, entry/exit criteria, roles, decision rights, and escalation paths.
  • Build the process: Turn the model into practical workflows across NOC, SRE, Solution Engineering, Engineering, Product, Support, and Account teams.
  • Run the response: Own the pager and participate in the Incident Commander rotation, driving clear owners, actions, decisions, communication, and resolution.
  • Build the IC bench: Train, coach, and certify Incident Commanders so the organization does not depend on a few individuals.
  • Establish customer communication: Define communication cadence, executive updates, customer RCAs, and post-incident follow-through.
  • Drive accountability: Every escalation has an owner, next action, and date. Challenge stalled work and escalate when commitments are missed.
  • Close the loop: Ensure post-incident actions have owners and deadlines and are driven to completion.
  • Build the feedback loop: Identify recurring customer and operational issues and drive them into Engineering and Product priorities.
  • Measure the function: Establish meaningful metrics around response, resolution, escalations, RCA SLAs, repeat incidents, and corrective-action closure.
  • Evolve the model: Use automation and AI where appropriate to improve detection, triage, investigation, and incident response while maintaining clear human accountability.

What You Bring

  • 8–12+ years in incident management, escalation management, SRE operations, technical support escalation, or technical program management.
  • Experience building or transforming an incident/escalation program, not simply operating within one.
  • Strong ability to turn an ambiguous vision into process, ownership, tooling, training, and measurable outcomes.
  • Experience building and developing Incident Commanders.
  • Strong technical understanding of distributed systems, cloud, networking, observability, or enterprise platforms.
  • Calm under pressure, decisive with incomplete information, and willing to challenge unclear ownership or stalled execution.
  • Strong written and verbal communication, including executive updates and customer-facing RCAs.
  • Familiarity with PagerDuty, Jira, Slack, and incident management workflows.
  • AIOps, NOC, networking, telecom, or AI-assisted operations experience is a strong plus.
  • Willingness to participate in the Incident Commander rotation, including off-hours coverage while the function is being established.

What Success Looks Like

You will build a predictable, scalable Incident & Escalation capability where the right people engage quickly, ownership is clear, communication is consistent, resolution is driven, and lessons become measurable improvements.

This is an opportunity to build and lead an important operational capability at a growing AI company.

Compensation:
165K–200K base

Perks: discretionary PTO, health insurance, 401k, bonus potential, and more.

Skills Required

  • 8-12+ years in incident management, escalation management, SRE operations, technical support escalation, or technical program management
  • Experience building or transforming an incident or escalation program
  • Ability to convert ambiguous vision into process, ownership, tooling, training, and measurable outcomes
  • Experience building and developing Incident Commanders
  • Strong technical understanding of distributed systems, cloud, networking, observability, or enterprise platforms
  • Calm under pressure, decisive with incomplete information, and willing to challenge unclear ownership or stalled execution
  • Strong written and verbal communication, including executive updates and customer-facing RCAs
  • Familiarity with PagerDuty, Jira, Slack, and incident management workflows
  • AIOps, NOC, networking, telecom, or AI-assisted operations experience
  • Willingness to participate in Incident Commander rotation, including off-hours coverage
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
104 Employees
Year Founded: 2019

What We Do

Selector AI is the industry leading AIOps platform designed to provide instant, real-time actionable insights for managing multi-domain network and application infrastructures. By bringing together multiple sources of data into one easy to use platform, IT teams can troubleshoot network issues faster, avoid downtime, reduce MTTR and improve efficiency.

Similar Jobs

Flex Logo Flex

Production Manager

Hardware • Other • Appliances
In-Office
Milpitas, CA, USA
52479 Employees
100K-137K Annually

Braze Logo Braze

Consultant

Marketing Tech • Mobile • Software
Easy Apply
Hybrid
San Francisco, CA, USA
2000 Employees
102K-204K Annually

Pluralsight Logo Pluralsight

Senior Project Manager

Edtech • Information Technology • Software
Remote or Hybrid
USA
1000 Employees
99K-130K Annually

Superhuman Logo Superhuman

Senior Manager, BU Finance

Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
Hybrid
San Francisco, CA, USA
1500 Employees
171K-236K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account