Production Support Lead (Incident management/Problem management)

Posted 9 Days Ago
Be an Early Applicant
Toronto, ON, CAN
In-Office
90K-93K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Consulting
The Role
Leads critical production incident response from triage through restoration and closure, coordinating technology, security, product, business, and vendor teams. Owns SWAT queue health, ticket prioritization, SLA performance, stakeholder communications, operational reporting, and root-cause reviews. Facilitates Scrum ceremonies, coaches teams on Agile practices, removes impediments, and drives corrective actions and continuous improvements across incident management, support workflows, monitoring, resilience, and service governance.
Summary Generated by Built In

About the Job: We are seeking an experienced Production Support Lead with Scrum Master capabilities to lead the response, coordination, and governance of production incidents across cross-functional technology teams. The successful candidate will own critical incident execution, SWAT queue health, stakeholder communication, and service restoration while applying Agile practices to improve team flow, accountability, and continuous improvement.

Office Location: Toronto

Employment Type: Permanent

Role Type: New position - current requirement

Work Arrangement: Hybrid (2 days in office per week)

Position Responsibilities:

Incident Leadership & Response Management

· Lead the end-to-end management of critical production incidents from initial triage through service restoration, stakeholder communication, root-cause review, and closure.

· Establish incident command, confirm severity and business impact, assign clear ownership, and coordinate application, engineering, infrastructure, security, product, and vendor teams.

· Drive timely resolution of critical tickets within agreed SLAs and escalate risks, blockers, and resource constraints appropriately.

· Maintain accurate incident timelines, decisions, actions, dependencies, and recovery updates throughout the incident lifecycle.

· Remove production support bottlenecks and enable rapid decision-making during high-priority incidents.


Ticket Triage & SWAT Queue Management

· Own daily ticket triage and the SWAT queue, ensuring incidents and support tickets are correctly categorized, prioritized, assigned, and progressed.

· Monitor ticket ageing, stalled work, recurring issues, capacity constraints, and ownership gaps to maintain a manageable backlog.

· Balance urgent restoration work with defects, service requests, technical debt, and preventive improvement initiatives.

· Improve ticket throughput and backlog hygiene while maintaining quality, compliance, and operational controls.


Scrum Master & Agile Delivery Responsibilities

· Facilitate daily SWAT stand-ups, sprint planning, backlog refinement, retrospectives, service reviews, and operational governance meetings.

· Coach support and engineering teams on Scrum and Agile practices suited to production support and interrupt-driven work.

· Partner with product owners and service owners to maintain a prioritized, transparent backlog with clear acceptance criteria and ownership.

· Identify and remove team impediments, manage dependencies, support capacity planning, and improve delivery flow across teams.

· Use retrospectives and operational data to implement measurable improvements in incident response and support delivery.


Operational Metrics, Reporting & Governance

· Track and report SLA compliance, mean time to acknowledge, mean time to resolution (MTTR), ticket ageing, throughput, backlog health, critical incident volume, and recurrence trends.

· Prepare dashboards and scorecards that provide leadership with clear visibility into service performance, operational risks, bottlenecks, and improvement actions.

· Facilitate incident and operational governance reviews, ensuring decisions, escalations, risks, and action items are documented and closed on time.

· Promote cross-team accountability through clear owners, target dates, escalation paths, and transparent follow-through.


Problem Management & Operational Excellence

· Lead post-incident reviews and root-cause analysis for major and recurring incidents without creating a blame-focused environment.

· Ensure corrective and preventive actions are prioritized, tracked, and implemented to reduce recurring incidents.

· Identify trends and systemic weaknesses, then partner with technology teams to improve resilience, monitoring, automation, and support readiness.

· Drive continuous improvement in incident processes, escalation models, runbooks, communications, and service management practices.



Requirements

Required Qualifications:

· 8+ years of experience leading production support and incident management teams, including coordinating the triage, prioritization, and resolution of software incidents in an enterprise technology environment.

· Demonstrated Scrum Master experience, including facilitation of Agile ceremonies, backlog governance, impediment removal, coaching, and continuous improvement.

· Proven ability to coordinate high-severity incidents across application, engineering, infrastructure, security, product, business, and vendor teams.

· Hands-on experience with ticket triage, incident queues, escalation management, root-cause analysis, and corrective-action tracking.

· Working knowledge of SLA, MTTR, ticket ageing, throughput, backlog health, and other production support metrics.

· Strong stakeholder communication, facilitation, decision-making, conflict-resolution, and executive reporting skills.

· Ability to remain composed, establish accountability, and drive outcomes in high-pressure and time-sensitive situations.

· Experience managing cross-functional and geographically distributed teams.

Preferred Qualifications:

· Experience supporting enterprise applications, microservices, integrations, and cloud environments such as AWS, Microsoft Azure, or Google Cloud Platform.

· Familiarity with ITIL practices, DevOps, CI/CD pipelines, observability, monitoring, and modern production support workflows.

· Experience building operational dashboards and scorecards using data from service management and delivery platforms.

· Proficiency with tools such as ServiceNow, Confluence, or similar incident and collaboration platforms.

· Preferred certifications include ITIL, Certified Scrum Master (CSM), Professional Scrum Master (PSM), SAFe Scrum Master, PMP, or PRINCE2.



Benefits

Salary Range: $90,000 to $93,000 CAD/ year

 

The final compensation offered will depend on local market conditions and geographic location, as well as job-related factors such as the candidate’s knowledge, skills, qualifications, relevant experience, and education/training. Compensation may also include additional components such as benefits, and/or other incentives, where applicable. In accordance with new employment standards requirements, we retain copies of this job posting and applicant information for three (3) years after the posting is removed. We do not use AI technology; all applications are also reviewed by our recruitment team.

Infoya is an equal opportunity employer committed to diversity and inclusion. We welcome applications from all qualified individuals, regardless of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, protected veteran status, aboriginal status, or any other legally protected factors.



Skills Required

  • 8+ years leading production support and incident management teams in an enterprise technology environment
  • Experience coordinating triage, prioritization, and resolution of software incidents
  • Demonstrated Scrum Master experience, including Agile ceremonies, backlog governance, impediment removal, coaching, and continuous improvement
  • Experience coordinating high-severity incidents across application, engineering, infrastructure, security, product, business, and vendor teams
  • Hands-on experience with ticket triage, incident queues, escalation management, root-cause analysis, and corrective-action tracking
  • Working knowledge of SLA, MTTR, ticket ageing, throughput, backlog health, and production support metrics
  • Strong stakeholder communication, facilitation, decision-making, conflict-resolution, and executive reporting skills
  • Ability to remain composed, establish accountability, and drive outcomes in high-pressure situations
  • Experience managing cross-functional and geographically distributed teams
  • Experience supporting enterprise applications, microservices, integrations, and cloud environments such as AWS, Microsoft Azure, or Google Cloud Platform
  • Familiarity with ITIL practices, DevOps, CI/CD pipelines, observability, monitoring, and modern production support workflows
  • Experience building operational dashboards and scorecards using service management and delivery platform data
  • Proficiency with ServiceNow, Confluence, or similar incident and collaboration platforms
  • ITIL, Certified Scrum Master, Professional Scrum Master, SAFe Scrum Master, PMP, or PRINCE2 certification
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Oakville
Year Founded: 2017

What We Do

Infoya is a global IT solutions and consulting firm specializing in business transformation, digital innovation, and advanced engineering services, including AI and cloud enablement.

Similar Jobs

NBCUniversal Logo NBCUniversal

Executive Assistant

AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Remote or Hybrid
Toronto, ON, CAN
55K-70K Annually

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Sales Support Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Ottawa, ON, CAN
16000 Employees
18-22 Hourly

Block Logo Block

Senior Analytics Engineer

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
139K-245K Annually

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Sales Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Ottawa, ON, CAN
16000 Employees
18-22 Hourly

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account