Recovery and Incident Manager with AI-Ops

Posted Yesterday
Be an Early Applicant
8 Locations
In-Office
73K-171K Annually
Senior level
Information Technology
The Role
Lead incident management, operational resilience, and AIOps-driven automation across enterprise environments. Drive SRE practices, observability (Dynatrace), ServiceNow workflow automation, root-cause analysis, SLAs/SLOs, and build AI-powered self-healing workflows to reduce MTTD/MTTR and improve reliability.
Summary Generated by Built In

 

We currently have a career opportunity for a Recovery and Incident Manager with AI-Ops to join our team in Charlotte, NC or Tempe, AZ or NYC, NY. 

 

Job Overview:

 

We are seeking a Recovery and Incident Manager with AI-Ops to lead incident management, operational resilience, and intelligent automation initiatives across enterprise technology environments. This role will partner with Network Operations Center (NOC), Infrastructure Operations, Cloud Engineering, DevOps, and Application Support teams to proactively detect, respond to, and prevent technology incidents. The ideal candidate combines hands-on incident management expertise with experience implementing observability, automation, and AI-driven operational solutions to improve system reliability, reduce operational overhead, and enhance customer experience. The candidate should possess deep expertise in AIOps, ITSM, ITIL, SRE, Incident Management, Cloud Operations, and Enterprise Infrastructure. 

 Perficient is always looking for the best and brightest talent and we need you! We’re a quickly-growing, global digital consulting leader, and we’re transforming the world’s largest enterprises and biggest brands. You’ll work with the latest technologies, expand your skills, and become a part of our global community of talented, diverse, and knowledgeable colleagues.

 

Responsibilities

Incident & Recovery Management

  • Monitor, document, and analyze major incident response efforts and service recovery activities.
  • Serve as a senior escalation point for Tier 1 and Tier 2 operational incidents.
  • Conduct incident reviews, root cause analysis, and corrective action planning.
  • Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).

Site Reliability Engineering

  • Implement SRE practices to improve platform reliability, scalability, and resiliency.
  • Define and monitor SLAs, SLOs, and operational KPIs.
  • Develop proactive reliability and availability strategies.

AIOps & Automation

  • Implement AIOps solutions to automate incident detection, diagnosis, remediation, and prevention.
  • Build and optimize AI-powered operational agents and self-healing workflows.
  • Reduce operational effort through intelligent automation.

Observability & Monitoring

  • Lead enterprise monitoring initiatives using Dynatrace and related observability platforms.
  • Improve visibility across cloud, infrastructure, applications, and user experiences.
  • Enable predictive monitoring and anomaly detection.

ITSM & Service Operations

  • Develop and enhance incident, problem, change, and event management frameworks aligned with ITIL and ITSM best practices.  
  • Leverage ServiceNow workflow automation to improve service delivery.

Cross-Functional Leadership

  • Partner with Infrastructure, DevOps, Cloud, Security, Application Development, and NOC teams.
  • Mentor operational teams and promote an automation-first culture. 

 

Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field (or equivalent experience).
  • 8+ years of experience in IT Operations, Site Reliability Engineering, Infrastructure Operations, Network Operations, or Production Support environments.
  • 5+ years of experience leading incident management, operational transformation, or reliability engineering initiatives.
  • Strong experience with:  
    • Site Reliability Engineering (SRE)
    • IT Service Management (ITSM)
    • ITIL Framework
    • Incident, Problem, Change, and Event Management
    • Network Operations Center (NOC)
    • Infrastructure Operations
    • Service Desk Operations
    • Application Production Support
    • Cloud Platforms (AWS, Azure, or GCP)
    • DevOps Practices and Toolchains
  • Hands-on experience with Dynatrace, monitoring platforms, and observability solutions.
  • Experience using ServiceNow for ticketing, workflow automation, and service management.
  • Strong understanding of infrastructure, networking, cloud architecture, and enterprise application ecosystems.
  • Proven experience conducting root cause analysis and implementing preventive controls.
  • Experience leading enterprise AIOps implementations.
  • Experience building AI-powered operational agents and intelligent automation solutions.
  • Certifications such as:  
    • ITIL Foundation or ITIL Managing Professional
    • Certified Site Reliability Engineer (SRE)
    • AWS, Azure, or Google Cloud certifications
    • ServiceNow certifications
  • Experience with workflow orchestration and enterprise automation platforms.
  • Familiarity with predictive analytics, machine learning operations, and autonomous operations frameworks. 

ABOUT THE TEAM

Our Automation team empowers organizations to work smarter by connecting digital process automation (DPA), robotic process automation (RPA), and AI into seamless, intelligent workflows. We help leading brands streamline operations, enhance efficiency, and unlock new business potential. By embedding advanced AI models into automation strategies, we enable smarter decision-making, adaptive processes, and continuous optimization at scale. 

ADDITIONAL INFORMATION

Perficient, Inc. proudly provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, gender, sexual orientation, national origin, age, disability, genetic information, marital status, amnesty, or status as a protected veteran in accordance with applicable federal, state and local laws. Perficient, Inc. complies with applicable state and local laws governing nondiscrimination in employment in every location in which the company has facilities. This policy applies to all terms and conditions of employment, including, but not limited to, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training. Perficient, Inc. expressly prohibits any form of unlawful employee harassment based on race, color, religion, gender, sexual orientation, national origin, age, genetic information, disability, or covered veterans. Improper interference with the ability of Perficient, Inc. employees to perform their expected job duties is absolutely not tolerated.

 Disability Accommodations: Perficient is committed to providing a barrier-free employment process with reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or accommodation due to a disability, please contact us.

 Applications will be accepted until the position is filled or the posting is removed.

 The salary range for this position takes into consideration a variety of factors, including but not limited to skill sets, level of experience, applicable office location, training, licensure and certifications, and other business and organizational needs. The new hire salary range displays the minimum and maximum salary targets for this position across all US locations, and the range has not been adjusted for any specific state differentials. It is not typical for a candidate to be hired at or near the top of the range for their role, and compensation decisions are dependent on the unique facts and circumstances regarding each candidate. A reasonable estimate of the current salary range for this position is $ 73,008 to $ 170,640. Please note that the salary range posted reflects the base salary only and does not include benefits or any potential variable compensation programs. Information regarding the benefits available for this position are in our benefits overview.


#LI-MG1#

 

 

 

About UsPerficient is the global AI and technology consulting firm disrupting the traditional consulting model. Powered by our 7,000+ advisors, engineers, and designers, Perficient implements AI-first solutions that break conventions and deliver outcomes that matter. Proudly serving clients that represent the world’s most innovative brands, and in collaboration with our powerful technology partner ecosystem, we bring deep industry expertise and data-driven design to redefine how businesses run and succeed. Perficient is different. For real. Learn more at perficient.com.

Skills Required

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field (or equivalent experience)
  • 8+ years of experience in IT Operations, Site Reliability Engineering, Infrastructure Operations, Network Operations, or Production Support
  • 5+ years of experience leading incident management, operational transformation, or reliability engineering initiatives
  • Site Reliability Engineering (SRE) experience
  • IT Service Management (ITSM) experience
  • Familiarity with ITIL Framework and ITIL-based processes
  • Incident, Problem, Change, and Event Management experience
  • Network Operations Center (NOC) and Infrastructure Operations experience
  • Service Desk Operations and Application Production Support experience
  • Experience with cloud platforms (AWS, Azure, or GCP)
  • DevOps practices and toolchain experience
  • Hands-on experience with Dynatrace and observability/monitoring platforms
  • Experience using ServiceNow for ticketing, workflow automation, and service management
  • Proven experience conducting root cause analysis and implementing preventive controls
  • Experience leading enterprise AIOps implementations
  • Experience building AI-powered operational agents and intelligent automation solutions
  • Experience with workflow orchestration and enterprise automation platforms
  • Familiarity with predictive analytics, machine learning operations (MLOps), and autonomous operations frameworks
  • Certifications such as ITIL Foundation or ITIL Managing Professional, Certified Site Reliability Engineer (SRE), AWS/Azure/GCP certifications, or ServiceNow certifications

Perficient Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Perficient and has not been reviewed or approved by Perficient.

  • Retirement Support Retirement offerings are positioned as robust, including a 401(k) with company match, an Employee Stock Purchase Plan, and an option for after-tax contributions (mega backdoor Roth). Eligibility details are described as clear for core benefits, supporting confidence in plan access timing.
  • Parental & Family Support Parental benefits are described as structured, with paid maternity recovery time and paid parental leave for all new parents. Company-paid disability coverage is also highlighted, strengthening the overall family support posture.
  • Fair & Transparent Compensation Compensation is characterized as generally market-aligned for a portion of roles, with examples of pay being viewed as fair or decent in certain contexts (such as remote or region-specific situations). Variable pay potential appears stronger in some tracks, improving perceived competitiveness for those roles.

Perficient Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Saint Louis, MO
3,295 Employees
Year Founded: 1997

What We Do

Perficient is a leading global digital consultancy. We imagine, create, engineer, and run digital transformation solutions that help our clients exceed customers’ expectations, outpace competition, and grow their business. With unparalleled strategy, creative, and technology capabilities, we bring big thinking and innovative ideas, along with a practical approach to help the world’s largest enterprises and biggest brands succeed.

Similar Jobs

Narmi Logo Narmi

Sales Engineer

Enterprise Web • Fintech • Payments • Software • Financial Services
Remote or Hybrid
2 Locations
140 Employees
80K-95K Annually
Hybrid
2 Locations
289097 Employees
Hybrid
3 Locations
289097 Employees
Hybrid
2 Locations
289097 Employees

Similar Companies Hiring

Scrunch  Thumbnail
Artificial Intelligence • Information Technology • Marketing Tech • Software • SEO
Salt Lake City, Utah
Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account