Lead AI Infrastructure Operations Engineer

Posted 4 Hours Ago
Be an Early Applicant
Hiring Remotely in Office, Machaze, Manica, MOZ
Remote
130K-150K Annually
Senior level
Insurance
The Role
Lead infrastructure operations for production AI solutions, translating AI designs into cloud, networking, security, observability, capacity, and operational requirements. Coordinate AWS managed services, infrastructure changes, production readiness, monitoring, incident resolution, and root-cause analysis. Establish reusable infrastructure patterns, dashboards, runbooks, SLIs, SLOs, and governance controls while monitoring AI quality, reliability, cost, drift, and performance.
Summary Generated by Built In

Department:

Information Technology

Job Description:

The Lead AI Infrastructure Operations Engineer is responsible for enabling, operating, and continuously improving the infrastructure and operational capabilities required by production AI solutions.

As TMG expands its AI capabilities, AI CoE will develop large language model, retrieval-augmented generation, agent-based, machine-learning, and other AI-enabled use cases. This role will work within the Infrastructure team to help those teams define and coordinate their infrastructure, environment, connectivity, observability, security, capacity, and production-readiness needs.

The engineer in collaboration with other Infrastructure resources will focus on the AI application, platform, and operational layers. TMG’s managed services provider operates the underlying AWS cloud infrastructure. This role will translate AI solution requirements, coordinate infrastructure services and changes, monitor delivery, diagnose issues, and validate that environments meet reliability, security, performance, and operational expectations.

This is a hands-on role that works closely with AI engineering, application development, data, architecture, cybersecurity, infrastructure operations, and external managed-services partners.

Work Arrangement:

  • Employees who live within 30 miles of the TMG home office are expected to follow a hybrid or in-office schedule. The initial training period may require additional in‑office days.

Accountabilities:

Enable AI Infrastructure and Environments

  • Partner with AI engineering, application, data, security, architecture, and platform teams to understand the infrastructure and operational needs of new AI use cases.

  • Translate AI solution designs into requirements for environments, compute, storage, networking, connectivity, identity, security, observability, capacity, and supporting cloud services.

  • In Collaboration with other Infrastructure resources and managed service provider, coordinate infrastructure provisioning, configuration, access, and changes with TMG’s AWS managed-services provider and other technology partners.

  • Support operational readiness for AI solutions transitioning into production.

  • Identify infrastructure dependencies, constraints, risks, costs, and lead times early in the delivery lifecycle.

  • Establish reusable infrastructure patterns, operational standards, dashboards, runbooks, and production-readiness requirements across AI use cases.

  • Validate that environments and supporting services are appropriately configured, monitored, secured, scalable, and ready for production use.

  • Implement new cloud functionality requirements in collaboration with architecture and AI engineering within approved architecture and guardrails.

  • Follow and enforce established security, compliance, and operational controls across AI platforms and infrastructure.

Operate and Improve Production AI Solutions

  • Work with managed services to Implement and maintain observability, logging, tracing, monitoring, dashboards, and alerting for production AI applications.

  • Define and track operational metrics covering AI quality, reliability, latency, cost, utilization, adoption, and business outcomes.

  • Work with AI engineering team to diagnose production issues across AI applications, models, prompts, retrieval systems, data, integrations, and supporting infrastructure.

  • Support root-cause analysis and coordinate resolution with AI engineers, application teams, platform teams, and managed-services providers.

  • Support AI-related incident management, problem management, operational reviews, release validation, and production-readiness activities.

  • Create and maintain service-health dashboards, runbooks, troubleshooting guidance, support procedures, and escalation paths.

  • Define and track service health indicators, SLIs, SLOs, and operational KPIs for production AI solutions.

  • Use telemetry, evaluations, and production data to validate fixes, releases, configuration changes, and system improvements.

Support AI Risk and Operational Governance

  • Partner with AI and IT governance team to operationalize applicable controls and monitoring requirements.

  • Operationalize evaluation processes and support AI Governance team to monitor AI performance, regressions, drift, grounding, retrieval quality, and overall effectiveness.

  • Support the collection and retention of operational evidence, including model and prompt versions, evaluation results, incidents, exceptions, and corrective actions.

  • Identify material changes in AI behavior and help ensure they are evaluated, documented, and appropriately addressed.

  • Analyze trends and proactively identify degradation, drift, capacity constraints, reliability risks, and quality issues before they become production incidents

Key Outcomes

  • AI teams receive timely and consistent infrastructure and operational support.

  • AI use cases move efficiently from experimentation to reliable production operation.

  • Reusable infrastructure and operational patterns are applied across AI initiatives.

  • End-to-end visibility exists across AI applications and supporting services.

  • AI quality, reliability, cost, usage, and business impact are consistently measured.

  • Production issues are detected, diagnosed, and resolved more quickly.

  • Releases result in fewer regressions and operational disruptions.

  • Secure, compliant, and audit-ready AI platforms meeting enterprise governance and risk requirements

Qualifications

  • Bachelor’s degree in computer science, engineering, information technology, or a related field, or equivalent practical experience.

  • 8+ years of overall information technology experience, including 5+ years in infrastructure operations, cloud operations, site reliability engineering, DevOps, platform operations, application operations, or production engineering.

  • Experience supporting cloud-hosted, distributed, data-intensive, or AI-enabled production applications.

  • Experience implementing and supporting observability, logging, tracing, monitoring, dashboards, alerting, and operational reporting.

  • Working knowledge of cloud infrastructure, networking, APIs, integrations, identity and access management, security, and data pipelines.

  • Experience coordinating infrastructure services, changes, dependencies, and issue resolution across internal teams, technology partners, and managed-services providers.

  • Familiarity with AWS services and capabilities related to monitoring, logging, networking, security, identity, and infrastructure operations.

  • Experience with automation, CI/CD pipelines, infrastructure-as-code, configuration management, and release-management practices.

  • Experience operating or supporting LLM, generative AI, machine-learning, RAG, or agent-based applications in production is strongly preferred.

  • Familiarity with AI observability, evaluation, model monitoring, hallucination detection, grounding, retrieval quality, model or data drift, and prompt-related risks.

  • Familiarity with vector databases, model APIs, AI gateways, prompt-management platforms, or agent orchestration frameworks.

  • Strong analytical, documentation, communication, and cross-functional collaboration skills, with the ability to manage multiple priorities across concurrent technology initiatives.

  • Experience in insurance, financial services, or another regulated industry, along with relevant AWS, cloud, infrastructure, DevOps, SRE, security, or AI certifications, is preferred.

Pay Range:

Anticipated Hiring Range:

  • $130,000 - $150,000 annual base salary depending on experience, qualifications, and geographic location 

Benefits:

We are proud to offer our full-time regular employees a robust benefits suite that includes:

  • Competitive base salary plus incentive plans for eligible team members

  • 401(K) retirement plan that includes a company match of up to 6% of your eligible salary

  • Free basic life and AD&D, long-term disability and short-term disability insurance

  • Medical, dental and vision plans to meet your unique healthcare needs

  • Wellness incentives

  • Generous time off program that includes personal, holiday and volunteer paid time off

  • Flexible work schedules and hybrid/remote options for eligible positions

  • Educational assistance

Equal Opportunity Employer

The Mutual Group is an Equal Opportunity Employer. It is our policy to recruit, hire, train and promote individuals in all job classifications without regard to race, color, religion, sex, national origin, age, veteran status, disability, sexual orientation, gender identity or any other characteristic protected by law.

  • Know Your Rights: Workplace Discrimination is Illegal

  • Your Rights Under USERRA

Applicants requiring a reasonable accommodation due to a disability at any stage of the employment application process should contact [email protected].

Employment Verification

The Mutual Group participates in the E-Verify program and will provide the federal government with your Form I-9 information to confirm that you are authorized to work in the U.S. You are protected from employment discrimination based on your citizenship status and national origin.

E-Verify Program Overview

 

E-Verify Participation Poster


All offers of employment are contingent upon the successful completion of a background check.

#TMG

Skills Required

  • Bachelor's degree in computer science, engineering, information technology, or a related field, or equivalent practical experience
  • 8+ years of overall information technology experience, including 5+ years in infrastructure operations, cloud operations, site reliability engineering, DevOps, platform operations, application operations, or production engineering
  • Experience supporting cloud-hosted, distributed, data-intensive, or AI-enabled production applications
  • Experience implementing and supporting observability, logging, tracing, monitoring, dashboards, alerting, and operational reporting
  • Working knowledge of cloud infrastructure, networking, APIs, integrations, identity and access management, security, and data pipelines
  • Experience coordinating infrastructure services, changes, dependencies, and issue resolution across internal teams, technology partners, and managed-services providers
  • Familiarity with AWS services and capabilities related to monitoring, logging, networking, security, identity, and infrastructure operations
  • Experience with automation, CI/CD pipelines, infrastructure-as-code, configuration management, and release-management practices
  • Experience operating or supporting LLM, generative AI, machine-learning, RAG, or agent-based applications in production
  • Familiarity with AI observability, evaluation, model monitoring, hallucination detection, grounding, retrieval quality, model or data drift, and prompt-related risks
  • Familiarity with vector databases, model APIs, AI gateways, prompt-management platforms, or agent orchestration frameworks
  • Strong analytical, documentation, communication, and cross-functional collaboration skills, with the ability to manage multiple priorities across concurrent technology initiatives
  • Experience in insurance, financial services, or another regulated industry
  • Relevant AWS, cloud, infrastructure, DevOps, SRE, security, or AI certifications
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
400 Employees
Year Founded: 2024

What We Do

The Mutual Group is an insurance services platform and strategic partner for independent mutual insurance carriers. It provides permanent capital, shared infrastructure, technology, and specialized expertise across underwriting, claims, corporate services, and technology. Through a membership-based model, the company helps mutuals improve operational efficiency, scale, strengthen competitiveness, and innovate while preserving their independence, identity, brands, balance sheets, and policyholder ownership.

Similar Jobs

Invenergy Logo Invenergy

Senior PI Administrator

Greentech • Real Estate • Social Impact • Energy • Industrial • Solar • Renewable Energy
Remote or Hybrid
18 Locations
2500 Employees
125K-155K Annually

Auror Logo Auror

Senior Engineering Lead (Risk Detection)

Artificial Intelligence • Big Data • Retail • Security • Social Impact • Software • Business Intelligence
Remote or Hybrid
Office, Machaze, Manica, MOZ
212 Employees
143K-190K Annually

Compa Logo Compa

Software Engineer

Artificial Intelligence • HR Tech • Software • Business Intelligence
Remote or Hybrid
4 Locations
75 Employees
125K-180K Annually

CrowdStrike Logo CrowdStrike

Growth Development Representative (Hybrid)

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
Office, Machaze, Manica, MOZ
11000 Employees

Similar Companies Hiring

Globe Life Thumbnail
Insurance • Financial Services
McKinney, TX
3000 Employees
MassMutual India Thumbnail
Big Data • Fintech • Information Technology • Insurance • Financial Services
Hyderabad, Telangana
Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account