Department:
Information TechnologyJob Description:
The Lead AI Infrastructure Operations Engineer is responsible for enabling, operating, and continuously improving the infrastructure and operational capabilities required by production AI solutions.
As TMG expands its AI capabilities, AI CoE will develop large language model, retrieval-augmented generation, agent-based, machine-learning, and other AI-enabled use cases. This role will work within the Infrastructure team to help those teams define and coordinate their infrastructure, environment, connectivity, observability, security, capacity, and production-readiness needs.
The engineer in collaboration with other Infrastructure resources will focus on the AI application, platform, and operational layers. TMG’s managed services provider operates the underlying AWS cloud infrastructure. This role will translate AI solution requirements, coordinate infrastructure services and changes, monitor delivery, diagnose issues, and validate that environments meet reliability, security, performance, and operational expectations.
This is a hands-on role that works closely with AI engineering, application development, data, architecture, cybersecurity, infrastructure operations, and external managed-services partners.
Work Arrangement:
Employees who live within 30 miles of the TMG home office are expected to follow a hybrid or in-office schedule. The initial training period may require additional in‑office days.
Accountabilities:
Enable AI Infrastructure and Environments
Partner with AI engineering, application, data, security, architecture, and platform teams to understand the infrastructure and operational needs of new AI use cases.
Translate AI solution designs into requirements for environments, compute, storage, networking, connectivity, identity, security, observability, capacity, and supporting cloud services.
In Collaboration with other Infrastructure resources and managed service provider, coordinate infrastructure provisioning, configuration, access, and changes with TMG’s AWS managed-services provider and other technology partners.
Support operational readiness for AI solutions transitioning into production.
Identify infrastructure dependencies, constraints, risks, costs, and lead times early in the delivery lifecycle.
Establish reusable infrastructure patterns, operational standards, dashboards, runbooks, and production-readiness requirements across AI use cases.
Validate that environments and supporting services are appropriately configured, monitored, secured, scalable, and ready for production use.
Implement new cloud functionality requirements in collaboration with architecture and AI engineering within approved architecture and guardrails.
Follow and enforce established security, compliance, and operational controls across AI platforms and infrastructure.
Operate and Improve Production AI Solutions
Work with managed services to Implement and maintain observability, logging, tracing, monitoring, dashboards, and alerting for production AI applications.
Define and track operational metrics covering AI quality, reliability, latency, cost, utilization, adoption, and business outcomes.
Work with AI engineering team to diagnose production issues across AI applications, models, prompts, retrieval systems, data, integrations, and supporting infrastructure.
Support root-cause analysis and coordinate resolution with AI engineers, application teams, platform teams, and managed-services providers.
Support AI-related incident management, problem management, operational reviews, release validation, and production-readiness activities.
Create and maintain service-health dashboards, runbooks, troubleshooting guidance, support procedures, and escalation paths.
Define and track service health indicators, SLIs, SLOs, and operational KPIs for production AI solutions.
Use telemetry, evaluations, and production data to validate fixes, releases, configuration changes, and system improvements.
Support AI Risk and Operational Governance
Partner with AI and IT governance team to operationalize applicable controls and monitoring requirements.
Operationalize evaluation processes and support AI Governance team to monitor AI performance, regressions, drift, grounding, retrieval quality, and overall effectiveness.
Support the collection and retention of operational evidence, including model and prompt versions, evaluation results, incidents, exceptions, and corrective actions.
Identify material changes in AI behavior and help ensure they are evaluated, documented, and appropriately addressed.
Analyze trends and proactively identify degradation, drift, capacity constraints, reliability risks, and quality issues before they become production incidents
Key Outcomes
AI teams receive timely and consistent infrastructure and operational support.
AI use cases move efficiently from experimentation to reliable production operation.
Reusable infrastructure and operational patterns are applied across AI initiatives.
End-to-end visibility exists across AI applications and supporting services.
AI quality, reliability, cost, usage, and business impact are consistently measured.
Production issues are detected, diagnosed, and resolved more quickly.
Releases result in fewer regressions and operational disruptions.
Secure, compliant, and audit-ready AI platforms meeting enterprise governance and risk requirements
Qualifications
Bachelor’s degree in computer science, engineering, information technology, or a related field, or equivalent practical experience.
8+ years of overall information technology experience, including 5+ years in infrastructure operations, cloud operations, site reliability engineering, DevOps, platform operations, application operations, or production engineering.
Experience supporting cloud-hosted, distributed, data-intensive, or AI-enabled production applications.
Experience implementing and supporting observability, logging, tracing, monitoring, dashboards, alerting, and operational reporting.
Working knowledge of cloud infrastructure, networking, APIs, integrations, identity and access management, security, and data pipelines.
Experience coordinating infrastructure services, changes, dependencies, and issue resolution across internal teams, technology partners, and managed-services providers.
Familiarity with AWS services and capabilities related to monitoring, logging, networking, security, identity, and infrastructure operations.
Experience with automation, CI/CD pipelines, infrastructure-as-code, configuration management, and release-management practices.
Experience operating or supporting LLM, generative AI, machine-learning, RAG, or agent-based applications in production is strongly preferred.
Familiarity with AI observability, evaluation, model monitoring, hallucination detection, grounding, retrieval quality, model or data drift, and prompt-related risks.
Familiarity with vector databases, model APIs, AI gateways, prompt-management platforms, or agent orchestration frameworks.
Strong analytical, documentation, communication, and cross-functional collaboration skills, with the ability to manage multiple priorities across concurrent technology initiatives.
Experience in insurance, financial services, or another regulated industry, along with relevant AWS, cloud, infrastructure, DevOps, SRE, security, or AI certifications, is preferred.
Pay Range:
Anticipated Hiring Range:
$130,000 - $150,000 annual base salary depending on experience, qualifications, and geographic location
Benefits:
We are proud to offer our full-time regular employees a robust benefits suite that includes:
Competitive base salary plus incentive plans for eligible team members
401(K) retirement plan that includes a company match of up to 6% of your eligible salary
Free basic life and AD&D, long-term disability and short-term disability insurance
Medical, dental and vision plans to meet your unique healthcare needs
Wellness incentives
Generous time off program that includes personal, holiday and volunteer paid time off
Flexible work schedules and hybrid/remote options for eligible positions
Educational assistance
Equal Opportunity Employer
The Mutual Group is an Equal Opportunity Employer. It is our policy to recruit, hire, train and promote individuals in all job classifications without regard to race, color, religion, sex, national origin, age, veteran status, disability, sexual orientation, gender identity or any other characteristic protected by law.
Know Your Rights: Workplace Discrimination is Illegal
Your Rights Under USERRA
Applicants requiring a reasonable accommodation due to a disability at any stage of the employment application process should contact [email protected].
Employment Verification
The Mutual Group participates in the E-Verify program and will provide the federal government with your Form I-9 information to confirm that you are authorized to work in the U.S. You are protected from employment discrimination based on your citizenship status and national origin.
E-Verify Program Overview
E-Verify Participation Poster
All offers of employment are contingent upon the successful completion of a background check.
#TMG
Skills Required
- Bachelor's degree in computer science, engineering, information technology, or a related field, or equivalent practical experience
- 8+ years of overall information technology experience, including 5+ years in infrastructure operations, cloud operations, site reliability engineering, DevOps, platform operations, application operations, or production engineering
- Experience supporting cloud-hosted, distributed, data-intensive, or AI-enabled production applications
- Experience implementing and supporting observability, logging, tracing, monitoring, dashboards, alerting, and operational reporting
- Working knowledge of cloud infrastructure, networking, APIs, integrations, identity and access management, security, and data pipelines
- Experience coordinating infrastructure services, changes, dependencies, and issue resolution across internal teams, technology partners, and managed-services providers
- Familiarity with AWS services and capabilities related to monitoring, logging, networking, security, identity, and infrastructure operations
- Experience with automation, CI/CD pipelines, infrastructure-as-code, configuration management, and release-management practices
- Experience operating or supporting LLM, generative AI, machine-learning, RAG, or agent-based applications in production
- Familiarity with AI observability, evaluation, model monitoring, hallucination detection, grounding, retrieval quality, model or data drift, and prompt-related risks
- Familiarity with vector databases, model APIs, AI gateways, prompt-management platforms, or agent orchestration frameworks
- Strong analytical, documentation, communication, and cross-functional collaboration skills, with the ability to manage multiple priorities across concurrent technology initiatives
- Experience in insurance, financial services, or another regulated industry
- Relevant AWS, cloud, infrastructure, DevOps, SRE, security, or AI certifications
What We Do
The Mutual Group is an insurance services platform and strategic partner for independent mutual insurance carriers. It provides permanent capital, shared infrastructure, technology, and specialized expertise across underwriting, claims, corporate services, and technology. Through a membership-based model, the company helps mutuals improve operational efficiency, scale, strengthen competitiveness, and innovate while preserving their independence, identity, brands, balance sheets, and policyholder ownership.








