Infrastructure Engineer Lead – Cloud AI

Sorry, this job was removed at 04:47 a.m. (UTC) on Saturday, Oct 03, 2026
Be an Early Applicant
Gahanna, OH, USA
In-Office
137K-178K Annually
Expert/Leader
Energy • Renewable Energy
The Role
Leads the design, deployment, and operation of secure, scalable AWS and on-premises infrastructure for AI and machine learning workloads. Builds cloud foundations, Kubernetes platforms, networking, identity, observability, automation, and FinOps practices. Supports disconnected and edge environments, security governance, incident response, cost optimization, and platform lifecycle management. Provides technical direction, mentoring, documentation, and recommendations while collaborating with cybersecurity, architecture, networking, application, project, and vendor teams.
Summary Generated by Built In

Job Posting End Date

10-03-2026

Please note the job posting will close on the day before the posting end date.

Job Summary

This hands-on role builds and operates secure, scalable infrastructure for AI and machine learning workloads across AWS and approved on-premises environments. Responsibilities include cloud foundations, networking, containers, compute, identity, cost management, and observability.
The engineer partners with AI, architecture, cybersecurity, network, and application teams to deliver secure, cost-effective, production-ready infrastructure while supporting core cloud engineering services.

Job Description

What you’ll do:


Essential Job Functions & Tasks


AI Cloud Foundations

  • Design, build and maintain the AWS account structure, landing zones, and reference patterns that host AI and machine learning workloads, leveraging AWS Control Tower, Landing Zone Accelerator, and Terraform.
  • Engineer the compute, storage, and networking foundations required for AI workloads, including GPU and accelerated instance families, high-throughput storage, and data access paths to enterprise data platforms.
  • Enable and operate AI platform services such as Amazon Bedrock and SageMaker, including private connectivity, model access provisioning, logging, and quota management.
  • Build reusable infrastructure-as-code modules, templates and pipelines so that AI teams can deploy quickly within approved guardrails.
  • Design, build and maintain on-premises AI infrastructure for edge and specialized use cases.

Support AI in disconnected or intermittently connected environments, including local compute, model distribution, patching, monitoring, backup, and recovery.


Create deployment standards, automation, runbooks, and support processes for on-premises and edge AI.

Integrate on-premises AI with enterprise identity, security, networking, monitoring, and governance where feasible.


Controls, Security & Governance

  • Partner with Cybersecurity, Enterprise Architecture and Compliance to align AI infrastructure with AEP security standards, regulatory obligations, and responsible AI guardrails.
  • Review and remediate configuration drift, vulnerabilities and audit findings across the AI cloud estate; support evidence requests and control attestations.
  • Contribute to onboarding and intake processes for new AI use cases, ensuring workloads are provisioned into the right accounts with the right controls from day one.

FinOps & Cost Management

  • Establish and operate FinOps practices for AI workloads, including tagging standards, showback/chargeback, budgets, anomaly detection, and forecasting.
  • Analyze and optimize spend on GPU compute, inference and token consumption, storage, and data transfer; recommend commitment strategies such as Savings Plans and Reserved Instances.
  • Provide cost transparency and consumption reporting to business stakeholders and technology leadership, and identify optimization opportunities before they become budget issues.

Monitoring & Observability

  • Partner with monitoring team to create logging, alerting and dashboards for AI infrastructure and workloads.
  • Define service-level objectives and operational thresholds for AI platforms, including model endpoint availability, latency, throughput, and error rates.
  • Support incident response, root cause analysis and problem management for AI-related infrastructure events; drive preventive actions and automation to reduce recurrence.

Core Cloud Engineering & Platform Services

  • Perform standard cloud engineering functions across the AEP AWS environment including account provisioning, environment builds, platform upgrades, patching, automation and lifecycle management.
  • Design, deploy and operate Kubernetes environments (Amazon EKS, ROSA/OpenShift) including cluster architecture, autoscaling, node group and GPU scheduling, ingress, service mesh, RBAC, and cluster security hardening.
  • Engineer and support platform services including load balancing, traffic management, API gateways, and integration patterns across cloud, on-premises, edge, and disconnected environments.
  • Apply networking fundamentals, VPC design, subnetting, routing, Transit Gateway, Direct Connect, VPN, firewalls, TLS and certificate management to deliver secure, performant connectivity for AI and general workloads.
  • Prepare cost estimates, justifications, alternative solutions and technical recommendations; produce technical documentation, runbooks and standards.
  • Collaborate with Project Managers, Architects, Solution Engineers, Business Analysts and vendor partners to deliver consistent, reliable solutions that leverage AEP's technology standards, architectures and best practices.
  • Adhere to and advocate for change, incident and problem management processes; participate in on-call rotation and after-hours support as required.
  • Provide training, mentoring and technical work direction to other engineers on the team.

Required Skills & Experience

  • Demonstrated hands-on engineering experience in AWS, including IAM, VPC networking, compute, storage, encryption/KMS, and account/organization structure.
  • Strong Kubernetes expertise including cluster design, operations, troubleshooting and security in EKS, ROSA or OpenShift.
  • Solid networking fundamentals, including routing, DNS, load balancing, traffic management, firewalls, hybrid connectivity, and network design for disconnected or intermittently connected environments.
  • Experience with API gateways, integration patterns, and exposing and securing services across environments.
  • Proficiency with infrastructure-as-code and automation (Terraform required; Ansible, Python or PowerShell preferred) and CI/CD pipelines.
  • Working knowledge of cloud security principles, identity and access management, and compliance requirements.
  • Experience implementing monitoring and cost management for cloud environments, with the ability to establish local monitoring, logging, patching, backup, and recovery processes for on-premises and disconnected environments.
  • Strong analytical, troubleshooting and problem-solving skills, with the ability to work independently on complex assignments.
  • Effective written and verbal communication skills, including the ability to present technical recommendations clearly to management and non-technical stakeholders.

Preferred Qualifications

  • Architecture experience including designing end-to-end cloud solutions and influencing platform direction is highly desirable.
  • Direct experience building or operating AI/ML or GenAI workloads across cloud or on-premises infrastructure (Amazon Bedrock, SageMaker, vector databases, retrieval-augmented generation patterns, GPU-based training or inference, edge AI, or disconnected environments).
  • Experience with FinOps tooling and practices for high-variability workloads.
  • Experience in a regulated industry (utility, energy, financial services) or with FedRAMP/GovCloud environments.
  • Familiarity with multi-cloud environments (Azure, OCI, GCP) and enterprise data platforms such as Snowflake.
  • Certifications: AWS Solutions Architect – Associate or Professional; AWS Certified Machine Learning; Certified Kubernetes Administrator (CKA).

Other Requirements

  • Adhere to policies, procedures, standards, codes and regulations relevant to assignments.
  • Demonstrate in-depth knowledge of AEP infrastructure, environment and components to enable efficient, comprehensive responses to projects and problems.
  • Participate in on-call rotation, after-hours maintenance windows, and storm/emergency response support as required.

What We're Looking For:


Education requirements are listed below:

  • Bachelor's degree in computer science, engineering, or related technical field is required.

Work Experience requirement listed below:

  • 12 years of relevant work experience required. An equivalent combination of education and related experience may be considered.

What You'll Get:

  • Base salary
  • Annual bonus
  • Long-term incentive
  • 401(k) match
  • AEP Pension
  • Comprehensive benefits package designed to support and enhance the overall well-being of employees.

At AEP, we’re more than just an energy company — we’re a team of dedicated professionals committed to delivering safe, reliable, and innovative energy solutions. Guided by our mission to put the customer first, we strive to exceed expectations by listening, responding, and continuously improving the way we serve our communities. If you're passionate about making a meaningful impact and being part of a forward-thinking organization, this is the company for you!


Compensation Data

Compensation Grade:

SP20-010

Compensation Range:

$136,539.00 - $177,503.00

The Physical Demand Level for this job is: S – Sedentary Work: Exerting up to 10 pounds of force occasionally (Occasionally: activity or condition exists up to 1/3 of the time) and/or a negligible amount of force frequently. (Frequently: activity or condition exists from 1/3 to 2/3 of the time) to lift, carry, push, pull or otherwise move objects, including the human body. Sedentary work involves sitting most of the time but may involve walking or standing for brief periods of time. Jobs are sedentary if walking and standing are required only occasionally, and all other sedentary criteria are met.  

Hear about it first!   Get job alerts by email.  Log in to your Candidate Home Account today!  If you don't have an account, you can create one.

It is hereby reaffirmed that it is the policy of American Electric Power (AEP) to provide Equal Employment Opportunity in all respects of the employer-employee relationship including recruiting, hiring, upgrading and promotion, conditions and privileges of employment, company sponsored training programs, educational assistance, social and recreational programs, compensation, benefits, transfers, discipline, layoffs and termination of employment to all employees and applicants without discrimination because of race, color, religion, sex (including pregnancy, gender identity, and sexual orientation), national origin, age, veteran or military status, disability, genetic information, or any other basis prohibited by applicable law. When required by law, we might record certain information or applicants for employment may be invited to voluntarily disclose protected characteristics.

Skills Required

  • Bachelor's degree in computer science, engineering, or a related technical field
  • 12 years of relevant work experience, or equivalent education and related experience
  • Hands-on AWS engineering experience with IAM, VPC networking, compute, storage, encryption/KMS, and account or organization structures
  • Strong Kubernetes expertise, including cluster design, operations, troubleshooting, and security in EKS, ROSA, or OpenShift
  • Networking experience with routing, DNS, load balancing, traffic management, firewalls, hybrid connectivity, and disconnected environments
  • Experience with API gateways, integration patterns, and securing services across environments
  • Proficiency with Terraform and infrastructure-as-code, automation, and CI/CD pipelines
  • Working knowledge of cloud security, identity and access management, and compliance requirements
  • Experience implementing cloud monitoring and cost management, including monitoring, logging, patching, backup, and recovery for on-premises or disconnected environments
  • Strong analytical, troubleshooting, problem-solving, written communication, and verbal communication skills
  • Architecture experience designing end-to-end cloud solutions and influencing platform direction
  • Experience operating AI/ML or generative AI workloads across cloud or on-premises infrastructure
  • Experience with FinOps tooling and practices for high-variability workloads
  • Experience in a regulated industry or with FedRAMP/GovCloud environments
  • Familiarity with Azure, OCI, GCP, and Snowflake
  • AWS Solutions Architect, AWS Certified Machine Learning, or Certified Kubernetes Administrator certification

Similar Jobs

Hybrid
3 Locations
289097 Employees

EXL Logo EXL

Artificial Intelligence Engineer

Information Technology • Database • Consulting
Remote or Hybrid
United States
30246 Employees
150K-170K Annually

PNC Bank Logo PNC Bank

System Reliability & Support Lead-Application Support

Machine Learning • Payments • Security • Software • Financial Services
Hybrid
Cincinnati, OH, USA
55000 Employees
75K-125K Annually

MetLife Logo MetLife

Principal Data & Analytics Lead

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
140K-210K Annually
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Columbus, OH
12,632 Employees

What We Do

Our team at American Electric Power is committed to improving our customers' lives with reliable, affordable power. We are investing $54 billion from 2025 through 2029 to enhance service for customers and support the growing energy needs of our communities. Our nearly 16,000 employees operate and maintain the nation's largest electric transmission system with 40,000 line miles, along with more than 225,000 miles of distribution lines to deliver energy to 5.6 million customers in 11 states. AEP also is one of the nation's largest electricity producers with approximately 29,000 megawatts of diverse generating capacity. We are focused on safety and operational excellence, creating value for our stakeholders and bringing opportunity to our service territory through economic development and community engagement. Our family of companies includes AEP Ohio, AEP Texas, Appalachian Power (in Virginia and West Virginia), AEP Appalachian Power (in Tennessee), Indiana Michigan Power, Kentucky Power, Public Service Company of Oklahoma, and Southwestern Electric Power Company (in Arkansas, Louisiana, east Texas and the Texas Panhandle). AEP also owns AEP Energy, which provides innovative competitive energy solutions nationwide. AEP is headquartered in Columbus, Ohio. For more information, visit aep.com.

Similar Companies Hiring

UL Solutions Thumbnail
Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Chicago, IL
15000 Employees
Runwise Thumbnail
Greentech • Hardware • Real Estate • Software • Energy • PropTech
New York, NY
199 Employees
Energy CX Thumbnail
Greentech • Professional Services • Business Intelligence • Consulting • Energy • Financial Services • Utilities
Chicago, IL
150 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account