AI & High Performance Computing Infrastructure Engineering Manager, Senior Director

Posted Yesterday
Be an Early Applicant
New York, NY, USA
In-Office
Senior level
Fintech • Financial Services
The Role
Leads the engineering strategy and delivery of enterprise AI and HPC infrastructure. Manages AI infrastructure engineers building, operating, automating, and scaling GPU platforms for model training, inference, RAG, agentic AI, and production workloads. Oversees Kubernetes, cloud and hybrid infrastructure, automation, observability, capacity planning, reliability, governance, vendor relationships, and executive communication. Partners with AI/ML teams to improve model deployment and platform adoption while ensuring security, compliance, resilience, and operational readiness.
Summary Generated by Built In

AI & High Performance Computing (HPC) Infrastructure Engineering Manager, Senior Director

We’re seeking a future team member for the role of AI & High Performance Computing Infrastructure Engineering Manager, Senior Director to join our Enterprise Infrastructure Delivery organization. This role is in New York, NY.

The AI and HPC Infrastructure Engineering Manager will lead the engineering and strategic evolution of the bank’s AI, machine learning, and high-performance computing infrastructure platforms. Responsible for managing a team of AI Infrastructure Engineers resources who design, operate, automate, and scale GPU-based infrastructure supporting model training, inference, agentic AI, data pipelines, and production AI workloads.

 In this role, you'll make an impact in the following ways

  • Lead and develop a team of AI Infrastructure Engineers responsible for building, operating, and scaling the firm's AI, machine learning, and high-performance computing (HPC) platforms.
  • Define and execute the strategic roadmap for AI infrastructure, partnering across Engineering, Architecture, Security, Risk, Compliance, Production Services, and AI/ML teams to deliver secure, scalable, and resilient platforms.
  • Drive the design, deployment, and optimization of GPU-based infrastructure supporting model training, inference, retrieval-augmented generation (RAG), agentic AI, data pipelines, and other production AI workloads across on-premises, hybrid, and cloud environments.
  • Serve as the senior technical leader for AI infrastructure, guiding decisions related to Kubernetes, distributed systems, GPU orchestration, infrastructure automation, observability, performance engineering, capacity planning, and operational resilience.
  • Develop and mature enterprise platform capabilities including AI model serving, scalable data and compute infrastructure, vector and graph database ecosystems, AI orchestration frameworks, and microservices architectures.
  • Partner closely with AI/ML engineering teams to accelerate model deployment, improve platform adoption, optimize performance, and enable the delivery of innovative AI solutions.
  • Advance automation across infrastructure provisioning, configuration management, monitoring, incident response, workload onboarding, and platform lifecycle management.
  • Establish and govern operating models, service standards, production readiness requirements, support processes, and reliability objectives to ensure exceptional platform availability and user experience.
  • Manage strategic technology vendor relationships, influencing product roadmaps, driving technical evaluations and proof-of-concepts, and ensuring alignment with business objectives.
  • Communicate platform strategy, investment priorities, operational performance, risks, and roadmap progress to executive leadership and technology governance forums.

To be successful in this role, we're seeking the following

  • Proven leadership experience managing infrastructure engineering, platform engineering, DevOps, SRE, or production operations teams in large-scale enterprise environments.
  • Deep expertise designing and operating AI, machine learning, GPU, HPC, or distributed computing platforms in production environments.
  • Strong experience with Kubernetes, containerized platforms, GPU orchestration technologies, and modern infrastructure automation practices.
  • Expertise with NVIDIA GPU ecosystems, including GPU lifecycle management, workload optimization, and large-scale compute environments.
  • Solid understanding of AI infrastructure patterns, including model training, inference, model serving, RAG architectures, data pipelines, and emerging agentic AI frameworks.
  • Experience deploying and supporting hybrid cloud infrastructure, with working knowledge of Azure, GCP, or similar cloud platforms.
  • Strong technical foundation in Linux, networking, storage, distributed systems, observability, performance engineering, and production operations.
  • Experience leveraging Infrastructure as Code, CI/CD, and automation tools such as Terraform, Ansible, Helm, ArgoCD, GitLab, Jenkins, or similar technologies.
  • Ability to influence senior stakeholders and effectively communicate complex technical concepts to both engineering and executive audiences.
  • Demonstrated success developing technology strategies, operating models, governance frameworks, and long-term infrastructure roadmaps.
  • Strong vendor management experience, including strategic partnerships, architecture reviews, technical assessments, and escalation management.

Preferred Qualifications

  • Experience within financial services or other highly regulated industries.
  • Expertise supporting AI platforms subject to stringent security, compliance, audit, and operational resilience requirements.
  • Experience with large-scale GPU environments, DGX platforms, SuperPOD architectures, InfiniBand, high-performance storage, or enterprise AI infrastructure.
  • Familiarity with modern AI serving and orchestration technologies such as Triton, vLLM, NVIDIA NIM, KServe, Seldon, LangGraph, vector databases, graph databases, and multi-agent frameworks.
  • Experience with capacity planning, usage analytics, cost optimization, GPU quota management, and chargeback/showback models.
  • 10+ years of experience in a related field with a 6-8 years’ experience managing staff 
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline; advanced STEM degrees or relevant cloud, Kubernetes, AI, or infrastructure certifications are a plus.
About Us

At BNY, our culture allows us to run our company better and enables employees’ growth and success. As a leading global financial services company at the heart of the global financial system, we influence nearly 20% of the world’s investible assets. Every day, our teams harness cutting-edge AI and breakthrough technologies to collaborate with clients, driving transformative solutions that redefine industries and uplift communities worldwide.

Recognized as a top destination for innovators, BNY is where bold ideas meet advanced technology and exceptional talent. Together, we power the future of finance – and this is what #LifeAtBNY is all about. Join us and be part of something extraordinary. About the Team

At BNY, our culture speaks for itself, check out the latest BNY news at BNY Newsroom & BNY LinkedIn

 Here’s a few of our recent awards:

  • America’s Most Innovative Companies, Fortune, 2025
  • World’s Most Admired Companies, Fortune 2025
  • “Most Just Companies”, Just Capital and CNBC, 2025

    Our Benefits and Rewards:

    BNY offers highly competitive compensation, benefits, and wellbeing programs rooted in a strong culture of excellence and our pay-for-performance philosophy. We provide access to flexible global resources and tools for your life’s journey. Focus on your health, foster your personal resilience, and reach your financial goals as a valued member of our team, along with generous paid leaves, including paid volunteer time, that can support you and your family through moments that matter.

    BNY is an Equal Employment Opportunity/Affirmative Action Employer - Underrepresented racial and ethnic groups/Females/Individuals with Disabilities/Protected Veterans.

    BNY assesses market data to ensure a competitive compensation package for our employees. The expected base salary for this position when employment commences can be found in the Job Info section at the bottom of the posting. 

    Base salary offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience. Base salary is only part of the total rewards package, which may include eligibility for an annual discretionary incentive award. Subject to the terms and conditions of the applicable plans then in effect, eligible employees may enroll in a 401(k) plan as well as participate in Company-sponsored medical, dental, vision, and basic life insurance plans for the employee and the employee’s eligible dependents. Eligible employees also may receive other benefits (including various paid time off benefits, such as vacation and sick time), dependent on the position offered. Details of participation in these benefit plans will be provided if an employee receives an offer of employment.

    If hired, the employee will be in an “at will” position and the Company reserves the right to modify base salary (as well as any other discretionary payments or compensation programs) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

    Skills Required

    • Leadership experience managing infrastructure engineering, platform engineering, DevOps, SRE, or production operations teams in large-scale enterprise environments
    • Deep expertise designing and operating AI, machine learning, GPU, HPC, or distributed computing platforms in production
    • Strong experience with Kubernetes, containerized platforms, GPU orchestration technologies, and infrastructure automation
    • Expertise with NVIDIA GPU ecosystems, GPU lifecycle management, workload optimization, and large-scale compute environments
    • Understanding of AI infrastructure patterns, including model training, inference, model serving, RAG architectures, data pipelines, and agentic AI frameworks
    • Experience deploying and supporting hybrid cloud infrastructure, with working knowledge of Azure, GCP, or similar platforms
    • Strong technical foundation in Linux, networking, storage, distributed systems, observability, performance engineering, and production operations
    • Experience with Infrastructure as Code, CI/CD, and automation tools such as Terraform, Ansible, Helm, ArgoCD, GitLab, or Jenkins
    • Ability to influence senior stakeholders and communicate complex technical concepts to engineering and executive audiences
    • Success developing technology strategies, operating models, governance frameworks, and long-term infrastructure roadmaps
    • Strong vendor management experience, including strategic partnerships, architecture reviews, technical assessments, and escalation management
    • Experience within financial services or another highly regulated industry
    • Experience supporting AI platforms subject to stringent security, compliance, audit, and operational resilience requirements
    • Experience with large-scale GPU environments, DGX platforms, SuperPOD architectures, InfiniBand, high-performance storage, or enterprise AI infrastructure
    • Familiarity with Triton, vLLM, NVIDIA NIM, KServe, Seldon, LangGraph, vector databases, graph databases, and multi-agent frameworks
    • Experience with capacity planning, usage analytics, cost optimization, GPU quota management, and chargeback or showback models
    • 10 or more years of experience in a related field
    • 6 to 8 years of experience managing staff
    • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline
    • Advanced STEM degree or relevant cloud, Kubernetes, AI, or infrastructure certifications

    BNY Compensation & Benefits Highlights

    The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about BNY and has not been reviewed or approved by BNY.

    • Healthcare Strength Health coverage includes comprehensive options with a $0‑premium plan for eligible lower earners, expanded mental‑health support with personalized therapy, and strong income protection through short‑ and long‑term disability. These features have been recently enhanced and are paired with dental and vision coverage.
    • Parental & Family Support Parental leave provides 16 weeks of fully paid time for all parents, with added support such as adoption assistance. This breadth offers strong coverage for major family events.
    • Retirement Support The 401(k) program includes a company match and Roth options to support long‑term savings. Additional financial programs like tuition assistance and savings vehicles complement retirement readiness.

    BNY Insights

    Am I A Good Fit?
    beta
    Get Personalized Job Insights.
    Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

    The Company
    HQ: New York, NY
    41,739 Employees

    What We Do

    We help make money work for the world — managing it, moving it and keeping it safe. As a leading global financial services company at the center of the world’s financial system, we touch nearly 20% of the world’s investable assets. Today we help over 90% of Fortune 100 companies and nearly all the top 100 banks globally access the money they need. For 240 years we have partnered alongside our clients to create solutions that benefit businesses, communities and people everywhere.

    Similar Jobs

    MetLife Logo MetLife

    Actuarial Intern

    Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
    Hybrid
    New York, NY, USA
    43000 Employees
    28-35 Hourly

    Wells Fargo Logo Wells Fargo

    Personal Banker Cold Spring

    Fintech • Financial Services
    Hybrid
    Cold Spring, NY, USA
    205000 Employees
    23-31 Hourly
    Hybrid
    New York, NY, USA
    205000 Employees
    215K-355K Annually

    Cox Enterprises Logo Cox Enterprises

    Human Resources Business Partner

    Artificial Intelligence • Automotive • Greentech • Information Technology • Machine Learning • Software • Cybersecurity
    Remote or Hybrid
    New York, NY, USA
    30000 Employees
    81K-122K Annually

    Similar Companies Hiring

    Hanover Park Thumbnail
    Artificial Intelligence • Fintech • Software • Financial Services
    New York, New York
    42 Employees
    Kepler  Thumbnail
    Artificial Intelligence • Fintech • Software
    New York, New York
    9 Employees
    Onshore Thumbnail
    Artificial Intelligence • Fintech • Software • Financial Services
    New York, New York
    60 Employees

    Sign up now Access later

    Create Free Account

    Please log in or sign up to report this job.

    Create Free Account