Senior AI Compute Engineer

Posted 3 Days Ago
Be an Early Applicant
Mumbai, Maharashtra, IND
In-Office
Senior level
Artificial Intelligence • Cloud • Infrastructure as a Service (IaaS)
The Role
Designs, deploys, operates, and optimizes NVIDIA and AMD GPU clusters for LLM training, inference, and HPC workloads. Responsibilities include Linux administration, kernel and driver tuning, Kubernetes orchestration, CUDA and GPU networking configuration, Slurm and MPI operations, infrastructure automation, monitoring, incident analysis, and customer support. The role owns deployments from architecture through production acceptance and requires strong cross-layer debugging, documentation, and collaboration skills.
Summary Generated by Built In
About the Role

We are building next-generation AI infrastructure powering LLM training, inference clusters, and HPC workloads. As a Senior AI Compute Engineer, you will design, deploy, and operate GPU clusters based on NVIDIA and AMD GPU infrastructure—managing everything from hardware configuration and Linux optimization to Kubernetes orchestration and customer success. You will work with team end-to-end: from architecture planning and production rollout through optimization and technical support. This is hands-on infrastructure engineering at neo cloud.

What you will be doing:

·       Deploy and manage AI GPU clusters (NVIDIA and AMD) for enterprise and cloud customers—end-to-end ownership from planning to production acceptance

·       Manage advanced Linux systems (RHEL, Ubuntu, Rocky) with expertise in kernel tuning, driver optimization, and system performance at scale

·       Build and optimize GPU infrastructure: configure CUDA, NVIDIA drivers, GPU Operator, GPUDirect RDMA, NVLink, and NVSwitch

·       Deploy and operate Kubernetes clusters with GPU support using Helm, Docker, and Containerd for AI workload orchestration

·       Configure and optimize Slurm, MPI, and parallel file systems for distributed AI training and HPC workloads

·       Perform root cause analysis on production incidents and proactively reduce cluster issues through validation and monitoring


What we need to see:Core Compute (8+ years)

·       8+ years of hands-on Linux systems administration and data center infrastructure deployment

·       3+ years of HPC infrastructure experience with job schedulers (Slurm/PBS) and parallel computing

·       Expertise in Linux administration (RHEL, Ubuntu, Rocky)—kernel tuning, driver management, PCIe troubleshooting, performance optimization

·       Proficiency with GPU infrastructure (NVIDIA GPUs, CUDA, GPUDirect RDMA, NVLink, DCGM monitoring and troubleshooting)

·       Experience with Kubernetes and container orchestration (Helm, Docker, Containerd, GPU Operator, CSI drivers)

Automation & Infrastructure-as-Code

·       3+ years of infrastructure automation using Python, Bash, Ansible, Terraform, or SaltStack

·       Ability to develop provisioning workflows, CI/CD pipelines, and version control with Git

·       Strong scripting skills to automate deployment, validation, and operational tasks at scale


Ways to stand out from the rest:

·       NVIDIA certifications (AI Infrastructure, AI Operations, Certified Associate/Professional)

·       Kubernetes certifications (CKA, CKS) or Red Hat Certified Engineer (RHCE)

·       Experience with AI Factory deployments, LLM training clusters, or GPU cloud platforms

·       Background with NVIDIA DGX SuperPOD, HGX clusters, or NVIDIA Spectrum-X networking

·       Experience with monitoring stacks (Prometheus, Grafana, DCGM, ELK, Loki) and observability in distributed systems

·       Hands-on experience with advanced storage systems (Ceph, GPFS, Weka, VAST) or bare-metal provisioning (MAAS, Foreman)

Minimum Qualifications:

·       Bachelor's degree in Computer Science, Electrical Engineering, Electronics, Information Technology, or equivalent professional experience

·       8+ years of Linux systems administration and data center deployment

·       4+ years of consulting or customer-success engineering roles


Soft Skills:

·       Strong problem-solving and debugging abilities across hardware, kernel, and application layers

·       Ownership mindset with accountability for deployment quality and customer success

·       Cross-functional collaboration with other teams

·       Proactive approach to continuous learning and staying current with AI infrastructure trends

·       Strong documentation and presentation skills, able to defend design decisions amongst peers.

Skills Required

  • 8+ years of hands-on Linux systems administration and data center infrastructure deployment
  • 3+ years of HPC infrastructure experience with Slurm or PBS and parallel computing
  • Expertise with RHEL, Ubuntu, or Rocky Linux administration, kernel tuning, driver management, PCIe troubleshooting, and performance optimization
  • Proficiency with NVIDIA GPUs, CUDA, GPUDirect RDMA, NVLink, and DCGM monitoring and troubleshooting
  • Experience with Kubernetes and container orchestration, including Helm, Docker, Containerd, GPU Operator, and CSI drivers
  • 3+ years of infrastructure automation using Python, Bash, Ansible, Terraform, or SaltStack
  • Ability to develop provisioning workflows, CI/CD pipelines, and Git-based version control processes
  • Strong scripting skills for automated deployment, validation, and operational tasks at scale
  • Bachelor's degree in Computer Science, Electrical Engineering, Electronics, Information Technology, or equivalent professional experience
  • 4+ years of consulting or customer-success engineering experience
  • NVIDIA AI Infrastructure, AI Operations, or Certified Associate/Professional certification
  • Kubernetes CKA or CKS certification, or RHCE certification
  • Experience with AI Factory deployments, LLM training clusters, or GPU cloud platforms
  • Experience with NVIDIA DGX SuperPOD, HGX clusters, or NVIDIA Spectrum-X networking
  • Experience with Prometheus, Grafana, DCGM, ELK, Loki, or distributed-systems observability
  • Experience with Ceph, GPFS, Weka, VAST, MAAS, or Foreman
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
141 Employees
Year Founded: 2023

What We Do

Neysa is an India-based AI acceleration cloud provider focused on democratizing enterprise AI adoption. Co-founded by Sharad Sanghi and Anindya Das, it combines cloud infrastructure, AI systems, and cybersecurity expertise. Its flagship Neysa Velocis platform supports AI training, fine-tuning, inference, compute orchestration, security, and observability, helping organizations across high-growth markets such as India and beyond deploy and scale AI more quickly, safely, and cost-effectively.

Similar Jobs

Mastercard Logo Mastercard

Lead Product Manager

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Pune, Maharashtra, IND
38800 Employees

Coursera + Udemy  Logo Coursera + Udemy

Content Marketing Manager

Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
Remote or Hybrid
India
1500 Employees
106K-143K Annually

Mastercard Logo Mastercard

Software Engineer

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Pune, Maharashtra, IND
38800 Employees

Mastercard Logo Mastercard

Senior Software Engineer

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Pune, Maharashtra, IND
38800 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account