Technology Platform Engineer

Posted Yesterday
Be an Early Applicant
Gurugram, Haryana, IND
In-Office
Senior level
Information Technology
The Role
Designs, deploys, and manages large-scale HPC and AI infrastructure across on-premises, cloud, and hybrid environments. Responsibilities include architecting GPU/XPU clusters, administering Linux and Slurm/Kubernetes platforms, optimizing performance and cost, integrating AI workloads and data pipelines, monitoring resiliency, automating infrastructure with Terraform and Ansible, and supporting users running advanced models and simulations.
Summary Generated by Built In
Project Role : Technology Platform Engineer
Project Role Description : Creates production and non-production cloud environments using the proper software tools such as a platform for a project or product. Deploys the automation pipeline and automates environment creation and configuration.
Must have skills : Linux Architecture
Good to have skills : Machine Learning (ML), Cloud Technology Architecture, Docker Kubernetes Administration, Edge Computing, Microsoft Agentic AI Architecture
Minimum 7.5 year(s) of experience is required
Educational Qualification : 15 years full time education
Key Responsibilities:
1) Design and implement HPC and AI infrastructure solutions, aligning system architecture and deployment roadmaps to industry-specific performance and scalability needs
2) Deploy, configure, and manage XPU-based clusters (CPU/GPU/accelerators) using schedulers, VM/K8s orchestration platforms, Slurm, and containerized platforms in scalable designs to provide Metal as a Service (MaaS), GPUaaS, AIaaS, and other offerings
3) Optimize cluster performance, scalability, energy, and cost efficiency across on-premises, cloud, and hybrid environments
4) Integrate AI and HPC platforms with existing IT systems, data pipelines, and security frameworks
5) Monitor, troubleshoot, and tune infrastructure to ensure high availability, low-latency networking, and workload resiliency
6) Develop and maintain documentation including architecture diagrams, configuration baselines, and operational runbooks
7) Provide Provide technical guidance and support to users, enabling efficient execution of HPC/AI workloads, large-scale models, and simulations
Required Skills and Qualifications:
1) Experience in enterprise-wide HPC strategy and architecting next-gen supercomputing environments across the full stack.
2) Primary Skills: Linux Administrator/Architect, Advanced CUDA/GPU on H100/A100, HPC cluster design (SLURM/PBS Pro)
3) Secondary Skills: Parallel programming: MPI, OpenMP, CUDA, SYCL, NVIDIA DGX SuperPOD & BCM, InfiniBand NDR 400Gb/s & RoCEv2 design, MLOps & HPC integration (Kubeflow), Containerization: Singularity, Kubernetes, DevOps: Ansible, Terraform for HPC, UFM management, RunAI, Azure ML integration: distributed training, MLflow, Terraform / Bicep IaC for Azure HPC, FinOps: Reserved Instances, Spot VM strategies, Hybrid cloud HPC: on-prem to Azure/AWS burst
4) Proven hands-on experience designing, deploying, and managing HPC and AI infrastructure across on-premises, cloud, and hybrid environments in 2 or more segments: hyperscaler, neocloud, large Enterprise, Telco/Mobile, supporting key industries such as Financial Services, Life Sciences, Manufacturing, and Retail
5) Deep knowledge of accelerated computing architectures (GPUs, XPUs, DPUs), high-performance fabrics (InfiniBand, Ethernet), SONiC, networking, and modern storage/data platforms (e.g. NVMe-oF, Lustre, GPFS, BeeGFS, VAST, DDN, Weka) to build robust solutions
6) Proficiency with cluster management and orchestration (e.g. Slurm, Run:ai, Kubernetes, Docker), real-time performance monitoring, and observability frameworks
7) Hands-on experience with cloud and virtualization platforms (e.g. AWS, Azure, GCP, VMware, Nutanix) and expertise in automation and optimization using scripting (Python, AI tools) with foundational Infrastructure-as-Code tools such as Terraform and Ansible.
Preferred Skills and Qualifications:
1) Experience managing the deployment of 1,000+ GPU clusters for HPC and AI workloads with various infrastructure services enabled
2) Experience with GPU computing libraries and accelerators (e.g., NVIDIA CUDA, Dynamo, AMD ROCm).
3) Experience with AI and HPC Networking (e.g., RoCE, InfiniBand, muti-planar/multi-rail designs, platform buffer architectures)
4) Knowledge of Machine Learning and AI frameworks (e.g., TensorFlow, PyTorch, JAX), Jupyter notebooks / Google Colab environments
5) Familiarity with DevOps practices and tools (e.g., Ansible, Terraform) for infrastructure automation
6) Experience with AgenticAI and associated technologies to leverage and build agents for workflow automation and observability
7) Industry certifications in NVIDIA infrastructure, public cloud providers, Data Science, etc. are a plus
8) Strong problem-solving, troubleshooting, communication, and collaboration skills to deliver reliable, scalable, and high-performance infrastructure solutions in fast-paced, dynamic environments that reward technical talent

15 years full time education

About Accenture

Accenture is a leading global professional services company that helps the world’s leading businesses, governments and other organizations build their digital core, optimize their operations, accelerate revenue growth and enhance citizen services—creating tangible value at speed and scale. We are a talent- and innovation-led company with approximately 791,000 people serving clients in more than 120 countries. Technology is at the core of change today, and we are one of the world’s leaders in helping drive that change, with strong ecosystem relationships. We combine our strength in technology and leadership in cloud, data and AI with unmatched industry experience, functional expertise and global delivery capability. Our broad range of services, solutions and assets across Strategy & Consulting, Technology, Operations, Industry X and Song, together with our culture of shared success and commitment to creating 360° value, enable us to help our clients reinvent and build trusted, lasting relationships. We measure our success by the 360° value we create for our clients, each other, our shareholders, partners and communities.

Visit us at www.accenture.com 

Equal Employment Opportunity Statement

We believe that no one should be discriminated against because of their differences. All employment decisions shall be made without regard to age, race, creed, color, religion, sex, national origin, ancestry, disability status, military veteran status, sexual orientation, gender identity or expression, genetic information, marital status, citizenship status or any other basis as protected by applicable law. Our rich diversity makes us more innovative, more competitive, and more creative, which helps us better serve our clients and our communities.

Skills Required

  • Minimum 7.5 years of professional experience
  • 15 years of full-time education
  • Linux administration or architecture experience
  • Experience developing enterprise-wide HPC strategies and next-generation supercomputing environments
  • Advanced CUDA/GPU experience with NVIDIA H100 or A100 systems
  • HPC cluster design experience using Slurm or PBS Pro
  • Experience designing, deploying, and managing HPC and AI infrastructure across on-premises, cloud, or hybrid environments
  • Knowledge of accelerated computing, high-performance networking, storage, and data platforms
  • Proficiency with Slurm, Run:ai, Kubernetes, Docker, monitoring, and observability tools
  • Experience with AWS, Azure, GCP, VMware, or Nutanix
  • Infrastructure automation experience with Python, Terraform, or Ansible
  • Experience deploying 1,000 or more GPU clusters
  • Experience with NVIDIA CUDA, Dynamo, or AMD ROCm
  • Experience with RoCE, InfiniBand, multi-rail designs, or related AI/HPC networking
  • Knowledge of TensorFlow, PyTorch, JAX, Jupyter, or Google Colab
  • Experience with Agentic AI technologies
  • Industry certifications in NVIDIA infrastructure, public cloud, or data science
  • Strong problem-solving, troubleshooting, communication, and collaboration skills

Accenture Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Accenture and has not been reviewed or approved by Accenture.

  • Healthcare Strength — Pay is considered competitive when paired with robust insurance options and other perks that compare well with large consulting and IT services peers. Multiple national medical plan options plus dental and vision are positioned as a core strength of the overall package.
  • Retirement Support — Retirement support is positioned as a standout feature through a 401(k) dollar-for-dollar match up to a set percentage after eligibility. The package is reinforced by additional financial programs such as savings tools and related resources.
  • Parental & Family Support — Parental and caregiving supports are presented as a meaningful benefit differentiator through substantial paid parental leave and multiple caregiver-oriented programs. Backup care and fertility/adoption/surrogacy navigation and reimbursements add breadth to family support beyond leave alone.

Accenture Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Dublin
456,553 Employees
Year Founded: 1989

What We Do

Accenture is a global professional services company with leading capabilities in digital, cloud and security. Combining unmatched experience and specialized skills across more than 40 industries, we offer Strategy and Consulting, Interactive, Technology and Operations services—all powered by the world’s largest network of Advanced Technology and Intelligent Operations centers. Our 500,000+ people deliver on the promise of technology and human ingenuity every day, serving clients in more than 120 countries. We embrace the power of change to create value and shared success for our clients, people, shareholders, partners and communities. Visit us at www.accenture.com.

Similar Jobs

Accenture Logo Accenture

Platform Engineer

Information Technology
In-Office
Gurugram, Haryana, IND
456553 Employees

Mastercard Logo Mastercard

Senior Specialist, Product Management, Custom Analytics

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Gurugram, Haryana, IND
38800 Employees

Circle (circle.so) Logo Circle (circle.so)

Lead Engineer, AI Platform

Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
In-Office or Remote
18 Locations
250 Employees

Circle (circle.so) Logo Circle (circle.so)

Senior Quality Engineer

Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
In-Office or Remote
18 Locations
250 Employees
120K-130K Annually

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account