Senior Platform Engineer

Posted Yesterday
Hiring Remotely in USA
Remote
Senior level
Artificial Intelligence • Big Data • Machine Learning
The Role
Design, optimize, and scale multi-GPU infrastructure for GenAI/LLM workloads. Perform GPU profiling and benchmarking, manage Slurm and Kubernetes/OpenShift clusters, enable NVIDIA GPU stack, build GenAI pipelines (fine-tuning, RAG, multi-modal), create IaC templates, and support production deployments and client engagements.
Summary Generated by Built In

While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth.
If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!

About Quantiphi:

Quantiphi is an award-winning, AI-First digital engineering and consulting company focused on delivering high-impact Services and Solutions that help organizations solve what truly matters. We partner with enterprises to reimagine their businesses through intelligent, scalable, and transformative AI driving measurable outcomes at the very core of their operations.

Since our founding in 2013, Quantiphi has tackled some of the world’s most complex business challenges by combining deep industry expertise, disciplined cloud and data engineering practices, and cutting-edge applied AI research. Our work is rooted in delivering accelerated, quantifiable business value, not just technology for technology’s sake.

Headquartered in Boston, Quantiphi is a global organization with 4,000+ professionals serving clients across key industry verticals, including BFSI, Healthcare & Life Sciences, CPG, MFG, TME etc. As an Elite and Premier partner to leading cloud and AI platforms such as NVIDIA, Google Cloud, AWS, and Snowflake, we build and deliver enterprise-grade AI services and solutions that create real-world impact. 

We’ve been recognized with:

  • 21x Google Cloud Partner of the Year awards in the last 8 years.

  • 3x AWS AI/ML award wins.

  • 3x NVIDIA Partner of the Year titles.

  • 2x Snowflake Partner of the Year awards.

  • We have also garnered top analyst recognitions from Gartner, ISG, and Everest Group.

  • We offer first-in-class industry solutions across Healthcare, Financial Services, Consumer Goods, Manufacturing, and more, powered by cutting-edge Generative AI and Agentic AI accelerators.

  • We have been certified as a Great Place to Work for the third year in a row- 2021, 2022, 2023.

Be part of a trailblazing team that’s shaping the future of AI, ML, and cloud innovation. 

Your next big opportunity starts here!

For more details, visit: Website or LinkedIn Page.

Role: Senior Platform Engineer

Experience Level: 10+ yrs

Work Location: US/Canada [ET & CT] 

Role Overview:

We are looking for a highly skilled Senior Platform Engineer to design, optimize, and scale infrastructure for GenAI and LLM workloads. This role is ideal for someone with deep hands-on experience in GPU profiling, distributed training, and high-performance compute environments.

You’ll play a key role in building out GenAI platform foundations, supporting production-grade deployments, and partnering closely with data science, MLOps, and application teams to bring cutting-edge AI solutions to life.

Key Responsibilities:

  • Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments

  • Perform GPU profiling, benchmarking, and performance optimization for distributed training workloads

  • Manage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environments

  • Enable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS, etc.)

  • Collaborate with cross-functional teams to deploy models in research and production environments

  • Build and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps)

  • Develop reusable infrastructure templates using tools like Terraform and Helm

  • Contribute to internal innovation (PoCs, workshops) and support client-facing delivery engagements

Basic Qualifications:

  • Strong experience with Slurm and distributed training environments

  • Hands-on expertise with Red Hat OpenShift and/or Kubernetes

  • Deep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT)

  • Strong foundation in Linux systems, performance tuning, and multi-GPU optimization

  • Experience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems)

  • Familiarity with Infrastructure-as-Code tools (Terraform, Ansible)

  • Experience with cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clusters

Other Qualifications (OQs):

  • Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containers

  • Knowledge of LLMOps frameworks and MLOps integration

  • Familiarity with vector databases and retrieval systems for RAG architectures

  • Comfortable working in client-facing environments and collaborating with AI solution teams

Healthcare Domain Experience (Nice to Have):

  • Experience working with FHIR R4, HL7 v2, or SMART on FHIR

  • Integration with EHR systems (e.g., Epic)

  • Understanding of HIPAA compliance and healthcare data privacy

  • Exposure to clinical workflows, CDS Hooks, or patient-facing applications

  • Experience building clinical decision support systems or healthcare interoperability solutions

What’s in it for YOU at Quantiphi:

  • Make an impact at one of the world’s fastest-growing AI-first digital engineering companies.

  • Upskill and discover your potential as you solve complex challenges in cutting-edge areas of technology alongside passionate, talented colleagues.

  • Work where innovation happens - work with disruptive innovators in a research-focused organization with 60+ patents filed across various disciplines.

  • Stay ahead of the curve, immerse yourself in breakthrough AI, ML, data, and cloud technologies and gain exposure working with Fortune 500 companies.

If you like wild growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us!

Skills Required

  • Strong experience with Slurm and distributed training environments
  • Hands-on expertise with Red Hat OpenShift and/or Kubernetes
  • Deep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT)
  • Strong foundation in Linux systems, performance tuning, and multi-GPU optimization
  • Experience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems)
  • Familiarity with Infrastructure-as-Code tools (Terraform, Ansible)
  • Experience with cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clusters
  • Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containers
  • Knowledge of LLMOps frameworks and MLOps integration
  • Familiarity with vector databases and retrieval systems for RAG architectures
  • Comfortable working in client-facing environments and collaborating with AI solution teams
  • Healthcare domain experience (FHIR, HL7, EHR integration, HIPAA)

Quantiphi Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Quantiphi and has not been reviewed or approved by Quantiphi.

  • Wellbeing & Lifestyle Benefits Wellbeing initiatives such as monthly meeting-free AMA-Zen Days, health check-ups, and wellness counseling are designed to reduce burnout and support day-to-day balance. Broader wellness programs reinforce both physical and mental health.
  • Flexible Benefits Remote/hybrid options with flexible working hours provide meaningful autonomy over where and when work gets done. Flexible leave constructs, including sabbaticals and special day leaves, add practical adaptability to the package.
  • Parental & Family Support Paid parental leave in the U.S., alongside maternity and childcare support, signals solid backing for families. These family-oriented policies integrate with a wider health and wellness focus.

Quantiphi Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Marlborough, MA
3,494 Employees
Year Founded: 2013

What We Do

Quantiphi is an award-winning AI-first digital engineering company driven by the desire to solve transformational problems at the heart of business. Quantiphi solves the toughest and complex business problems by combining deep industry experience, disciplined cloud, and data-engineering practices, and cutting-edge artificial intelligence research to achieve quantifiable business impact at unprecedented speed.

Similar Jobs

Vannevar Logo Vannevar

Senior Software Engineer

Artificial Intelligence • Machine Learning • Software • Defense
Remote
USA
225 Employees
150K-215K Annually

Samsara Logo Samsara

Senior Software Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
6 Locations
4000 Employees
131K-198K Annually

Optum Logo Optum

Platform Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
135K-231K Annually

General Motors Logo General Motors

Machine Learning Engineer

Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Remote or Hybrid
4 Locations
165000 Employees
171K-261K Annually

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account