Sr. Technology Enablement Engineer

Reposted One Month Ago
Ann Arbor, MI, USA
In-Office
130K-190K Annually
Senior level
Hardware
The Role
Design, develop, and deploy production-grade AI/ML models and scalable architectures. Lead model training, validation, deployment, monitoring, and cross-functional integration. Conduct research, document methods, and share knowledge to drive business impact and operational reliability.
Summary Generated by Built In

To make electronics, you need chips, wafers, transistors, reticles, and... To make these, you must see, test and manufacture them at scale—faster and better than ever before. That's where KLA comes in. Whether you're early in your career or an experienced professional, you'll solve complex challenges, work alongside brilliant minds and help shape the future of technology.


Group/Division

KLA's IT group supports business growth and productivity by connecting people, process and technology around the world. We work to improve the technology that drives our business to thrive and focus on empowering employee use of technology. This integrated approach to customer service, creativity and technological excellence enhances employee productivity, business analytics and process excellence.

What You'll Do

In this role, you will play a key part in advancing business priorities by delivering high-impact work across your area of expertise.


We are seeking a highly skilled Sr. Software Engineer with specialized expertise

Design, build, and operate production-grade AI infrastructure for large-scale GPU training and inference. The role owns cluster architecture from ground zero, distributed workload orchestration, accelerator management, performance engineering, and open-source integration. TPU experience is optional. 

Key Responsibilities 

  • Design and deploy scalable, multi-node GPU clusters on Kubernetes, including compute, networking, storage, scheduling, security, and observability. 

  • Build distributed training and reinforcement learning platforms using Ray and NVIDIA NeMo RL, supporting frameworks such as PyTorch and JAX. 

  • Deploy and optimize high-throughput LLM inference using vLLM, SGLang, and NVIDIA Dynamo. 

  • Implement GPU scheduling, quotas, isolation, autoscaling, health monitoring, and capacity management for multi-tenant environments. 

  • Profile and troubleshoot GPU workloads using NVIDIA Nsight Systems, Nsight Compute, DCGM, CUDA, NCCL, and related diagnostics. 

  • Automate cluster provisioning, upgrades, workload deployment, and operational recovery through infrastructure-as-code and GitOps practices. 

  • Contribute to or actively participate in relevant open-source AI infrastructure communities and bring upstream best practices into the platform. 

Minimum Qualifications

  • ​Bachelor's Degree and eight (8) years of Software Engineering experience

  • Four (4) years in software, cloud, platform, HPC, or infrastructure engineering, including two (2) years supporting distributed AI/ML workloads. 

  • Proven hands-on experience building GPU clusters from the ground up and operating them at production scale. 

  • Deep Kubernetes expertise, including operators, CRDs, Helm, networking, storage, scheduling, and cluster lifecycle management. 

  • Production experience with Ray and at least two of the following: NVIDIA NeMo RL, vLLM, SGLang, or NVIDIA Dynamo. 

  • Strong knowledge of NVIDIA GPUs, CUDA, NCCL, GPU Operator, DCGM, MIG, RDMA, and multi-node collective communication. 

  • Experience debugging training and inference performance across GPU, CPU, memory, network, and storage layers. 

  • Proficiency in Python and Linux, plus experience with containers, CI/CD, GitOps, infrastructure-as-code, and platform observability. 

  • Demonstrated open-source contribution, maintainership, or meaningful participation in an AI infrastructure project. 

Preferred Qualifications 

  • Experience with Google TPUs and TPU-oriented frameworks or distributed workloads. 

  • Experience with large language model training, fine-tuning, RLHF or agentic reinforcement learning. 

  • Knowledge of TensorRT-LLM, Triton Inference Server, DeepSpeed, Megatron-LM, or similar performance-oriented frameworks. 

  • Experience operating secure, multi-tenant AI platforms in enterprise or regulated environments. 

About KLA

We provide advanced inspection tools, metrology systems, process solutions, and computational analytics that make electronics possible, tackling complex challenges. From electron and photon optics to machine learning and data analytics, we seek perfection at the most fundamental level of matter in the universe. If you want to make electronics that push industries forward and make the world a better place, join us.


Total Rewards

Base Pay Range: $129,600.00 - $190,067.00 Annually Primary Location: USA-MI-Ann Arbor-KLA

KLA’s total rewards package for employees may also include participation in performance incentive programs and eligibility for additional benefits including but not limited to: medical, dental, vision, life, and other voluntary benefits, 401(K) including company matching, employee stock purchase program (ESPP), student debt assistance, tuition reimbursement program, development and career growth opportunities and programs, financial planning benefits, wellness benefits including an employee assistance program (EAP), paid time off and paid company holidays, and family care and bonding leave.
Interns are eligible for some of the benefits listed. Our pay ranges are determined by role, level, and location. The range displayed reflects the pay for this position in the primary location identified in this posting. Actual pay depends on several factors, including state minimum pay wage rates, location, job-related skills, experience, and relevant education level or training. We are committed to complying with all applicable federal and state minimum wage requirements where applicable. If applicable, your recruiter can share more about the specific pay range for your preferred location during the hiring process.


Use of AI Statement 

At KLA, our interviews seek to understand your individual skills, problem-solving approach and authentic thinking. To ensure a fair and consistent evaluation, the use of AI, recording tools or other technologies to generate, suggest or provide responses during interviews—whether virtual or in person—is not permitted unless explicitly approved in advance as part of a reasonable accommodation or invited by the interviewer. Use of these tools may interfere with our ability to evaluate your individual qualifications and affect your candidacy. KLA is committed to advancing innovation through responsible AI, and we value candidates who share this mindset.


Equal Opportunity Statement

KLA is proud to be an Equal Opportunity Employer. We will ensure that qualified individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us at [email protected] or at +1-408-352-2808 to request accommodation.


For additional information, view the US Know Your Rights poster on the U.S. Equal Employment Opportunity Commission website.

Skills Required

  • Bachelor's degree in Computer Science, Data Science, AI, or related field (Master's/PhD preferred).
  • Minimum eight (8) years software engineering experience with 3-5 years in AI/ML and production deployments.
  • Proficiency in Python or similar languages.
  • Experience with machine learning libraries such as TensorFlow and Scikit-learn.
  • Experience with Retrieval-Augmented Generation (RAG).
  • Strong background in statistics, machine learning algorithms, deep learning, and neural networks.
  • Strong understanding of system architecture and scalable system design for large data/AI workloads.
  • Familiarity with cloud platforms, DevOps practices, and data processing technologies.
  • Experience with software engineering best practices including version control (Git), CI/CD pipelines, and agile methodologies.
  • Demonstrated ability to lead technical projects and work effectively in matrixed environments.
  • Ability to communicate technical concepts to non-technical stakeholders.
  • Hybrid work based in Ann Arbor, MI (role location requirement).
  • Experience with specialized AI domains or industry-specific applications.
  • Contributions to open-source projects or research publications in AI/ML.
  • Knowledge of AI ethics and responsible AI practices.

KLA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about KLA and has not been reviewed or approved by KLA.

  • Retirement Support — U.S. programs include a formulaic 401(k) employer match with clearly described plan designs in company materials. Consistent plan structure and employer matching signal strong support for long‑term savings.
  • Equity Value & Accessibility — A broad base of employees is eligible for RSUs and can participate in an ESPP, extending ownership beyond executives. Company materials position equity as a core component of total compensation.
  • Strong & Reliable Incentives — A quarterly profit‑sharing program and performance bonuses are outlined in filings, providing regular variable pay opportunities. These incentives complement base pay and equity to lift overall earnings in strong periods.

KLA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Milpitas, CA
10,001 Employees

What We Do

KLA develops industry-leading equipment and services that enable innovation throughout the electronics industry. We provide advanced process control and process-enabling solutions for manufacturing wafers and reticles. In close collaboration with leading customers across the globe, our expert teams of physicists, engineers, data scientists and problem-solvers design solutions that move the world forward.

Similar Jobs

Octus Logo Octus

Legal Workflow Strategist

Fintech • News + Entertainment • Software • Database • Financial Services
Easy Apply
Remote or Hybrid
United States
808 Employees

Samsara Logo Samsara

Senior Software Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
155K-260K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
14 Locations
40000 Employees
45K-85K Annually

CrowdStrike Logo CrowdStrike

Consultant

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
110K-160K Annually

Similar Companies Hiring

Blissway Thumbnail
Computer Vision • Fintech • Hardware • Internet of Things • Machine Learning • Software • Transportation
Denver, CO
24 Employees
Fairly Even Thumbnail
Hardware • Robotics • Sales • Software • Hospitality
New York, NY
30 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account