Senior HPC Architect

Posted 2 Days Ago
Be an Early Applicant
Santa Clara, CA
Senior level
Artificial Intelligence • Hardware • Robotics • Software • Metaverse
The Role
The Senior HPC Architect role involves deploying and managing large-scale GPU compute clusters, providing engineering solutions for GPU computing products, and collaborating with researchers and developers. Responsibilities include system administration, large-scale performance architecture, and fostering technical relationships within the engineering community.
Summary Generated by Built In

NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a “learning machine” that constantly evolves by adapting to new opportunities that are hard to solve, that only we can tackle, and that matter to the world. This is our life’s work, to amplify human imagination and intelligence. Make the choice, join our diverse team today!

We are looking for an outstanding hands-on architect/engineer for a Senior HPC architect role to support deployment and bringup of large-scale GPU compute clusters. Be a key player to enable the most exciting computing hardware and software and contribute to the latest breakthroughs in artificial intelligence and GPU computing. Provide insights on and implement at-scale system administration and tuning mechanisms for large-scale compute runs. You will work with the latest accelerated computing and Deep Learning software and hardware platforms, and with many scientific researchers, developers, and customers to craft improved workflows and develop new, leading differentiated solutions. You will interact with HPC, OS, GPU compute, and systems specialist to architect, develop and bring up large scale performance platforms.

What you’ll be doing:

Provide engineering solutions to operationalize the latest GPU Computing products and software stacks, ensure technical relationships with internal and external engineering teams, and assisting systems, machine learning/deep learning engineers in building creative solutions based on NVIDIA technology. Be an internal reference for system administration, at-scale system analysis, and other datacenter and large-scale GPU-accelerated system solutions among the NVIDIA technical community.

What we need to see:

  • 5+ years of experience using in accelerated computing for datacenter/HPC solutions.

  • Solid understanding of accelerated computing scheduling and I/O stacks.

  • Experience using and handling HPC-based Enterprise computing architectures.

  • C/C++/Python/Bash programming/scripting experience.

  • Experience working with engineering or academic research community supporting high performance computing or deep learning.

  • Background with scheduling and resource management systems.

  • Experience with parallel filesystems.

  • Strong verbal and written communication skills.

  • Strong teamwork and communication skills.

  • Ability to multitask effectively in a dynamic environment.

  • Action driven with strong analytical and analytical skills.

  • Desire to be involved in multiple diverse and innovative projects.

  • BS (or equivalent experience) in Engineering, Mathematics, Physics, or Computer Science. MS or PhD desirable.

Ways to stand out from the crowd:

  • Deep Learning framework skills.

  • Exposure to using and deploying telemetry and visualization pipelines

  • Exposure to container technology and Linux performance tools.

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/ We have some of the most brilliant and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you.

The base salary range is 148,000 USD - 276,000 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Top Skills

C
C++
Python
The Company
HQ: Santa Clara, CA
21,960 Employees
On-site Workplace
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

BlackLine Logo BlackLine

Senior Database Administrator (FedRamp)

Cloud • Fintech • Information Technology • Machine Learning • Software • App development • Generative AI
Pleasanton, CA, USA
1810 Employees
145K-193K Annually

BlackLine Logo BlackLine

Staff I Systems Engineer (FedRamp)

Cloud • Fintech • Information Technology • Machine Learning • Software • App development • Generative AI
Pleasanton, CA, USA
1810 Employees
160K-213K Annually

BlackLine Logo BlackLine

Staff I Reliability Engineer (FedRamp)

Cloud • Fintech • Information Technology • Machine Learning • Software • App development • Generative AI
Pleasanton, CA, USA
1810 Employees
160K-213K Annually

BlackLine Logo BlackLine

Senior Site Reliability Engineer (FedRamp)

Cloud • Fintech • Information Technology • Machine Learning • Software • App development • Generative AI
Pleasanton, CA, USA
1810 Employees
145K-193K Annually

Similar Companies Hiring

TrainingPeaks (A Peaksware Company) Thumbnail
Software • Fitness
Louisville, CO
69 Employees
bet365 Thumbnail
Software • Gaming • eSports • Digital Media • Automation
Denver, Colorado
6100 Employees
Jobba Trade Technologies, Inc. Thumbnail
Software • Professional Services • Productivity • Information Technology • Cloud
Chicago, IL
45 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account