HPC/AI Systems Administrator

Posted Yesterday
Be an Early Applicant
Hiring Remotely in United States
Remote
Senior level
Hardware • Software
The Role
Administer, provision, and maintain HPC/AI compute, storage, networking, and software stacks. Develop automation for provisioning, configuration management, and monitoring. Install, configure, and optimize job schedulers (e.g., Slurm), deploy MPI and containerized HPC applications, perform benchmarking, tuning, capacity planning, security patching, troubleshooting, and vendor coordination while supporting researchers and documenting procedures.
Summary Generated by Built In
Description

NextSilicon is revolutionizing high-performance computing. Our innovative coprocessor technology dramatically accelerates supercomputers, propelling them into a new era. Our software-defined hardware architecture empowers HPC/AI to deliver groundbreaking discoveries across all areas of advanced research. We're seeking a dynamic and results-oriented HPC/AI Systems Administrator to join our team.

At NextSilicon, everything we do is guided by three core values:

  • Professionalism: We strive for exceptional results through professionalism and unwavering dedication to quality and performance. 
  • Unity: Collaboration is key to success. That's why we foster a work environment where every employee can feel valued and heard. 
  • Impact: We're passionate about developing technologies that make a meaningful impact on industries, communities, and individuals worldwide.

Join our Field Deployment & Systems team as an HPC/AI Systems Administrator.

As an HPC/AI Systems Administrator at NextSilicon, you will be central to sustaining the successful operation of HPC/AI systems. You will stand-up and maintain HPC/AI hardware and software resources. You will tune and configure systems for high-quality benchmarking efforts. You will ensure that the health and accessibility of the HPC/AI systems is top-notch via cluster management tools and capacity planning efforts.

This is a highly technical, execution-focused individual contributor role with no people management or leadership responsibilities at this time.

Location: Hybrid in either our Austin, TX or Minneapolis, MN offices preferred but Remote considered for exceptional candidates.

Requirements
  • Bachelor’s degree in engineering, mathematics, computer science, related field, or equivalent experience. Advanced degree is a plus.
  • 5-10+ years of experience with HPC/AI system administration.
  • Deep understanding of HPC & AI technologies and software ecosystems
  • Experience in a fast-paced, entrepreneurial environment is a plus
  • Ability to travel within the USA approx. 4 times per year
  • US citizenship with eligibility to visit US government research facilities
Responsibilities
  • Administer, install, monitor, and maintain HPC/AI systems, including compute nodes, storage, networking, and software stacks.
  • Develop and maintain automation tools for system provisioning, configuration management, and monitoring. 
  • Install, configure, and optimize job scheduling and resource management tools (e.g., Slurm). 
  • Assist in system security, patch management, and troubleshooting operational issues. 
  • Contribute to performance benchmarking, system tuning, and capacity planning. 
  • Deploy and maintain commonly used HPC/AI applications, software stacks, and technologies (e.g., MPI, containers, spack, modules)
  • Document system administration procedures and contribute to knowledge-sharing initiatives.
  • Support researchers by providing technical expertise and resolving escalated support tickets.
  • Participate in vendor coordination, system procurement, and hardware/software lifecycle management.

Skills Required

  • Bachelor's degree in engineering, mathematics, computer science, related field, or equivalent experience
  • Advanced degree
  • 5-10+ years of experience with HPC/AI system administration
  • Deep understanding of HPC & AI technologies and software ecosystems
  • Ability to travel within the USA approximately 4 times per year
  • US citizenship with eligibility to visit US government research facilities
  • Experience installing, configuring, and optimizing job scheduling/resource management tools (e.g., Slurm)
  • Experience deploying and maintaining HPC/AI applications and software stacks (e.g., MPI, containers, spack, modules)
  • Experience developing and maintaining automation for system provisioning, configuration management, and monitoring
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Tel Aviv-Yafo
280 Employees
Year Founded: 2017

What We Do

We believe in a smarter future and want to create new opportunities for innovation. In order to achieve this, we’re rethinking compute architectures for the future of computer processing.

Similar Jobs

General Dynamics Information Technology Logo General Dynamics Information Technology

Senior Systems Engineer

Aerospace • Information Technology • Professional Services • Security • Software
Remote
United States
21625 Employees
136K-184K Annually

Eve Logo Eve

Consultant

Legal Tech • Software • Generative AI
Easy Apply
Remote or Hybrid
United States
180 Employees
90K-110K Annually

Lob Logo Lob

Account Executive

Logistics • Marketing Tech • Software
Easy Apply
Remote
United States
125 Employees
71K-85K Annually

Lob Logo Lob

Customer Success Manager

Logistics • Marketing Tech • Software
Easy Apply
Remote
United States
125 Employees
7K-80K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account