System Administrator

Posted Yesterday
Be an Early Applicant
2 Locations
In-Office
Mid level
Information Technology • Professional Services • Consulting • App development
The Role
Manage and maintain Linux-based HPC CPU/GPU clusters (hardware, schedulers, parallel storage), support researchers with application builds and runtime debugging, troubleshoot CUDA/GPU and scheduler issues, manage configuration (Git, Ansible, MS DevOps), package management, and document processes while reporting progress to technical authority.
Summary Generated by Built In

MGIS is seeking a System Administrator, Level 2, to manage High Performance Computing (HPC) clusters and support the scientists who rely on them. This role blends HPC system administration with hands-on user support — helping researchers install, run, and debug applications on HPC infrastructure so they can focus on their science instead of IT issues.

HPC environments in scope include clustered CPU/GPU systems with job schedulers and attached parallel storage (e.g., Lustre, GPFS).


What you'll be doing

HPC Administrator duties

  • Maintain the HPC cluster — hardware, image management, local networking, scheduler, and backups

  • Troubleshoot environment incidents to ensure a quick return to normal operations

HPC Analyst duties

  • Meet with scientists to evaluate their HPC support requirements

  • Develop task plans to meet researchers' needs, consulting the technical authority for approval

  • Support application builds, installs, and runtime troubleshooting (GNU, Intel, Fortran, Nvidia)

  • Support open-source and commercial software, including Python/Anaconda installs, Bash scripting, build/make tools, EasyBuild, Spack, and MPI implementations (MPICH, OpenMPI, IntelMPI, HPMPI)

  • Assist with compilation and runtime of in-house developed applications

General systems management

  • Manage Linux OS patching schedules and reliability

  • Manage user accounts (creation, deletion) and environment modules

  • Manage configuration via Git, MS DevOps, and Ansible Playbooks

  • Manage RPM/DEB packages and troubleshoot ThinLinc

Troubleshooting & hardware

  • Troubleshoot jobs on schedulers (PBS Pro/Torque, SLURM, SGE)

  • Ensure reliable CUDA installs; troubleshoot GPU failures and CUDA software/driver issues

  • Provide hardware support — memory upgrades, storage arrays, power/network cabling, ILO

Documentation

  • Document every process and task to support enterprise knowledge continuity

  • Submit weekly progress reports to the Technical Authority


Requirements

What we're looking for

  • Solid experience administering Linux-based HPC clusters (CPU/GPU nodes, schedulers, parallel storage)

  • Hands-on experience with job schedulers such as PBS Pro/Torque, SLURM, or SGE

  • Experience troubleshooting CUDA installations, GPU failures, and driver issues

  • Familiarity with scientific computing toolchains — compilers (GNU, Intel), MPI implementations, EasyBuild, and Spack

  • Experience supporting researchers or end-users with application builds and runtime issues

  • Working knowledge of configuration management tools (Git, Ansible, MS DevOps)

  • Comfortable working independently and producing clear technical documentation

  • Eligible to obtain and maintain a Secret-level security clearance


Skills Required

  • Experience administering Linux-based HPC clusters (CPU/GPU nodes, schedulers, parallel storage)
  • Hands-on experience with job schedulers such as PBS Pro/Torque, SLURM, or SGE
  • Experience troubleshooting CUDA installations, GPU failures, and driver issues
  • Familiarity with scientific computing toolchains: compilers (GNU, Intel), MPI implementations, EasyBuild, Spack
  • Experience supporting researchers or end-users with application builds and runtime issues
  • Working knowledge of configuration management tools (Git, Ansible, MS DevOps)
  • Ability to produce clear technical documentation and work independently
  • Eligible to obtain and maintain a Secret-level security clearance
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
11 Employees
Year Founded: 2001

What We Do

MGIS Inc. is an Ottawa-based professional-services and information-technology company serving government and international corporate clients. It provides recruitment and project-planning support, consulting resources, web and mobile application development, geospatial application development, project management, database management, and business analysis. The company builds and maintains projects ranging from short-term initiatives to multi-year, multimillion-dollar ventures, including work supporting Government of Canada clients.

Similar Jobs

Social Discovery Group Logo Social Discovery Group

Senior System Administrator

Artificial Intelligence • Fintech • Machine Learning • Software • App development • Conversational AI • Generative AI
In-Office or Remote
7 Locations

CTC Global Corporation Logo CTC Global Corporation

IT System Administrator

Greentech • Energy • Utilities • Manufacturing
In-Office or Remote
7 Locations
500 Employees
In-Office
Montréal, QC, CAN
21450 Employees

BellatRx Inc. Logo BellatRx Inc.

Administrateur Informatique / System Administrator

Robotics • Industrial • Automation • Manufacturing
In-Office
Pointe-Claire, QC, CAN
81 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account