Senior Infrastructure Engineer

Posted 2 Days Ago
Be an Early Applicant
Hiring Remotely in Québec, QC, CAN
Remote
Senior level
Artificial Intelligence • Blockchain • Cloud • Software
The AI Factory. Accelerating the Future.
The Role
Own the design, deployment, operation, and improvement of globally scaled OpenStack and Kubernetes infrastructure for GPU workloads. Build automated infrastructure using infrastructure-as-code and GitOps, optimize GPU scheduling, implement monitoring and alerting, lead incident response, and maintain security controls including RBAC, network policies, and tenant isolation. The role also requires hands-on Linux administration, server and rack assembly, data center work, and collaboration across technical and customer-facing teams.
Summary Generated by Built In

ABOUT NEXGEN CLOUD:

NexGen Cloud is the company behind Hyperstack, a full-stack AI cloud serving tens of thousands of customers from AI researchers to enterprises running the world's most compute-intensive workloads. We deliver on-demand and private GPU infrastructure to teams who treat performance as a requirement, not a feature.

We're a tight-knit, fast-moving team working at the cutting edge of AI cloud infrastructure. We practice what we preach, equipping our people with AI at every level so we can solve harder problems, ship faster, and keep raising the bar for what enterprise GPU infrastructure looks like.

THE ROLE: Senior Infrastructure Engineer

This role exists because our platform is scaling quickly — and complexity comes with it. As we expand our OpenStack and Kubernetes environments globally, we need engineers who can take real ownership of how the platform is designed, operated, and improved. You'll have direct ownership over business-critical infrastructure that impacts performance, reliability, and customer experience.

This is not a maintenance role. If you like solving hard problems, owning systems end-to-end, and seeing the impact of your work immediately — you'll enjoy this.

WHAT YOU'LL BE DOING:

Rather than a long checklist, here's what success in this role looks like:

  • Own the design, deployment, and operation of OpenStack and Kubernetes environments — ensuring platform performance, scalability, and resilience for GPU workloads
  • Build and improve infrastructure using infrastructure-as-code and GitOps practices, driving automation across provisioning, deployment, and operational workflows
  • Optimise GPU workload scheduling using Kubernetes and NVIDIA tooling, and implement monitoring, logging, and alerting to ensure platform stability
  • Lead incident response and drive continuous improvement of reliability across the platform
  • Maintain strong security controls across infrastructure and container layers — RBAC, network policies, and tenant isolation
  • Work closely with Platform, DevOps, AI, Product, and Support teams to align infrastructure capabilities with customer and platform requirements
ABOUT YOU:

We're more interested in how you think and work than in a perfect CV. You'll likely bring a combination of the following:

Essential

  • Extensive hands-on Linux systems administration skills and knowledge — genuine depth, not surface-level familiarity
  • Strong, proven experience building servers and racks — you've physically assembled, cabled, and commissioned hardware, not just specified or overseen it
  • Direct hands-on experience physically working in data centres — you've personally stacked and racked hardware on-site
  • A willingness and ability to travel to Quebec sites as required
  • A solid understanding of networking and storage systems

Nice to Have

  • Experience installing, racking, and configuring GPU hardware specifically, ideally including NVIDIA platforms
  • Production experience running OpenStack and/or Kubernetes at scale
  • Experience with infrastructure automation, CI/CD, and Git-based workflows
  • Broader exposure to HPC or large-scale compute environments
  • Contributions to open-source projects
WHAT WE OFFER:
  • Competitive salary and annual discretionary bonus scheme
  • Employee wellbeing benefits
  • 25 days of holiday, plus public holidays
  • Flexible working arrangements (remote or hybrid, depending on role and location)
  • Real ownership and autonomy, with the trust to take initiative and experiment
  • The opportunity to make a visible, meaningful impact as we scale
  • Clear career progression and growth opportunities in a fast-growing company
  • A collaborative, international culture built on trust, transparency, and ownership
  • The chance to help shape NexGen Cloud's team, culture, and future alongside ambitious, mission-driven colleagues
MORE INFORMATION

Head over to our NexGen Cloud careers page to view current openings and follow us on LinkedIn and X to learn more about our journey, newest releases and hear exciting news in the neocloud space.



Skills Required

  • Extensive hands-on Linux systems administration experience
  • Strong hands-on experience building servers and racks, including physical assembly, cabling, and commissioning
  • Direct hands-on experience working physically in data centers and racking hardware on-site
  • Willingness and ability to travel to Quebec sites as required
  • Solid understanding of networking and storage systems
  • Experience installing, racking, and configuring GPU hardware, ideally NVIDIA platforms
  • Production experience running OpenStack and/or Kubernetes at scale
  • Experience with infrastructure automation, CI/CD, and Git-based workflows
  • Exposure to HPC or large-scale compute environments
  • Contributions to open-source projects
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
92 Employees
Year Founded: 2020

What We Do

NexGen Cloud, founded in 2020, is a global leader in sustainable AI Cloud solutions, offering Data Sovereignty to its clients. Powered by 100% renewable energy, NexGen Cloud's expertise is rooted in the deployment and management of advanced AI infrastructure and cloud services. NexGen Cloud’s solutions are tailored to meet the diverse needs of AI enterprises and practitioners through a suite of specialised products and services, including the AI Supercloud for large-scale bespoke environments, and Hyperstack, a service for on-demand enterprise GPU access.

Similar Jobs

Samsara Logo Samsara

Senior Software Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
Canada
4000 Employees
143K-185K Annually

Planet Logo Planet

Senior Software Engineer

Aerospace • Big Data • Greentech • Hardware • Social Impact
Remote
Canada
747 Employees
117K-146K Annually

Affirm Logo Affirm

Senior Software Engineer

Big Data • Fintech • Mobile • Payments • Financial Services
Easy Apply
Remote
Canada
2200 Employees
153K-213K Annually

Very Good Security Logo Very Good Security

Infrastructure Engineer

Security • Database • Cybersecurity
Remote
2 Locations
223 Employees
185K-290K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account