HPC Network Engineer

Posted 2 Days Ago
Be an Early Applicant
Hiring Remotely in Tokyo, JPN
In-Office or Remote
Mid level
Software
The Role
Designs, deploys, configures, and maintains InfiniBand and Ethernet networks for AI/HPC infrastructure. Troubleshoots connectivity, latency, routing, and performance issues; optimizes HCAs, switches, and fabrics; implements network security (FortiGate, VPNs); monitors and tunes performance; collaborates with compute/storage teams; participates in incident response and on-call rotations; documents architectures and automates workflows with Ansible/Terraform.
Summary Generated by Built In
Company Description

About Mirantis

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.

Job Description

We are looking for an HPC Network Engineer to join our global team responsible for managing and supporting high-performance networking environments for large-scale AI infrastructure. You will help ensure the performance, reliability, and security of critical network fabrics, working closely with distributed compute and storage teams to support modern AI workloads.

This is an opportunity to work hands-on with some of the most advanced InfiniBand and Ethernet networking infrastructure in production today, while developing deep expertise in AI infrastructure and high-performance computing (HPC) technologies.

Responsibilities:

  • Support the design, deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructures.

  • Troubleshoot complex network issues, including connectivity, latency, routing, and performance degradation across hybrid environments.

  • Manage and optimize high-performance network components, such as switches, Host Channel Adapters (HCAs), subnet managers, and fabric configurations.

  • Implement, manage, and troubleshoot network security and firewall technologies (e.g., Fortinet solutions like FortiGate, VPNs).

  • Monitor network health, perform performance tuning, and assist in capacity planning for HPC and AI networking systems.

  • Collaborate with compute, storage, and platform teams to seamlessly support HPC and AI workloads.

  • Participate in incident response, on-call support activities, and drive long-term operational improvements.

  • Document network architectures, configurations, and operational procedures, and adopt automation tools (e.g., Ansible, Terraform) to streamline workflows.

Qualifications

  • Proven experience in networking, system engineering, or data center infrastructure roles.

  • Solid understanding of networking fundamentals, including TCP/IP, routing protocols (BGP, OSPF), switching, VLANs, QoS, and network design.

  • Hands-on experience or strong familiarity with Linux operating systems and command-line tools.

  • Effective verbal and written communication skills in English, with the ability to collaborate in a global team environment.

  • Strong analytical and problem-solving skills to diagnose and resolve complex network challenges.

Nice to have:

  • Exposure to or hands-on experience with InfiniBand fabrics and high-performance networking concepts (e.g., NVIDIA/Mellanox).

  • Familiarity with or certifications in enterprise firewall technologies (e.g., Fortinet FCSS/FCNSP, CCNP/CCIE).

  • Experience with large-scale HPC clusters, AI/ML infrastructure, or distributed systems environments.

  • Knowledge of RDMA, MPI, and low-latency networking concepts.

  • Proficiency in scripting languages (e.g., Python, Bash) and Infrastructure-as-Code (IaC) tools for automation.

Additional Information

What does Mirantis offer you?

- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.

It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to [email protected]

By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.

We are a Leader for Container Management in G2 (#2 after AWS)!

Skills Required

  • Proven experience in networking, system engineering, or data center infrastructure roles
  • Solid understanding of networking fundamentals including TCP/IP, routing protocols (BGP, OSPF), switching, VLANs, QoS, and network design
  • Hands-on experience or strong familiarity with Linux operating systems and command-line tools
  • Effective verbal and written communication skills in English and ability to collaborate in a global team
  • Strong analytical and problem-solving skills to diagnose and resolve complex network challenges
  • Exposure to or hands-on experience with InfiniBand fabrics and high-performance networking (NVIDIA/Mellanox)
  • Familiarity with or certifications in enterprise firewall technologies (e.g., Fortinet FCSS/FCNSP, CCNP/CCIE)
  • Experience with large-scale HPC clusters, AI/ML infrastructure, or distributed systems environments
  • Knowledge of RDMA, MPI, and low-latency networking concepts
  • Proficiency in scripting languages (Python, Bash) and Infrastructure-as-Code tools for automation
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Campbell, CA
729 Employees
Year Founded: 1999

What We Do

We are dedicated to helping organizations increase developer productivity and ship code faster on public and private clouds. We provide a ZeroOps experience to remove the stress of managing cloud native infrastructure by combining software and automation tools with our cloud native expertise to deliver the industry's leading secure cloud platforms. Our capabilities allow us to provide a secure and reliable cloud native platform that includes validated FIPS-140-2 Encryption and DISA STIG ready capabilities. Who do we serve? We serve a wide range of industries, building on our extensive customer experience to provide distinct value in specific verticals including Financial Services, Government & Education, Healthcare, Manufacturing, and Telecommunications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Inmarsat, PayPal, Reliance Jio, Societe Generale, Splunk, and S&P Global. Learn more at www.mirantis.com.

Similar Jobs

PureSpectrum Logo PureSpectrum

Sales Director, North Asia

Big Data • Marketing Tech • Sales • Software • Analytics • Big Data Analytics
Remote or Hybrid
Tokyo, JPN
283 Employees
22M-22M Annually

UL Solutions Logo UL Solutions

Senior Sales Executive

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Remote or Hybrid
3 Locations
15000 Employees

MongoDB Logo MongoDB

Regional Vice President, Enterprise, Acquisition

Big Data • Cloud • Software • Database
Easy Apply
Remote or Hybrid
Japan
5550 Employees

GitLab Logo GitLab

Sales Manager

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
Japan
2500 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account