Senior Lead Network Engineer - HPC

Posted 6 Days Ago
Be an Early Applicant
Cairo, EGY
In-Office
Senior level
Software
The Role
Lead architecture, operations, and performance tuning of Ethernet and InfiniBand fabrics in an HPC/AI data center. Own fabric design, lifecycle management, monitoring, incident response, diagnostics, automation, and documentation. Collaborate with platform, storage, and software teams, participate in on-call rotation, and lead design reviews and knowledge-sharing to ensure low-latency, high-bandwidth network performance.
Summary Generated by Built In
Overview

Integrant is seeking a Senior Lead Network & Infrastructure Engineer with 14+ years of experience to provide technical leadership in a fast-paced, complex HPC and AI environment. This is a multi-disciplinary senior role: the core is deep network engineering across Ethernet and InfiniBand fabrics, ideally complemented by hands-on experience in high-performance storage and/or Linux systems operations (SysOps). The Senior Lead owns fabric architecture and performance, acts as the highest technical escalation point, and works across engineering, platform, storage, and client teams. Adaptability, ownership, and clear communication are key to success in an environment where network performance is critical.

Responsibilities
  • Develop network configurations and architectures
  • Operate, maintain, and support Ethernet and InfiniBand networks in a high-performance computing (HPC) and AI environment.
  • Perform ongoing maintenance, upgrades, and lifecycle management of network equipment.
  • Monitor network health, performance, and capacity to ensure reliable, low-latency data flow.
  • Respond to and resolve network and server-related incidents in a timely manner.
  • Run hardware diagnostics and coordinate replacement of failing network components.
  • Support and maintain Linux-based HPC and AI platforms across a wide range of technologies.
  • Collaborate with senior network engineers, software teams, and platform teams on network efficiency, reliability, and security.
  • Assist with configuration, deployment, and operational support of InfiniBand and Ethernet fabrics.
  • Develop and maintain operational documentation, including configuration examples, build guides, and best practices.
  • Support on-site staff during hardware updates, card replacements, and infrastructure changes.
  • Stay current with advancements in data center networking, HPC interconnects, and AI infrastructure technologies.
  • Work within the client ticketing / IT service management system (e.g., TopDesk) to manage incidents and service requests to SLA.
  • Build and maintain automation and tooling (scripting, monitoring integrations, infrastructure-as-code) to improve operational efficiency.
  • Collaborate with software, platform, storage, and client teams on efficiency, reliability, and security.
  • Own the quality of operational documentation: configuration examples, build guides, runbooks, and best practices.
  • Lead design reviews and knowledge-sharing.
Work Conditions

Participate in a weekly on-call rotation and respond to network and infrastructure issues after hours when required.


Requirements
  • 14+ years of hands-on experience supporting enterprise or data center-scale networks.
  • Experience working in HPC, AI/ML, or performance-sensitive environments.
  • Practical experience administering InfiniBand (Mellanox/NVIDIA) and Ethernet (Cumulus, SONiC) networks.
  • Strong understanding of data center networking concepts, including servers, storage, and high-speed interconnects.
  • Solid knowledge of Layer 2 and Layer 3 networking, including routing and switching fundamentals.
  • Installing, monitoring, and maintaining very large-scale data center networks.
  • Low-latency, high-bandwidth fabric support and performance tuning for distributed compute and GPU workloads.
  • VXLAN/EVPN architectures and routing protocols such as BGP and OSPF.
  • Exposure to communication libraries such as NCCL, UCX, and MPI.
  • Network management and monitoring tools: UFM, OpenSM, NetQ, or similar.
  • Ability to troubleshoot and resolve network issues in complex, distributed environments.
  • Strong documentation and communication skills.
  • Proven ability to work effectively as part of a team and provide operational support.
Preferred (Multi-Skill) Qualifications
  • Storage: Hands-on experience with high-performance / parallel storage environments (e.g., Lustre, GPFS/Spectrum Scale, BeeGFS, Ceph, NVMe-oF), including storage networking and I/O performance troubleshooting.
  • SysOps / Linux systems: Production Linux systems administration at scale — provisioning, configuration management (Ansible/Salt), kernel/network stack tuning, schedulers (Slurm), containerization.

Benefits
  • Salary paid in USD
  • Six-month career advancing opportunities
  • Supportive and friendly work environment
  • Premium medical insurance [employee +family]
  • English language development courses
  • Interest-free loans paid over 2.5 years
  • Technical development courses
  • Employment referral program
  • Premium location in Maadi
  • Social insurance

Skills Required

  • 14+ years hands-on experience supporting enterprise or data center-scale networks
  • Experience working in HPC, AI/ML, or performance-sensitive environments
  • Practical experience administering InfiniBand (Mellanox/NVIDIA) networks
  • Practical experience administering Ethernet networks (Cumulus, SONiC)
  • Strong understanding of data center networking, servers, storage, and high-speed interconnects
  • Solid knowledge of Layer 2 and Layer 3 networking, routing and switching fundamentals
  • Experience installing, monitoring, and maintaining very large-scale data center networks
  • Low-latency, high-bandwidth fabric support and performance tuning for distributed compute and GPU workloads
  • Experience with VXLAN/EVPN architectures and routing protocols such as BGP and OSPF
  • Exposure to communication libraries such as NCCL, UCX, and MPI
  • Familiarity with network management and monitoring tools (UFM, OpenSM, NetQ, or similar)
  • Ability to troubleshoot and resolve network issues in complex, distributed environments
  • Strong documentation and communication skills
  • Proven ability to work effectively as part of a team and provide operational support
  • Participate in weekly on-call rotation and respond to after-hours incidents
  • Hands-on experience with high-performance / parallel storage environments (Lustre, GPFS/Spectrum Scale, BeeGFS, Ceph, NVMe-oF)
  • Production Linux systems administration at scale; provisioning and configuration management (Ansible, Salt)
  • Kernel/network stack tuning, scheduler experience (Slurm), and containerization experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Diego, CA
263 Employees
Year Founded: 1992

What We Do

Integrant, Inc. is a custom software development company focused on providing tailor made software solutions to fit your needs to a tee. We strive to uncover your pain points and identify how our team can seamlessly integrate with you and your business for a one-team approach. Our guiding principle is to always do the right thing for our customers and employees. Some days this means happy news of a “hit on the mark” demo, successful launch, or challenging problem solved. Other days this means making hard decisions, asking tough questions, or working more than we planned. Every day, it means doing our best to provide the highest quality service to each of our customers. We do that by investing our people in you and inspiring a people-to-people connection so when we say, “we share your goals,” we truly mean it.

Similar Jobs

Pfizer Logo Pfizer

Regulatory Data Coordinator

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
2 Locations
121990 Employees

Pfizer Logo Pfizer

Team Lead

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
2 Locations
121990 Employees

Pfizer Logo Pfizer

LOM - International Associate

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote or Hybrid
Cairo, EGY
121990 Employees

Mastercard Logo Mastercard

Director Specialist Sales - Egypt

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Cairo, EGY
38800 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account