HPC Network Architect

Posted 11 Days Ago
Be an Early Applicant
Dallas, TX, USA
In-Office
Entry level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
The Role
Designs, deploys, and optimizes high-performance networking architectures for HPC, AI/ML, and data-intensive workloads. Advises customers throughout solution lifecycles, leads proof-of-concepts and benchmarking, tunes network performance, and integrates compute, storage, orchestration, and security layers. The role develops observability frameworks, collaborates with engineering and vendors, influences product roadmaps, and presents architectures to technical and executive stakeholders.
Summary Generated by Built In

The Company

NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients' research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision-making, accelerating discovery and driving faster innovation.

The Position

As an HPC Network Solutions Architect, you will design, integrate, and optimize high-performance networking architectures that form the backbone of HPC, AI/ML, and data-intensive workloads. You will act as a trusted advisor to customers, guiding them across the entire solution lifecycle — from requirements gathering and design, through proof-of-concept and deployment, to optimization and long-term adoption.

This is a customer-facing, technically focused role. You will collaborate closely with customers to align low-latency, high-bandwidth networking designs with their workload requirements, while also working with internal engineering and product teams to influence roadmap priorities. Your role will bridge the gap between cutting-edge networking technologies (InfiniBand, RoCE, EVPN, VXLAN) and real-world HPC adoption at scale.

This position offers the opportunity to shape the future of HPC networking, deliver measurable impact for customers, and influence vendor ecosystems by incorporating emerging innovations into enterprise-ready solutions.

Responsibilities

  • Act as the primary networking SME for customers adopting or scaling HPC environments.

  • Partner with customers to capture network performance goals, scalability requirements, and integration constraints.

  • Design and document end-to-end HPC network architectures, including Ethernet, InfiniBand, RoCE, EVPN, and VXLAN fabrics.

  • Lead proof-of-concept and benchmarking engagements, validating low-latency and high-throughput designs against workload requirements.

  • Optimize multi-vendor, multi-protocol data center and HPC interconnects, addressing scaling challenges such as data gravity and throughput bottlenecks.

  • Define integration strategies across compute, storage, orchestration, and security layers to deliver resilient, workload-aware solutions.

  • Conduct network performance assessments and tuning, identifying bottlenecks and recommending enhancements.

  • Build observability frameworks for HPC networks at scale using tools like Prometheus, Grafana, and vendor telemetry.

  • Collaborate with engineering, product, and operations teams to refine architecture blueprints and ensure consistent delivery.

  • Partner with ecosystem vendors (e.g., NVIDIA, Mellanox, Cisco, Arista) to integrate cutting-edge features and influence roadmap evolution.

  • Stay current with emerging HPC networking technologies and protocols, providing future insight to customers on adoption strategies.

  • Represent the organization at customer design sessions, workshops, and industry events, building strong technical relationships.

Requirements

  • Demonstrated experience in HPC networking solution architecture, systems design, or data center network engineering.

  • Strong expertise in InfiniBand and RoCE protocols, including deployment and tuning at scale.

  • Hands-on experience designing and implementing large-scale Ethernet networks, including BGP, OSPF, EVPN, and VXLAN.

  • Deep understanding of GPU communication frameworks such as MPI and NCCL, and their integration with HPC interconnects.

  • Proficiency with Linux-based environments and scripting (e.g., Python, Bash, PowerShell) for automation.

  • Experience supporting multi-vendor environments and evaluating new networking platforms.

  • Ability to translate complex networking requirements into clear solution architectures and present them effectively to customers.

  • Strong customer-facing communication skills, including the ability to engage executives and technical stakeholders alike.

Preferred Experience

  • Experience delivering HPC or AI/ML workloads across large-scale, low-latency network infrastructures.

  • Familiarity with CNI plugins (Multus, Cilium, NVIDIA CNI) for HPC/Kubernetes environments.

  • Exposure to automation and infrastructure-as-code practices for network provisioning (Terraform, Ansible).

  • Experience in vendor collaboration, including influencing feature roadmaps and participating in joint evaluations.

  • Contributions to open-source HPC networking or infrastructure projects.

  • Bachelor’s or Master’s degree in Computer Science, Networking, Engineering, or a related technical field.

  • Relevant Networking and systems certifications such as Cisco CCNP/CCIE, Juniper JNCIP, AWS Advanced Networking Specialty, or Red Hat RHCE.

It is impossible to list every requirement for, or responsibility of, any position.  Similarly, we cannot identify all the skills a position may require since job responsibilities and the Company’s needs may change over time.  Therefore, the above job description is not comprehensive or exhaustive.  The Company reserves the right to adjust, add to or eliminate any aspect of the above description.  The Company also retains the right to require all employees to undertake additional or different job responsibilities when necessary to meet business needs.

Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future.

Benefits & Perks:

  • Company-Paid Lunch Stipend: Lunch is provided via GrubHub

  • Company-Paid Benefits: 100% Employer-Paid Medical in our High Deductible Health Plan, Dental and Vision benefits for employees and their families, 16 weeks of Paid Parental Leave, Employee Assistance Program, Life insurance, Short-Term Disability and Long-Term Disability

  • 401(k): Company will match 100% of your contributions up to 6%

  • Optional Employee-Paid Benefits: Medical insurance in our PPO plan and a variety of other benefits such as Health Savings Accounts (with Company Contribution!), Flexible Spending Accounts, Supplemental Life Insurance, Wellhub and more.

  • Time Off:  25 days of Paid Time Off plus 12 company holidays

EQUAL OPPORTUNITY EMPLOYER

NORTHMARK STRATEGIES LLC IS AN EQUAL EMPLOYMENT OPPORTUNITY EMPLOYER. THE COMPANY'S POLICY IS NOT TO DISCRIMINATE AGAINST ANY APPLICANT OR EMPLOYEE BASED ON RACE, COLOR, RELIGION, NATIONAL ORIGIN, GENDER, AGE, SEXUAL ORIENTATION, GENDER IDENTITY OR EXPRESSION, MARITAL STATUS, MENTAL OR PHYSICAL DISABILITY, AND GENETIC INFORMATION, OR ANY OTHER BASIS PROTECTED BY APPLICABLE LAW. THE FIRM ALSO PROHIBITS HARASSMENT OF APPLICANTS OR EMPLOYEES BASED ON ANY OF THESE PROTECTED CATEGORIES.

Skills Required

  • Experience in HPC networking solution architecture, systems design, or data center network engineering
  • Expertise in InfiniBand and RoCE deployment and tuning at scale
  • Experience designing and implementing large-scale Ethernet networks using BGP, OSPF, EVPN, and VXLAN
  • Understanding of GPU communication frameworks including MPI and NCCL
  • Proficiency with Linux-based environments and scripting using Python, Bash, or PowerShell
  • Experience supporting multi-vendor environments and evaluating networking platforms
  • Ability to translate networking requirements into solution architectures and present them to customers
  • Strong customer-facing communication skills with technical and executive stakeholders
  • Experience delivering HPC or AI/ML workloads across large-scale, low-latency networks
  • Familiarity with Multus, Cilium, or NVIDIA CNI plugins for HPC or Kubernetes environments
  • Experience with Terraform or Ansible for network provisioning and infrastructure as code
  • Experience collaborating with vendors and influencing feature roadmaps
  • Contributions to open-source HPC networking or infrastructure projects
  • Bachelor's or Master's degree in Computer Science, Networking, Engineering, or a related technical field
  • Relevant certifications such as Cisco CCNP or CCIE, Juniper JNCIP, AWS Advanced Networking Specialty, or Red Hat RHCE
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
157 Employees

What We Do

NorthMark Strategies is a strategic capital firm that combines investment capital with engineering and technology to build enduring businesses. The firm operates a High-Performance Computing platform and supports simulation, AI/ML-enabled engineering and data-driven design to accelerate portfolio companies. NorthMark deploys capital, operates complex businesses, and builds infrastructure (including compute and cloud services) to drive long‑term innovation and operational outcomes.

Similar Jobs

Hybrid
Plano, TX, USA
289097 Employees

Rapid7 Logo Rapid7

Director, Product Management, Threat Detection

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
United States
2400 Employees
206K-278K Annually

Rapid7 Logo Rapid7

Director, Product Management, Threat Response

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
United States
2400 Employees
206K-278K Annually

Rapid7 Logo Rapid7

Consultant

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
United States
2400 Employees
89K-121K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account