HPC Network Architect

Posted One Month Ago
Be an Early Applicant
Dallas, TX, USA
In-Office
Entry level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
The Role
Designs, deploys, and optimizes high-performance networking architectures for HPC, AI/ML, and data-intensive workloads. Advises customers throughout solution lifecycles, leads proof-of-concepts and benchmarking, tunes network performance, and integrates compute, storage, orchestration, and security layers. The role develops observability frameworks, collaborates with engineering and vendors, influences product roadmaps, and presents architectures to technical and executive stakeholders.
Summary Generated by Built In

THE COMPANY 

NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients' research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision-making, accelerating discovery and driving faster innovation. 

 

THE POSITION 

The Architecture team's mission is to build a sound, coherent architecture for the HPC platform. The team takes in product requirements and operational constraints, incorporates validated new technologies and produces reference architectures ready for implementation and rollout. The team is comprised of senior architects with subject matter expertise spanning HPC compute, storage, networking, and orchestration. 

As the HPC Network Architect, you will own the architecture of highly scalable, highly available, and globally distributed networks that support large-scale research and AI/ML workloads. You will assess existing and emerging network technologies against real business and product needs, and turn architectural decisions into clear direction that engineering teams can execute against with confidence. 

You will work closely with Storage, Compute, and Orchestration architects, along with engineering teams and hardware vendors, to ensure every networking decision fits into a coherent, future-proof HPC platform. You will also establish a data-driven approach to continuously evaluating operational needs and constraints, proactively identifying opportunities to improve performance and resiliency. 

Success in this role looks like bringing rigor and clarity to some of the most demanding networking challenges in HPC, while staying ahead of a fast-moving technology landscape and mentoring engineering teams along the way. 

 

RESPONSIBILITIES 

  • Design the architecture of highly scalable, highly available, and globally distributed networks to support large-scale research workloads. 
  • Assess existing and emerging networking technologies against business and product needs. 
  • Guide and drive execution by bringing clarity of the architecture to engineering teams and ensuring high-quality engineering standards are followed. 
  • Collaborate with storage, compute, and orchestration architects to propose coherent, complete, and future-proof HPC solutions. 
  • Establish a data-driven approach to continuously evaluate operational needs and constraints and proactively propose improvements. 
  • Define and evangelize network architecture standards and best practices across engineering teams. 
  • Stay current with emerging technologies and approaches and apply new knowledge across disciplines. 

 

REQUIREMENTS 

  • Bachelor's or Master's degree in Computer Science, Engineering, Physics, or a related technical field, with 10-12 years of experience. 
  • Strong working knowledge of network technologies and architectures, including Ethernet and InfiniBand, with hands-on experience applying them in an HPC context. 
  • Hands-on experience with high-speed fabric solutions, particularly InfiniBand and NVLink, and ideally Ethernet (RoCE) and Omni-Path. 
  • Experience with network offloading and hardware acceleration technologies and their role in large-scale / exascale systems. 
  • Solid grasp of packet switching, routing algorithms, flow control, and congestion management, with the judgment to apply the right solution for optimal network performance. 
  • Solid grasp of networking security and network virtualization, and their integration with IaaS and Kubernetes. 
  • Hands-on lab experience running benchmarks and test jobs on pre-production / unproven hardware in a testing environment. 
  • Experience designing or contributing to HPC clusters and parallel computing environments, with proficiency in Linux kernel tuning, system-level optimization, or performance profiling. 
  • Experience with emerging network trends such as CXL, PCIe Gen 6, or DPUs is a plus. 

It is impossible to list every requirement for, or responsibility of, any position.  Similarly, we cannot identify all the skills a position may require since job responsibilities and the Company’s needs may change over time.  Therefore, the above job description is not comprehensive or exhaustive.  The Company reserves the right to adjust, add to or eliminate any aspect of the above description.  The Company also retains the right to require all employees to undertake additional or different job responsibilities when necessary to meet business needs.

Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future.

Benefits & Perks:

  • Company-Paid Lunch Stipend: Lunch is provided via GrubHub

  • Company-Paid Benefits: 100% Employer-Paid Medical in our High Deductible Health Plan, Dental and Vision benefits for employees and their families, 16 weeks of Paid Parental Leave, Employee Assistance Program, Life insurance, Short-Term Disability and Long-Term Disability

  • 401(k): Company will match 100% of your contributions up to 6%

  • Optional Employee-Paid Benefits: Medical insurance in our PPO plan and a variety of other benefits such as Health Savings Accounts (with Company Contribution!), Flexible Spending Accounts, Supplemental Life Insurance, Wellhub and more.

  • Time Off:  25 days of Paid Time Off plus 12 company holidays

EQUAL OPPORTUNITY EMPLOYER

NORTHMARK STRATEGIES LLC IS AN EQUAL EMPLOYMENT OPPORTUNITY EMPLOYER. THE COMPANY'S POLICY IS NOT TO DISCRIMINATE AGAINST ANY APPLICANT OR EMPLOYEE BASED ON RACE, COLOR, RELIGION, NATIONAL ORIGIN, GENDER, AGE, SEXUAL ORIENTATION, GENDER IDENTITY OR EXPRESSION, MARITAL STATUS, MENTAL OR PHYSICAL DISABILITY, AND GENETIC INFORMATION, OR ANY OTHER BASIS PROTECTED BY APPLICABLE LAW. THE FIRM ALSO PROHIBITS HARASSMENT OF APPLICANTS OR EMPLOYEES BASED ON ANY OF THESE PROTECTED CATEGORIES.

Skills Required

  • Experience in HPC networking solution architecture, systems design, or data center network engineering
  • Expertise in InfiniBand and RoCE deployment and tuning at scale
  • Experience designing and implementing large-scale Ethernet networks using BGP, OSPF, EVPN, and VXLAN
  • Understanding of GPU communication frameworks including MPI and NCCL
  • Proficiency with Linux-based environments and scripting using Python, Bash, or PowerShell
  • Experience supporting multi-vendor environments and evaluating networking platforms
  • Ability to translate networking requirements into solution architectures and present them to customers
  • Strong customer-facing communication skills with technical and executive stakeholders
  • Experience delivering HPC or AI/ML workloads across large-scale, low-latency networks
  • Familiarity with Multus, Cilium, or NVIDIA CNI plugins for HPC or Kubernetes environments
  • Experience with Terraform or Ansible for network provisioning and infrastructure as code
  • Experience collaborating with vendors and influencing feature roadmaps
  • Contributions to open-source HPC networking or infrastructure projects
  • Bachelor's or Master's degree in Computer Science, Networking, Engineering, or a related technical field
  • Relevant certifications such as Cisco CCNP or CCIE, Juniper JNCIP, AWS Advanced Networking Specialty, or Red Hat RHCE
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
157 Employees

What We Do

NorthMark Strategies is a strategic capital firm that combines investment capital with engineering and technology to build enduring businesses. The firm operates a High-Performance Computing platform and supports simulation, AI/ML-enabled engineering and data-driven design to accelerate portfolio companies. NorthMark deploys capital, operates complex businesses, and builds infrastructure (including compute and cloud services) to drive long‑term innovation and operational outcomes.

Similar Jobs

In-Office
2 Locations
121228 Employees
175K-230K Annually

CrowdStrike Logo CrowdStrike

Senior Site Reliability Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Hybrid
4 Locations
11000 Employees
140K-215K Annually

CoreWeave Logo CoreWeave

Staff Product Quality Engineer

Cloud • Information Technology • Machine Learning
In-Office
6 Locations
1450 Employees
Remote or Hybrid
2 Locations
175633 Employees
65K-125K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account