HPC Orchestration Architect

Posted One Month Ago
Be an Early Applicant
Dallas, TX, USA
In-Office
Expert/Leader
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
The Role
Leads a multidisciplinary HPC Solutions Architecture team delivering secure, scalable compute, storage, networking, Kubernetes, and systems-integration solutions. Oversees customer engagements from requirements discovery through proof of concept, deployment, optimization, and adoption. Builds reference architectures, guides technical reviews, supports HPC and AI/ML workloads, and partners with product and engineering teams on platform roadmaps. Requires extensive HPC and distributed-systems expertise, customer-facing communication, and technical leadership.
Summary Generated by Built In

THE COMPANY 

NorthMark Compute & Cloud (NMC²) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients' research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision-making, accelerating discovery and driving faster innovation. 

 

THE POSITION 

The Architecture team's mission is to build a sound, coherent architecture for the HPC platform. The team takes in product requirements and operational constraints, incorporates validated technologies, and produces reference architectures ready for implementation and rollout. The team is comprised of senior architects with subject matter expertise spanning HPC compute, storage, networking, and orchestration. 

As the HPC Orchestration Architect, you will own the architecture of highly scalable, highly available, and globally distributed systems that support large-scale research workloads, with a focus on compute, virtualization, and Kubernetes-based orchestration. You will assess existing and emerging technologies against real business and product needs and drive continuous improvement in developer velocity across the platform. 

You will work closely with Storage, Networking, and Compute architects, along with engineering teams, to ensure every orchestration decision fits into a coherent, future-proof HPC platform. You will also establish a data-driven approach to continuously evaluating and improving operational needs and constraints. 

Success in this role looks like bringing rigor and clarity to some of the most demanding orchestration and resource-management challenges in HPC, while staying ahead of a fast-moving technology landscape. 

 

RESPONSIBILITIES 

  • Design the architecture of highly scalable, highly available, and globally distributed systems to support large-scale research workloads. 
  • Assess existing and emerging HPC technologies against business and product needs. 
  • Guide and drive execution by bringing clarity of the architecture to engineering teams and ensuring high-quality engineering standards are followed. 
  • Collaborate with storage, networking, and compute architects to propose coherent, complete, and future-proof HPC solutions. 
  • Drive continuous improvement in developer velocity. 
  • Establish a data-driven approach to continuously evaluating and improving operational needs and constraints. 
  • Stay current with emerging technologies and approaches and apply new knowledge across disciplines. 

 

REQUIREMENTS 

  • Bachelor's or Master's degree in Computer Science, Engineering, Physics, or a related technical field, with 10-12 years of experience. 
  • Strong working knowledge of compute architectures, including CPUs, GPUs, and domain-specific accelerators, and their role in large-scale / exascale systems. 
  • Solid grasp of virtualization technologies, containers, container orchestration, and Kubernetes resource management, with demonstrated understanding of scale, Reliability/Availability/Serviceability, and security consideration 
  • Hands-on experience building, deploying, and scaling Kubernetes clusters. 
  • A working understanding of how resource placement tradeoffs impact performance. 
  • Experience designing or contributing to HPC clusters and parallel computing environments, with proficiency in Linux kernel tuning, system-level optimization, or performance profiling. 
  • Familiarity with Kubernetes CNI and CSI plugins. 
  • Experience working directly with clients to capture technical requirements and deliver tailored, scalable system designs across AI/ML and scientific computing. 

It is impossible to list every requirement for, or responsibility of, any position.  Similarly, we cannot identify all the skills a position may require since job responsibilities and the Company’s needs may change over time.  Therefore, the above job description is not comprehensive or exhaustive.  The Company reserves the right to adjust, add to or eliminate any aspect of the above description.  The Company also retains the right to require all employees to undertake additional or different job responsibilities when necessary to meet business needs.

Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future.

Benefits & Perks:

  • Company-Paid Lunch Stipend: Lunch is provided via GrubHub

  • Company-Paid Benefits: 100% Employer-Paid Medical in our High Deductible Health Plan, Dental and Vision benefits for employees and their families, 16 weeks of Paid Parental Leave, Employee Assistance Program, Life insurance, Short-Term Disability and Long-Term Disability

  • 401(k): Company will match 100% of your contributions up to 6%

  • Optional Employee-Paid Benefits: Medical insurance in our PPO plan and a variety of other benefits such as Health Savings Accounts (with Company Contribution!), Flexible Spending Accounts, Supplemental Life Insurance, Wellhub and more.

  • Time Off:  25 days of Paid Time Off plus 12 company holidays

EQUAL OPPORTUNITY EMPLOYER

NORTHMARK STRATEGIES LLC IS AN EQUAL EMPLOYMENT OPPORTUNITY EMPLOYER. THE COMPANY'S POLICY IS NOT TO DISCRIMINATE AGAINST ANY APPLICANT OR EMPLOYEE BASED ON RACE, COLOR, RELIGION, NATIONAL ORIGIN, GENDER, AGE, SEXUAL ORIENTATION, GENDER IDENTITY OR EXPRESSION, MARITAL STATUS, MENTAL OR PHYSICAL DISABILITY, AND GENETIC INFORMATION, OR ANY OTHER BASIS PROTECTED BY APPLICABLE LAW. THE FIRM ALSO PROHIBITS HARASSMENT OF APPLICANTS OR EMPLOYEES BASED ON ANY OF THESE PROTECTED CATEGORIES.

Skills Required

  • Bachelor's degree or equivalent experience
  • 10+ years of experience in HPC, solutions architecture, or large-scale systems design
  • 3+ years of technical team leadership managing architects or engineers across domains
  • Leadership experience managing multidisciplinary architecture or engineering teams in HPC, cloud, or large-scale distributed systems
  • Strong architectural expertise across compute, storage, networking, Kubernetes, and security
  • Hands-on experience with GPU acceleration, CUDA, NVIDIA technologies, workload schedulers, and distributed storage
  • Experience designing secure and compliant architectures involving identity management, encryption, and regulatory frameworks
  • Customer-facing communication and technical requirements gathering experience
  • Ability to design scalable HPC systems for AI/ML, scientific computing, simulation, and data-intensive workloads
  • Track record of driving solution adoption and measurable customer success
  • Experience delivering or supporting HPC or AI/ML workloads at scale
  • Familiarity with Kafka, Spark, or similar data and analytics platforms
  • Knowledge of CI/CD, GitLab, Jenkins, Terraform, and Ansible
  • Experience leading customer workshops, technical reviews, or industry presentations
  • Knowledge of next-generation GPUs, InfiniBand, RDMA, and HPC container runtimes
  • Advanced degree in Computer Science, Engineering, Physics, or a related technical field
  • Relevant AWS, Azure, GCP, Cisco, Red Hat, or Kubernetes certifications
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
157 Employees

What We Do

NorthMark Strategies is a strategic capital firm that combines investment capital with engineering and technology to build enduring businesses. The firm operates a High-Performance Computing platform and supports simulation, AI/ML-enabled engineering and data-driven design to accelerate portfolio companies. NorthMark deploys capital, operates complex businesses, and builds infrastructure (including compute and cloud services) to drive long‑term innovation and operational outcomes.

Similar Jobs

ServiceNow Logo ServiceNow

Principal Customer Success Executive-Financial Services-ServiceNow

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
Austin, TX, USA
29000 Employees

Samsara Logo Samsara

Data Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
99K-167K Annually

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Katy, TX, USA
16000 Employees
15-20 Hourly

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Katy, TX, USA
16000 Employees
15-20 Hourly

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account