Group Lead - Infrastructure Services

Posted Yesterday
Be an Early Applicant
Hiring Remotely in United States
Remote
Expert/Leader
Artificial Intelligence • Cloud • Software
The Role
Leads the Infrastructure Services team of Solution Engineers and Solutions Architects delivering AI/HPC infrastructure solutions. Oversees customer engagements from pre-sales architecture and proof-of-concept execution through deployment, benchmarking, and operations. Acts as the senior technical escalation point, partners with Sales, Product, and Engineering, guides performance optimization, develops deployment processes, manages technology partnerships, creates technical content, and recruits and mentors team members.
Summary Generated by Built In
Description

Title: Group Lead - DriveNets Infrastructure Services (DIS)

#LI-Remote

Remote US - East Coast and Central Time zone preferred

About the Company 

DriveNets is a leader in high-scale networking software for AI infrastructure and service providers. The company pioneered a disaggregated networking architecture that transforms the economics of large-scale networks while maximizing performance, utilization, and operational efficiency. DriveNets-powered networks are deployed by global leaders, including AT&T and Comcast, supporting more than 30% of total U.S. internet traffic. DriveNets AI Fabric delivers full-stack networking for AI infrastructures, providing the highest-performance, Ethernet-based alternative to InfiniBand. The solution is deployed by hyperscalers, NeoClouds, and enterprises worldwide. With over $1B raised, DriveNets continues to push the boundaries of modern networking infrastructure.

The Role

DriveNets is seeking a Group Lead for its Infrastructure Services (DIS) team to be a key member of our customer-facing technical organization. Join a dynamic and forward-thinking company at the forefront of AI infrastructure. We leverage advanced technologies to develop innovative solutions that drive efficiency, scalability, and exceptional compute performance. Collaborate with the industry's best as we partner with hyperscalers, emerging NeoClouds, and enterprises building large-scale AI/HPC clusters, shaping the future of disaggregated AI networking and compute infrastructure. Our environment fosters creativity, teamwork, and growth, and offers you the opportunity to make a meaningful impact while leading a high-performing team on groundbreaking deployments.

As Group Lead for DIS, you will manage and develop a team of Solution Engineers and Solutions Architects responsible for designing, deploying, and optimizing DriveNets' AI/HPC infrastructure solutions at customer sites. You will provide technical leadership across the full customer lifecycle - from pre-sales architecture and POC execution through deployment, performance benchmarking, and ongoing operations. You will work cross-functionally with Sales, Product Management, and Engineering to ensure customer success, drive product feedback, and continuously raise the bar for technical delivery quality across the team.

Responsibilities

  • Lead and develop the DIS team - a group of Solution Engineers and Solutions Architects - setting technical direction, managing execution, and fostering a culture of ownership, learning, and customer focus.
  • Oversee end-to-end customer engagement for DIS - from pre-sales technical support and solution architecture through POC planning, deployment execution, and post-deployment operations.
  • Serve as the senior technical escalation point for customer infrastructure challenges, including AI cluster performance issues, networking design trade-offs, and operational reliability concerns.
  • Partner with Sales Account Managers to support business opportunities, lead technical responses to RFP/RFQs, and influence technical decision-makers at the VP and CxO level.
  • Guide the team in conducting performance benchmarking activities - including NCCL/RCCL, RDMA, and LLM benchmarks - and ensure results are translated into actionable product and deployment insights.
  • Work with Product Management and Engineering to funnel customer requirements, field observations, and performance data into the product roadmap and development backlog.
  • Define and drive internal processes for deployment planning, operational readiness, monitoring standards, and technical documentation across the DIS team.
  • Build and maintain relationships with compute, NIC, and storage partners to support joint POCs, reference deployments, and solution validation.
  • Represent DriveNets at industry events and conferences, and contribute to external technical content including white papers, blogs, and design guides.
  • Recruit, mentor, and grow team members, and establish clear performance goals aligned with business objectives.
Requirements

What we need to see:

  • 10+ years of experience in AI/HPC infrastructure, data center networking, or solutions architecture, with at least 2-3 years in a technical leadership or team lead capacity.
  • Hands-on technical depth across both compute infrastructure (GPU clusters, Linux systems, AI workloads) and data center networking (routing, switching, fabric design), with the ability to engage credibly across both disciplines.
  • Proven experience leading customer-facing technical teams through complex deployment and POC cycles in AI/HPC or data center environments.
  • Strong understanding of AI cluster architecture - including GPU platforms (NVIDIA, AMD), RDMA networking, storage connectivity, and the interaction between compute, network, and storage layers.
  • Experience with performance benchmarking methodologies (NCCL/RCCL, RDMA, LLM workloads) and the ability to interpret and act on results at a system level.
  • Demonstrated ability to work cross-functionally with Sales, Product Management, and Engineering teams, translating customer feedback into product improvements and go-to-market strategy.
  • Excellent communication and presentation skills, with proven ability to influence technical and executive stakeholders at customer organizations.
  • Ability to write extensive technical content (white papers, technical briefs, design guides, etc.) for external audiences with a balance of technical accuracy and clear messaging.
  • Ability to travel domestic and international.

Ways to stand out from the crowd:

  • Deep familiarity with AI-relevant infrastructure technologies - InfiniBand, RoCEv2, lossless Ethernet (PFC, ECN), GPU, NIC, DPU, and accelerated computing platforms.
  • Hands-on experience deploying and operating large-scale AI/HPC clusters, including GPU resource scheduling (Slurm, Kubernetes), monitoring (Prometheus, Grafana, DCGM), and operational tooling.
  • Understanding of scale-up (NVLink, UALink) and scale-out (Enhanced Ethernet, UEC, InfiniBand) interconnect technologies and their design trade-offs.
  • Experience with CCL tuning (NCCL/RCCL), GPU environment setup, and performance optimization across large multi-node GPU clusters.
  • Familiarity with AI/ML frameworks (PyTorch, TensorFlow) and how workload characteristics interact with infrastructure design decisions.
  • Proven experience with one or more Tier-1 Clouds (AWS, Azure, GCP, or OCI) or emerging NeoClouds, and cloud-native architectures and software.
  • Background in data center operations fundamentals - networking, cooling, power, and rack-level design.
  • Experience engaging compute, NIC, or storage vendors on joint solution definition, reference architecture development, or benchmarking programs.

EDUCATION

BS/MS/PhD in Electrical/Computer Engineering, Computer Science, Physics, or other Engineering fields, or equivalent experience.

More About DriveNets  

Based in Israel with locations in Romania, US, India and Japan as well as extended teams, DriveNets operations cover more than twelve countries. With recognition by industry analysts and through partnerships with market leaders such as AMD, Broadcom, Dell and others, DriveNets is pushing market momentum, delivering the scale and efficiency that modern AI workloads demand. Visit our website:  https://drivenets.com/company

 If your experience is close but doesn’t fulfil all requirements, please submit your application. DriveNets is on a mission to build a special company comprised of individuals with different backgrounds, perspectives, and experiences. 

 DriveNets is an equal opportunity employer. We do not discriminate based on upon race, religion, national origin, sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with disability, or other applicable legally protected characteristics. 

  

Skills Required

  • 10+ years of experience in AI/HPC infrastructure, data center networking, or solutions architecture
  • 2-3 years of experience in a technical leadership or team lead capacity
  • Hands-on technical expertise in GPU clusters, Linux systems, AI workloads, routing, switching, and fabric design
  • Experience leading customer-facing technical teams through complex AI/HPC or data center deployments and proof-of-concept cycles
  • Strong understanding of GPU platforms, RDMA networking, storage connectivity, and compute-network-storage interactions
  • Experience with NCCL, RCCL, RDMA, and LLM performance benchmarking
  • Experience translating customer feedback into product improvements and go-to-market strategy
  • Excellent communication and presentation skills with the ability to influence technical and executive stakeholders
  • Ability to write extensive external technical content, including white papers, technical briefs, and design guides
  • Ability to travel domestically and internationally
  • BS, MS, or PhD in Electrical Engineering, Computer Engineering, Computer Science, Physics, another engineering field, or equivalent experience
  • Familiarity with InfiniBand, RoCEv2, lossless Ethernet, PFC, ECN, GPU, NIC, DPU, and accelerated computing platforms
  • Experience deploying and operating large-scale AI/HPC clusters with Slurm, Kubernetes, Prometheus, Grafana, or DCGM
  • Understanding of NVLink, UALink, Enhanced Ethernet, UEC, and InfiniBand interconnect technologies
  • Experience with CCL tuning, GPU environment setup, and multi-node GPU cluster optimization
  • Familiarity with PyTorch and TensorFlow
  • Experience with AWS, Azure, GCP, OCI, or emerging NeoClouds
  • Knowledge of data center operations, including networking, cooling, power, and rack-level design
  • Experience engaging compute, NIC, or storage vendors on joint solutions, reference architectures, or benchmarking programs
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Raanana
349 Employees
Year Founded: 2015

What We Do

DriveNets is a rapidly growing software company that has created a radical new way for service providers and hyperscalers to build their networking infrastructure. DriveNets Network Cloud and DriveNets Network Cloud-AI are new innovative networking solutions that apply the cloud architectural approach to high-scale networking. They bring together the scalability of standard Ethernet Clos architecture with the high performance and reliability of service provider networking, delivering optimal networking performance, scale and cost structure for service providers and hyperscalers. Founded by Ido Susan and Hillel Kobrinsky, two successful telco entrepreneurs, DriveNets Network Cloud is the leading open disaggregated networking solution based on cloud-native software running over standard white boxes. Over three funding rounds, DriveNets raised $587 million. Its solutions are used by tens of service providers globally and are in proof-of-concept and lab trials at dozens of operators and hyperscalers, consistently ranking #1 in trials for breadth of capabilities and solution quality. AT&T, the largest backbone in the US, deployed DriveNets Network Cloud across its core network, and DriveNets is currently transporting more than 52% of AT&T’s core network traffic. DriveNets is engaged with over 100 Tier-1 operators and cloud-providers on large projects in North America, Asia and Europe.

Similar Jobs

Samsara Logo Samsara

Senior Data Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
118K-179K Annually

Samsara Logo Samsara

Technical Support

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
71K-96K Annually

Samsara Logo Samsara

Customer Support Specialist

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
43K-58K Annually

Samsara Logo Samsara

Senior Data Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
120K-201K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account