Senior Network Engineer (Amsterdam)

Reposted 5 Days Ago
Be an Early Applicant
Amsterdam, NLD
In-Office
Senior level
Artificial Intelligence • Information Technology
The Role
Design, deploy, and maintain large-scale hybrid HPC/data-center networks. Troubleshoot and analyze network issues, recommend technologies, develop automation, ensure reliability, manage vendor relationships, and lead complex network projects.
Summary Generated by Built In
About the Role

Together AI is looking for a Senior Network Engineer to design, deploy, and operate the global network infrastructure supporting our production services and high-performance AI compute environments.

This is a hands-on engineering role for someone with deep networking expertise who can also troubleshoot across Linux, Kubernetes, automation, and application boundaries. You will work on large-scale, multi-vendor data center networks and help ensure they remain highly available, reliable, scalable, and performant.

The ideal candidate has strong networking fundamentals, experience operating complex networks at scale, and a structured, evidence-based approach to troubleshooting. You should be comfortable owning problems from initial investigation through root cause and resolution, including situations where the issue may extend beyond the network itself.


Requirements

  • 8+ years of professional experience designing, building, and supporting large-scale production data center, cloud, service-provider, or high-performance computing networks (excluding enterprise networks).
  • Deep understanding of TCP/IP and strong experience with technologies such as BGP, OSPF, VXLAN, EVPN, ECMP, and QoS.
  • Experience designing and supporting multi-tenant network environments using technologies such as VRFs, VLANs, overlays, and policy-based segmentation.
  • Hands-on experience deploying and troubleshooting network platforms from vendors such as Arista, Cisco, Juniper, and NVIDIA.
  • Strong troubleshooting skills using tools such as Wireshark, tcpdump, MTR, curl, nmap, and standard Linux networking utilities.
  • Ability to diagnose connectivity, latency, packet-loss, routing, and performance issues across the network, host, and application layers.
  • Experience developing or maintaining network automation using Python, Ansible, or similar tools.
  • Experience working through a Git-based software development lifecycle, including branching, code review, validation, linting, testing, CI/CD, deployment, and rollback.
  • Working knowledge of Kubernetes networking, including pods, services, CNIs, and basic connectivity troubleshooting.
  • Foundational knowledge of RDMA networking and technologies such as RoCE or InfiniBand.
  • Experience with cloud networking in AWS, GCP, or Azure.
  • Strong Linux administration and troubleshooting skills.

Responsibilities

  • Design, deploy, operate, and maintain global, multi-vendor, multi-protocol networks supporting high-performance AI compute infrastructure.
  • Troubleshoot complex network and application-connectivity issues, identify root causes, and drive problems through resolution.
  • Analyze telemetry, packet captures, logs, and performance data to identify network degradation, congestion, packet loss, and capacity constraints.
  • Participate in architecture and design reviews to ensure solutions meet requirements for performance, availability, scalability, security, and operational supportability.
  • Develop and maintain automation, validation, and operational tooling that improves network reliability and reduces manual effort.
  • Evaluate network hardware, software, optics, and emerging technologies for use in production environments.
  • Establish standards and operational best practices for network design, deployment, monitoring, change management, and incident response.
  • Lead projects addressing complex technical challenges and contribute directly to the network engineering roadmap.
  • Partner with infrastructure, systems, security, and application teams to troubleshoot issues that cross traditional ownership boundaries.

Requirements

  • Must have hands-on experience deploying or operating RoCE and/or InfiniBand fabrics.
  • Experience supporting GPU clusters, HPC environments, distributed storage, or other high-bandwidth and latency-sensitive workloads.
  • Understanding of AI training and inference traffic patterns and the demands they place on network infrastructure.
  • Experience operating networks spanning thousands of devices, multiple data centers, and multiple geographic regions.
  • Familiarity with AI-assisted engineering tools and the ability to validate, test, and safely deploy AI-generated automation or code.
About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy  

Skills Required

  • 8+ years professional experience building, managing, and supporting large-scale hybrid data center networks (excluding enterprise networks)
  • High proficiency with TCP/IP networking architecture and technologies such as BGP, OSPF, VXLAN, EVPN, and QoS
  • Experience developing network automation pipelines using Python, Ansible, or other infrastructure automation tools
  • Proficient using Wireshark, tcpdump, nmap, MTR, and curl to diagnose connectivity and latency issues
  • Experience designing and supporting multi-tenant networks
  • Hands-on experience deploying and supporting network devices from Cisco, Arista, Juniper, and Mellanox
  • Experience working with cloud networks such as AWS, GCP, and Azure
  • Solid experience working in and troubleshooting within a Linux environment
  • Must have either RoCE or Infiniband real world experience
  • Evaluate and recommend network technologies, hardware, and software solutions
  • Participate in design reviews to align architecture with business needs and optimize performance, scalability, and reliability
  • Manage relationships with external vendors and partners to test and verify hardware and software selections
  • Establish and implement industry best practices and ensure compliance with IT governance standards
  • Lead projects addressing complex technical challenges and contribute directly to roadmaps
  • Experience with Docker, Kubernetes, or Slurm
  • Understanding of AI training workloads and their network demands
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Surry Hills
84 Employees
Year Founded: 2022

What We Do

Together AI is a research-driven artificial intelligence company. We contribute leading open-source research, models, and datasets to advance the frontier of AI. Our decentralized cloud services empower developers and researchers at organizations of all sizes to train, fine-tune, and deploy generative AI models. We believe open and transparent AI systems will drive innovation and create the best outcomes for society

Similar Jobs

FareHarbor Logo FareHarbor

Account Executive

Sales • Software • Travel
Easy Apply
Hybrid
Amsterdam, NLD
960 Employees
45K-45K Annually

Xero Logo Xero

Software Engineer

Cloud • Fintech • Information Technology • Machine Learning • Software
Remote or Hybrid
Wormer, NLD
4500 Employees

Tulip Logo Tulip

Marketing Manager

Enterprise Web • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
27 Locations
310 Employees

Tulip Logo Tulip

DACH Regional Sales Lead – Enterprise Manufacturing

Enterprise Web • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
27 Locations
310 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account