AI Infrastructure Solutions Engineer

Posted 4 Days Ago
Be an Early Applicant
Hiring Remotely in Office, Machaze, Manica, MOZ
Remote
Senior level
Artificial Intelligence • Machine Learning • Software • Analytics
The Role
Deploy and customize DDN’s AI infrastructure solutions for strategic customers across on-premises, hybrid, and cloud-adjacent environments. Integrate NVIDIA AI Enterprise services, vector databases, RAG workflows, parallel and object storage, networking, and monitoring tools. Develop automation and deployment scripts, optimize AI and HPC workloads, troubleshoot Linux and distributed storage systems, and serve as the primary technical customer contact while coordinating with internal teams, vendors, and partners.
Summary Generated by Built In

DDN is expanding our Enterprise AI offerings to include the integration of industry leading technologies with DDN Infinia and DDN EXAScaler storage. These solutions will be optimized for inference and RAG workloads and require integration into the customer’s environment. Our support organization is deep on storage (Infinia, EXAScaler); we are now hiring an AI Infrastructure Solutions Engineer to deploy our complete AI solutions. This implementation will include NVIDIA AI Enterprise services (NIMs, NeMo, Triton, GPU Operator, licensing), vector databases (initially Milvus), RAG/agentic workflows, and the high‑performance storage and networking fabric that underpins them.

 

In this role, you will either remotely or onsite in some cases deploy the DDN AI solutions and work to customize this to the end user requirements. You will work with DDN internal teams, vendors and other partners as needed to successfully deploy these solutions.

 
Key Responsibilities
  • Serve as the primary technical point of contact for assigned strategic customers, that are deploying DDN AI solutions

  • Work with Pre-sales to interpret design considerations during solution deployment

  • Drive operational efficiency through automation, tooling, documentation, and repeatable deployment workflows

  • Develop and deploy scripts and tools to support customer environments (DevOps-focused)

  • Be prepared to develop scripting to deploy system monitoring and other metrics based tools to integrate with customer infrastructure

  • Support AI/ML, data‑intensive, and HPC workloads running at scale in on‑prem, hybrid, and cloud‑adjacent environments

  • Work closely with customers to optimize the their AI applications to better work with DDN technology

Required Qualifications
  • 5+ years of experience in a senior technical role deploying complex, customer‑facing production systems

  • Experience administering and operating Lustre or similar parallel file systems in large‑scale environments

  • Experience with object storage and S3‑compatible systems

  • Strong Linux systems knowledge, including performance tuning and troubleshooting

  • Solid understanding of distributed storage architectures, networking fundamentals, and data movement at scale

  • Proven ability to work directly with customers, communicate clearly, and build trusted technical relationships

  • Ability to work effectively across cross‑functional teams including Engineering, Product Management, Support, and Field Services

Preferred Qualifications
  • Experience with additional parallel file systems such as IBM Spectrum Scale or StorNext

  • Experience developing and debugging automation using shell scripting, Python, Bash, or similar languages

  • Strong understanding of networking technologies including InfiniBand, Ethernet, TCP/IP, and routing

  • Knowledge of NVAIE services (e.g., NIMs, NeMo, Triton, TensorRT/TensorRT‑LLM, GPU Operator, licensing/NLS) and vector databases (e.g., Milvus)

  • Familiarity with NAS and data transfer protocols (NFS, SMB/CIFS, SFTP, rsync, etc.)

  • Experience with authentication and identity systems (LDAP, Active Directory, Kerberos, OAuth2/OIDC, SAML)

  • Experience using network diagnostics and troubleshooting tools (tcpdump, Wireshark, LLDP, etc.)

  • Exposure to AI/ML infrastructure operations, GPU‑accelerated environments, or large‑scale data pipelines

  • Experience with deployment and orchestration of large scale compute systems (Kubernetes, SLURM, BCM etc)

  • Experience working in globally distributed or remote teams

Additional Information
  • Occasional physical tasks related to hardware setup may be required, with appropriate tools and support

Why Join DDN
  • Work on real, production‑scale AI and HPC systems that power world‑class innovation

  • Influence product direction through direct customer engagement

  • Collaborate with highly skilled engineers across storage, networking, and distributed systems

  • Grow your career into senior technical leadership, architecture, or product‑facing roles

Skills Required

  • 5+ years of experience in a senior technical role deploying complex, customer-facing production systems
  • Experience administering and operating Lustre or similar parallel file systems in large-scale environments
  • Experience with object storage and S3-compatible systems
  • Strong Linux systems knowledge, including performance tuning and troubleshooting
  • Understanding of distributed storage architectures, networking fundamentals, and data movement at scale
  • Ability to work directly with customers, communicate clearly, and build trusted technical relationships
  • Ability to work effectively across Engineering, Product Management, Support, and Field Services
  • Experience with IBM Spectrum Scale or StorNext
  • Experience developing and debugging automation using shell scripting, Python, Bash, or similar languages
  • Understanding of InfiniBand, Ethernet, TCP/IP, and routing
  • Knowledge of NVIDIA AI Enterprise services and vector databases such as Milvus
  • Familiarity with NAS and data transfer protocols including NFS, SMB/CIFS, SFTP, and rsync
  • Experience with LDAP, Active Directory, Kerberos, OAuth2/OIDC, or SAML
  • Experience using tcpdump, Wireshark, LLDP, or similar network diagnostics tools
  • Exposure to AI/ML infrastructure operations, GPU-accelerated environments, or large-scale data pipelines
  • Experience deploying and orchestrating large-scale compute systems using Kubernetes, SLURM, BCM, or similar technologies
  • Experience working in globally distributed or remote teams
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Chatsworth, CA
706 Employees
Year Founded: 1998

What We Do

DDN is the world’s largest private data storage company and the leading provider of intelligent technology and infrastructure solutions for Enterprise At Scale, AI and analytics, HPC, government and academia customers. Through its DDN and Tintri divisions, the company delivers AI, Data Management software and hardware solutions, and unified analytics frameworks to solve complex business challenges for data-intensive, global organizations. DDN provides its enterprise customers with the most flexible, efficient and reliable data storage solutions for on-premises and multi-cloud environments at any scale. Over the last two decades, DDN has established itself as the data management provider of choice for over 11,000 enterprises, government, and public-sector customers, including many of the world’s leading financial services firms, life science organizations, manufacturing and energy companies, research facilities, and web and cloud service providers.

Similar Jobs

Mondelēz International Logo Mondelēz International

Senior Product Manager

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
3 Locations
90000 Employees

Mondelēz International Logo Mondelēz International

o9 Data & Integration Lead

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
2 Locations
90000 Employees

Compa Logo Compa

Enterprise Account Executive

Artificial Intelligence • HR Tech • Software • Business Intelligence
Remote or Hybrid
3 Locations
75 Employees
200K-225K Annually

Invenergy Logo Invenergy

Senior PI Administrator

Greentech • Real Estate • Social Impact • Energy • Industrial • Solar • Renewable Energy
In-Office or Remote
18 Locations
2500 Employees
125K-155K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account