Role Summary
Responsible for the setup, maintenance, security, and day-to-day operation of a technical/research lab environment (servers, storage systems, networking equipment, and lab-issued workstations), ensuring high availability, performance, and compliance with organizational policies.
Key Responsibilities
Install, configure, and maintain lab hardware (servers, storage arrays, NICs, switches) and software (OS images, drivers, monitoring tools)
Manage user access, accounts, and permissions across lab systems
Monitor system health, performance, and capacity (CPU, memory, storage, network utilization)
Troubleshoot hardware/software issues and coordinate with vendors for support/RMAs
Maintain documentation: network diagrams, asset inventories, configuration baselines, SOPs
Implement and enforce security policies (patching, firewall rules, access controls)
Manage backups, disaster recovery procedures, and data retention policies
Support researchers/engineers with environment setup for experiments, benchmarks, or testing (e.g., provisioning compute nodes, storage volumes, network configs)
Track licensing, warranties, and hardware lifecycle (procurement to decommissioning)
Coordinate lab scheduling/resource allocation if shared across teams
Required Skills/Qualifications
Strong Linux administration experience (Ubuntu/RHEL/CentOS)
Networking fundamentals (TCP/IP, VLANs, bonding/LACP, basic troubleshooting)
Experience with storage systems (SAN/NAS, parallel filesystems like Lustre/GPFS a plus)
Scripting ability (Bash, Python) for automation
Familiarity with virtualization/containerization (KVM, Docker) is a plus
Experience with monitoring tools (Prometheus/Grafana, Nagios, Zabbix)
Understanding of hardware components (CPUs, NUMA architecture, PCIe, NICs/RDMA)
Good documentation and communication skills
Nice to Have
Experience with HPC/AI infrastructure (InfiniBand, RDMA, GPU clusters)
Experience with configuration management (Ansible, Puppet)
Skills Required
- Strong Linux administration experience with Ubuntu, RHEL, or CentOS
- Networking fundamentals, including TCP/IP, VLANs, bonding/LACP, and troubleshooting
- Experience with storage systems such as SAN/NAS
- Bash and Python scripting ability for automation
- Experience with monitoring tools such as Prometheus/Grafana, Nagios, or Zabbix
- Understanding of hardware components, including CPUs, NUMA, PCIe, NICs, and RDMA
- Good documentation and communication skills
- Experience with parallel filesystems such as Lustre or GPFS
- Familiarity with virtualization or containerization using KVM or Docker
- Experience with HPC or AI infrastructure, InfiniBand, RDMA, or GPU clusters
- Experience with configuration management tools such as Ansible or Puppet
What We Do
DDN is the world’s largest private data storage company and the leading provider of intelligent technology and infrastructure solutions for Enterprise At Scale, AI and analytics, HPC, government and academia customers. Through its DDN and Tintri divisions, the company delivers AI, Data Management software and hardware solutions, and unified analytics frameworks to solve complex business challenges for data-intensive, global organizations. DDN provides its enterprise customers with the most flexible, efficient and reliable data storage solutions for on-premises and multi-cloud environments at any scale. Over the last two decades, DDN has established itself as the data management provider of choice for over 11,000 enterprises, government, and public-sector customers, including many of the world’s leading financial services firms, life science organizations, manufacturing and energy companies, research facilities, and web and cloud service providers.







