Staff Engineer

Posted 2 Days Ago
Be an Early Applicant
Santa Clara, Tarapacá, Amazonas, COL
In-Office
185K-275K Annually
Senior level
Artificial Intelligence • Machine Learning • Software • Analytics
The Role
Own complex customer escalations for DDN Infinia, diagnosing distributed-systems, storage, and performance issues through resolution, mitigation, and root-cause analysis. Lead incident response, customer communications, cross-functional investigations, and reliability improvements. Develop troubleshooting guidance, runbooks, and automation; mentor engineers and influence architecture. The role requires deep storage, Linux, coding, observability, and performance-debugging expertise, customer engagement, and participation in an on-call rotation.
Summary Generated by Built In

DDN is seeking a Staff Engineer to join our Infinia Core team. This is a hands-on technical role combining deep distributed-systems engineering with direct engagement with customers running Infinia in production.

You'll own complex technical escalations end-to-end, from root-cause analysis and incident response through to mitigation, customer communication and product improvements. You'll also help shape engineering best practice, mentor other engineers and drive the use of AI and automation to improve reliability and diagnostics.

If you love the technical depth but want to stay behind the curtain, this probably isn't the right fit but if you want to combine serious engineering with real customer impact, read on.

About Infinia

Infinia is DDN's next-generation, software-defined storage platform, built from the ground up for AI and accelerated computing. It combines separate control and data planes, all-flash performance, sub-millisecond latency and multi-tenancy for demanding enterprise and hyperscale AI and GPU workloads.

What You'll Do

  • Communicate technical issues clearly to customers, engineers and senior stakeholders, including executive audiences.

  • Own complex customer escalations from diagnosis through to resolution, mitigation and RCA.

  • Lead live incident response, war rooms and cross-functional investigations with Engineering, QA and Field teams.

  • Debug complex distributed-systems, storage and performance issues across the system, protocol and application layers.

  • Reproduce customer issues and feed findings into product and reliability improvements.

  • Develop runbooks, troubleshooting guidance and performance-tuning practices.

  • Act as a technical authority on Infinia internals, mentoring engineers and influencing architectural best practice.

  • Partner with Field CTOs, Solutions Architects and Sales Engineers on strategic customer issues.

  • Use AI, automation and observability to improve diagnostics, reliability and MTTR.

  • Communicate technical issues clearly to customers, engineers and senior stakeholders, including executive audiences.

  • This position requires participation in an on-call rotation to provide after-hours support as needed.

What You'll Bring

Must-Haves

  • Significant experience in enterprise storage, distributed systems or cloud infrastructure, with technical leadership at Senior or Staff level.

  • Deep understanding of file systems and storage technologies, including S3, POSIX, NFS and storage performance.

  • Strong Linux systems knowledge, including kernel-level troubleshooting and debugging.

  • Strong coding ability in Python or C++.

  • Proven ability to diagnose complex issues using tools such as strace, tcpdump and perf.

  • Genuine interest in working directly with customers and taking ownership of complex problems through to resolution.

Nice-to-Haves

  • Experience with DDN, VAST, Weka or similar scale-out storage/file systems.

  • Familiarity with observability platforms such as Prometheus, Grafana, ELK or OpenTelemetry.

  • Knowledge of replication, consistency models and data integrity mechanisms.

  • Experience supporting AI/ML, LLM training or other high-performance computing environments.

  • Experience using AI tools for log analysis, troubleshooting, automated RCA or reducing MTTR.

Skills Required

  • Significant experience in enterprise storage, distributed systems, or cloud infrastructure, with technical leadership at Senior or Staff level.
  • Deep understanding of file systems and storage technologies, including S3, POSIX, NFS, and storage performance.
  • Strong Linux systems knowledge, including kernel-level troubleshooting and debugging.
  • Strong coding ability in Python or C++.
  • Proven ability to diagnose complex issues using strace, tcpdump, and perf.
  • Willingness and ability to work directly with customers and own complex problems through resolution.
  • Experience with DDN, VAST, Weka, or similar scale-out storage/file systems.
  • Familiarity with Prometheus, Grafana, ELK, or OpenTelemetry.
  • Knowledge of replication, consistency models, and data integrity mechanisms.
  • Experience supporting AI/ML, LLM training, or high-performance computing environments.
  • Experience using AI tools for log analysis, troubleshooting, automated root-cause analysis, or reducing MTTR.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Chatsworth, CA
706 Employees
Year Founded: 1998

What We Do

DDN is the world’s largest private data storage company and the leading provider of intelligent technology and infrastructure solutions for Enterprise At Scale, AI and analytics, HPC, government and academia customers. Through its DDN and Tintri divisions, the company delivers AI, Data Management software and hardware solutions, and unified analytics frameworks to solve complex business challenges for data-intensive, global organizations. DDN provides its enterprise customers with the most flexible, efficient and reliable data storage solutions for on-premises and multi-cloud environments at any scale. Over the last two decades, DDN has established itself as the data management provider of choice for over 11,000 enterprises, government, and public-sector customers, including many of the world’s leading financial services firms, life science organizations, manufacturing and energy companies, research facilities, and web and cloud service providers.

Similar Jobs

DDN Storage Logo DDN Storage

Staff Engineer

Artificial Intelligence • Machine Learning • Software • Analytics
Remote or Hybrid
Santa Clara, Tarapacá, Amazonas, COL
706 Employees
225K-275K Annually

DDN Storage Logo DDN Storage

Senior Staff Storage Engineer

Artificial Intelligence • Machine Learning • Software • Analytics
Hybrid
Santa Clara, Tarapacá, Amazonas, COL
706 Employees

DDN Storage Logo DDN Storage

Security Engineer

Artificial Intelligence • Machine Learning • Software • Analytics
Hybrid
Santa Clara, Tarapacá, Amazonas, COL
706 Employees

LawnStarter Logo LawnStarter

Staff Software Engineer

Marketing Tech • Software
In-Office or Remote
6 Locations
366 Employees
85K-125K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account