Data and Platform Engineer

Posted An Hour Ago
Be an Early Applicant
4 Locations
In-Office or Remote
200K-322K Annually
Expert/Leader
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
The Role
Owns architecture, roadmap, development, and production operations for major components of NVIDIA’s distributed data platform. Builds batch and streaming pipelines, data products, shared platform capabilities, quality standards, and operational practices using Python, SQL, Databricks, and Spark. Leads cross-team technical delivery, production investigations, architectural changes, mentoring, and adoption of reliable, secure, scalable data systems supporting GPU fleet health, capacity, utilization, cost, and operational decisions.
Summary Generated by Built In

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions.

You will be responsible for the architecture, technical plan, and production results of a major platform area like ingestion and orchestration, data quality and reconciliation, or data serving and consumption. You will clarify requirements with customers, define technical objectives, guide design and development among engineers and partner teams, and stay actively engaged in coding, debugging, and production tasks. Successful candidates have already led complex technical work across team boundaries and delivered improvements that other groups adopted. Our primary implementation environment is Python, SQL, Databricks, and Spark.

What you’ll be doing:

  • You will own a major platform component and its roadmap. For example, define its architecture, interfaces, technical goals, and evolution. Anticipate capacity, compatibility, and operational needs over a multi-year horizon, and translate them into achievable breakthroughs that balance immediate delivery with long-term maintainability.

  • Lead technical delivery across teams. Work with customers and interested parties to clarify vague requirements. Break down design and implementation work for contributing engineers. Establish release turning points and manage dependencies and delivery risks. Guide the work process, revise plans when requirements shift, and keep management and partner teams informed and aligned.

  • Build data pipelines and products. Plan and carry out batch and streaming ingestion, transformation, reconciliation, and serving processes for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Establish data models and agreements that remain stable as sources, consumers, and scale progress.

  • Develop shared platform capabilities. Direct the creation and adoption of libraries, workflow and DAG or comparable experience abstractions, deployment tools, and standard implementation approaches. Partner with related teams to solve shared challenges and evaluate progress in onboarding time, engineering effort, reliability, and cost.

  • Lead complex production investigations. Serve as the technical point of accountability for issues spanning pipelines, applications, SQL engines, Spark, storage, networks, and cloud services. Coordinate investigations across owners, drive resolution of release blockers and critical issues from partners, and implement preventive measures.

  • Define quality, security, and operational expectations. Establish and implement testing, data-quality, reconciliation, lineage, SLO, and release-readiness standards for your platform area. Partner with security and infrastructure teams on trust boundaries, service identities, least privilege, secrets, environment isolation, and auditability, and drive adoption across contributing teams.

  • Make trusted data usable. Deliver well-modeled tables, APIs, automation, dashboards, and focused internal applications. Align with consumers on semantics, access patterns, freshness, compatibility, and ownership so that shared capabilities support dependable operational decisions.

  • Provide technical leadership through others. Guide design reviews, mentor engineers taking on larger ownership, and resolve technical disagreements using evidence and clear tradeoffs. Partner with leadership on priorities and explain how technical investments support DGXC objectives.

What we need to see:

  • BS or MS in Computer Science, Engineering, or a related field (or equivalent experience), and at least 12+ years of equivalent experience

  • A sustained record of building and operating production software, data platforms, databases, or distributed systems. This includes owning a major component or complex project from requirements and architecture through release and ongoing operation.

  • Proven ability to outline a component’s technical plan, establish objectives for engineers, assign design and implementation tasks, and guide delivery within your team and nearby teams with little supervision.

  • Extensive practical experience in one or more of these areas: distributed processing using Spark or a similar system; relational, distributed, or analytical databases; production ETL, change-data capture, streaming, or event handling; or backend and cloud platforms managing large data volumes. You must grasp the interfaces and failure modes of nearby layers thoroughly to inform solid architectural choices.

  • Strong software-engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL. Experience designing reusable abstractions, reviewing substantial changes, and personally implementing and debugging critical code paths.

  • Strong SQL and data-modeling skills, with practical depth in query execution, incremental processing, schema evolution, consistency, and analytical consumption. Ability to reason about idempotency, replay, late-arriving data, partial failure, and correctness across system boundaries.

  • Experience leading complex investigations involving multiple components and teams. Ability to use logs, metrics, traces, query plans, profiles, and controlled experiments to establish root cause, coordinate resolution, and prevent recurrence.

  • Demonstrated architectural judgment: evaluating alternatives, anticipating future requirements, and balancing reliability, performance, cost, security, compatibility, and maintainability. Experience leading significant migrations or architectural changes while preserving production service.

  • Experience establishing production quality and operational practices that other engineers adopt, including testing, CI/CD, monitoring, alerting, rollback, incident response, and secure deployment.

  • Proven success in influencing technical decisions without official authority, advising engineers outside your immediate project, and advancing workflow improvements across closely related teams. Ability to simplify complex issues, offer a course of action, and communicate decisions and delivery risks clearly.

Ways to stand out from the crowd:

  • Proven expertise in building, refining, and running Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, or Unity Catalog workloads along with shared platform features.

  • Experience designing and operating Kafka or comparable streaming systems, including partitioning, consumer behavior, offset management, backpressure, replay, and schema compatibility.

  • Experience scaling, migrating, or tuning relational, distributed, time-series, object-storage, or information retrieval systems, including Elasticsearch or OpenSearch.

  • Background operating compute or GPU clusters, or working with Kubernetes, Slurm, cloud infrastructure, and fleet telemetry across AWS, Azure, GCP, or other providers.

  • Experience building production agentic systems or agent harnesses, including tool integration, context management, evaluation, permissions, observability, and failure recovery. Evidence of measurable improvements in engineering productivity or operational outcomes.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 200,000 USD - 322,000 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 25, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Skills Required

  • BS or MS in Computer Science, Engineering, or a related field, or equivalent experience
  • At least 12 years of equivalent experience
  • Experience building and operating production software, data platforms, databases, or distributed systems
  • Experience owning major components or complex projects from requirements and architecture through release and ongoing operation
  • Experience defining technical plans, objectives, design tasks, and delivery across teams
  • Practical experience with distributed processing, databases, production ETL, change-data capture, streaming, event handling, backend systems, or cloud platforms managing large data volumes
  • Strong software engineering fundamentals and production proficiency in Python or another backend or systems language
  • Ability and willingness to work primarily in Python and SQL
  • Experience designing reusable abstractions, reviewing substantial changes, and implementing and debugging critical code paths
  • Strong SQL and data-modeling skills, including query execution, incremental processing, schema evolution, consistency, and analytical consumption
  • Understanding of idempotency, replay, late-arriving data, partial failure, and correctness across system boundaries
  • Experience leading complex investigations involving multiple components and teams
  • Experience using logs, metrics, traces, query plans, profiles, and controlled experiments for root-cause analysis
  • Demonstrated architectural judgment balancing reliability, performance, cost, security, compatibility, and maintainability
  • Experience leading significant migrations or architectural changes while preserving production service
  • Experience establishing testing, CI/CD, monitoring, alerting, rollback, incident response, and secure deployment practices
  • Ability to influence technical decisions without official authority and advise engineers outside the immediate project
  • Ability to communicate complex issues, decisions, priorities, and delivery risks clearly
  • Expertise with Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, or Unity Catalog
  • Experience designing and operating Kafka or comparable streaming systems
  • Experience scaling, migrating, or tuning relational, distributed, time-series, object-storage, or information-retrieval systems
  • Experience with Elasticsearch or OpenSearch
  • Experience operating compute or GPU clusters, Kubernetes, Slurm, cloud infrastructure, or fleet telemetry
  • Experience with AWS, Azure, GCP, or other cloud providers
  • Experience building production agentic systems or agent harnesses

NVIDIA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about NVIDIA and has not been reviewed or approved by NVIDIA.

  • Equity Value & Accessibility Equity awards and a discounted ESPP are highlighted as core parts of total compensation, enabling employees to share in the company’s success. Stock-based compensation and the two-year lookback ESPP are consistently described as especially valuable.
  • Healthcare Strength Health coverage is portrayed as robust, with comprehensive medical, dental, and vision options alongside mental health support and on-site care resources. Employer HSA contributions and wellness perks reinforce the depth of the offering.
  • Retirement Support Retirement programs are depicted as strong, featuring a meaningful 401(k) match with Roth options and support for Mega Backdoor Roth contributions. These elements position long-term savings as a notable advantage of the total rewards package.

NVIDIA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
21,960 Employees
Year Founded: 1993

What We Do

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as “the AI computing company.”

Similar Jobs

CrowdStrike Logo CrowdStrike

Data Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
120K-180K Annually
Easy Apply
Remote
United States
900 Employees
130K-160K Annually

NVIDIA Logo NVIDIA

Platform Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office or Remote
4 Locations
21960 Employees
140K-270K Annually

Peraton Logo Peraton

Data Scientist

Aerospace • Information Technology • Security • Cybersecurity • Defense
Remote
United States
18000 Employees
80K-128K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account