Principal Engineer - Data Engineering

Posted 2 Days Ago
Be an Early Applicant
Singapore, SGP
In-Office
Entry level
Big Data • Cloud • Hardware • Software
The Role
Build and maintain scalable data pipelines for scientific, sensor, and operational data used in machine learning. Responsibilities include feature engineering, data quality monitoring, versioning, lineage, drift detection, data contracts, governance, sensor ingestion, synthetic data infrastructure, and MLOps data-layer components such as feature stores and dataset registries.
Summary Generated by Built In
Company Description

WD is building the infrastructure behind the AI-driven data economy.

As AI scales, so does data. Every interaction, every model, every system generates data that must be stored, managed, and made accessible over time. That’s where we come in.

We combine deep engineering expertise with global-scale manufacturing to deliver the storage systems that make AI possible, powering hyperscale data centers, cloud platforms, and enterprise infrastructure worldwide.

This isn’t theoretical work. It’s real systems, at real scale, people solving some of the hardest challenges in technology today.

We’re looking for people who want to build, solve, and operate at that level.

Join us and let’s shape the future of data.

Job Description

About This Role — The Mission

The data you will build pipelines for is not transactional data or clickstream data. It is experimental measurement data from precision product development instruments — each data point costs real time and resources to generate. Getting the data infrastructure right for this kind of scientific data is a genuinely different engineering challenge from standard web-scale or financial data work. You will develop rare expertise in ML-ready scientific data pipelines that very few data engineers in Singapore or globally have built.

Key Responsibilities

  • Feature Engineering Pipelines: Build and maintain reliable, versioned feature engineering pipelines that transform raw engineering, sensor, and operational data into structured ML-ready feature sets — delivered to the specification defined
  • Data Quality Frameworks: Design and operate data quality checks covering completeness, schema consistency, statistical distribution stability, and label accuracy across all AI training datasets. Alert the ML team when data quality degrades before it impacts model training. Collaborate with team who performs final downstream validation.
  • Data Versioning, Lineage & Drift Detection: Build and maintain training data versioning and lineage tracking — ensuring full reproducibility of all model training runs and early alerting when deployment data diverges from training distributions.
  • Data Contracts & Governance — Guided Implementation: Implement and maintain agreed data contracts between upstream data producers and downstream ML consumers, following governance standards established with guidance from ML Engineer. Establish access control and retention practices for all AI data assets.
  • Real-Time Streaming — Sensor Data Ingestion: Contribute to real-time sensor data ingestion pipelines under technical direction. Develops operational ownership progressively over 6–12 months. Not a solo day-1 requirement.
  • Synthetic Data Pipeline Support: Build pipeline infrastructure to operationalize synthetic data generation workstreams. With generative model methodology provided, builds ingestion, storage, and versioning infrastructure.
  • MLOps Data Layer: Build and maintain the training dataset registry, feature store, and model input validation — tightly integrated with the AI platform (AWS Kubernetes, PortKey, Agent Gateway, LangFuse, AWS Guardrails, Elastic Search etc.).

Qualifications

Requirements:

Education

  • Bachelor's or Master's degree in AI, Computer Science, Data Engineering, Electrical Engineering, Applied Mathematics, or related field. AI major preferred; strong data engineering fundamentals required.

Experience

  • Fresh to 1 year. Demonstrated project experience building end-to-end data pipelines — academic, personal, or internship contexts — is the primary evaluation criterion. Python, SQL, and pipeline design fundamentals must be solid and demonstrable through project evidence.

Must Have Skills:

  • Python: Strong proficiency — primary pipeline development language
  • SQL: Strong proficiency — complex queries, window functions, data transformation logic
  • Scalable Pipeline Design: Batch pipeline architecture; reliability, schema management, fault tolerance; Pipeline orchestration (Airflow, Prefect, AWS Glue Jobs)
  • Data Quality Principles: Completeness checks, schema validation, distribution stability monitoring
  • Data Versioning & Lineage: Reproducibility of training data; ability to trace data origin and transformations
  • ML Data Lifecycle Awareness: Basic understanding of how data pipelines connect to ML model training. Awareness that data quality and pipeline design affect model performance downstream — specifically, awareness of risks like train/test data leakage and label quality impact on model accuracy. Does not require prior ML work experience; requires curiosity and conceptual understanding.

Good to have Skills:

  • Feature store (Feast, Tecton) · Synthetic data generation (VAE, GAN, diffusion models) · Data Built Tool (DBT) · Data Load Tool DLT ·  Annotation platform integration (Label Studio, CVAT) · MES / LIMS system integration · Active learning data loop design · Data lakehouse (Iceberg, Dremio, AWS Glue, AWS Lake Formation) / Redhsift

Additional Information

#LI-FN1 

WD thrives on the power and potential of diversity. As a global company, we believe the most effective way to embrace the diversity of our customers and communities is to mirror it from within. We believe the fusion of various perspectives results in the best outcomes for our employees, our company, our customers, and the world around us. We are committed to an inclusive environment where every individual can thrive through a sense of belonging, respect and contribution.

WD is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at [email protected] to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.

Notice To Candidates: Please be aware that WD and its subsidiaries will never request payment as a condition for applying for a position or receiving an offer of employment. Should you encounter any such requests, please report it immediately to WD Ethics Helpline or email [email protected].

WD thrives on the power and potential of diversity. As a global company, we believe the most effective way to embrace the diversity of our customers and communities is to mirror it from within. We believe the fusion of various perspectives results in the best outcomes for our employees, our company, our customers, and the world around us. We are committed to an inclusive environment where every individual can thrive through a sense of belonging, respect and contribution.

WD is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at [email protected] to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.

Notice To Candidates: Please be aware that WD and its subsidiaries will never request payment as a condition for applying for a position or receiving an offer of employment. Should you encounter any such requests, please report it immediately to WD Ethics Helpline or email [email protected].

Skills Required

  • Bachelor's or Master's degree in AI, Computer Science, Data Engineering, Electrical Engineering, Applied Mathematics, or a related field
  • Fresh to 1 year of experience
  • Demonstrated project experience building end-to-end data pipelines through academic, personal, or internship projects
  • Strong proficiency in Python
  • Strong proficiency in SQL, including complex queries, window functions, and data transformation logic
  • Understanding of scalable batch pipeline architecture, reliability, schema management, fault tolerance, and pipeline orchestration
  • Knowledge of data quality principles, including completeness checks, schema validation, and distribution stability monitoring
  • Understanding of data versioning, lineage, and reproducibility of training data
  • Basic understanding of the ML data lifecycle, including data quality, train/test leakage, and label quality
  • Experience with feature stores such as Feast or Tecton
  • Knowledge of synthetic data generation, including VAE, GAN, or diffusion models
  • Experience with dbt
  • Experience with DLT
  • Experience integrating annotation platforms such as Label Studio or CVAT
  • Experience integrating MES or LIMS systems
  • Knowledge of active learning data loop design
  • Experience with data lakehouse technologies such as Iceberg, Dremio, AWS Glue, AWS Lake Formation, or Redshift

Western Digital Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Western Digital and has not been reviewed or approved by Western Digital.

  • Strong & Reliable Incentives Strong & Reliable Incentives: Incentive structures in variable‑pay roles are portrayed as well‑designed, and annual or quarterly bonuses are commonly part of total compensation.
  • Healthcare Strength Healthcare Strength: Company materials highlight comprehensive medical, dental, vision, and mental‑health resources, complemented by options like HSA/FSA and disability coverage.
  • Parental & Family Support Parental & Family Support: Caregiving support across life stages and children’s behavioral health resources are featured, with programs such as Bright Horizons referenced for U.S. employees.

Western Digital Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Bengaluru, Karnataka
25,132 Employees

What We Do

At Western Digital we create data storage solutions that power the technology of today and inspire the innovations of tomorrow.

Similar Jobs

Micron Technology Logo Micron Technology

Senior Engineer

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
Singapore, SGP
45000 Employees

Micron Technology Logo Micron Technology

Intern - MSB Process & Equipment Engineer

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
Singapore, SGP
45000 Employees

Adyen Logo Adyen

Product Manager

Fintech • Payments • Financial Services
Easy Apply
Hybrid
Singapore, SGP
4771 Employees

Mastercard Logo Mastercard

Senior Specialist, Product Management - Operational Intelligence Commercialisation

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Singapore, SGP
38800 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account