Machine Learning Data Engineer (DataOps), Materra

Posted Yesterday
Mountain View, CA, USA
In-Office
166K-244K Annually
Mid level
Artificial Intelligence • Greentech • Hardware • Internet of Things • Transportation • Cybersecurity • Automation
The Role
Build and consolidate data infrastructure for ML training: design automated ETL/ELT pipelines, implement DataOps practices (validation, monitoring, anomaly detection), integrate annotation workflows, manage dataset versioning and storage, and collaborate with ML engineers and operations to produce high-quality, reproducible training datasets.
Summary Generated by Built In

About the team:

Materra is on a mission to radically reduce global waste and move to a true circular economy. The team has developed technology that identifies waste material at the molecular level—starting with plastics. Materra works with industry partners to improve the way recycling centers process plastics using AI and robotics, to make recycling more affordable and scalable.
About the Role
We are looking for a Machine Learning Data Engineer (DataOps) to build and unify the data infrastructure that powers our model training pipelines. In this role, you will lead the effort to consolidate fragmented data sources into a cohesive, high-quality data foundation.
Your primary focus will be designing automated ingestion pipelines, establishing data quality validation frameworks, and managing dataset versioning to support our machine learning training loops. You will bridge the gap between operations, remote annotation teams, and machine learning engineers to ensure our models are trained on reliable, well-structured data.
Key Responsibilities

  • Architect and build automated ETL (Extract, Transform, Load) and ELT (Extract, Load, Transform) data pipelines to aggregate, clean, and harmonize data from disparate sources, databases, and operational ingestion flows.
  • Implement DataOps practices, including data quality monitoring, automated schema validation, and anomaly detection to catch corrupt or mislabeled data early.
  • Standardize and integrate third-party annotation workflows and remote labeling feeds into unified datasets ready for model training.
  • Design and maintain dataset versioning and storage systems to allow reproducible machine learning experiments and seamless data retrieval.
  • Collaborate with machine learning engineers and operations teams to translate raw material, form factor, and sensor metadata into structured training features.

Requirements

  • Education: Degree in Computer Science, Data Engineering, Software Engineering, or a related technical field.
  • Data Engineering & Architecture: 3+ years  experience building scalable data pipelines, managing relational and non-relational databases, and unifying fragmented data storage systems.
  • Modern Python Proficiency: Expertise in Python and data manipulation libraries (e.g., Pandas, NumPy, or SQL).
  • Data Quality & DataOps: Practical experience implementing automated data validation, quality control frameworks, and dataset versioning practices.
  • ML Data Lifecycle Understanding: Hands-on experience structuring datasets specifically for machine learning workflows, including handling annotations, metadata tracking, and training set curation.

Preferred Skills

  • Google Cloud Ecosystem: Hands-on experience with Google Cloud platform tools (e.g., BigQuery, Cloud Storage, Dataflow, Dataproc, Vertex AI Data Pipelines).
  • Workflow Orchestration: Experience managing pipelines using Google Cloud Composer or equivalent orchestration frameworks (e.g., Apache Airflow, Prefect, Dagster).
  • Multimodal / Unstructured Data: Experience handling mixed data types, including image datasets, sensor metadata, and unstructured physical property records.
  • Annotation Platform Integration: Familiarity with data labeling platforms, human-in-the-loop workflows, or integrating third-party annotation APIs.
  • Validation & Versioning Tooling: Exposure to data quality and ML versioning tools (e.g., Great Expectations, DVC, or TFX/Data Validation).

The US base salary range for this full-time position is $166,000 - $244,000 + bonus + equity + benefits. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your location during the hiring process.
Please note that the compensation details listed in US role postings reflect the base salary only, and do not include bonus, equity, or benefits.

Skills Required

  • Degree in Computer Science, Data Engineering, Software Engineering, or related technical field.
  • 3+ years experience building scalable data pipelines and unifying fragmented data storage systems (relational and non-relational).
  • Expertise in Python and data manipulation libraries (Pandas, NumPy) and proficiency with SQL.
  • Practical experience implementing automated data validation, data quality frameworks, and dataset versioning practices (DataOps).
  • Hands-on experience structuring datasets for machine learning workflows, including annotation handling, metadata tracking, and training set curation.

X, The Moonshot Factory Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about X, The Moonshot Factory and has not been reviewed or approved by X, The Moonshot Factory.

  • Fair & Transparent Compensation Pay is considered competitive for core technical and senior roles, with employer-posted ranges and clear statements that total compensation includes base, bonus, equity, and benefits. Feedback suggests posted bands and explicit structure provide clarity on how pay is constructed.
  • Parental & Family Support Family support is described as generous, including paid parental leave, baby bonding, and transitional support for parents returning to work. Fertility treatments and maternity care are also covered, indicating depth in family-focused provisions.
  • Retirement Support Retirement programs include a 401(k) with a notable company match and immediate vesting of matched funds. Additional financial supports such as student loan reimbursement and coaching strengthen long-term financial security.

X, The Moonshot Factory Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Mountain View, CA
2,277 Employees
Year Founded: 2010

What We Do

We create breakthrough technologies to help solve some of the world’s biggest problems. Born at Google, we got our start creating self-driving cars and smart glasses. Since then, we’ve continued to bring sci-fi ideas into reality.

Similar Jobs

CrowdStrike Logo CrowdStrike

Program Manager

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Hybrid
Sunnyvale, CA, USA
11000 Employees
120K-180K Annually

Mondelēz International Logo Mondelēz International

TA Advisor, Sales

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
United States
90000 Employees
84K-115K Annually

Square Logo Square

Account Manager

eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Hybrid
San Francisco, CA, USA
12000 Employees
80K-160K Annually

MongoDB Logo MongoDB

Senior Director, HR Business Partnering

Big Data • Cloud • Software • Database
Easy Apply
Hybrid
15 Locations
5550 Employees
151K-297K Annually

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account