Data Infra - Signals and Evaluations (IC)

Posted 5 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
Entry level
Aerospace • Artificial Intelligence • Robotics • Defense
The Role
Build versioned datasets, benchmark suites, simulation environments, and evaluation infrastructure for AI models and agents. Develop reproducible testing across accuracy, robustness, safety, calibration, latency, cost, and operational outcomes. Convert field and customer failures into durable test cases, adversarial evaluations, and environment scenarios. Establish release gates and communicate evidence, uncertainty, capability, and residual risk. The role requires strong statistical and experimental foundations, Python and machine-learning framework fluency, and software engineering discipline.
Summary Generated by Built In
About Matter Intelligence

Welcome to Matter, where we are building the future of vision AI: pairing a world-first sensor that sees molecular chemistry, temperature, and 3D shape with a Large World Model that will be the most powerful intelligence engine for the physical world. This system doesn't just see what something looks like; it understands everything from a single pixel. We call this Superintelligent Vision.

Our team has delivered technologies to Mars for NASA/JPL, designed advanced sensors for U.S. Defense, and frontier artificial intelligence systems. We are now building the next generation of space- and airborne-based sensing systems.

About the Role

Matter is hiring a Signal, Data, and Evaluation Engineer to build the datasets, benchmarks, environments, and evidence that make our models and AI systems scientifically defensible. Reporting to Ignacio Cases Martin, this individual contributor will work across data curation, simulation and learning environments, model and agent evaluation, and release-quality evidence.

Key Responsibilities
  • Build versioned dataset workflows for collection, curation, filtering, labeling, ground truth, synthetic data, hard-negative mining, contamination controls, and train-evaluation isolation.

  • Design benchmark suites across perception, multimodal reasoning, physics-informed prediction, retrieval, planning, tool use, world modeling, and customer tasks.

  • Create simulation and environment infrastructure for reinforcement learning, imitation learning, offline learning, model-based learning, and agent training.

  • Define repeatable evaluation methods across accuracy, calibration, generalization, robustness, safety, latency, cost, and operational outcomes.

  • Turn field, mission, and customer failures into versioned datasets, benchmark cases, adversarial tests, and environment scenarios.

  • Build automated promotion gates and communicate clearly what has been tested, what remains unproven, and what evidence supports release.

QualificationsRequired
  • Experience in ML evaluation, dataset engineering, reinforcement-learning environments, simulation, scientific ML, test infrastructure, or production AI systems.

  • Strong foundations in statistics, experimental design, benchmark validity, distribution shift, calibration, reward design, and evidence-based acceptance.

  • Fluency in Python and modern machine-learning frameworks, with the software engineering discipline to build tested modules, data contracts, or services.

  • Experience converting qualitative model or agent failures into reproducible datasets, tests, or environment scenarios.

  • Ability to reason across the complete data path and identify how acquisition, transformation, indexing, orchestration, and presentation affect correctness.

Preferred
  • Experience with dataset curation, benchmark platforms, deep reinforcement learning, world models, synthetic data, multimodal evaluation, or scientific validation.

  • Experience with PyTorch or JAX, distributed evaluation, experiment tracking, environment frameworks, or model and agent observability.

  • Experience with uncertainty quantification, calibration, out-of-distribution detection, conformal methods, or selective prediction.

  • Experience evaluating models deployed to aircraft, satellites, robotics, industrial systems, or other constrained environments.

What Success Looks Like
  • Model and agent releases are supported by reproducible datasets, benchmarks, and clearly stated evidence.

  • Failures from experiments and deployments become durable test cases that improve future systems.

  • Evaluation results preserve scientific meaning and make capability, uncertainty, and residual risk legible to the team.

Location

This role is based in San Francisco, CA, and requires onsite work.

ITAR Requirements

To comply with U.S. export regulations, applicants must be one of the following:

  • A U.S. citizen or national

  • A lawful permanent resident (green card holder)

  • Eligible to obtain required authorizations from the U.S. Department of State

Employee Offerings and Benefits

At Matter, we believe in rewarding high performance and providing the support you need to thrive. Our compensation and benefits package includes:

  • Competitive compensation based on experience

  • Early-stage equity package

  • 100% employer-paid health, dental, and vision coverage

  • Opportunity to work on novel sensing, data, and AI systems with real-world deployment paths to the largest industries in the world

Matter Intelligence is an equal opportunity employer. We welcome candidates from all backgrounds who can raise the ambition and performance of the team.

Skills Required

  • Experience in ML evaluation, dataset engineering, reinforcement-learning environments, simulation, scientific ML, test infrastructure, or production AI systems.
  • Strong foundations in statistics, experimental design, benchmark validity, distribution shift, calibration, reward design, and evidence-based acceptance.
  • Fluency in Python and modern machine-learning frameworks.
  • Software engineering discipline to build tested modules, data contracts, or services.
  • Experience converting qualitative model or agent failures into reproducible datasets, tests, or environment scenarios.
  • Ability to reason across the complete data path, including acquisition, transformation, indexing, orchestration, and presentation.
  • Experience with dataset curation, benchmark platforms, deep reinforcement learning, world models, synthetic data, multimodal evaluation, or scientific validation.
  • Experience with PyTorch or JAX, distributed evaluation, experiment tracking, environment frameworks, or model and agent observability.
  • Experience with uncertainty quantification, calibration, out-of-distribution detection, conformal methods, or selective prediction.
  • Experience evaluating models deployed to aircraft, satellites, robotics, industrial systems, or other constrained environments.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
25 Employees
Year Founded: 2023

What We Do

Matter Intelligence develops advanced sensors and geospatial AI platforms that capture detailed, beyond-visible data of natural and artificial materials from space to surface. Their technology accelerates computer vision and geospatial modeling to understand and predict real-world events.

Similar Jobs

Capital One Logo Capital One

Full-stack Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
4 Locations
55000 Employees
230K-286K Annually

Micron Technology Logo Micron Technology

Artificial Intelligence Engineer

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
2 Locations
45000 Employees
110K-234K Annually

Square Logo Square

Software Engineer

eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Remote or Hybrid
8 Locations
12000 Employees
185K-327K Annually

Wipfli Logo Wipfli

Tax Manager

Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Remote or Hybrid
United States
2900 Employees
106K-160K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account