Data Infra - Telemetry and Observability (IC)

Posted 6 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
Entry level
Aerospace • Artificial Intelligence • Robotics • Defense
The Role
Build telemetry and observability across AI training, data, infrastructure, models, agents, and missions. The role establishes shared telemetry schemas, instruments distributed workloads and agent workflows, connects mission context to model behavior, and develops service objectives, alerts, reliability mechanisms, capacity and cost visibility, and incident tooling. Success requires correlating operational, scientific, and AI-system signals to diagnose failures and support safe deployments.
Summary Generated by Built In
About Matter Intelligence

Welcome to Matter, where we are building the future of vision AI: pairing a world-first sensor that sees molecular chemistry, temperature, and 3D shape with a Large World Model that will be the most powerful intelligence engine for the physical world. This system doesn't just see what something looks like; it understands everything from a single pixel. We call this Superintelligent Vision.

Our team has delivered technologies to Mars for NASA/JPL, designed advanced sensors for U.S. Defense, and frontier artificial intelligence systems. We are now building the next generation of space- and airborne-based sensing systems.

About the Role

Matter is hiring a Telemetry and Observability Engineer to make datasets, training runs, models, environments, agents, and missions observable as one AI system. Reporting to Ignacio Cases Martin, this individual contributor will connect flight and mission reality with infrastructure, model, and product behavior through shared telemetry, traces, alerts, reliability mechanisms, and operational context.

Key Responsibilities
  • Build a common telemetry model linking mission state, datasets, code, checkpoints, accelerators, model and prompt versions, environment episodes, agent steps, tools, decisions, cost, latency, and feedback.

  • Instrument distributed training, data loading, GPU utilization, checkpointing, experiment health, batch and online inference, and model-serving behavior.

  • Instrument agent workflows across prompts, context, retrieval, memory, planning, tool calls, graph state, human intervention, evidence, outcomes, and safety controls.

  • Connect planned-versus-actual mission state, hardware context, time, geometry, and geolocation to relevant datasets and model results.

  • Establish actionable service objectives, alerts, incident interfaces, rollout and rollback signals, and escalation paths across platform, model, environment, and agent failures.

  • Build capacity and cost visibility across storage, processing, accelerators, training, evaluation, inference, retrieval, and agent execution.

QualificationsRequired
  • Experience with ML observability, telemetry, time-series systems, distributed training or inference, cloud infrastructure, or production AI platforms.

  • Strong hands-on engineering skills in observability, automation, APIs, infrastructure, incident tooling, and systems debugging.

  • Understanding of metrics, logs, traces, events, model and data monitoring, service objectives, capacity, release safety, and incident diagnosis.

  • Ability to reason about AI failure modes including bad datasets, training instability, stale checkpoints, drift, version mismatch, environment bugs, and agent-tool failure.

  • Experience designing telemetry schemas and correlation workflows across multiple systems, time domains, and ownership boundaries.

Preferred
  • Experience with MLOps, GPU workloads, model serving, agent observability, reinforcement-learning environments, mission telemetry, or scientific instrumentation.

  • Experience operating high-throughput inference, streaming systems, geospatial pipelines, or mixed cloud and edge deployments.

  • Experience with reliability platforms or observability products used by multiple engineering teams.

  • Familiarity with aerospace, defense, regulated operations, or other settings requiring formal change control and incident evidence.

What Success Looks Like
  • Teams can correlate mission, data, infrastructure, model, environment, and agent behavior during normal operation and incidents.

  • Alerts and service objectives identify actionable failures without obscuring scientific or operational context.

  • Capacity, cost, version, and reliability signals support safe deployments and faster diagnosis across the AI stack.

Location

This role is based in San Francisco, CA, and requires onsite work.

ITAR Requirements

To comply with U.S. export regulations, applicants must be one of the following:

  • A U.S. citizen or national

  • A lawful permanent resident (green card holder)

  • Eligible to obtain required authorizations from the U.S. Department of State

Employee Offerings and Benefits

At Matter, we believe in rewarding high performance and providing the support you need to thrive. Our compensation and benefits package includes:

  • Competitive compensation based on experience

  • Early-stage equity package

  • 100% employer-paid health, dental, and vision coverage

  • Opportunity to work on novel sensing, data, and AI systems with real-world deployment paths to the largest industries in the world

Matter Intelligence is an equal opportunity employer. We welcome candidates from all backgrounds who can raise the ambition and performance of the team.

Skills Required

  • Experience with ML observability, telemetry, time-series systems, distributed training or inference, cloud infrastructure, or production AI platforms.
  • Strong hands-on engineering skills in observability, automation, APIs, infrastructure, incident tooling, and systems debugging.
  • Understanding of metrics, logs, traces, events, model and data monitoring, service objectives, capacity, release safety, and incident diagnosis.
  • Ability to reason about AI failure modes, including bad datasets, training instability, stale checkpoints, drift, version mismatch, environment bugs, and agent-tool failure.
  • Experience designing telemetry schemas and correlation workflows across multiple systems, time domains, and ownership boundaries.
  • Experience with MLOps, GPU workloads, model serving, agent observability, reinforcement-learning environments, mission telemetry, or scientific instrumentation.
  • Experience operating high-throughput inference, streaming systems, geospatial pipelines, or mixed cloud and edge deployments.
  • Experience with reliability platforms or observability products used by multiple engineering teams.
  • Familiarity with aerospace, defense, regulated operations, or settings requiring formal change control and incident evidence.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
25 Employees
Year Founded: 2023

What We Do

Matter Intelligence develops advanced sensors and geospatial AI platforms that capture detailed, beyond-visible data of natural and artificial materials from space to surface. Their technology accelerates computer vision and geospatial modeling to understand and predict real-world events.

Similar Jobs

MongoDB Logo MongoDB

Head of Executive Operations - Office of CEO

Big Data • Cloud • Software • Database
Easy Apply
Hybrid
3 Locations
5550 Employees
162K-318K Annually

Alaffia Health Logo Alaffia Health

Implementation Lead

Artificial Intelligence • Healthtech • Insurance • Machine Learning • Payments
Remote or Hybrid
United States
80 Employees
120K-160K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Engineering Manager

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Hybrid
6 Locations
40000 Employees
156K-281K Annually

ServiceNow Logo ServiceNow

Sr Enterprise Account Exec

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
San Diego, CA, USA
29000 Employees
131K-216K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account