Senior Software Engineer, Data Platform

Posted 2 Days Ago
Be an Early Applicant
Somerville, MA, USA
Hybrid
Senior level
Big Data • Machine Learning • Software • Analytics • Biotech
AI-Powered Tools to Engineer Biology
The Role
Design, build, and scale the data platform that converts raw mass-spectrometry and biochemical data into validated, labeled datasets. Own pipelines, data contracts, serving, cost optimization, quality gates, SDKs and interfaces, and operational SLAs. Collaborate with ML researchers, scientists, and product to enable model development and product features.
Summary Generated by Built In
About Us

Most of the molecules driving human biology are invisible to us. Mass spectrometers already detect metabolites, lipids, and peptides, but the vast majority of those signals never get identified. A typical experiment names a small fraction of its features and discards the rest. We call this biology's dark matter. It’s signal-rich and mechanism-defining, yet almost entirely opaque.

Matterworks is building the foundation models that make that dark matter legible. Our Large Spectral Models do for biochemical biology what AlphaFold and ESM did for proteins: turning a library-bound discipline into something predictable and generative, and embedding it at every stage of R&D.

Come build the future of biological discovery with us.

Position Overview

As a Senior Software Engineer you’ll work to build the connective tissue of our data platform, in both what you build and how you build it. Design, build and scale systems to enrich data from raw samples and information into readily usable datasets enriched with biological context. The data produced by you and the team will serve our customers through both our ML research and model development activities as well as our product.

You will report to the Head of Engineering and work daily with our machine learning researchers, scientists, and product team.

Key Responsibilities

  • Build and Scale Data Contracts: Own the pipelines and system other teams consume from. Design and implement systems that scale to multiple petabytes of data in an effective way.

  • Serving and Cost at Scale: You will build and scale systems to acquire, store and serve data in a fast and affordable as it grows: data layout, featurization throughput, Kubernetes-native orchestration, and cost surfaced before it adds up.

  • Labels and Enrichment: Turn raw data into datasets people can use, with consistent schemas, trustworthy metadata, and documented definitions. Scale scientific labels from studies down to their spectra and underlying features.

  • Quality Gates: Automate quality checks that enable increasing capability without regression and promote only on a pass.

  • Interfaces People Use: Own the surfaces AI, chemistry, product, and agents call, from the SDK used to build datasets to the tools that expose platform capabilities.

  • Operations and Data Rights: Ensure effective operations of our data needs meeting our designed service level agreements, while providing high quality, provenance, and secure data processing in line with our customer needs.

About You

  • Significant professional experience building production data systems and pipelines. We level on scope and judgment rather than years.

  • Proficient in Python and SQL for large-scale data processing.

  • Proficient in Kubernetes-native batch orchestration and modern data lake technologies (Argo Workflows, Metaflow, EKS, Glue, Athena, Apache Iceberg, Parquet, DuckDB, Terraform). Airflow or Dagster experience transfers fine.

  • Demonstrated experience designing stable identifiers for a large, changing corpus, and building validation that gates a publish rather than reporting on it after the fact.

  • Experience putting an LLM or agent component into a production data path, including the eval loop, the gold set, and cost per record.

  • Daily use of AI coding tools, paired with healthy skepticism about their output on questions of production data correctness.

  • A track record of owning work through to a running, validated system, including the unglamorous parts: fixing the malformed dataset, writing the backfill, debugging last night's bad publish.

  • Comfort with messy scientific formats and toolchains (mzML, RDKit, ProteoWizard or similar). Engineering depth is the requirement.

  • A passion for contributing to an early-stage startup where autonomy, eagerness to learn, and enthusiasm for solving novel scientific challenges prevail over rigid processes and egos.

Working at Matterworks

Given the cross-disciplinary and innovative nature of our work, effective collaboration and communication are critical to our progress. We operate in a flexible hybrid model that accommodates both fully remote team members and those who work full-time from our Somerville, MA office. While some positions may require regular in-person presence for hands-on work or local collaboration, many roles can be performed remotely with team members distributed across various locations.

Compensation and Benefits

Matterworks offers full-time employees a competitive base salary, stock options, and benefits (health & dental, vision, long- and short-term disability, life insurance, 401k with company match). Employees enjoy a flexible work & unlimited time away policy, commuter benefits and parking, regular team meals and outings, and company support for continued education/coursework and conference participation.

Matterworks, Inc. is an equal opportunity employer. All candidates for employment at Matterworks are considered without regard to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, sexual orientation, or any other category protected by law.

Skills Required

  • Significant professional experience building production data systems and pipelines
  • Proficient in Python for large-scale data processing
  • Proficient in SQL for large-scale data processing
  • Experience with Kubernetes-native batch orchestration (Argo Workflows, Metaflow); Airflow or Dagster experience transferable
  • Experience with modern data lake technologies and tooling (EKS, AWS Glue, Athena, Apache Iceberg, Parquet, DuckDB, Terraform)
  • Experience designing stable identifiers for large, changing corpora and building publish-time validation gates
  • Experience putting an LLM or agent component into a production data path, including eval loop and gold set
  • Proven track record of owning work through to running, validated systems (backfills, debugging, malformed data fixes)
  • Comfort working with messy scientific formats and toolchains (mzML, RDKit, ProteoWizard or similar)
  • Daily use of AI coding tools and critical evaluation of their outputs
  • Ability and enthusiasm to contribute in an early-stage startup environment
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Somerville, MA
25 Employees
Year Founded: 2019

What We Do

Matterworks is developing advanced AI-powered software and proprietary chemistry reagents to enable real-time, quantitative metabolomics to accelerate the research, development, and manufacturing of biologic therapeutics. Metabolomics measures all the small molecules a cell needs to function. In spite of its far-reaching potential impact across life science disciplines, adoption of metabolomics remains limited and lags behind other ‘omic techniques due to a number of inherent challenges. Matterworks is developing a novel alternative methodology that activates deep learning (DL) for analytical chemistry, to remove the constraints imposed by the requirement for human-interpretable data. Matterworks is on a mission to democratize metabolomics and catalyze its widespread adoption and integration into the life sciences. Our view is very simple: solving metabolomics will advance all of the life sciences industries.

Why Work With Us

We are a passionate team of scientists and engineers committed to building better tools to interrogate biology and accelerate the rate of scientific progress. In pursuit of our ambitious goals, we rely on our core company values to guide our work: Mission First, Communication & Trust, Growth Through Challenges, and Scientific Excellence.

Gallery

Gallery

Similar Jobs

Klaviyo Logo Klaviyo

Lead Software Engineer

Consumer Web • eCommerce • Marketing Tech • Retail • Software • Analytics • Generative AI
Easy Apply
Hybrid
Boston, MA, USA
2400 Employees
216K-324K Annually

Samsara Logo Samsara

Senior Software Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
131K-220K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
9 Locations
40000 Employees
45K-85K Annually

PwC Logo PwC

Security Risk & Engineering - Tech and Cyber Risk & Compliance -Manager

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
9 Locations
370000 Employees
99K-232K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account