Data Engineer

Posted 24 Days Ago
San Francisco, CA, USA
In-Office
Senior level
Artificial Intelligence • Fintech • Payments • Software
The Role
Own and scale end-to-end data infrastructure: build production ETL pipelines, design scalable schemas, enforce data quality/governance/security, create self-serve data models for analytics and ML, instrument observability, and partner cross-functionally while participating in on-call incident response.
Summary Generated by Built In

About Sapiom

Sapiom is the end-to-end platform that removes barriers to ship and scale agentic products.

We unify everything an agent needs to act in the world: compute and sandboxes, memory, identity, domains and DNS, spend controls, browser automation, web search and deep research, databases, storage, queues, messaging, image generation, voice, enrichment, verification, and monitoring provisioned together as one thing, not handed over as a framework for builders to assemble themselves. Pricing is just as simple: a plan, a generous free tier, pay for what you use when you use it.

We have assembled a world-class team with deep infrastructure and payments DNA to build the operating system for machines. Backed by a $15.75M investment from Accel, Menlo, and Anthropic, we are moving with relentless focus to allow builders to ship and scale agentic products.

About the Role

This is a foundational infrastructure role at a company where the data layer isn't a back-office function — it's the nervous system of a payments platform processing every agent transaction, policy decision, and risk signal in real time. The right person thrives on ownership, has strong opinions about data quality and governance, and moves with the urgency of someone who knows that bad data costs more than bad code. As an early data engineer, you'll define not just the pipelines but the standards, architecture, and culture of data at Sapiom.

What You Will Do

You'll own Sapiom's data infrastructure end-to-end — designing and scaling ETL pipelines, defining schemas that survive 10x growth, and building the governance and quality frameworks that make data trustworthy across the company. You'll architect standardized data models that enable self-serve AI-powered insights, giving Analytics, Data Science, and product teams the visibility they need to move fast without coming to you for every query. The mandate is broad: pipelines, quality, security, observability, and the cross-functional partnerships that keep it all running.

Responsibilities

  • Build, scale, and optimize production-quality ETL pipelines — owning the full lifecycle from ingestion through availability, with clear quality and SLA standards

  • Design data schemas and architect for scale — anticipating 10x data growth and building models that don't require rework when it arrives

  • Own data quality, governance, security, and schema design across the platform — setting the standards and making sure they hold

  • Develop standardized, self-serve data models that enable AI-powered analytics — reducing friction for partner teams and eliminating one-off data pulls

  • Instrument pipeline observability and surface key health metrics to Analytics, Data Science, and DevOps — proactively surfacing issues before they become incidents

  • Partner closely with Data Science, Analytics, and DevOps — operating as a force multiplier across teams, not a bottleneck

Requirements

  • Demonstrated track record — 5+ years — transforming raw data into governed, well-documented, production-ready datasets that business teams can trust and use

  • Deep hands-on experience building and deploying production data pipelines using SQL, Python, Spark, AWS Glue, EMR, DBT, and Airflow

  • Strong command of MPP databases — Snowflake, AWS Redshift, or Teradata — with 3+ years of hands-on production use

  • Proven partnership record with Engineering, Analytics, Data Science, and DevOps teams — someone who treats cross-functional relationships as core to the job, not peripheral to it

  • Architectural instincts — able to design schemas and systems that scale gracefully, not just handle today's load

  • Comfort operating in an on-call rotation — including incident response outside regular working hours when the pipeline demands it

  • Clear communicator who can translate complex data infrastructure decisions into plain-language insights for both technical and non-technical stakeholders

Skills Required

  • 5+ years transforming raw data into governed, production-ready datasets
  • Production data pipelines using SQL
  • Production data pipelines using Python
  • Production data pipelines using Spark
  • Production data pipelines using AWS Glue
  • Production data pipelines using EMR
  • Production data pipelines using DBT
  • Production data pipelines using Airflow
  • 3+ years hands-on production use of Snowflake, AWS Redshift, or Teradata (MPP databases)
  • Proven partnership with Engineering, Analytics, Data Science, and DevOps teams
  • Architectural instincts to design scalable schemas and systems
  • Comfort operating in an on-call rotation, including incident response outside regular hours
  • Clear communicator able to translate complex technical decisions for non-technical stakeholders
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
10 Employees
Year Founded: 2025

What We Do

Sapiom is a fintech startup building the financial infrastructure and operating system for the agentic economy. It provides the essential payment and identity rails, including KYA protocols, that enable AI agents to autonomously and securely transact with real-world digital services such as APIs, software, and compute, while abstracting the complexities of authentication, billing, and compliance for developers.

Similar Jobs

PwC Logo PwC

Data Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
66 Locations
370000 Employees
124K-280K Annually

PwC Logo PwC

Data Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
65 Locations
370000 Employees
99K-232K Annually

PwC Logo PwC

Data Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
68 Locations
370000 Employees
77K-202K Annually

PwC Logo PwC

Data Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
68 Locations
370000 Employees
77K-202K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account