Data Engineer, Red Tape Index - Labrynth

Posted Yesterday
Be an Early Applicant
11 Locations
Remote
Mid level
Artificial Intelligence • Financial Services
Amplifying founders and building companies with exponential potential, founded by Invisible with a focus on AI services
The Role
Build and operate the data platform behind regulatory indices: source messy public and commercial data, implement idempotent medallion ingestion pipelines, design Postgres schemas and migrations, construct auditable index methodology (normalization, weighting, sensitivity testing), and run pipelines on Prefect/ECS with strong observability and test coverage.
Summary Generated by Built In
About Labrynth

Labrynth accelerates progress by streamlining regulatory complexity. We build AI-powered platforms that navigate complex regulations, generate audit-level documentation, and provide certainty, not shortcuts. Our technology serves clients across heavily regulated industries including energy, compliance, and government regulations.

We operate as a forward-deployed engineering organization: small, high-velocity teams embedded directly with clients to rapidly discover needs and ship production-quality solutions.

About the Role

We are hiring a Data Engineer to build the data platform behind our regulatory indices: acquiring fragmented public data, transforming it into clean, auditable datasets, and constructing the index methodology that turns it into published rankings.

This is a data platform role more than a pure pipeline or backend role. You will sit close to the raw sources and close to the math. The work spans three modes:

  • Acquire: source data from fragmented and often hostile places, including government open-data portals, APIs, HTML, PDFs, legacy Excel formats, login-protected portals, and commercial sites behind anti-bot protection.

  • Transform: normalize inconsistent jurisdictional data through bronze → silver → gold pipelines with idempotent ingestion, content hashing, and run-level lineage.

  • Construct: turn clean data into transparent, auditable indices through winsorization, percentile ranks, weighting, composites, and sensitivity testing.

What You'll Do
  • Ship scrapers and ingestion flows against messy, sometimes adversarial sources, using HTTP/2 clients, TLS-fingerprint evasion, and browser automation fallbacks, and keep them resilient as sources change

  • Own Postgres schema design and migrations end to end across per-country and per-domain schemas

  • Build and maintain medallion (bronze → silver → gold) transforms that are idempotent, content-hashed, and lineage-tracked

  • Implement and defend index methodology: normalization, weighting, and composite construction where the math verifiably says what it claims (our scoring core is held to 100% test coverage)

  • Assess data feasibility early, clarify requirements with partners, and convert ambiguous index ideas into executable plans

  • Take an index end to end: sourcing, validation, methodology, publication, and refresh planning

  • Operate pipelines on our orchestration stack (Prefect dispatching per-flow ECS Fargate tasks) with observability everywhere

What We're Looking For

Our stack is deliberately modern (Python 3.14, uv, ruff, ty, polars, Prefect 3, marimo). We don't filter on those exact tools; we hire for Python and data depth and expect a short ramp.

  • Strong Python and SQL; you have designed Postgres schemas and owned migrations (SQLAlchemy and Alembic, or equivalents) in production

  • Data pipeline experience with a lakehouse/medallion mindset: idempotent ingestion, content hashing, and lineage are habits, not aspirations

  • Web scraping beyond requests: anti-bot evasion, browser automation, and resilience against messy or hostile sources

  • Statistics literacy for index methodology: winsorization, normalization, weighting, and sensitivity testing, and you can reason about whether an index's math supports its claims

  • Comfort with modern Python tooling and CI discipline: typing, linting, coverage gates, and conventional commits

  • Product discovery instincts: you talk with partners in plain language, assess data feasibility before committing, and flag what is proven versus assumed

  • End-to-end ownership: you are a pragmatic generalist who moves across data, backend, infrastructure, and basic product decisions in an uncertain environment

Nice to Have
  • Prefect experience, or Airflow/Dagster with willingness to switch

  • AWS (ECS, S3) and Terraform

  • polars, pyarrow, and marimo or a Jupyter background

  • LLM-in-pipeline experience (pydantic-ai, AWS Bedrock, evals)

  • Actuarial, quantitative research, or data science background in ranking or index construction

  • Experience with government open data (permits, energy, environmental, or economic datasets)

  • Comfort working alongside AI tooling; our repos are agent-forward (Claude agent teams, spec-driven docs)

What We Offer
  • High-impact work at the intersection of AI and critical infrastructure regulation

  • End-to-end ownership of indices, from raw source to published methodology

  • Small team with outsized influence; your feasibility calls shape what we build

  • Modern AI-native development environment (Claude Code, Cursor, multi-model orchestration)

  • Remote-first

  • Competitive compensation

Values We Hire For
  • Character: integrity and trustworthiness above all

  • Competency: evoking trust and reliably delivering

  • Togetherness: family-level support and alignment

  • Impact: meaningful outcomes over activity

  • Commitment: ownership and follow-through

Equal Opportunity Statement

We’re an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability, or veteran status, or any other basis protected by law.

Skills Required

  • Strong Python (3.x) and SQL skills
  • Design and ownership of Postgres schemas and migrations (SQLAlchemy, Alembic or equivalent)
  • Data pipeline experience with lakehouse/medallion mindset: idempotent ingestion, content hashing, lineage
  • Web scraping against adversarial sources: HTTP/2 clients, TLS-fingerprint evasion, browser automation fallbacks
  • Statistics literacy for index methodology: winsorization, normalization, weighting, percentile ranks, sensitivity testing
  • Comfort with modern Python tooling and CI: typing, linting, test coverage gates, conventional commits
  • Product discovery and partner communication: assess data feasibility and convert ambiguous requirements into plans
  • End-to-end ownership across data, backend, infrastructure, and product decisions
  • Prefect experience (or Airflow/Dagster) for orchestration
  • AWS experience (ECS, S3) and Terraform
  • Experience with polars, pyarrow, marimo, or similar data tooling / Jupyter background
  • LLM-in-pipeline experience (pydantic-ai, AWS Bedrock, evals) or familiarity with AI tooling
  • Actuarial, quantitative research, or data science background in ranking/index construction
  • Experience with government open data (permits, energy, environmental, economic datasets)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
17 Employees
Year Founded: 2023

What We Do

Infinity incubates companies focused on AI service businesses, combining repeat founders with world class applied AI engineers creating the next generation of service industries.

Similar Jobs

InterSystems Logo InterSystems

Technical Specialist

Artificial Intelligence • Big Data • Healthtech • Machine Learning • Software • Database • Analytics
Easy Apply
Remote
Chile
2100 Employees

Deepgram Logo Deepgram

Research Staff, LLMs

Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
In-Office or Remote
49 Locations
150 Employees
150K-250K Annually
Easy Apply
Remote
37 Locations
55 Employees
140K-178K Annually

InterSystems Logo InterSystems

Executive Assistant

Artificial Intelligence • Big Data • Healthtech • Machine Learning • Software • Database • Analytics
Easy Apply
Remote
Chile
2100 Employees

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account