Data Engineer

Posted 3 Days Ago
Be an Early Applicant
6 Locations
Remote
Senior level
Cloud • Information Technology • Software • Consulting • Web3
The Role
Build and operate production data infrastructure, including batch and streaming pipelines, warehouses and lakehouses, retrieval systems, and data-quality controls. Design for schema evolution, reliability, cost, security, compliance, and observability using cloud platforms and modern orchestration tools. Partner directly with clients, explain technical trade-offs, and deploy containerized systems with CI/CD and infrastructure as code. Support AI applications through embedding, indexing, and retrieval pipelines.
Summary Generated by Built In

Azumo builds and operates production AI systems for companies ranging from seed-stage startups to Meta. We are hiring a Data Engineer to own the layer everything else depends on: ingestion and transformation pipelines, storage and warehouse design, and the retrieval infrastructure that AI systems query. The role is fully remote across Latin America, aligned to your client's working day.

You will not be building pipelines that work until the schema changes. Azumo has shipped production data systems since 2016, and the work here is judged downstream: whether a model can be trained on what you deliver, whether a retrieval query returns the right passage, whether a number survives being questioned by the client.

Where this role sits

Azumo's engineering organization is built around four lanes. The Data Scientist lane owns the question and the method. The AI Engineer lane owns production behavior. The Software Engineer lane owns AI-augmented product delivery. This role is the Data Engineer lane, and it owns pipelines, storage, and the retrieval layer.
One question places the boundary: when the output is wrong, whose problem is it? "The data was missing, stale, or wrong by the time it arrived" is yours."The system did the wrong thing with data that was correct" is the AI Engineer's.

Not quite your profile? Check our other openings:

- If you decide what to measure and which method answers it — Data Scientist

- If you own how an AI system behaves in production — AI Engineer

- If you ship product software with agents in your toolchain — AI-Augmented Software Engineer

- If you've done all of the above and answered to the client directly — Forward Deployed Engineer

What you will build

- Ingestion and transformation pipelines. Batch and streaming ingestion on Spark, Kafka, dbt and Airflow, with idempotency, backfills, schema evolution and late-arriving data handled by design rather than by hand.

- Storage and modeling. Warehouse and lakehouse design on Snowflake, BigQuery, Redshift or Databricks, with partitioning, file layout and query cost treated as engineering decisions.

- The retrieval layer. The chunking, embedding and indexing pipelines that feed RAG systems on pgvector, Pinecone, Qdrant or Azure AI Search, and the freshness, deduplication and permission problems that come with them.

- Data quality as a contract. Tests, expectations, lineage and alerting. If a pipeline is wrong, the people downstream should hear it from you and not from the client.

- Sensitive data by default. PII classification, masking, row and column level access, retention and deletion, and an audit trail that holds up when a client asks who read what.

- Production operation. Containerized deployment on Azure or AWS, CI/CD, orchestration, observability, and explicit cost and runtime budgets that you own rather than discover after the invoice.

- Work inside the client's environment. Their repositories, their standups, sometimes their customer calls. Azumo is SOC 2 certified, client code stays in client repositories, and some engagements carry additional requirements such as HIPAA.

How we work

Our engineers build with AI every day. Claude Code, Codex, and similar tools are part of the standard toolchain here, not an experiment. We run an automated audit across the whole codebase on day one and every day after, grading security, cost, and architecture findings by severity with the exact file and line, so a small team can move quickly without quality drifting. We stay vendor-neutral across OpenAI, Anthropic, and open-weight models, and we run Valkyrie, our own production layer, when a single interface to any model is the right call.

About Azumo

Azumo is a San Francisco based software development company that has been building intelligent applications since 2016. We provide nearshore AI engineering teams to organizations that need production AI faster than they can hire for it: as an embedded engineering team, as AI staff augmentation alongside an existing team, or as a full project build. Our engineers work from Latin America, aligned to United States time zones, and have delivered for Twitter, Meta, Discovery Channel, Omnicom, UnitedHealth, and CENTEGIX.

We hire for seniority and test for it before anyone joins a client team. We support engineers in going deep on the modern AI stack, and we give time back to open-source work, community teaching, and philanthropy.

Apply at https://azumo.com/join-our-team or write to us at [email protected].


Requirements

Basic qualifications

- 5+ years building and operating production data pipelines, with Python and SQL as your primary languages, plus the engineering fundamentals that go with it: testing, code review, CI/CD, Git, containers, and orchestration.

- Deep expertise in designing and building data warehouses or lakehouses, including dimensional modeling, incremental processing, and the cost and performance trade-offs behind each choice.

- Distributed processing at production scale with Spark, Kafka, Flink or equivalent, including the failure modes that only appear under load.

- Orchestration as an engineering discipline rather than a cron replacement: Airflow, Dagster or Prefect, with retries, idempotency and backfill strategy you can defend.

- Transformation under version control, with tests and lineage: dbt or something you built yourself.

- Cloud deployment experience, Azure preferred and AWS acceptable, with Docker, CI/CD pipelines, and infrastructure as code (GitHub Actions, Terraform, or Bicep).

- Working discipline around pipeline cost, runtime and throughput. You can explain what a pipeline costs to run and what you did about it.

- Active use of AI-assisted coding tools such as Claude Code, Cursor, or GitHub Copilot in real delivery work.

- Clear written and spoken English, C1 or above, and the confidence to explain a technical trade-off directly to a client.

- Bachelor's degree in Computer Science, Data Science, or a related field, or equivalent professional experience.

Preferred qualifications

- Vector and retrieval infrastructure: pgvector, Pinecone, Qdrant, FAISS or Azure AI Search, and the retrieval-quality problems that come with it.

- Streaming, real-time or high-throughput workloads.

- Experience with cloud-based managed services like Airflow, Glue, Elastic stack, Amazon Redshift, Snowflake, BigQuery, Azure SQL Db, EMR, Databricks.

- Prior experience with notebooks using Jupyter, Google Collab, or similar.

- Delivery under a compliance regime such as SOC 2 or HIPAA.

- Contributions to open-source data libraries, published technical writing, or active participation in the data engineering community.


Benefits
  • Paid time off (PTO)
  • U.S. Holidays
  • AI Training
  • Mentored career development
  • Profit sharing
  • $US remuneration

Skills Required

  • 5+ years building and operating production data pipelines
  • Python and SQL experience
  • Testing, code review, CI/CD, Git, containers, and orchestration experience
  • Deep expertise designing and building data warehouses or lakehouses
  • Experience with dimensional modeling, incremental processing, and cost and performance trade-offs
  • Production-scale distributed processing using Spark, Kafka, Flink, or equivalent
  • Experience with Airflow, Dagster, Prefect, or equivalent orchestration tools
  • Experience with retries, idempotency, and backfill strategies
  • Experience with dbt or equivalent version-controlled transformation tooling
  • Cloud deployment experience with Azure preferred or AWS acceptable
  • Docker, CI/CD pipelines, and infrastructure as code using GitHub Actions, Terraform, or Bicep
  • Ability to manage and explain pipeline cost, runtime, and throughput
  • Active use of AI-assisted coding tools such as Claude Code, Cursor, or GitHub Copilot
  • Clear written and spoken English at C1 level or above
  • Confidence explaining technical trade-offs directly to clients
  • Bachelor's degree in Computer Science, Data Science, or related field, or equivalent professional experience
  • Experience with vector and retrieval infrastructure such as pgvector, Pinecone, Qdrant, FAISS, or Azure AI Search
  • Experience with streaming, real-time, or high-throughput workloads
  • Experience with managed cloud services such as Airflow, Glue, Elastic, Redshift, Snowflake, BigQuery, Azure SQL Database, EMR, or Databricks
  • Experience using Jupyter, Google Colab, or similar notebooks
  • Experience delivering under SOC 2, HIPAA, or similar compliance regimes
  • Contributions to open-source data libraries, published technical writing, or active data engineering community participation
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
79 Employees
Year Founded: 2016

What We Do

Based in San Francisco, Azumo is a top-rated nearshore software development company. Customers work with us to scale their software development efforts and build high quality web, mobile, data and cloud applications. At Azumo, we build intelligent applications. We are passionate about technology to solve complex problems for our customers around the globe. From our proprietary AI-based solutions like, HealthyScreen.ai, Baneka NeuralDB, and myNLU to custom software development solutions that have helped our customers scale their businesses, we are focused on innovation through our nearshore development capabilities. For more information please visit www.azumo.com

Similar Jobs

NTT DATA Services Logo NTT DATA Services

Data Engineer

Big Data • Cloud • Information Technology • Analytics • Consulting
Remote
11 Locations
24200 Employees

NTT DATA Services Logo NTT DATA Services

Data Engineer

Big Data • Cloud • Information Technology • Analytics • Consulting
Remote
11 Locations
24200 Employees
22-32 Hourly

Nortal Logo Nortal

Data Engineer

Big Data • Blockchain • Software • Business Intelligence • App development • Big Data Analytics • Automation
Remote
11 Locations
2800 Employees

Hilbert's AI Logo Hilbert's AI

Data Engineer

Artificial Intelligence • Information Technology • Machine Learning • Software
In-Office or Remote
14 Locations
20 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account