Data Lead- Dallas, TX

Reposted One Month Ago
Be an Early Applicant
2 Locations
In-Office or Remote
46K-162K Annually
Expert/Leader
Agency • Information Technology
The Role
Design and scale end-to-end data pipelines and RAG architectures for LLM consumption: ETL/ELT, vector databases, chunking/embeddings, data cleaning/PII removal, metadata engineering, low-latency streaming, and evaluation/versioning to support agentic AI.
Summary Generated by Built In

We are seeking a Lead Data Engineer to build and scale the data infrastructure powering our Agentic AI products. You will be responsible for the "Ingestion-to-Insight" pipeline that allows autonomous agents to access, search, and reason over vast amounts of proprietary and public data.

Your role is critical: you will design the RAG (Retrieval-Augmented Generation) architectures and data pipelines that ensure our agents have the right context at the right time to make accurate decisions.

Key Responsibilities

  • AI-Ready Data Pipelines: Design and implement scalable ETL/ELT pipelines that process both structured (SQL, logs) and unstructured (PDFs, emails, docs) data specifically for LLM consumption.
  • Vector Database Management: Architect and optimize Vector Databases (e.g., Pinecone, Weaviate, Milvus, or Qdrant) to ensure high-speed, relevant similarity searches for agentic retrieval.
  • Chunking & Embedding Strategies: Collaborate with AI Engineers to optimize data chunking strategies and embedding models to improve the "recall" and "precision" of the agent's knowledge retrieval.
  • Data Quality for AI: Develop automated "Data Cleaning" workflows to remove noise, PII (Personally Identifiable Information), and toxicity from training/context datasets.
  • Metadata Engineering: Enrich raw data with advanced metadata tagging to help agents filter and prioritize information during multi-step reasoning tasks.
  • Real-time Data Streaming: Build low-latency data streams (using Kafka or Flink) to provide agents with "fresh" data, enabling them to act on real-time market or operational changes.
  • Evaluation Frameworks: Construct "Gold Datasets" and versioned data snapshots to help the team benchmark agent performance over time.

Required Skills & Qualifications

  • Experience: 10+ years in Data Engineering, with at least 2 years focusing on data for LLMs or AI/ML applications.
  • Python Mastery: Deep expertise in Python (Pandas, Pydantic, FastAPI) for data manipulation and API integration.
  • Data Tooling: Strong experience with modern data stack tools (e.g., dbt, Airflow, Dagster, Snowflake, or Databricks).
  • Vector Expertise: Hands-on experience with at least one major Vector Database and knowledge of similarity search algorithms (HNSW, Cosine Similarity).
  • Search Knowledge: Familiarity with hybrid search techniques (combining semantic search with traditional keyword search like Elasticsearch/BM25).
  • Cloud Infrastructure: Proficiency in managing data workloads on AWS, Azure, or GCP.

Preferred Qualifications

  • Experience with LlamaIndex or LangChain for data ingestion.
  • Knowledge of Graph Databases (e.g., Neo4j) to help agents understand complex relationships between data points.
  • Familiarity with "Data-Centric AI" principles—prioritizing data quality over model size.

Compensation, Benefits and Duration

Minimum Compensation: USD 46,000
Maximum Compensation: USD 162,000
Compensation is based on actual experience and qualifications of the candidate. The above is a reasonable and a good faith estimate for the role.
Medical, vision, and dental benefits, 401k retirement plan, variable pay/incentives, paid time off, and paid holidays are available for full time employees.
This position is not available for independent contractors
No applications will be considered if received more than 120 days after the date of this post

Skills Required

  • 10+ years in Data Engineering with at least 2 years focused on data for LLMs or AI/ML applications
  • Deep expertise in Python (Pandas, Pydantic, FastAPI)
  • Strong experience with modern data stack tools (dbt, Airflow, Dagster, Snowflake, Databricks)
  • Hands-on experience with at least one major Vector Database (Pinecone, Weaviate, Milvus, Qdrant) and familiarity with similarity search algorithms (HNSW, Cosine Similarity)
  • Familiarity with hybrid search techniques combining semantic search and keyword search (e.g., Elasticsearch/BM25)
  • Proficiency managing data workloads on cloud platforms (AWS, Azure, or GCP)
  • Experience building low-latency data streaming pipelines (Kafka or Flink)
  • Experience with LlamaIndex or LangChain for data ingestion
  • Knowledge of graph databases (e.g., Neo4j)
  • Familiarity with Data-Centric AI principles (prioritizing data quality)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
5,017 Employees
Year Founded: 2007

What We Do

Photon.com has emerged as one of the world’s largest and fastest-growing Digital Agencies. We work with 40% of the Fortune 100 on their Digital initiatives and are known for our ability to integrate Strategy Consulting, Creative Design, and Technology at scale. Please visit www.photon.com to learn more about us, how we work, and our customer case studies. Digital Transformation Starts Here.

Similar Jobs

Capital One Logo Capital One

Artificial Intelligence Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
5 Locations
55000 Employees
286K-392K Annually

Optum Logo Optum

Director, Executive and Enterprise Storytelling - Remote

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
135K-231K Annually

Optum Logo Optum

LVN or LPN, Case Manager - Remote

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Cypress, CA, USA
160000 Employees
20-36 Hourly

Optum Logo Optum

Account Manager

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Remote
Pennsylvania, USA
160000 Employees
20-36 Hourly

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account