Data Engineer — ML Training Data Pipeline

Posted 6 Days Ago
Be an Early Applicant
Hyderabad, Telangana, IND
In-Office
Senior level
Big Data • Cloud • Analytics • Consulting
The Role
Build and maintain AWS-based pipelines that transform raw production traces into high-quality LLM training datasets. Responsibilities include ingestion, deduplication, format conversion, quality filtering, identity-aware train/test splitting, schema validation, sampling, and continuous retraining-data processing. The role requires strong Python data engineering skills, large-scale JSONL processing, Hugging Face Datasets, Arrow-based storage, AWS services, and familiarity with conversational and tokenizer-specific data formats.
Summary Generated by Built In
Job Title:Data Engineer - ML Training Data Pipeline
Notice period: 0-30 Days
Experience : 5+ Years
Location: Hyderabad OR Pune 

We are looking for Data Engineer - ML Training Data Pipeline who can Build and maintain the data pipeline that transforms raw production traces into high-quality training datasets for LLM fine-tuning-ingestion, deduplication, format conversion, quality filtering, and train/test splitting at scale on AWS.


What We Expect:

  • Build end-to-end data pipelines: raw trace ingestion → dedup → format conversion → quality gating → training-ready datasets
  • Process large-scale JSONL data on AWS S3 (tens of thousands of traces per batch)
  • Convert between chat-completion formats (e.g., OpenAI → Llama 3.1 tool-calling format)
  • Implement smart deduplication and sampling to balance training distribution
  • Design identity-aware train/test splits that measure true generalization
  • Build data validation gates to detect schema drift and format anomalies
  • Create a continuous pipeline that auto-processes new production traces for retraining


Requirements
  • Experience: 6+ years data engineering focused on ML data pipelines
  • Python: Strong — pandas, pyarrow, JSONL processing at scale
  • ML Data Libraries: HuggingFace Datasets, Arrow-based storage
  • Data Formats: Multi-turn conversation/chat data structures and tokenizer-specific formatting
  • Deduplication: Content hashing, identity-based grouping strategies
  • AWS: S3, EC2, batch processing workflows

Preferred (Not Required): LLM training data prep (chat templates, tool-calling schemas); Axolotl or similar dataset formats; data versioning (DVC, LakeFS); browser-automation trace data or Playwright.



Benefits
  • Comprehensive Medical Coverage:
    Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
  • Robust Protection Plans:
    Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
  • Retirement Benefits:
    PF and Gratuity provided as per standard government regulations.
  • Flexible Work Options:
    Enjoy hybrid work arrangements & flexible working hours.
  • Generous Leave Policy:
    21 days of annual leave, in addition to 10 company-declared holidays.
  • Employee Well-being Spaces:
    Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.


Skills Required

  • 5+ years of professional experience
  • 6+ years of data engineering experience focused on ML data pipelines
  • Strong Python experience, including pandas, PyArrow, and large-scale JSONL processing
  • Experience with Hugging Face Datasets and Arrow-based storage
  • Experience processing multi-turn conversation and chat data structures with tokenizer-specific formatting
  • Experience with content hashing and identity-based deduplication strategies
  • Experience with AWS S3, EC2, and batch-processing workflows
  • Experience preparing LLM training data, including chat templates and tool-calling schemas
  • Experience with Axolotl or similar dataset formats
  • Experience with DVC or LakeFS data versioning
  • Experience with browser-automation trace data or Playwright
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
414 Employees
Year Founded: 2018

What We Do

DATAECONOMY is a global, cloud-first data and AI consultancy delivering enterprise-grade solutions through an innovative intellectual-property suite. Its work spans data and BI platform modernization, self-service AI, data mesh and fabric, master data management, governance, cloud enablement, digital engineering, knowledge graphs, and machine lakes supporting cybersecurity and financial-crime use cases for enterprise clients.

Similar Jobs

Optum Logo Optum

Full-stack Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Senior Data Analyst

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Technical Product Manager

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Senior Quality Engineer I - ACCELQ Selenium

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Similar Companies Hiring

Northslope Thumbnail
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
London, GB
100 Employees
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account