AI Pipeline Engineer

Posted 2 Days Ago
Be an Early Applicant
Tempe, AZ, USA
In-Office
100K-150K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
The Role
Design, build, and operate petabyte-scale AI data pipelines for training and evaluation across multimodal data. Implement ingestion, cleaning, lineage, versioning, high-throughput loading, labeling/active-learning workflows, privacy controls, observability, and cost/performance optimizations while collaborating with ML researchers and documenting systems.
Summary Generated by Built In
AI Pipeline Engineer- Remote 
 
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. 
 
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. 

 
Job Title:
AI Pipeline Engineer

Location: 100% Remote (U.S.) 
Position Type: Full-time, Direct W2 
Salary Range: $100,000–$150,000 Annually 
Experience Required: 6+ years 
 
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position. 
 

Job Summary 
We are seeking an AI Pipeline Engineer to build and operate the large-scale data systems that power modern AI training and evaluation pipelines. The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities. The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and clear understanding of how data infrastructure choices propagate into model quality and training efficiency.
Key Responsibilities
  • Design and operate large-scale data pipelines supporting AI training, evaluation, and continual improvement workflows.
  • Build ingestion systems for diverse modalities including text, image, audio, video, and structured signals.
  • Implement data cleaning, deduplication, filtering, and quality assurance at petabyte scale.
  • Develop dataset versioning, lineage, and provenance tracking systems suitable for reproducible training.
  • Build high-throughput data loading systems that maximize GPU utilization during training.
  • Implement labeling workflows, active learning pipelines, and human-in-the-loop data improvement systems.
  • Design storage architectures balancing cost, throughput, and latency across data tiers.
  • Build evaluation dataset construction pipelines with strict integrity and contamination controls.
  • Implement data privacy, redaction, and consent enforcement throughout the pipeline.
  • Collaborate with ML researchers and engineers to align data systems with model development needs.
  • Drive observability of data quality, drift, and pipeline health across the AI data estate.
  • Optimize cost and performance through compression, format selection, and caching strategies.
  • Document data systems, schemas, and operational procedures for broad internal use.
  • Stay current with AI data infrastructure research and emerging open-source tools.
Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Six or more years of data engineering experience, with significant work supporting ML or AI workloads.
  • Strong proficiency in Python and at least one JVM or systems language.
  • Deep experience with modern data processing frameworks such as Spark, Ray, or Beam.
  • Hands-on experience operating petabyte-scale storage and pipeline systems.
  • Strong understanding of distributed systems, data modeling, and storage formats.
  • Experience with dataset versioning, lineage, and reproducibility for ML workflows.
  • Familiarity with high-throughput data loading for accelerator-based training.
  • Strong software engineering practices including testing, CI/CD, and code review.
  • Excellent communication and cross-functional collaboration skills.
Preferred Qualifications
  • Experience with multimodal datasets at large scale.
  • Familiarity with data quality tooling and dataset evaluation methodology.
  • Exposure to privacy-preserving data systems and regulated data handling.
  • Open-source contributions to data infrastructure projects.
  • Experience supporting frontier model training pipelines.

How to Apply 
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3544. Learn more about Bright Vision Technologies at www.bvteck.com
 
Bright Vision Technologies is an Equal Opportunity Employer. 
 

Skills Required

  • Bachelor's or Master's degree in Computer Science or a related field
  • Six or more years of data engineering experience supporting ML or AI workloads
  • Strong proficiency in Python
  • Proficiency in at least one JVM or systems language
  • Deep experience with modern data processing frameworks such as Spark, Ray, or Beam
  • Hands-on experience operating petabyte-scale storage and pipeline systems
  • Strong understanding of distributed systems, data modeling, and storage formats
  • Experience with dataset versioning, lineage, and reproducibility for ML workflows
  • Familiarity with high-throughput data loading for accelerator-based training
  • Strong software engineering practices including testing, CI/CD, and code review
  • Excellent communication and cross-functional collaboration skills
  • Experience with multimodal datasets at large scale
  • Familiarity with data quality tooling and dataset evaluation methodology
  • Exposure to privacy-preserving data systems and regulated data handling
  • Open-source contributions to data infrastructure projects
  • Experience supporting frontier model training pipelines
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
53 Employees
Year Founded: 2020

What We Do

Bright Vision Technologies is a minority-owned organization founded in July 2020 and based in New Jersey, USA. The company specializes in delivering top-tier staffing and IT consulting services, including custom computer programming and systems design. Additionally, they are a product engineering firm with a flagship AI-powered talent intelligence and enterprise automation platform called Lumina, which helps transform IT into a strategic asset for their valued partners.

Similar Jobs

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
10 Locations
40000 Employees
44K-100K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
9 Locations
40000 Employees
45K-85K Annually

Wipfli Logo Wipfli

Manager, Financial Reporting - FQHC Industry Clients

Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Remote or Hybrid
United States
3000 Employees
97K-145K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
11 Locations
40000 Employees
45K-85K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account