AI Data Engineer

Posted Yesterday
Be an Early Applicant
2 Locations
In-Office
Senior level
Biotech
The Role
Build domain data products and pipelines for AI consumption: create retrieval-ready, semantically annotated, contract-governed assets; deploy agents for metadata, entity resolution and classification; ensure data quality, lineage, certification, and reuse; collaborate with domain stewards, business teams and IT to model, store and serve structured and unstructured data for retrieval and modeling.
Summary Generated by Built In
Job Description

Agilent inspires and supports discoveries that advance the quality of life by providing life science, diagnostic and applied market laboratories worldwide with instruments, services, consumables, application and both measurement and asset management expertise.

Builds the data products and pipelines a pod runs on; turns raw domain data into model-ready assets. Pods do not wait for the Fabric to be complete; they build the Fabric through execution, and this role is where that happens for the data plane. Every domain data product built in a pod is constructed to certification standards from the start, because the second consumer of the asset is the point, not an afterthought. 

The role goes beyond the conventional pipeline engineering. AI consumption changes what "model-ready" means: data must be retrieval-ready, semantically annotated, contract-governed, and quality-scored, and the AI Data Engineer often uses agents to do the building, generating metadata, resolving entities, and classifying unstructured domain content rather than hand-curating at a scale that cannot hold. 

Responsible for 

  • Domain data products and pipelines serving the pod's use case, built to the data plane's certification standards: semantic definition, data contract, entitlement metadata including agent identity, lineage, and certification tier from day one. 

  • Data quality and model-readiness, including quality scoring against the defined dimensions and the quality signals that feed the evals spine; when quality falls below threshold, this role is the one who says so before the agent does something embarrassing with it. 

  • Working relationships with the domain's data owners and stewards, so that steward-validated definitions and certified sources ground the pod's retrieval rather than whatever was easiest to reach. 

  • Retrieval foundations for the pod: structured and unstructured grounding, vector and graph assets where the use case requires relationship reasoning, built on the platform estate rather than parallel infrastructure. 

  • AI-built curation in practice: deploying metadata generation, entity resolution, and content classification agents against the domain's data, contributing those outputs to the registry. 

  • Reusable assets back to the Fabric: every domain data product is a candidate for certification and enterprise reuse, documented and handed to the data plane for promotion, not a point integration that dies with the pod. 

  • Works with the business teams and IT to develop domain data models.

  • Applying broad understanding of on premise and cloud deployment topologies, creates data collection frameworks to capture, manage, store and utilize large sets of structured, unstructured and/or disconnected data from a wide variety of internal and external sources.

  • Oversees the establishment of data set processes and builds data structures based on business and technical requirements to funnel data from multiple sources and store using any combination of storage forms (e.g. cloud, local databases) as required.

  • Consults in design standards and assurance processes for software, systems and applications development to ensure compatibility and operability of data connections, flows and storage requirements.

  • Creates data tools for analytics and data scientist team members that assist them in building and optimizing our analytics products.

  • Prepares and manipulates data for predictive and prescriptive modeling. Uses data to discover tasks that can be automated. Identify ways to improve data reliability, efficiency and quality.

  • Reviews internal and external business and product requirements for data operations and activity and suggests changes and upgrades to systems and storage to accommodate ongoing needs. 

What success looks like in year one 

  • The pod's use case running entirely on contract-governed, quality-scored data products, with no undocumented side channels into source systems. 

  • Multiple domain data products from the pod certified into the registry and consumed or queued for consumption by multiple use cases. 

  • Quality signals from the pod's domain flowing into the evals spine, with at least one instance where a quality threshold correctly gated an agent behavior. 

  • Measurably faster data-to-build time driven by reuse and process optimization

Qualifications
  • Strong data engineering with experience building for AI consumption: retrieval-ready, semantically annotated, contract-governed data products, not only warehouse tables and dashboards. 

  • Hands-on familiarity with the platform estate (Microsoft Fabric, Snowflake, vector and graph stores) and with operating under data contracts, lineage, and certification requirements. 

  • Experience with RAG data foundations: chunking, embedding, hybrid retrieval, and the failure modes that show up in agent behavior rather than in pipeline monitoring. 

  • The disposition to work inside a business domain: you interview stewards, read the pipeline code that produces the data to understand what it actually means, and treat domain knowledge as the scarce input. 

  • Curiosity about AI, its potential and its pitfalls. The field moves monthly, and the people who thrive here are genuinely curious about both sides of it: what these systems can newly do, and where they fail, mislead, or quietly degrade. We want people who read the failure analyses as eagerly as the launch posts, who experiment on their own initiative, and who hold excitement and skepticism at the same time without letting either one win permanently. 

  • Lifelong learners. Whatever expertise a candidate arrives with will be partially obsolete within a year, and that is not a defect of the candidate; it is the condition of the field. We hire people who have reinvented their toolkit before and expect to do it again, who learn in public, and who treat being wrong as information rather than injury. A history of deliberate self-reinvention counts for more than any single credential. 

  • Excellent communication and the ability to influence. Nothing in this organization ships by authority alone. Every AI CoE role here persuades: domain experts to engage, stewards to share what they know, sponsors to stay honest about value, and functions like Legal, Quality, and Security to move from gatekeeping to partnership. We look for people who write and speak clearly, who adapt their register from bench scientist to Board, and who change minds through credibility and clarity rather than escalation. 

  • The instinct to generalize: you build for the second consumer of the asset, not only the first. 

  • Bachelor's or Master's Degree or equivalent.

  • Typically, at least 8+ years relevant experience for entry to this level.

Additional Details

This job has a full time weekly schedule.

Our pay ranges are determined by role, level, and location. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. During the hiring process, a recruiter can share more about the specific pay range for a preferred location. Pay and benefit information by country are available at: https://careers.agilent.com/locations

Agilent Technologies Inc. is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability or any other protected categories under all applicable laws.Travel Required: 10% of the TimeShift: DayDuration: No End DateJob Function: Administration

Skills Required

  • Strong data engineering experience building retrieval-ready, semantically annotated, contract-governed data products
  • Hands-on familiarity with Microsoft Fabric
  • Hands-on familiarity with Snowflake
  • Experience with vector stores and graph stores
  • Experience with RAG data foundations including chunking, embeddings, and hybrid retrieval
  • Experience operating under data contracts, lineage, and certification requirements
  • Experience deploying agents for metadata generation, entity resolution, and content classification
  • Ability to work within a business domain and engage stewards and data owners
  • Curiosity about AI, awareness of AI failure modes and ongoing learning
  • Excellent communication and influencing skills across technical and business stakeholders
  • Instinct to generalize assets for reuse and enterprise certification
  • Bachelor's or Master's degree or equivalent
  • Typically at least 8+ years relevant experience

Agilent Technologies Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Agilent Technologies and has not been reviewed or approved by Agilent Technologies.

  • Retirement Support The core U.S. package highlights a generous 401(k) match as a strength. Retirement programs are positioned as competitive within the company’s total rewards.
  • Equity Value & Accessibility An Employee Stock Purchase Plan at a discount provides accessible equity and augments total compensation. Ownership opportunities are presented as a notable advantage alongside retirement benefits.
  • Leave & Time Off Breadth Flexible Time Off, company holidays, a personal holiday, and paid volunteer time create a broad leave offering. Time off can accrue into multiple weeks in the first year, supporting flexibility.

Agilent Technologies Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
17,369 Employees
Year Founded: 1999

What We Do

Analytical scientists and clinical researchers worldwide rely on Agilent to help fulfill their most complex laboratory demands. Our instruments, software, services and consumables address the full range of scientific and laboratory management needs—so our customers can do what they do best: improve the world around us. Whether a laboratory is engaged in environmental testing, academic research, medical diagnostics, pharmaceuticals, petrochemicals or food testing, Agilent provides laboratory solutions to meet their full spectrum of needs. We work closely with customers to help address global trends that impact human health and the environment, and to anticipate future scientific needs. Our solutions improve the efficiency of the entire laboratory, from sample prep to data interpretation and management. Customers trust Agilent for solutions that enable insights...for a better world.

Similar Jobs

In-Office
Barcelona, Cataluña, ESP
10549 Employees
Remote or Hybrid
12 Locations
2449 Employees
172K-200K Annually
Hybrid
Barcelona, Cataluña, ESP
99 Employees

Enverus Logo Enverus

Sr. Application Services Engineer -- 26146

Big Data • Information Technology • Software • Analytics • Energy
In-Office or Remote
2 Locations
1800 Employees

Similar Companies Hiring

Formation Bio Thumbnail
Artificial Intelligence • Big Data • Healthtech • Biotech • Pharmaceutical
New York, NY
150 Employees
SOPHiA GENETICS Thumbnail
Software • Healthtech • Biotech • Big Data • Artificial Intelligence
Boston, MA
450 Employees
Pfizer Thumbnail
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
New York, NY
121990 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account