Agentic AI, Data Engineer COA Accelerator

Posted Yesterday
Be an Early Applicant
7 Locations
In-Office or Remote
Entry level
Healthtech
The Role
Build and maintain data infrastructure for AI-enabled Clinical Outcomes Assessment capabilities. Develop pipelines for structured and unstructured clinical, regulatory, scientific, and proprietary sources; transform content into searchable, traceable, governed, retrieval-ready assets; and implement parsing, OCR, metadata, indexing, versioning, quality controls, audit trails, and hybrid retrieval. Collaborate with AI engineers, scientists, product, software, legal, security, and compliance teams to support accurate, secure, evidence-backed AI workflows.
Summary Generated by Built In
Job Description Summary

IQVIA provides scientific services spanning clinical trials, real world evidence, and consulting in all areas of the product lifecycle. Our Clinical Outcomes Assessments (COAs) organisation leads the industry in generating data to ensure that the patient voice is incorporated into the development and commercialisation of medication and other drug/non-drug interventions.

We focus on understanding and meeting the needs of our clients – mostly life science/pharmaceutical companies – through the application of broad consulting expertise and technical scientific knowledge to conduct scientifically rigorous research. This research is broad and includes qualitative, quantitative, and passive approaches to understand patient, caregiver, and healthcare professional experiences and expectations of disease and treatment.

The COA Accelerator ecosystem is expanding its AI-enabled capabilities to help internal teams and external clients generate evidence-backed COA strategy recommendations. These capabilities depend on high-quality, traceable, well-governed, AI-ready data assets spanning COA instrument information, psychometric evidence, regulatory precedent, clinical trial records, scientific publications, internal consulting deliverables, and related sources. The Data Engineer will play a critical role in transforming fragmented clinical, regulatory, scientific, and proprietary content into structured, searchable, and secure knowledge assets that power AI-enabled COA strategy workflows.

To meet our client expectations and retain the excellent reputation built up over time, the IQVIA COAs team is committed to recruiting, training and supporting driven individuals who have life science, consulting, product development, data engineering, and/or AI-enablement skills that can be applied to COA research and technology-enabled offerings.

Individuals joining us are assured of a rewarding and progressive career in patient-focused research. You’ll have the opportunity to address challenging client issues, across multiple geographies, with a hands-on influence in developing and delivering innovative solutions. We operate in a truly multi-cultural, collegial and collaborative work environment that is rich in development and growth.


Role & Responsibilities
  • Design, build, and maintain data infrastructure that supports IQVIA’s AI-enabled COA strategy and COA Accelerator capabilities.
  • Own the ingestion, transformation, normalisation, enrichment, indexing, versioning, and governance of public, proprietary, and client-specific data sources.
  • Build ingestion pipelines for structured and unstructured sources, including PDFs, Word documents, slide decks, spreadsheets, databases, APIs, clinical trial registries, regulatory documents, scientific publications, and internal repositories.
  • Transform raw source material into standardised, searchable, AI-ready formats that support evidence retrieval, source citation, recommendation generation, and expert review workflows.
  • Develop repeatable processes for document parsing, OCR, text extraction, metadata enrichment, chunking, deduplication, versioning, indexing, and quality control.
  • Provide technical support to the teams building and maintaining the platform’s core knowledge layer, including COA instrument metadata, instrument versions, translations, modes of administration, usage rights, psychometric evidence, therapeutic area mappings, endpoint usage, regulatory precedent, and related scientific evidence.
  • Support integration of public data sources such as clinical trial registries, FDA labels, EMA EPARs, HTA records, scientific literature, FDA guidance, public qualification documents, and other relevant evidence repositories.
  • Support ingestion of proprietary internal knowledge, publications, thought leadership, other expert-authored content.
  • Prepare data for retrieval-augmented generation workflows through high-quality chunking, embeddings, indexes, metadata filters, and source reference structures.
  • Collaborate with AI engineers to improve retrieval precision, recall, relevance, and citation accuracy.
  • Implement hybrid retrieval approaches combining semantic search, keyword search, structured database queries, and metadata filtering.
  • Maintain traceability between AI-generated outputs and source documents.
  • Implement data quality controls to identify incomplete, outdated, duplicated, poorly parsed, incorrectly tagged, or otherwise unreliable content.
  • Maintain audit trails for source ingestion, transformation, updates, deletions, access rights, and downstream use.
  • Work with legal, security, compliance, product, and domain stakeholders to ensure data use aligns with contractual, licensing, privacy, intellectual property, and governance requirements.
  • Collaborate across AI Engineering, Product, COA Science and Software Engineering.

Skills & Qualifications
  • Degree in computer science, data engineering, data science, information systems, bioinformatics, computational biology, engineering, or a related technical field.
  • Experience designing, building, and maintaining data pipelines for structured and unstructured data.
  • Strong Python and SQL skills.
  • Experience with APIs, relational databases, document stores, search indexes, cloud data platforms, and ETL/ELT workflows.
  • Experience handling large volumes of text-heavy documents such as PDFs, slide decks, reports, publications, regulatory files, scientific literature, or knowledge repositories.
  • Strong understanding of data cleaning, normalisation, metadata management, document parsing, indexing, lineage, versioning, and auditability.
  • Familiarity with data modelling approaches for complex knowledge domains, including scientific, clinical, regulatory, or healthcare content.
  • Ability to translate domain expert requirements into practical data structures, metadata models, retrieval-ready content, and maintainable pipelines.
  • Strong attention to detail and ability to identify data quality issues that could affect AI outputs, source traceability, or regulatory credibility.
  • Ability to collaborate effectively with AI engineers, product managers, COA scientists, software engineers, security stakeholders, legal teams, and commercial teams.
  • Strong documentation skills, including ability to document data sources, transformations, lineage, access rules, data quality checks, and known limitations.

Additional Requirements
  • Experience in life sciences, clinical research, healthcare, regulatory data, scientific publishing, HEOR, clinical outcome assessments, patient-reported outcomes, or medical evidence management is strongly preferred.
  • Experience with clinical trial registries, regulatory labels, HTA reports, scientific literature databases, medical knowledge repositories, or similar evidence sources.
  • Experience with vector databases, embeddings, semantic search, Elasticsearch/OpenSearch, Azure AI Search, Pinecone, Weaviate, Milvus, Qdrant, or similar technologies.
  • Experience with cloud data platforms such as Azure, AWS, or GCP.
  • Experience with document AI, OCR, layout-aware parsing, table extraction, metadata enrichment, taxonomy development, or controlled vocabularies is desirable.
  • Familiarity with ontology development, biomedical terminologies, controlled vocabularies, evidence classification, and structured knowledge representation.
  • Understanding of GDPR, data privacy, intellectual property constraints, licensed content, confidential client data, access control, and secure data handling.
  • Ability to work independently in a remote or hybrid environment while collaborating across global teams.
  • Significant experience leveraging AI tools for work.
  • Fluency in English.

IQVIA is a leading global provider of clinical research services, commercial insights, and healthcare intelligence to the life sciences and healthcare industries. We create intelligent connections to accelerate the development and commercialization of innovative medical treatments to help improve patient outcomes and population health worldwide. Learn more at https://jobs.iqvia.com.

IQVIA is committed to integrity in our hiring process and maintains a zero tolerance policy for candidate fraud. All information and credentials submitted in your application must be truthful and complete. Any false statements, misrepresentations, or material omissions during the recruitment process will result in immediate disqualification of your application, or termination of employment if discovered later, in accordance with applicable law. We appreciate your honesty and professionalism.

At IQVIA, we believe that diversity, inclusion, and belonging empower our mission to accelerate innovation for a healthier world. We create a culture of belonging by valuing the perspectives of all talented employees worldwide and providing them with the opportunity to power smarter healthcare for everyone, everywhere. When our talented employees bring their authentic selves and their diverse experiences to work, they enable us to accomplish extraordinary things. Multifaceted thought processes spark innovation. Multi-talented collaboration harnesses innovation to deliver superior outcomes. Likewise, as part of this culture, IQVIA is committed to ensuring effective equality between women and men, integrating it as a strategic principle in its corporate and human resources policies.

Skills Required

  • Degree in computer science, data engineering, data science, information systems, bioinformatics, computational biology, engineering, or a related technical field
  • Experience designing, building, and maintaining data pipelines for structured and unstructured data
  • Strong Python and SQL skills
  • Experience with APIs, relational databases, document stores, search indexes, cloud data platforms, and ETL/ELT workflows
  • Experience handling large volumes of text-heavy documents, including PDFs, presentations, reports, publications, regulatory files, scientific literature, or knowledge repositories
  • Strong understanding of data cleaning, normalization, metadata management, document parsing, indexing, lineage, versioning, and auditability
  • Familiarity with data modeling for complex scientific, clinical, regulatory, or healthcare knowledge domains
  • Ability to translate domain expert requirements into data structures, metadata models, retrieval-ready content, and maintainable pipelines
  • Strong attention to detail and ability to identify data quality issues affecting AI outputs, traceability, or regulatory credibility
  • Ability to collaborate with AI engineers, product managers, scientists, software engineers, security, legal, and commercial teams
  • Strong documentation skills covering data sources, transformations, lineage, access rules, quality checks, and limitations
  • Experience in life sciences, clinical research, healthcare, regulatory data, scientific publishing, HEOR, clinical outcome assessments, patient-reported outcomes, or medical evidence management
  • Experience with clinical trial registries, regulatory labels, HTA reports, scientific literature databases, or medical knowledge repositories
  • Experience with vector databases, embeddings, semantic search, Elasticsearch/OpenSearch, Azure AI Search, Pinecone, Weaviate, Milvus, Qdrant, or similar technologies
  • Experience with Azure, AWS, or GCP cloud data platforms
  • Experience with document AI, OCR, layout-aware parsing, table extraction, metadata enrichment, taxonomy development, or controlled vocabularies
  • Familiarity with ontology development, biomedical terminologies, controlled vocabularies, evidence classification, and structured knowledge representation
  • Understanding of GDPR, data privacy, intellectual property, licensed content, confidential client data, access control, and secure data handling
  • Ability to work independently in a remote or hybrid environment with global teams
  • Significant experience leveraging AI tools for work
  • Fluency in English

IQVIA Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about IQVIA and has not been reviewed or approved by IQVIA.

  • Healthcare Strength — Healthcare coverage is positioned as comprehensive, spanning medical/dental/vision plus programs like telemedicine, EAP resources, and additional insurance options. Feedback suggests the health offering is a meaningful part of the overall rewards package, though details can vary by location and plan design.
  • Retirement Support — Retirement benefits include an employer match structure that supports employee contributions through a defined formula. This adds steady long-term value to total rewards beyond base salary.
  • Leave & Time Off Breadth — Time off offerings include vacation/paid time off, holidays, and flexibility themes, with some roles described as having discretionary or unlimited time-off models. This can make the package feel more attractive even when cash compensation is viewed as only mid-range.

IQVIA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Durham, NC
61,500 Employees
Year Founded: 2016

What We Do

IQVIA (NYSE:IQV) is a leading global provider of advanced analytics, technology solutions, and clinical research services to the life sciences industry. IQVIA creates intelligent connections across all aspects of healthcare through its analytics, transformative technology, big data resources and extensive domain expertise. IQVIA Connected Intelligence™ delivers powerful insights with speed and agility — enabling customers to accelerate the clinical development and commercialization of innovative medical treatments that improve healthcare outcomes for patients. With approximately 70,000 employees, IQVIA conducts operations in more than 100 countries. To learn more, visit www.iqvia.com.

Similar Jobs

Mastercard Logo Mastercard

Content Manager

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Athens, GRC
38800 Employees

Pfizer Logo Pfizer

Senior Manager, HTA, Value and Evidence (HV&E), Genitourinary Cancer

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
30 Locations
121990 Employees
139K-232K Annually

Mastercard Logo Mastercard

Consultant

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Athens, GRC
38800 Employees

Mondelēz International Logo Mondelēz International

Brand Manager

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
Athens, GRC
90000 Employees

Similar Companies Hiring

Sailor Health Thumbnail
Healthtech • Social Impact • Telehealth
New York City, NY
20 Employees
Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account