Senior Data & Document Ingestion Engineer (OCR / RAG) - REMOTE

Posted 25 Days Ago
Be an Early Applicant
6 Locations
Remote
Senior level
Artificial Intelligence • Information Technology • Professional Services • Consulting
The Role
Build scalable pipelines to ingest and process high-volume unstructured insurance documents, including PDFs, scans, emails, Word, Excel, and PowerPoint files. Responsibilities include OCR integration, document parsing, text normalization, semantic chunking, metadata extraction, enterprise connectors, retrieval-ready data preparation for RAG systems, validation, monitoring, logging, testing, and secure cloud deployment using software engineering best practices.
Summary Generated by Built In
About Gramian

Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.

About the Role

Our client is a Big4 Consultancy group that works with leading financial institutions on AI-driven transformation, automation, advanced analytics, and financial crime prevention. Their work spans intelligent fraud detection, AML/KYC modernization, autonomous workflows, enterprise AI platforms, and the secure industrialization of AI in highly regulated environments.

We are looking for a Senior Data & Document Ingestion Engineer to build robust pipelines for processing high-volume unstructured insurance content. The role focuses on OCR, document parsing, ingestion pipelines, text normalization, semantic chunking, metadata extraction, and retrieval-ready data preparation for downstream AI systems.

CONTRACT: Contractor assignment, expected October 2026 – July 2027, with extension to other projects (and retention rate)

COMMITMENT: Full-time

LOCATIONS: REMOTE 100%, Europe-based

PROCESS: Initial qualification followed by technical and client interviews

NOTES: Fluent English is required. Must be able to work in EU.

Responsibilities
  • Design and build scalable document ingestion pipelines for PDFs, scans, emails, and office documents.
  • Integrate and optimize OCR and document extraction technologies for high-accuracy text and layout extraction.
  • Build workflows for text cleaning, normalization, semantic chunking, and metadata tagging.
  • Process unstructured formats including PDF, Word, Excel, and PowerPoint.
  • Develop connectors for enterprise sources such as SharePoint and email systems.
  • Design data schemas and retrieval mechanisms for downstream AI and RAG use cases.
  • Build validation and monitoring loops to detect low-confidence OCR or extraction results.
  • Ensure ingestion pipelines meet enterprise security, reliability, and latency requirements.
  • Implement logging, testing, and operational monitoring across data-processing workflows.
  • Apply Git, CI/CD, and software-engineering best practices to pipeline development.

Requirements
  • Approximately 5–10 years of professional data engineering or backend/data-platform experience.
  • Strong hands-on experience with Python and SQL.
  • Proven experience building data ingestion and document-processing pipelines.
  • Hands-on experience processing unstructured documents such as PDF, Word, Excel, PPT, scans, or emails.
  • Experience with OCR/document extraction tools such as AWS Textract or equivalent.
  • Professional experience building data-processing pipelines on public cloud platforms.
  • Experience with AWS services such as S3, Step Functions, and CloudWatch, or comparable cloud services.
  • Strong development practices including Git, CI/CD, and automated testing.

Preferred Qualifications

  • Experience with Azure, AWS, or Databricks in enterprise data environments.
  • Experience with vector databases, embeddings, or RAG architectures.
  • Experience designing connectors to SharePoint, email, or other enterprise content systems.
  • Background in insurance, financial services, or regulated-data environments.

Skills Required

  • Approximately 5–10 years of professional data engineering or backend/data-platform experience
  • Strong hands-on experience with Python and SQL
  • Experience building data ingestion and document-processing pipelines
  • Experience processing unstructured documents, including PDFs, Word, Excel, PowerPoint, scans, or emails
  • Experience with OCR or document extraction tools such as AWS Textract or equivalent
  • Professional experience building data-processing pipelines on public cloud platforms
  • Experience with AWS services such as S3, Step Functions, and CloudWatch, or comparable cloud services
  • Experience with Git, CI/CD, and automated testing
  • Experience with Azure, AWS, or Databricks in enterprise data environments
  • Experience with vector databases, embeddings, or RAG architectures
  • Experience designing connectors to SharePoint, email, or other enterprise content systems
  • Background in insurance, financial services, or regulated-data environments
  • Fluent English
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
6 Employees

What We Do

Gramian Consulting Group is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong foundation in software engineering and leadership, the firm helps organizations build high-performing teams by matching them with qualified professionals. They specialize in talent augmentation and recruiting, specifically focusing on connecting engineering and data/AI talent with organizations to unlock real business value.

Similar Jobs

Pfizer Logo Pfizer

Senior Manager, Scientific Learning, Medical Academy

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote or Hybrid
26 Locations
121990 Employees

Pfizer Logo Pfizer

Sr. Director, Product Intelligence & Marketing Lead

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
30 Locations
121990 Employees
215K-358K Annually

Pfizer Logo Pfizer

Director, AI Platform Product Management

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
30 Locations
121990 Employees
163K-272K Annually

Pfizer Logo Pfizer

Director, Build Engineer

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
30 Locations
121990 Employees
177K-294K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account