Senior Data & Document Ingestion Engineer (OCR / RAG)

Posted 11 Hours Ago
Be an Early Applicant
7 Locations
In-Office or Remote
Senior level
Artificial Intelligence • Information Technology • Professional Services • Consulting
The Role
Build scalable pipelines to ingest and process high-volume unstructured insurance documents, including PDFs, scans, emails, Word, Excel, and PowerPoint files. Responsibilities include OCR integration, document parsing, text normalization, semantic chunking, metadata extraction, enterprise connectors, retrieval-ready data preparation for RAG systems, validation, monitoring, logging, testing, and secure cloud deployment using software engineering best practices.
Summary Generated by Built In
About Gramian

Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.

About the Role

Our client is a Big4 Consultancy group that works with leading financial institutions on AI-driven transformation, automation, advanced analytics, and financial crime prevention. Their work spans intelligent fraud detection, AML/KYC modernization, autonomous workflows, enterprise AI platforms, and the secure industrialization of AI in highly regulated environments.

We are looking for a Senior Data & Document Ingestion Engineer to build robust pipelines for processing high-volume unstructured insurance content. The role focuses on OCR, document parsing, ingestion pipelines, text normalization, semantic chunking, metadata extraction, and retrieval-ready data preparation for downstream AI systems.

CONTRACT: Contractor assignment, expected October 2026 – July 2027, with potential extension

COMMITMENT: Full-time

LOCATIONS: Europe-based, preferably CEE; remote, with potential future hybrid work in Prague

PROCESS: Initial qualification followed by technical and client interviews

NOTES: Fluent English is required.

Responsibilities
  • Design and build scalable document ingestion pipelines for PDFs, scans, emails, and office documents.
  • Integrate and optimize OCR and document extraction technologies for high-accuracy text and layout extraction.
  • Build workflows for text cleaning, normalization, semantic chunking, and metadata tagging.
  • Process unstructured formats including PDF, Word, Excel, and PowerPoint.
  • Develop connectors for enterprise sources such as SharePoint and email systems.
  • Design data schemas and retrieval mechanisms for downstream AI and RAG use cases.
  • Build validation and monitoring loops to detect low-confidence OCR or extraction results.
  • Ensure ingestion pipelines meet enterprise security, reliability, and latency requirements.
  • Implement logging, testing, and operational monitoring across data-processing workflows.
  • Apply Git, CI/CD, and software-engineering best practices to pipeline development.

Requirements
  • Approximately 5–10 years of professional data engineering or backend/data-platform experience.
  • Strong hands-on experience with Python and SQL.
  • Proven experience building data ingestion and document-processing pipelines.
  • Hands-on experience processing unstructured documents such as PDF, Word, Excel, PPT, scans, or emails.
  • Experience with OCR/document extraction tools such as AWS Textract or equivalent.
  • Professional experience building data-processing pipelines on public cloud platforms.
  • Experience with AWS services such as S3, Step Functions, and CloudWatch, or comparable cloud services.
  • Strong development practices including Git, CI/CD, and automated testing.

Preferred Qualifications

  • Experience with Azure, AWS, or Databricks in enterprise data environments.
  • Experience with vector databases, embeddings, or RAG architectures.
  • Experience designing connectors to SharePoint, email, or other enterprise content systems.
  • Background in insurance, financial services, or regulated-data environments.

Skills Required

  • Approximately 5–10 years of professional data engineering or backend/data-platform experience
  • Strong hands-on experience with Python and SQL
  • Experience building data ingestion and document-processing pipelines
  • Experience processing unstructured documents, including PDFs, Word, Excel, PowerPoint, scans, or emails
  • Experience with OCR or document extraction tools such as AWS Textract or equivalent
  • Professional experience building data-processing pipelines on public cloud platforms
  • Experience with AWS services such as S3, Step Functions, and CloudWatch, or comparable cloud services
  • Experience with Git, CI/CD, and automated testing
  • Experience with Azure, AWS, or Databricks in enterprise data environments
  • Experience with vector databases, embeddings, or RAG architectures
  • Experience designing connectors to SharePoint, email, or other enterprise content systems
  • Background in insurance, financial services, or regulated-data environments
  • Fluent English
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
6 Employees

What We Do

Gramian Consulting Group is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong foundation in software engineering and leadership, the firm helps organizations build high-performing teams by matching them with qualified professionals. They specialize in talent augmentation and recruiting, specifically focusing on connecting engineering and data/AI talent with organizations to unlock real business value.

Similar Jobs

Samsara Logo Samsara

Entry Level Tech Sales - Benelux Market (work from the UK, Netherlands, Germany or France - Dutch speaking role)

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
4 Locations
4000 Employees

Inato Logo Inato

Product Engineer

Artificial Intelligence • Greentech • Healthtech • Social Impact • Software • Biotech • Pharmaceutical
In-Office or Remote
Paris, Île-de-France, FRA
63 Employees
65K-85K Annually

Samsara Logo Samsara

Sales Manager

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
France
4000 Employees

GitLab Logo GitLab

Commercial Account Executive

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
6 Locations
2500 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account