Data Engineer

Posted 3 Days Ago
Be an Early Applicant
Hiring Remotely in Poland
Remote
Mid level
Information Technology • Consulting
The Role
Build, test, and maintain data pipelines that process AST outputs, knowledge graphs, and vector embeddings. Ingest and normalize code dependency graphs into Neo4j and vector indexes in Qdrant, transform source code into structured Markdown/JSON, monitor pipeline runs, debug errors, and collaborate with senior engineers to optimize retrieval and pipeline efficiency.
Summary Generated by Built In

Ciklum is looking for a Data Engineer to join our team full-time in Poland.

We are a custom product engineering company that supports both multinational organizations and scaling startups to solve their most complex business challenges. With a global team of over 4,000 highly skilled developers, consultants, analysts and product owners, we engineer technology that redefines industries and shapes the way people live.

About the role:

As a Data Engineer, become a part of a cross-functional development team engineering experiences of tomorrow.

The Legacy Code Semantic Documentation Project is a 26-week enterprise initiative for a global industrial automation leader. The primary objective is to engineer an automated, AI-assisted pipeline to generate structured, system-level Markdown documentation directly from an undocumented ~400K LOC codebase (spanning IEC 61131-3 languages and ANSI C/C++).

Operating within a dedicated, zero-data-egress secure tenant, the technical architecture combines Tree-sitter AST parsing, Neo4j knowledge graphs, Qdrant vector search, and self-hosted open-weight LLMs (Llama 3.1 / Mixtral family) with NLI-based validation. Delivery is structured across 7 work packages executing over a 6-month period, governed by strict contractual KPIs: ≥95% code coverage, ≥92% NLI-verified factual precision, and ≥90% SME validation acceptance. All outputs align with EU Cyber Resilience Act (EU CRA) requirements for SBOM and source-level traceability.

Responsibilities:

  • Data Pipeline Execution: Develop, test, and maintain robust data pipelines that process AST outputs, knowledge graph structures, and vector embeddings
  • Knowledge Base Ingestion: Write data processing scripts to ingest, normalize, and update code dependency graphs in Neo4j and vector indexes in Qdrant
  • Data Cleansing & Transformation: Parse raw source code structures and raw metadata into structured Markdown assets and standardized JSON/RAG inputs
  • Pipeline Monitoring & Debugging: Monitor execution throughput, resolve batch processing errors, and ensure data state consistency across pipeline runs
  • Collaboration: Work closely with Senior Data Engineers and AI/ML Engineers to optimize data retrieval speeds and pipeline efficiency

Requirements:

  • Professional Experience: 3+ years of commercial Data Engineering experience building and maintaining production data pipelines
  • Core Technical Proficiency: Solid proficiency in Python, SQL, and standard data manipulation frameworks
  • Hands-on Stack Exposure: Direct working knowledge of graph databases (Neo4j), vector search engines (Qdrant), or static code parsing frameworks (Tree-sitter)
  • Engineering Best Practices: Proficiency with Git version control, Docker containerization, unit/integration testing for data pipelines, and CI/CD workflows
  • Problem-Solving & Detail Orientation: Strong analytical skills with a focus on data accuracy, schema consistency, and output validation
  • Language: Professional proficiency in English (B2+/C1)

What’s in it for you?

  • Strong community: Work alongside top professionals in a friendly, open-door environment
  • Growth focus: Take on large-scale projects with a global impact and expand your expertise
  • Tailored learning: Boost your skills with internal events (meetups, conferences, workshops), Udemy access, language courses, and company-paid certifications
  • Endless opportunities: Explore diverse domains through internal mobility, finding the best fit to gain hands-on experience with cutting-edge technologies
  • Flexibility: Enjoy flexibility – full remote working possibilities
  • Care: We've got you covered with company-paid premium medical package through Luxmed
  • Benefits: Access the MyBenefit cafeteria platform, allowing you to choose perks that best suit your lifestyle and needs

About us:

At Ciklum, we are always exploring innovations, empowering each other to achieve more, and engineering solutions that matter. With us, you’ll work with cutting-edge technologies, contribute to impactful projects, and be part of a One Team culture that values collaboration and progress.

With delivery centers in Wrocław and Gdańsk, our 300+ professionals in Poland drive forward-thinking solutions for global clients. Join a community where collaboration sparks innovation—and your impact reaches millions.

Explore, empower, engineer with Ciklum!

Interested already? We would love to get to know you! Submit your application. We can’t wait to see you at Ciklum.

#LI-RS1

Skills Required

  • 3+ years of commercial Data Engineering experience building and maintaining production data pipelines
  • Proficiency in Python, SQL, and standard data manipulation frameworks
  • Direct working knowledge of graph databases (Neo4j), vector search engines (Qdrant), or static code parsing frameworks (Tree-sitter)
  • Proficiency with Git version control, Docker containerization, unit/integration testing, and CI/CD workflows
  • Strong analytical skills with focus on data accuracy, schema consistency, and output validation
  • Professional proficiency in English (B2+/C1)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
2,995 Employees
Year Founded: 2002

What We Do

Ciklum is a global Digital Solutions Company for Fortune 500 and fast-growing organisations alike around the world. The company is headquartered in London and has software development centres and branch offices in the United States, Spain, Switzerland, Denmark, Israel, Poland, Ukraine, Czech Republic, Slovakia, Romania, UAE and Pakistan. Ciklum builds tailored digital solutions that leverage emerging technologies for such clients as Just Eat, Flixbus, Metro Markets, EFG International, Zurich Insurance, Lottoland and others. For more information about us visit www.ciklum.com

Similar Jobs

Xebia Logo Xebia

Senior Data Engineer

Artificial Intelligence • Cloud • Information Technology • Software • Consulting • Data Privacy
Remote
3 Locations
3254 Employees

Ruby Labs Logo Ruby Labs

Data Engineer

Information Technology • Software
In-Office or Remote
25 Locations
28 Employees

Ciklum Logo Ciklum

Data Engineer

Information Technology • Consulting
Remote
Poland
2995 Employees

Primer (UK) Logo Primer (UK)

Data Engineer

eCommerce • Fintech • Payments • Software • Financial Services
Remote
7 Locations
166 Employees

Similar Companies Hiring

Scrunch  Thumbnail
Artificial Intelligence • Information Technology • Marketing Tech • Software • SEO
Salt Lake City, Utah
Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account