Senior Data Engineer

Posted 5 Days Ago
Be an Early Applicant
Hiring Remotely in Poland
Remote
Senior level
Information Technology • Consulting
The Role
Design and maintain high-throughput data pipelines that integrate AST outputs, knowledge graphs, and vector indexes; build Neo4j graph schemas; integrate Tree-sitter and Qdrant embeddings; ensure data lineage, snapshot versioning, validation, and monitoring; collaborate with AI/ML engineers to optimize retrieval and LLM orchestration for automated code documentation.
Summary Generated by Built In

Ciklum is looking for a Senior Data Engineer to join our team full-time in Poland.

We are a custom product engineering company that supports both multinational organizations and scaling startups to solve their most complex business challenges. With a global team of over 4,000 highly skilled developers, consultants, analysts and product owners, we engineer technology that redefines industries and shapes the way people live.

About the role:

As a Senior Data Engineer, become a part of a cross-functional development team engineering experiences of tomorrow. 

The Legacy Code Semantic Documentation Project is a 26-week enterprise initiative for a global industrial automation leader. The primary objective is to engineer an automated, AI-assisted pipeline to generate structured, system-level Markdown documentation directly from an undocumented ~400K LOC codebase (spanning IEC 61131-3 languages and ANSI C/C++).

Operating within a dedicated, zero-data-egress secure tenant, the technical architecture combines Tree-sitter AST parsing, Neo4j knowledge graphs, Qdrant vector search, and self-hosted open-weight LLMs (Llama 3.1 / Mixtral family) with NLI-based validation. Delivery is structured across 7 work packages executing over a 6-month period, governed by strict contractual KPIs: ≥95% code coverage, ≥92% NLI-verified factual precision, and ≥90% SME validation acceptance. All outputs align with EU Cyber Resilience Act (EU CRA) requirements for SBOM and source-level traceability.

Responsibilities:

  • Data Pipeline Architecture: Build and optimize high-throughput data processing pipelines linking static code analysis outputs, graph databases, vector indexes, and LLM orchestration layers
  • Knowledge Graph Engineering: Design and maintain Neo4j graph schemas to map code dependencies, function calls, and module hierarchies across 400K+ LOC
  • AST Integration: Integrate Abstract Syntax Tree (AST) parser outputs (Tree-sitter) into downstream knowledge graphs and vector embeddings (Qdrant)
  • Data Lineage & Integrity: Implement bidirectional cross-reference tracking and graph snapshot versioning to ensure data consistency across batch documentation runs
  • Pipeline Testing & Monitoring: Establish data validation protocols and automated testing for pipeline health, throughput, and state persistence
  • Technical Collaboration: Partner closely with AI/ML Engineers and Architects to optimize retrieval quality and context assembly for RAG pipelines

Requirements:

  • Professional Experience: 5+ years of Data Engineering experience building complex, production-grade data pipelines and ETL/ELT workflows
  • Domain Specialization (Must cover at least ONE key domain):
    • Option 1 (Static Analysis Focus): Hands-on experience with Abstract Syntax Tree (AST) parsing, static code analysis (e.g., Tree-sitter, language grammars), and code structure modeling
    • Option 2 (Knowledge Graph Focus): Direct experience with graph database design and optimization (Neo4j/Cypher), documentation processing pipelines, and vector search integration (Qdrant)
  • Core Technical Stack: High proficiency in Python, graph query languages, vector storage, and modern data engineering toolkits
  • Infrastructure & Deployment: Familiarity with containerized execution (Docker, Kubernetes) in secure, private, or air-gapped environments
  • Quality & Engineering Standards: Strong focus on code quality, automated pipeline testing, CI/CD, and robust versioning
  • Language: Full professional proficiency in English (C1+)

What’s in it for you?

  • Strong community: Work alongside top professionals in a friendly, open-door environment
  • Growth focus: Take on large-scale projects with a global impact and expand your expertise
  • Tailored learning: Boost your skills with internal events (meetups, conferences, workshops), Udemy access, language courses, and company-paid certifications
  • Endless opportunities: Explore diverse domains through internal mobility, finding the best fit to gain hands-on experience with cutting-edge technologies
  • Flexibility: Enjoy flexibility – full remote working possibilities
  • Care: We've got you covered with company-paid premium medical package through Luxmed
  • Benefits: Access the MyBenefit cafeteria platform, allowing you to choose perks that best suit your lifestyle and needs

About us:

At Ciklum, we are always exploring innovations, empowering each other to achieve more, and engineering solutions that matter. With us, you’ll work with cutting-edge technologies, contribute to impactful projects, and be part of a One Team culture that values collaboration and progress.

With delivery centers in Wrocław and Gdańsk, our 300+ professionals in Poland drive forward-thinking solutions for global clients. Join a community where collaboration sparks innovation—and your impact reaches millions.

Explore, empower, engineer with Ciklum!

Interested already? We would love to get to know you! Submit your application. We can’t wait to see you at Ciklum.

#LI-MP1

Skills Required

  • 5+ years of Data Engineering experience building complex, production-grade data pipelines and ETL/ELT workflows
  • Hands-on experience with Abstract Syntax Tree (AST) parsing and static code analysis (e.g., Tree-sitter)
  • Experience with graph database design and optimization using Neo4j and Cypher
  • Experience integrating vector search/storage (Qdrant) and producing vector embeddings
  • High proficiency in Python and graph query languages
  • Familiarity with containerized execution and deployment (Docker, Kubernetes) in secure/private or air-gapped environments
  • Strong focus on code quality, automated pipeline testing, CI/CD, and robust versioning
  • Full professional proficiency in English (C1+)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
2,995 Employees
Year Founded: 2002

What We Do

Ciklum is a global Digital Solutions Company for Fortune 500 and fast-growing organisations alike around the world. The company is headquartered in London and has software development centres and branch offices in the United States, Spain, Switzerland, Denmark, Israel, Poland, Ukraine, Czech Republic, Slovakia, Romania, UAE and Pakistan. Ciklum builds tailored digital solutions that leverage emerging technologies for such clients as Just Eat, Flixbus, Metro Markets, EFG International, Zurich Insurance, Lottoland and others. For more information about us visit www.ciklum.com

Similar Jobs

In-Office or Remote
Kraków, Małopolskie, POL
1516 Employees

Xebia Logo Xebia

Senior Data Engineer

Artificial Intelligence • Cloud • Information Technology • Software • Consulting • Data Privacy
Remote
3 Locations
3254 Employees

Ciklum Logo Ciklum

Senior Data Engineer

Information Technology • Consulting
Remote
Poland
2995 Employees

Alpaca Logo Alpaca

Senior Software Engineer

Fintech • Information Technology
Remote
26 Locations
132 Employees

Similar Companies Hiring

Scrunch  Thumbnail
Artificial Intelligence • Information Technology • Marketing Tech • Software • SEO
Salt Lake City, Utah
Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account