Data Engineer

Posted Yesterday
Be an Early Applicant
Hiring Remotely in București, ROU
In-Office or Remote
Mid level
Information Technology
The Role
Build and maintain high-volume Kafka ingestion pipelines, develop Spark loaders and enrichments, design and optimize Iceberg lakehouse models, create dbt/Dremio transformations, manage Kubernetes/Helm deployments and CI/CD, and implement monitoring, PII handling, and governance for a multi-tenant telecom CDP.
Summary Generated by Built In

About VOX

VOX is a visionary company led by a single founder, currently leading the way in flashcall and telecom carrier services, transforming the way businesses communicate, authenticate and connect. As a hyper-growth company, VOX achieved over 25% YoY revenue growth last year and is aiming to reach $100M+ revenue this year. VOX is looking for a team of growth-driven individuals to take the company to new heights.

VOX's cutting-edge technology and dedicated customer service team ensure that telcos and enterprises maintain secure, fast, and reliable connections while protecting their networks. VOX's promise of a hassle-free experience and superior customer support enables telcos and enterprises to focus on success. As a company, VOX focuses on solutions that monetize the assets of mobile network operators. 

Joining VOX offers the opportunity to work with the industry's leading technologies and help them stay ahead and continue to innovate with a comprehensive suite of flashcall and telecom carrier services. VOX is highly committed to providing its employees with a dynamic, forward-thinking work environment, competitive compensation and benefits, vacation and time-off packages, and stock options. This is a once-in-a-lifetime opportunity for highly ambitious individuals, as VOX plans to expand its solutions portfolio and go public in the next 3-5 years.

 About the Role

VOX is building a multi-tenant Customer Data Platform for mobile network operators across multiple countries. Our platform ingests billions of events from telecom traffic  and transforms them into actionable insights, segmentation, and campaign activation.

As a Data Engineer on the VOX CDP team, you will work across Kafka ingestion, Spark processing, Iceberg/Nessie lake house modeling, Dremio/dbt transformations, and Kubernetes-based multi-tenant deployments.

This is a role for someone who wants to work deeply with high-volume event data and a modern, cloud-native analytics architecture.

 Responsibilities

 Event Ingestion & Streaming (Kafka KRaft)

  • Build and maintain Kafka ingestion pipelines

  • Define topic structures, partition strategies, retention policies, and consumer logic for multi-tenant setups

  • Manage data contracts and schema evolution

  • Develop idempotent ingestion services that land data into Iceberg tables

Lakehouse Architecture (Iceberg + Nessie)

  • Design and optimize Iceberg tables (partitioning, compaction, clustering, retention rules)

  • Work with Nessie branches/tags to manage multi-environment (dev/test/prod) and multi-MNO deployments

  • Implement Python/Spark loaders writing from Kafka → Iceberg

  • Manage Iceberg compaction, metadata pruning, snapshot control, and performance tuning

Distributed Processing & Enrichment (Spark)

  • Develop Spark jobs (batch + micro-batch where needed) for:
    Cleaning and normalizing events
    Categorizing senders
    Engagement signals
    Identity stitching and grouping
    Audience enrichment and behavioral metrics

  • Ensure Spark jobs scale efficiently across large volumes of event data

Query Layer & Transformations (Dremio + dbt)

  • Build dbt models on top of Iceberg datasets via Dremio and dbt

  • Deliver telecom-specific analytical models including:
    Descriptive Analytics
    Quality/quantity audience scoring
    Campaign performance metrics
    RFU relevance scoring
    Cohort segmentation pipelines

  • Optimise Dremio queries using reflections, column pruning, and Iceberg metadata

Kubernetes, Helm, and CI/CD

  • Maintain Helm charts for each VOX deployment (multiple clusters)

  • Build CI/CD pipelines (GitHub Actions/GitLab/Argo) for:
    Ingestion services
    Spark job deployment
    Kafka topic configs
    Dbt model updates
    Helm releases into customer clusters

  • Automate rollouts, config updates, and monitoring installation

Observe-ability, Quality & Governance

  • Implement monitoring for ingestion lag, consumer errors, Iceberg table health, Spark jobs, and Dremio performance

  • Implement custom Python-based data validation checks where needed

  • Handle all PII with strict tenant isolation and encryption (Vault)

  • Ensure compliant and governed data flows across all deployments.

Collaboration & Product Development

  • Work with product teams to translate telecom and marketing requirements into scalable data models

  • Support analysts and ML engineers with clean, enriched datasets from Iceberg

  • Collaborate with DevOps on cluster performance, scaling, and stability

Requirements

  • 3+ years of experience as a Data Engineer or equivalent, with documented experience in working with big datasets

  • Strong Python engineering skills (must-have)

  • Experience building pipelines on Apache Kafka (KRaft mode preferred)

  • Strong SQL + experience with Iceberg table design and optimization

  • Experience with Spark for large-scale processing

  • Experience with dbt and SQL modeling on lakehouse storage

  • Experience working with Dremio, Trino, or similar query engines

  • Experience with Kubernetes, Helm, and Git-based CI/CD

  • Understanding of PII handling, encryption, and compliance requirements

  • Ability to work in distributed, multi-environment setups (dev/test/prod + multi-deployments)

Nice to Have

  • Experience with telecom data structures (CDRs)

  • Experience with Nessie catalogs (branching, tagging, schema versioning)

  • Understanding of audience-building, scoring, or marketing activation models

  • Experience tuning object storage (S3/MinIO)

 Join the team and help shape the future of the telecom industry!

Skills Required

  • 3+ years experience as a Data Engineer or equivalent working with large datasets
  • Strong Python engineering skills
  • Experience building pipelines on Apache Kafka (KRaft mode preferred)
  • Strong SQL and experience with Apache Iceberg table design and optimization
  • Experience with Apache Spark for large-scale processing (batch and micro-batch)
  • Experience with dbt and SQL modeling on lakehouse storage
  • Experience with Dremio, Trino, or similar query engines
  • Experience with Kubernetes, Helm, and Git-based CI/CD (GitHub Actions/GitLab/Argo)
  • Understanding of PII handling, encryption, and compliance requirements (Vault mentioned)
  • Ability to work in distributed, multi-environment setups (dev/test/prod + multi-deployments)
  • Experience with Nessie catalogs (branching, tagging, schema versioning)
  • Experience with telecom data structures (CDRs) and audience/marketing modeling
  • Experience tuning object storage (S3/MinIO)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Hong Kong, Hong Kong
154 Employees

What We Do

VOX Solutions optimises, accelerates and simplifies International Voice and Messaging through innovating in technology, platforms and processes. We serve operators, carriers, aggregators and enterprises worldwide delivering an array of services such as A2P messaging, firewall managed services, service monetization and operational outsourcing. Our business is focused enabling partners to be successful in today’s Voice and messaging market through efficiently expanding reach, optimising operations, enhancing user experience, and mitigating fraud. We are a 10 ranked independent Telecom solution provider and the largest and most innovative businesses in the international Voice and messaging market trust us to deliver. Our award-winning VOX360 managed solution protect MNOs network from fraudulent routing, sms, voice and flash-call and illicit bypass attempts. We’re bringing Voice and messaging into the future and turning a pain point into new profitability.

Similar Jobs

Ruby Labs Logo Ruby Labs

Data Engineer

Information Technology • Software
In-Office or Remote
25 Locations
28 Employees

Primer (UK) Logo Primer (UK)

Data Engineer

eCommerce • Fintech • Payments • Software • Financial Services
Remote
7 Locations
166 Employees

Growe Talents Logo Growe Talents

Data Engineer

Agency • Gaming • HR Tech • Professional Services
In-Office or Remote
28 Locations
19 Employees

Miratech Logo Miratech

Data Engineer

Information Technology
In-Office or Remote
Bucharest, București, ROU
701 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account