We are seeking an experienced Data Engineer. The ideal candidate is self-motivated, a multitasker, and a demonstrated team player. You will be responsible for designing, developing, managing, and maintaining our open-source data platform, including our Data-Lakehouse (S3, Apache Iceberg, and ClickHouse), ETL processes, and orchestration tool (Temporal Workflow).
What You Will Do● Develop a scalable data platform integrating multiple sources for easy access.
● Design and enhance data tools (orchestration, governance, Data-Lakehouse, BI, etc.).
● Ensure smooth operation of data systems for analysts, scientists, and engineers.
● Optimize data pipelines (ingestion, processing, and output) in a microservices environment.
● Build, maintain, and monitor ETL/ELT processes and orchestrate workflows using Temporal.
● Troubleshoot and improve the performance, scalability, and reliability of the data infrastructure (S3, Apache Iceberg, ClickHouse).
● Collaborate cross-functionally with data scientists, analysts, and backend engineers to understand data needs and deliver solutions.
● Implement and champion data quality, governance, and security best practices across the platform.
Requirements● 3+ years of experience as a Data Engineer or in a similar data infrastructure role.
● Strong proficiency in SQL and hands-on experience with data modeling.
● Experience with data lake/lakehouse architectures (e.g., Apache Iceberg, S3, or similar).
● Experience with analytical / columnar databases (e.g., ClickHouse or similar).
● Experience building and orchestrating ETL/ELT pipelines (e.g., Temporal, Airflow, or similar).
● Strong programming skills in Python and/or Scala/Java.
● Experience working within a microservices architecture and cloud environments (AWS preferred).
● Self-motivated, strong multitasking skills, and a demonstrated team player.
● Excellent communication skills and the ability to work both independently and collaboratively.
● Hands-on experience with Apache Spark (or similar technologies) for large-scale data processing.
● Professional proficiency in written and spoken English.
● Note: this role is focused on batch data processing (not real-time streaming).
Nice to Have
● Experience working with and contributing to open-source data platforms and tools.
● Familiarity with BI and visualization tools (e.g., Superset, Looker, Tableau, Metabase, or similar).
● Experience with containerization and orchestration (Docker, Kubernetes).
● Experience with infrastructure-as-code and CI/CD practices.
● Experience with AWS EMR and running Apache Spark workloads in a cloud environment.
● Experience leveraging AI-assisted development tools (e.g., GitHub Copilot, Cursor, or similar) to boost engineering productivity.
Skills Required
- 3+ years experience as a Data Engineer or similar data infrastructure role
- Strong proficiency in SQL and data modeling
- Experience with data lake/lakehouse architectures (Apache Iceberg, S3)
- Experience with analytical/columnar databases (ClickHouse or similar)
- Experience building and orchestrating ETL/ELT pipelines (Temporal, Airflow, or similar)
- Strong programming skills in Python and/or Scala or Java
- Experience working within a microservices architecture
- Experience with cloud environments
- AWS experience (preferred)
- Hands-on experience with Apache Spark or similar technologies for large-scale data processing
- Professional proficiency in written and spoken English
- Self-motivated, strong multitasking skills, and demonstrated team player
- Excellent communication skills and ability to work independently and collaboratively
- Experience contributing to open-source data platforms and tools
- Familiarity with BI and visualization tools (Superset, Looker, Tableau, Metabase)
- Experience with containerization and orchestration (Docker, Kubernetes)
- Experience with infrastructure-as-code and CI/CD practices
- Experience with AWS EMR and running Apache Spark workloads in a cloud environment
- Experience leveraging AI-assisted development tools (GitHub Copilot, Cursor, or similar)
What We Do
Onebeat provides a cloud enterprise platform for retailers to manage merchandise and inventory as an integrated system. Its machine-learning capabilities address long-term forecasting inaccuracies and use short-term predictions to translate consumer behavior into daily execution. The platform supports product introduction, assortment management, smart allocation, liquidation, and discounting, helping retailers improve sell-through, margins, sales, and operational efficiency across global markets and optimize stock availability.








