As a Data Engineer, you will be responsible for building and maintaining scalable data pipelines and architectures that enable AI and analytics solutions. You’ll work closely with data scientists, software engineers, and product teams to ensure data is clean, reliable, and ready for advanced modeling and insights.
Data Pipeline Development – Design, build, and maintain robust ETL/ELT pipelines for structured and unstructured data.
Data Integration – Ingest data from various sources (APIs, databases, cloud storage) into centralized data platforms.
Performance Optimization – Enhance data processing speed, reliability, and scalability.
Cloud Data Engineering – Work with cloud platforms (AWS, Azure, or GCP) for data storage, processing, and orchestration.
Collaboration – Partner with data scientists and analysts to ensure data availability and quality for downstream use cases.
Best Practices – Ensure code quality, data security, and maintainability through testing, automation, and documentation.
2–3 years of experience in data engineering, data warehousing, or big data environments.
Strong programming skills in Python or Java (experience with PySpark or Scala is a plus).
Hands-on experience with ETL tools and data pipeline frameworks (e.g., Airflow, Spark, Kafka).
Proficiency in working with SQL and relational databases (PostgreSQL, MySQL, etc.).
Familiarity with cloud platforms such as AWS (Glue, S3, Redshift), Azure (Data Factory, Synapse), or GCP (BigQuery, Dataflow).
Exposure to containerization (Docker, Kubernetes) and CI/CD workflows is a plus.
Passionate about data, automation, and building systems that drive intelligent decision-making.
Based in Chennai, with flexibility to collaborate with global teams.
Java | Python | Spark | Kafka | Airflow |
Skills Required
- 3-5 years experience in data engineering or data-driven product development
- Strong programming skills in Python or Java
- Experience with PySpark or Scala
- Hands-on experience with ETL tools and data pipeline frameworks (Airflow, Spark, Kafka)
- Proficiency in SQL and relational databases (PostgreSQL, MySQL)
- Familiarity with cloud data platforms (AWS Glue, S3, Redshift; Azure Data Factory, Synapse; GCP BigQuery, Dataflow)
- Exposure to containerization (Docker, Kubernetes) and CI/CD workflows
- Based in Chennai and able to collaborate with global teams (on-site)
What We Do
Crayon Data is a leading provider of AI-led revenue acceleration solutions, headquartered in Singapore with a local presence in India and the UAE. The company was founded in 2012 with the vision of simplifying the world’s choices. Our flagship platform, maya.ai, helps enterprises across the Banking, Fintech, and Travel industries create and capture sustainable revenue streams by unlocking the value of data. maya.ai's capability is driven by four “as a Service” components - Data, Recommendation, Customer Experience, and Marketplace - that work individually and together to create tangible results. Crayon Data recently won the E50 awards organized by KPMG and the Business Times in Singapore. Crayon was featured in HFS Hot Vendors Compendium in 2021. They were also among the top 15 finalists at Emerging Enterprise Awards 2019, Singapore.








