We are looking for an experienced
Data Engineer to design, build, and maintain scalable data pipelines and
infrastructure that power analytics, reporting, and machine learning
initiatives across the organization. The ideal candidate has a strong foundation
in data modeling, distributed systems, and cloud-based data platforms, with a
track record of delivering reliable, high-performance data solutions.
● Design, build, and optimize scalable
ETL/ELT pipelines to ingest, transform, and load data from diverse sources
(databases, APIs, streaming platforms, third-party systems)
● Develop and maintain data
warehouse/data lake architectures, ensuring data quality, consistency, and
reliability
● Collaborate with data scientists,
analysts, and product teams to understand data requirements and deliver
well-structured, accessible datasets
● Build and maintain batch and
real-time streaming data pipelines using tools like Apache Spark, Kafka, or
Airflow
● Implement data quality checks,
monitoring, and alerting to ensure pipeline reliability and data integrity
● Optimize database and query
performance for large-scale datasets
● Design and maintain data models
(conceptual, logical, physical) and schemas that support analytics and
application needs
● Work with cloud platforms
(AWS/Azure/GCP) to manage data infrastructure, storage, and compute resources
● Implement and enforce data
governance, security, and compliance best practices
● Participate in code reviews, CI/CD
pipeline development, and infrastructure-as-code practices
● Document data pipelines,
architecture, and processes for team knowledge sharing
● Troubleshoot and resolve production
data pipeline issues in a timely manner
● Bachelor's degree in Computer
Science, Engineering, or a related field
● 4–6 years of hands-on experience as a
Data Engineer or in a similar role
● Strong programming skills in Python
and/or Scala; solid SQL expertise
● Experience with ETL/ELT tools and
orchestration frameworks (Apache Airflow, dbt, Luigi, or similar)
● Hands-on experience with big data
technologies (Apache Spark, Hadoop, Kafka)
● Proficiency with relational databases
(PostgreSQL, MySQL, SQL Server) and NoSQL databases (MongoDB, Cassandra,
DynamoDB)
● Experience working with cloud data
platforms and services (AWS Redshift/Glue/S3, Azure Data Factory/Synapse, GCP BigQuery/Dataflow)
● Solid understanding of data modeling
concepts — entity relationships, cardinality, normalization, and dimensional
modeling (star/snowflake schemas)
● Experience with data warehousing
solutions (Snowflake, Redshift, BigQuery, Databricks)
● Familiarity with version control
(Git) and CI/CD practices
● Understanding of data governance,
security, and privacy best practices (GDPR, data masking, access controls)
● Strong problem-solving skills and
ability to work with large, complex datasets
●
Skills Required
- Bachelor's degree in Computer Science, Engineering, or a related field
- 4-6 years of hands-on experience as a Data Engineer or in a similar role
- Strong programming skills in Python and/or Scala
- Solid SQL expertise
- Experience with ETL/ELT tools and orchestration frameworks such as Apache Airflow, dbt, or Luigi
- Hands-on experience with Apache Spark, Hadoop, and Kafka
- Proficiency with relational databases including PostgreSQL, MySQL, or SQL Server
- Proficiency with NoSQL databases including MongoDB, Cassandra, or DynamoDB
- Experience with cloud data platforms and services across AWS, Azure, or GCP
- Understanding of data modeling concepts, entity relationships, cardinality, normalization, and dimensional modeling
- Experience with data warehousing solutions such as Snowflake, Redshift, BigQuery, or Databricks
- Familiarity with Git and CI/CD practices
- Understanding of data governance, security, and privacy best practices, including GDPR, data masking, and access controls
- Strong problem-solving skills and ability to work with large, complex datasets
What We Do
Kavi Global is a data analytics and AI company that helps enterprises make intelligent, data-driven decisions. It provides analytics software, solutions, and services spanning strategy, design, development, implementation, and support. Its capabilities include business intelligence, data warehousing, big data, advanced analytics, machine learning, data management, and AI, serving organizations across multiple industries and supporting digital transformation initiatives.






