The Role
Designs, builds, maintains, and optimizes data pipelines, ETL processes, data warehouses, and data lakes. Collaborates with data scientists, analysts, software engineers, and stakeholders to develop scalable data models and infrastructure. Responsibilities include ensuring data quality, security, governance, reliability, and performance; testing and troubleshooting pipelines; documenting systems; and evaluating emerging data engineering technologies.
Summary Generated by Built In
Job Description:
As a Data Engineer, you will play a critical role in the development, implementation, and maintenance of data infrastructure and systems. Your primary responsibility will be to design, build, and optimize data pipelines and data warehouses, ensuring the efficient and reliable collection, storage, and processing of large volumes of data. You will collaborate with cross-functional teams, including data scientists, analysts, and software engineers, to understand data requirements and translate them into scalable solutions. Your work will enable the organization to extract valuable insights, drive data-based decision-making, and support various business initiatives.
Responsibilities:
- Design, develop, and maintain data pipelines and ETL processes to efficiently ingest, transform, and load data from various sources into data warehouses and data lakes.
- Collaborate with data scientists, analysts, and business stakeholders to understand data requirements and design data models that facilitate efficient data retrieval and analysis.
- Optimize data pipeline performance, ensuring scalability, reliability, and data integrity.
- Implement data governance and security measures to ensure compliance with data privacy regulations and protect sensitive information.
- Identify and implement appropriate tools and technologies to enhance data engineering capabilities and automate processes.
- Conduct thorough testing and validation of data pipelines to ensure data accuracy and quality.
- Monitor and troubleshoot data pipelines to identify and resolve issues, ensuring minimal downtime.
- Develop and maintain documentation, including data flow diagrams, technical specifications, and user guides.
- Collaborate with software engineers and infrastructure teams to optimize data infrastructure, including storage, processing, and retrieval systems.
- Stay up-to-date with emerging trends and technologies in the field of data engineering, and recommend innovative solutions to improve efficiency and performance.
Requirements:
- Bachelor's degree in Computer Science, Engineering, or a related field. A master's degree is a plus.
- Proven experience as a Data Engineer or in a similar role, with a strong understanding of data engineering concepts, practices, and tools.
- Proficiency in programming languages such as Python, Java, or Scala, and experience with data manipulation and transformation frameworks/libraries (e.g., Apache Spark, Pandas, SQL).
- Solid understanding of relational databases, data modeling, and SQL queries.
- Experience with distributed computing frameworks, such as Apache Hadoop, Apache Kafka, or Apache Flink.
- Knowledge of cloud platforms (e.g., AWS, Azure, GCP) and experience with cloud-based data engineering services (e.g., Amazon Redshift, Google BigQuery, Azure Data Factory).
- Familiarity with data warehousing concepts and technologies (e.g., dimensional modeling, columnar databases).
- Strong problem-solving skills and the ability to analyze complex data-related issues.
- Excellent communication and collaboration skills, with the ability to work effectively in cross-functional teams.
- Attention to detail and a commitment to delivering high-quality work within specified timelines.
Skills Required
- Bachelor's degree in Computer Science, Engineering, or a related field
- Proven experience as a Data Engineer or in a similar role
- Proficiency in Python, Java, or Scala
- Experience with Apache Spark, Pandas, SQL, or similar data manipulation and transformation tools
- Understanding of relational databases, data modeling, and SQL queries
- Experience with distributed computing frameworks such as Apache Hadoop, Apache Kafka, or Apache Flink
- Knowledge of AWS, Azure, or GCP cloud platforms
- Experience with cloud-based data engineering services such as Amazon Redshift, Google BigQuery, or Azure Data Factory
- Familiarity with data warehousing concepts and technologies, including dimensional modeling and columnar databases
- Strong problem-solving skills and ability to analyze complex data-related issues
- Excellent communication and cross-functional collaboration skills
- Attention to detail and commitment to delivering high-quality work within specified timelines
- Master's degree
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
ConveGenius is an education-technology company focused on making learning accessible and engaging for every child. It develops technology-enabled education solutions, including SwiftChat, SwiftPAL, Swift Insights, and Vidya Samiksha Kendra initiatives, and collaborates with governments and partners to support educational transformation. Its work combines digital learning, AI-powered tools, and large-scale outreach, with deployments spanning Indian states and a user base exceeding 150 million.







