Lead Data Engineer (Databricks, PySpark & GCP)

Posted Yesterday
Be an Early Applicant
Hyderabad, Telangana, IND
Hybrid
Expert/Leader
Artificial Intelligence • Big Data • Cloud • Information Technology • Machine Learning
The Role
Leads the design, architecture, development, testing, and support of scalable ETL and ELT data pipelines using Python, PySpark, Databricks, GCP, and Azure. Responsibilities include data ingestion, transformation, quality validation, Medallion and Delta Lake architecture, Spark optimization, cloud migrations, workflow orchestration, CI/CD, documentation, and direct customer requirements gathering. The role also collaborates with data scientists, analysts, and engineering teams to deliver reliable enterprise data solutions.
Summary Generated by Built In

Job Overview:

           
We are looking for a skilled and motivated Lead Data Engineer with strong experience in Python programming, PySpark, Databricks and Google Cloud Platform (GCP) to join our data engineering team. The ideal candidate will be responsible for requirements gathering, designing, architecting the solution, developing, and maintaining robust and scalable ETL (Extract, Transform, Load) & ELT data pipelines. The role involves working with customers directly, gathering requirements, discovery phase,  designing, architecting the solution, using various GCP services, implementing data transformations, data ingestion, data quality, and consistency across systems, and post post-delivery support.

Experience Level:

10 to 16 years of relevant IT experience

Key Responsibilities:

    • Design, develop, test, and maintain scalable ETL data pipelines using Python, PySpark, Databricks & GCP / Azure.

      • Architect the enterprise solutions with various technologies like GCP, Azure, Databricks, PySpark and Spark SQL.

      • Work extensively on Google Cloud Platform (GCP) services such as:

        • Dataflow for real-time and batch data processing

        • Cloud Functions for lightweight serverless compute

        • BigQuery for data warehousing and analytics

        • Cloud Composer for orchestration of data workflows (on Apache Airflow)

        • Google Cloud Storage (GCS) for managing data at scale

        • IAM for access control and security

        • Cloud Run for containerized applications

Should have experience in the following areas :

    • Develop production-grade Databricks notebooks and workflows.

    • Build data transformation pipelines using PySpark and Spark SQL.

    • Implement Delta Lake architecture.

    • Design Bronze, Silver, and Gold data layers using the Medallion Architecture.

    • Implement Databricks Workflows/Jobs and dependency management.

    • Tune Spark jobs for large-scale data processing.

    • Optimize cluster configuration and compute utilization.

    • Implement appropriate partitioning, caching, and file-size optimization strategies.

    • Perform data ingestion from various sources and apply transformation and cleansing logic to ensure high-quality data delivery.

      • Implement and enforce data quality checks, validation rules, and monitoring.

        • Collaborate with data scientists, analysts, and other engineering teams to understand data needs and deliver efficient data solutions.

          • Manage version control using GitHub and participate in CI/CD pipeline deployments for data projects.

            • Write complex SQL queries for data extraction and validation from relational databases such as SQL Server, Oracle, or PostgreSQL.

              • Document pipeline designs, data flow diagrams, and operational support procedures.

Required Skills:

    • 10+ years of hands-on experience in Python for backend or data engineering projects.

      • Strong understanding and working experience with GCP cloud services (especially Dataflow, BigQuery, Cloud Functions, Cloud Composer, etc.).

      • Working experience with Azure Data Factory (ADF), Azure Databricks, Azure Data Lake Storage Gen2 (ADLS).

        • Solid understanding of data pipeline architecture, data integration, and transformation techniques.

          • Experience in working with version control systems like GitHub and knowledge of CI/CD practices.

          • Experience in Apache Spark, Kafka, Redis, Fast APIs, Airflow, GCP Composer DAGs.

            • Strong experience in SQL with at least one enterprise database (SQL Server, Oracle, PostgreSQL, etc.).

            • Experience with PySpark is required.

            • Experience in data migrations from on-premise data sources to Cloud platforms.

            • Good to Have (Optional Skills):

              • Experience with AWS services.

              • Additional Details:

                • Excellent problem-solving and analytical skills.

                • Strong communication skills and ability to collaborate in a team environment.

                • Education:

                  ● Bachelor's degree in Computer Science, a related field, or equivalent experience.

Skills Required

  • 10+ years of hands-on Python experience in backend or data engineering projects
  • 10 to 16 years of relevant IT experience
  • Strong experience with GCP services, including Dataflow, BigQuery, Cloud Functions, and Cloud Composer
  • Experience with Azure Data Factory, Azure Databricks, and Azure Data Lake Storage Gen2
  • Experience designing data pipeline architecture, data integration, and transformation solutions
  • Experience with Databricks, PySpark, Apache Spark, and Spark SQL
  • Experience implementing Delta Lake and Bronze, Silver, and Gold Medallion Architecture layers
  • Experience with GitHub, version control, and CI/CD practices
  • Experience with Apache Kafka, Redis, FastAPI, Airflow, and GCP Composer DAGs
  • Strong SQL experience with an enterprise database such as SQL Server, Oracle, or PostgreSQL
  • Experience migrating data from on-premises sources to cloud platforms
  • Bachelor's degree in Computer Science, a related field, or equivalent experience
  • Experience with AWS services
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Naperville, IL
240 Employees
Year Founded: 2000

What We Do

Egen is a data engineering and cloud modernization firm partnering with leading Chicagoland companies to launch, scale, and modernize industry-changing technologies. We are catalysts for change who create digital breakthroughs at warp speed. Our team of cloud and data engineering experts are trusted by top clients in pursuit of the extraordinary. Our mission is to be an enabler of amazing possibilities for companies looking to use the power of cloud and data. We want to stand shoulder to shoulder with clients, as true technology partners, and make sure they succeed at what they have set out to do. We want to be disruptors, game-changers, and innovators who have played an important part in moving the world forward.

Similar Jobs

Optum Logo Optum

Assistant Manager

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Machine Learning Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Director AI/ML Engineering

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Senior Data Scientist

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account