Data Engineer- GCP/Databricks

Posted 4 Days Ago
Be an Early Applicant
Bengaluru North, Yelahanka, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Artificial Intelligence • Software • Consulting • Generative AI
The Role
Design, build, and optimize scalable batch and streaming data pipelines across GCP, Databricks, and lakehouse environments. Develop ingestion, transformation, analytical, and ML-supporting data layers using Spark, Delta Lake, BigQuery, and SQL. Implement governance, quality controls, schema evolution, observability, CI/CD, and security practices. Collaborate with architects, AI engineers, and data scientists while mentoring junior engineers and contributing to technical design, documentation, code reviews, and agile delivery.
Summary Generated by Built In

AuxoAI is seeking a Senior Data Engineer to lead the design, development, and optimisation of modern data pipelines and cloud-native platforms. This role is ideal for someone with deep experience building scalable batch and streaming data workflows across cloud and lakehouse environments, strong hands-on engineering skills, and a drive to mentor junior engineers.

You will work closely with AI engineers, solution architects, and cross-functional teams to build production-grade pipelines spanning ingestion, transformation, and curated data delivery — enabling high-quality data for AI and analytics use cases at scale.

Location: Bangalore / Mumbai / Hyderabad / Gurgaon (Hybrid — 3 days in office)

Responsibilities

  • Design and build scalable batch and streaming data pipelines across bronze, silver, and gold medallion layers.
  • Build historical and incremental ingestion using Auto Loader/Spark Structured Streaming/Kafka feeds, with GCS/Azure/AWS storage and Databricks Jobs orchestration.
  • Develop and maintain Databricks-based pipelines using Spark and Delta Lake for lakehouse architecture, including migration of legacy or on-premises data sources.
  • Design and maintain analytical data layers in BigQuery or Databricks SQL, applying best practices in partitioning, clustering, and performance tuning.
  • Implement SQL/PySpark transformations for wide and semi-structured data, including wide-to-long processing and typed or hybrid models suited to consumer requirements.
  • Collaborate with AI engineers and data scientists to build pipelines that feed ML models, AI agents, and analytical systems.
  • Implement data governance, quality controls, and security best practices including schema enforcement, lineage tracking, and access controls.
  • Drive engineering best practices across CI/CD, testing, monitoring, and pipeline observability.
  • Partner with solution architects to translate data requirements into technical designs.
  • Mentor junior data engineers and contribute to documentation, code reviews, and agile ceremonies.


Requirements
  • 5+ years of hands-on experience in data engineering, building and operating production-grade pipelines.
  • Hands-on experience with Databricks on GCP, including BigQuery, GCS, Databricks, Spark, Delta Lake, and structured streaming.
  • Hands-on experience with Databricks and Apache Spark, including Delta Lake and end-to-end lakehouse implementations.
  • Strong programming skills in Python and/or Scala, with solid SQL for modelling and transformation.
  • Experience with data modelling, ETL/ELT, pipeline orchestration, and data warehousing concepts, including experience working with large, evolving JSON/map/array payloads, wide-to-long transformations, event-time context joins and schema-change handling.
  • Familiarity with Git, CI/CD pipelines, and data quality monitoring frameworks.
  • Solid understanding of data architecture, schema design, and performance tuning.
  • Experience with Unity Catalog, source reconciliation, schema evolution, correction handling and replay/recovery testing.
  • Strong problem-solving and collaboration skills.

Bonus Skills

  • GCP Professional Data Engineer certification.
  • Experience with Vertex AI, Cloud Functions, Dataproc, or real-time streaming architectures.
  • Experience with factory or industrial data sources — MES systems, IoT sensor streams, or operational telemetry.
  • Familiarity with data governance and cataloguing tools such as Dataplex, Unity Catalog, Atlan, or Collibra.
  • Exposure to Docker, Kubernetes, API integration, and infrastructure-as-code (Terraform).


Skills Required

  • 5+ years of hands-on data engineering experience building and operating production-grade pipelines.
  • Hands-on experience with Databricks on GCP, including BigQuery, GCS, Databricks, Spark, Delta Lake, and structured streaming.
  • Hands-on experience with Databricks and Apache Spark, including Delta Lake and end-to-end lakehouse implementations.
  • Strong programming skills in Python and/or Scala, with solid SQL for modeling and transformation.
  • Experience with data modeling, ETL/ELT, pipeline orchestration, data warehousing, large evolving JSON/map/array payloads, wide-to-long transformations, event-time context joins, and schema-change handling.
  • Familiarity with Git, CI/CD pipelines, and data quality monitoring frameworks.
  • Solid understanding of data architecture, schema design, and performance tuning.
  • Experience with Unity Catalog, source reconciliation, schema evolution, correction handling, and replay/recovery testing.
  • Strong problem-solving and collaboration skills.
  • GCP Professional Data Engineer certification.
  • Experience with Vertex AI, Cloud Functions, Dataproc, or real-time streaming architectures.
  • Experience with factory or industrial data sources such as MES systems, IoT sensor streams, or operational telemetry.
  • Familiarity with Dataplex, Unity Catalog, Atlan, or Collibra.
  • Exposure to Docker, Kubernetes, API integration, and Terraform.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
Year Founded: 2022

What We Do

AuxoAI partners with enterprise leaders to build AI systems, enabling the creation of AI-first enterprises by moving from AI strategy to production-grade deployed systems in weeks.

Similar Jobs

Zscaler Logo Zscaler

Site Reliability Engineer

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Hybrid
Bangalore, Bengaluru, Karnataka, IND
8697 Employees

JPMorganChase Logo JPMorganChase

Client Data

Financial Services
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Firmware Validation Engineer (Python - AI, Automation)

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
85422 Employees
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account