The Role
Design, optimize, and support scalable batch and streaming data pipelines using Databricks, Spark, Delta Lake, Kafka, and SQL. Architect governed data lakes with Unity Catalog, implement data quality and access controls, and automate deployments using CI/CD, DevOps, and infrastructure-as-code tools. Collaborate on Generative AI and LLM data pipelines, troubleshoot production issues, and contribute to data platform architecture and scalability.
Summary Generated by Built In
Exp- 5+ yrs
Location- Remote (Preferred candidates to be in bangalore)
Notice- Looking candidates with to be joining within 30 Days
Key Responsibilities:
- Design, implement, and optimize scalable data pipelines using Databricks and Apache Spark.
- Architect data lakes using Delta Lake, ensuring reliable and efficient data storage.
- Manage metadata, security, and lineage through Unity Catalog for governance and compliance.
- Ingest and process streaming data using Apache Kafka and real-time frameworks.
- Collaborate with ML engineers and data scientists on LLM-based AI/GenAI project pipelines.
- Apply CI/CD and DevOps practices to automate data workflows and deployments (e.g., with GitHub Actions, Jenkins, Terraform).
- Optimize query performance and data transformations using advanced SQL.
- Implement and uphold data governance, quality, and access control policies.
- Support production data pipelines and respond to issues and performance bottlenecks.
- Contribute to architectural decisions around data strategy and platform scalability.
Requirements
Required Skills & Experience:
- 5+ years of experience in data engineering roles.
- Proven expertise in Databricks, Delta Lake, and Apache Spark (PySpark preferred).
- Deep understanding of Unity Catalog for fine-grained data governance and lineage tracking.
- Proficiency in SQL for large-scale data manipulation and analysis.
- Hands-on experience with Kafka for real-time data streaming.
- Solid understanding of CI/CD, infrastructure automation, and DevOps principles.
- Experience contributing to or supporting Generative AI / LLM projects with structured or unstructured data.
- Familiarity with cloud platforms (AWS, Azure, or GCP) and data services.
- Strong problem-solving, debugging, and system design skills.
- Excellent communication and collaboration abilities in cross-functional teams.
Skills Required
- 5+ years of experience in data engineering roles
- Expertise in Databricks, Delta Lake, and Apache Spark; PySpark preferred
- Deep understanding of Unity Catalog for data governance and lineage tracking
- Proficiency in SQL for large-scale data manipulation and analysis
- Hands-on experience with Kafka for real-time data streaming
- Understanding of CI/CD, infrastructure automation, and DevOps principles
- Experience supporting Generative AI or LLM projects with structured or unstructured data
- Familiarity with AWS, Azure, or GCP and cloud data services
- Strong problem-solving, debugging, and system design skills
- Excellent communication and cross-functional collaboration abilities
- Availability to join within 30 days
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Aptus Data Labs is a global data engineering and AI consulting company helping enterprises modernize their data landscapes, accelerate AI adoption, and build scalable digital and product capabilities. Its expertise spans artificial intelligence, generative AI, data engineering, cloud solutions, and industry platforms, supporting organizations in pharmaceuticals, financial services, manufacturing, supply chain, retail, consumer packaged goods, and technology.







