Senior Data Engineer / SSE

Posted 5 Hours Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Artificial Intelligence • Consumer Web • HR Tech • Information Technology
The Role
Build and operate Apna’s scalable data platform, including batch and real-time pipelines, lakehouse architecture, query platforms, orchestration workflows, curated datasets, and data marts. Improve reliability, observability, data quality, lineage, SLAs, query performance, and cost efficiency. Partner with product, analytics, ML, and backend teams to deliver scalable data solutions, establish engineering standards, own production systems, and mentor data engineers.
Summary Generated by Built In

Role: Senior Data Engineer / SSE

Requirement: 1

Team: Data Platform / Engineering

Location: Work from Office - Domlur, Bangalore (5 days / week)

Experience : 4-6 Years of Experience

Why Join Apna

At Apna, data is central to how we build products, understand users, improve employer outcomes, power recommendations, and scale decision-making. This role gives you the opportunity to build the backbone of Apna’s data platform and influence how data is used across the company.

You will work on real-world, high-scale problems across jobs, users, employers, communities, matching, growth, and AI-driven systems.

About the Role

Apna is looking for a Senior Software Engineer to build and scale our core data platform. This role will work on large-scale data pipelines, lakehouse architecture, query platforms, workflow orchestration, and data reliability systems that power analytics, product intelligence, machine learning, business dashboards, experimentation, and operational decision-making across Apna.

We are looking for someone who can think deeply about data architecture, design reliable pipelines, improve data quality, and help build a platform that can scale with Apna’s growth.


Requirements

What You’ll Own:

You will be responsible for designing, building, and operating critical parts of Apna’s data platform, including:

  • Building scalable batch and near-real-time data pipelines across product, business, growth, and ML use cases.
  • Designing and improving our lakehouse architecture using technologies likeApache Hudi.
  • Working with query engines such asPresto / Trinofor large-scale analytical workloads.
  • Building and maintaining orchestration workflows usingApache Airflow.
  • Creating reusable data models, curated datasets, and reliable data marts for analytics and product teams.
  • Improving data platform reliability, observability, SLA tracking, lineage, and data quality checks.
  • Optimizing storage, compute, query performance, and pipeline costs.
  • Partnering with product, analytics, ML, and backend engineering teams to understand data needs and convert them into scalable platform solutions.
  • Driving engineering standards around data modeling, schema evolution, partitioning, deduplication, backfills, replayability, and pipeline ownership.
  • Mentoring data engineers and influencing architecture decisions across teams.

What We’re Looking For

Must Have

  • Strong experience indata engineering, preferably at scale.
  • Hands-on experience withApache Airflowor similar orchestration systems.
  • Strong knowledge ofPresto / Trinoor other distributed query engines.
  • Good understanding ofApache Hudiconcepts such as:
    • Copy-on-write vs merge-on-read
    • Upserts and deletes
    • Incremental reads
    • Compaction
    • Clustering
    • Timeline and commits
    • Schema evolution
    • Partitioning strategy
  • Strong knowledge of distributed data processing and storage systems.
  • Ability to design and build reliable ETL / ELT pipelines.
  • Strong SQL skills and ability to debug complex data issues.
  • Good understanding of different data architectures, including:
    • Data warehouse
    • Data lake
    • Lakehouse
    • Lambda architecture
    • Kappa architecture
    • Medallion architecture
    • Event-driven data architecture
  • Experience with data modeling for analytics and reporting.
  • Strong programming skills in at least one language such asPython, Java, or Scala.
  • Ability to reason about trade-offs between freshness, cost, reliability, latency, and complexity.
  • Strong debugging and production ownership mindset.

Good to Have

  • Experience with Kafka, Spark, Flink, Hive, Iceberg, Delta Lake, or BigQuery.
  • Experience building internal data platforms or self-serve data infrastructure.
  • Experience with data quality frameworks such as Great Expectations, Deequ, Soda, or custom validation systems.
  • Exposure to ML feature pipelines or feature stores.
  • Experience with metadata management, data catalogs, lineage, and governance.
  • Experience with cloud infrastructure such as AWS, GCP, or Azure.
  • Understanding of privacy, compliance, PII handling, and access control in data systems.

What Success Looks Like
In this role, success means:

  • Critical business and product datasets are reliable, discoverable, and trusted.
  • Pipelines are observable, recoverable, and have clear SLAs.
  • Query performance improves across major analytical workloads.
  • Data freshness and quality issues reduce significantly.
  • Teams can build on top of the data platform faster without reinventing pipelines.
  • The platform can scale with Apna’s user, job, employer, and engagement data.

Skills Required

  • 4-6 years of professional experience
  • Strong data engineering experience, preferably at scale
  • Hands-on experience with Apache Airflow or similar orchestration systems
  • Strong knowledge of Presto, Trino, or other distributed query engines
  • Understanding of Apache Hudi, including upserts, deletes, incremental reads, compaction, clustering, schema evolution, and partitioning
  • Strong knowledge of distributed data processing and storage systems
  • Ability to design and build reliable ETL or ELT pipelines
  • Strong SQL skills and ability to debug complex data issues
  • Understanding of data warehouse, data lake, lakehouse, Lambda, Kappa, medallion, and event-driven architectures
  • Experience with data modeling for analytics and reporting
  • Strong programming skills in at least one of Python, Java, or Scala
  • Ability to evaluate trade-offs among freshness, cost, reliability, latency, and complexity
  • Strong debugging and production ownership mindset
  • Experience with Kafka, Spark, Flink, Hive, Iceberg, Delta Lake, or BigQuery
  • Experience building internal data platforms or self-service data infrastructure
  • Experience with data quality frameworks such as Great Expectations, Deequ, Soda, or custom validation systems
  • Exposure to ML feature pipelines or feature stores
  • Experience with metadata management, data catalogs, lineage, and governance
  • Experience with AWS, GCP, or Azure cloud infrastructure
  • Understanding of privacy, compliance, PII handling, and access control in data systems
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Year Founded: 2019

What We Do

Apna is India's largest professional networking and jobs platform, connecting job seekers, particularly blue and grey-collar workers, with employers. It facilitates job discovery, skill development, and professional networking.

Similar Jobs

Samsara Logo Samsara

Senior Saleforce Developer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
4000 Employees

Ericsson Logo Ericsson

Experienced Engineer - FPGA Verification

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
88000 Employees

Ericsson Logo Ericsson

System Verification Engineer II

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
88000 Employees
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account