Senior Data Engineer - Apache Spark and SQL - Vice President

Posted 4 Days Ago
Be an Early Applicant
Pune, Mahārāshtra, IND
In-Office
Senior level
Fintech • Financial Services
The Role
Modernize legacy Hadoop data platforms by building, optimizing, and simplifying Spark pipelines on Databricks and AWS. Design scalable data models and reusable pipeline components, improve batch performance, develop testing and validation frameworks, and support production releases. Collaborate with architects, platform teams, DevOps engineers, and global stakeholders on technical decisions, troubleshooting, and implementation guidance.
Summary Generated by Built In

We are looking for a highly skilled Senior Databricks Engineer to contribute to the engineering, modernization, and continuous evolution of data processing platform on Databricks on AWS. While supporting the transition from the legacy Cloudera Hadoop platform to Databricks on AWS, this role will continue to play a key part in enhancing performance, simplifying pipelines, and delivering new capabilities on the Databricks platform over the long term.

The ideal candidate is a strong hands‑on Spark engineer with solid design experience, capable of contributing to architectural decisions while leading complex implementation and optimization efforts.

Responsibilities:

1. Platform Engineering & Modernization

  • Refactor and modernize existing Spark pipelines to Databricks native architectures
  • Eliminate legacy Hadoop dependencies and adopt cloud native AWS patterns
  • Enhance and extend existing processing logic using optimized Spark (JavaSpark / PySpark) on Databricks

2. Databricks Native Development

  • Build and optimize solutions using Databricks features, including Delta Lake, Databricks Workflows for orchestration and Auto scaling and job clusters

3. Design & Solution Engineering

  • Contribute to low and mid level architecture and design
  • Translate high level architecture into detailed technical designs
  • Define data models, pipeline patterns, and reusable components
  • Ensure solutions are scalable, maintainable, and production ready

4. Performance Optimization & Simplification

  • Analyze, improve Spark job performance and simplify complex or over engineered pipelines into standardized, efficient patterns

5. Engineering Standards & Best Practices

  • Follow and contribute to Databricks and Spark engineering standards
  • Write clean, modular, and testable code
  • Contribute to shared frameworks, reusable libraries, and quality standards

6. Collaboration & Stakeholder Engagement

  • Work closely with senior architects, platform teams, and DevOps engineers
  • Provide technical inputs, troubleshooting support, and implementation guidance
  • Participate in design discussions and technical decision making

7. Testing & Quality Assurance

  • Develop unit, integration, and data validation tests
  • Support production releases and post deployment validation

Qualifications:

Core Technical Skills

  • 11+ years in data engineering or distributed systems
  • Strong expertise in Apache Spark (JavaSpark / PySpark), Databricks on AWS, and Delta Lake
  • Experience on SQL
  • Experience with AWS services and large‑scale distributed data processing

Modernization & Optimization Experience

  • Experience modernizing or refactoring legacy data platforms into cloud‑based architectures
  • Strong background in Spark performance tuning and large‑scale batch optimization

Design Capability

  • Ability to translate architecture into implementable designs
  • Understanding of data modeling and pipeline orchestration patterns

Behavioral Competencies

  • Strong problem‑solving mindset for complex distributed systems
  • Comfortable working in time‑bound, high‑impact environments
  • Proactive, accountable, and collaborative
  • Clear communication skills across global teams

Education:

  • Bachelor’s degree/University degree or equivalent experience

------------------------------------------------------

Job Family Group: Technology

------------------------------------------------------

Job Family:Applications Development

------------------------------------------------------

Time Type:Full time

------------------------------------------------------

Most Relevant Skills Please see the requirements listed above.

------------------------------------------------------

Other Relevant Skills For complementary skills, please see above and/or contact the recruiter.

------------------------------------------------------

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

 

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.
View Citi’s EEO Policy Statement and the Know Your Rights poster.

Skills Required

  • 11+ years of experience in data engineering or distributed systems
  • Strong expertise in Apache Spark, including JavaSpark or PySpark
  • Experience with Databricks on AWS
  • Experience with Delta Lake
  • Experience with SQL
  • Experience with AWS services and large-scale distributed data processing
  • Experience modernizing or refactoring legacy data platforms into cloud-based architectures
  • Strong background in Spark performance tuning and large-scale batch optimization
  • Ability to translate architecture into implementable technical designs
  • Understanding of data modeling and pipeline orchestration patterns
  • Strong problem-solving skills for complex distributed systems
  • Ability to work in time-bound, high-impact environments
  • Proactive, accountable, and collaborative working style
  • Clear communication skills across global teams
  • Bachelor's degree, university degree, or equivalent experience

Citi Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Citi and has not been reviewed or approved by Citi.

  • Healthcare Strength Benefits coverage is positioned as comprehensive, including health, dental, and vision insurance plus on-site clinics, prescription drug support, and disability coverage. Family-building support such as fertility assistance is described as a notable differentiator within the overall package.
  • Retirement Support Retirement benefits are framed as strong, highlighted by a 401(k) with matching and additional plan options like a Roth 401(k). Financial support is reinforced through discounts and broader financial guidance resources tied to the benefits ecosystem.
  • Wellbeing & Lifestyle Benefits Wellbeing support extends beyond insurance through programs like an Employee Assistance Program, counseling/legal resources, and gym or wellness reimbursement. These offerings increase the perceived total rewards value even when cash compensation sentiment varies by role.

Citi Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Kwun Tong, Kowloon
223,850 Employees

What We Do

Citi's mission is to serve as a trusted partner to our clients by responsibly providing financial services that enable growth and economic progress. Our core activities are safeguarding assets, lending money, making payments and accessing the capital markets on behalf of our clients. We have 200 years of experience helping our clients meet the world's toughest challenges and embrace its greatest opportunities. We are Citi, the global bank – an institution connecting millions of people across hundreds of countries and cities.

Similar Jobs

Zocdoc Logo Zocdoc

Senior Specialist, Provider Data Operations (US Healthcare Insurance)

Healthtech • Information Technology • Software • Telehealth
Easy Apply
Hybrid
Pune, Mahārāshtra, IND
900 Employees

Zocdoc Logo Zocdoc

Enterprise Support Associate

Healthtech • Information Technology • Software • Telehealth
Easy Apply
Hybrid
Pune, Mahārāshtra, IND
900 Employees

TransUnion Logo TransUnion

Lead Developer, C# .NET & APIs

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Hybrid
Pune, Mahārāshtra, IND
13000 Employees

Rubrik Logo Rubrik

Software Engineer

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Cybersecurity • Data Privacy
In-Office
Pune, Mahārāshtra, IND
3000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account