Data Engineer
About the role
We're looking for a Data Engineer to build and maintain the pipelines that power analytics and data products across Chubb. You'll design ETL/ELT workflows on Databricks, turn raw source data into reliable, well-modeled datasets, and keep those pipelines fast and cost-efficient as data volumes grow.
This role suits someone who is comfortable owning a pipeline end to end — from ingestion through transformation to the tables analysts and data scientists actually query.
What you'll do
- Design, build, and maintain batch and streaming ETL/ELT pipelines using Python, SQL, and PySpark on Databricks.
- Model data across raw, cleansed, and curated layers (medallion architecture) with Delta Lake.
- Ingest data from a range of sources — relational databases, APIs, files, and event streams — including incremental and change data capture patterns.
- Tune Spark jobs and SQL queries for performance and cost: partitioning, file sizing and compaction, caching, join strategies, and shuffle reduction.
- Build data quality checks, validation rules, and monitoring so problems are caught before downstream consumers see them.
- Orchestrate and schedule workflows (Databricks Workflows, Airflow, or similar), with proper retry, alerting, and dependency handling.
- Apply software engineering practices to data work: version control, code review, testing, and CI/CD for pipeline deployments.
- Partner with analysts, data scientists, and business stakeholders to translate requirements into usable data models.
- Document pipelines, data lineage, and design decisions.
Required qualifications
- [3]+ years of experience in a data engineering or comparable role.
- Strong Python for data processing, automation, and pipeline development.
- Advanced SQL: complex joins, window functions, aggregations, and query optimization.
- Hands-on experience with Databricks and PySpark in a production environment.
- Demonstrated experience designing and operating ETL/ELT pipelines at scale.
- Practical knowledge of performance optimization — able to diagnose a slow or expensive job and explain what you changed and why.
- Solid understanding of data warehousing and modeling concepts (dimensional modeling, slowly changing dimensions, normalization trade-offs).
- Experience with Git and collaborative development workflows.
Nice to have
- Delta Lake internals: OPTIMIZE, Z-ordering, liquid clustering, time travel, VACUUM.
- Databricks features such as Unity Catalog, Delta Live Tables / Lakeflow Declarative Pipelines, Auto Loader, or Databricks SQL.
- Cloud platform experience ([AWS / Azure / GCP]) and its storage and compute services.
- Streaming experience with Structured Streaming, Kafka, or Event Hubs.
- Infrastructure as code (Terraform) and CI/CD pipelines for data workloads.
- dbt or similar transformation frameworks.
- Databricks certification (Data Engineer Associate or Professional).
- Familiarity with data governance, access control, and PII handling.
Skills Required
- 3+ years of experience in data engineering or a comparable role
- Strong Python for data processing, automation, and pipeline development
- Advanced SQL, including complex joins, window functions, aggregations, and query optimization
- Hands-on production experience with Databricks and PySpark
- Experience designing and operating ETL/ELT pipelines at scale
- Practical experience diagnosing and optimizing slow or expensive data jobs
- Understanding of data warehousing and modeling concepts, including dimensional modeling, slowly changing dimensions, and normalization trade-offs
- Experience with Git and collaborative development workflows
- Knowledge of Delta Lake internals, including OPTIMIZE, Z-ordering, liquid clustering, time travel, and VACUUM
- Experience with Databricks Unity Catalog, Delta Live Tables or Lakeflow Declarative Pipelines, Auto Loader, or Databricks SQL
- Cloud platform experience with AWS, Azure, or GCP and related storage and compute services
- Streaming experience with Structured Streaming, Kafka, or Event Hubs
- Infrastructure as code experience with Terraform and CI/CD pipelines for data workloads
- Experience with dbt or similar transformation frameworks
- Databricks Data Engineer Associate or Professional certification
- Familiarity with data governance, access control, and PII handling
What We Do
Chubb is the world’s largest publicly traded property and casualty insurance company. With operations in 54 countries and territories, Chubb provides commercial and personal property and casualty insurance, personal accident and supplemental health insurance, reinsurance and life insurance to a diverse group of clients. As an underwriting company, we assess, assume and manage risk with insight and discipline. We service and pay our claims fairly and promptly. The company is also defined by its extensive product and service offerings, broad distribution capabilities, exceptional financial strength and local operations globally. Parent company Chubb Limited is listed on the New York Stock Exchange (NYSE: CB) and is a component of the S&P 500 index. Chubb maintains executive offices in Zurich, New York, London, Paris and other locations, and employs 31,000 people worldwide. Additional information can be found at: chubb.com.








