Senior Data Engineer (Databricks/AWS)

Posted Yesterday
Be an Early Applicant
Mexico City, Cuauhtémoc, Mexico City, MEX
In-Office
Senior level
Sharing Economy
The Role
Lead Databricks-based data engineering for a CRM migration from Veeva to Salesforce Life Sciences Cloud. Build and orchestrate ingestion pipelines (Airflow), implement dbt transformations and tests, optimize Delta Lake/Unity Catalog and Databricks clusters on AWS S3, enforce data quality, monitoring, and collaborate with architects and stakeholders on ET business hours.
Summary Generated by Built In

About Fusemachines

Fusemachines is a 12+ year old AI company, dedicated to delivering state-of-the-art AI products and solutions to a diverse range of industries. Founded by Sameer Maskey, Ph.D., an Adjunct Associate Professor at Columbia University, our company is on a steadfast mission to democratize AI and harness the power of global AI talent from underserved communities. With a robust presence in four countries and a dedicated team of over 400 full-time employees, we are committed to fostering AI transformation journeys for businesses worldwide. At Fusemachines, we not only bridge the gap between AI advancement and its global impact but also strive to deliver the most advanced technology solutions to the world.
Type: Full-time, Remote
 

About the role

This is a full-time, high-impact position for a Senior Data Engineer with expertise in Databricks, dbt, and Apache Airflow to support a critical CRM data architecture migration for a key client in the Life Sciences industry.

In this role, you will join an urgent initiative to backfill key engineering capabilities and maintain momentum during an ongoing CRM system transition. The project involves migrating enterprise customer data from Veeva CRM to Salesforce Life Sciences Cloud, integrated with an underlying AWS S3 cloud environment and Databricks data warehouse. Your main focus will be building out, configuring, and redirecting data ingestion pipelines out of Life Sciences Cloud into the data warehouse, while implementing dbt models and Airflow orchestrations to ensure complete data accuracy.

Candidates must be able to operate strictly on US East Coast business hours (location is flexible across North America, LATAM, or remote with full Eastern Time overlap).

Qualification / Skill Set Requirement:

  • Core Technical Expertise:

    • 5+ years of hands-on data engineering experience with deep expertise in AWS, Databricks, dbt, and Apache Airflow.

    • Strong programming proficiency in Python / PySpark and Advanced SQL (complex joins, analytical window functions).

    • Hands-on expertise in Databricks platform architecture, Lakehouse implementation, Delta Lake, Unity Catalog, and cluster performance tuning.

  • Architecture & Migration:

    • Proven track record of architecting and executing migrations.

    • Demonstrated experience scaling platform performance.

  • Pipeline Orchestration & Modeling:

    • Proven experience building scalable transformations pipelines using dbt for data transformation, testing, and documentation.

    • Solid background orchestrating complex workflow DAGs with Apache Airflow.

    • Experience working with AWS cloud infrastructure, specifically AWS S3 as an underlying data lake storage layer.

  • CRM Integration & Domain Knowledge:

    • Hands-on experience developing integrations and data ingestion pipelines for CRM platforms, specifically Salesforce, Salesforce Life Sciences Cloud, and/or Veeva CRM.

    • Understanding of data structures, customer master data, and analytics workflows within the Life Sciences.

  • DevOps & Governance:

    • Deep understanding of SDLC/Agile and DevOps for CI/CD and artifact management.

    • Knowledge of AWS and Databricks security best practices and compliance standards.

  • Certifications Preferred: Databricks Certified Data Engineer Associate/Professional, Databricks Spark Developer, and major cloud certifications in AWS.

  • Logistics & Shift Overlap:

    • Ability to maintain 100% full working time overlap with US East Coast business hours (ET). Flexible location (US, Canada, LATAM, or remote ET).

Responsibilities

  • Pipeline Development & Integration: Architect, build, and deploy data integration pipelines connecting Salesforce Life Sciences Cloud to the client’s Databricks warehouse environment.

  • CRM Migration Support: Execute pipeline modifications to transition legacy data feeds from Veeva CRM to Salesforce Life Sciences Cloud, updating warehouse models accordingly.

  • Transformation & Workflow Orchestration: Write clean, modular dbt transformation models and organize end-to-end DAG execution using Apache Airflow.

  • Data Warehouse & Storage Optimization: Manage Delta tables and optimize Databricks clusters and AWS S3 storage for high performance and cost efficiency.

  • Data Validation & Quality Assurance: Implement data quality testing, schemas, and verification rules in dbt and Python to guarantee accurate data delivery.

  • Monitoring & Alerting: Build and enforce proactive monitoring frameworks.

  • Agile Collaboration: Work closely with project leads, solution architects, and technical stakeholders during US East Coast hours to ensure rapid iteration and goal completion.

Equal Opportunity Employer: Race, Color, Religion, Sex, Sexual Orientation, Gender Identity, National Origin, Age, Genetic Information, Disability, Protected Veteran Status, or any other legally protected group status.

Skills Required

  • 5+ years data engineering experience with AWS, Databricks, dbt, and Apache Airflow
  • Strong programming proficiency in Python / PySpark
  • Advanced SQL (complex joins, analytical window functions)
  • Hands-on expertise in Databricks platform architecture, Lakehouse implementation, Delta Lake, and Unity Catalog
  • Experience with Databricks cluster performance tuning and Delta table management
  • Proven experience architecting and executing migrations and scaling platform performance
  • Proven experience building scalable transformation pipelines using dbt (models, testing, documentation)
  • Experience orchestrating complex DAGs with Apache Airflow
  • Experience working with AWS S3 as data lake storage
  • Hands-on experience integrating CRM platforms (Salesforce, Salesforce Life Sciences Cloud, and/or Veeva CRM)
  • Understanding of customer master data and analytics workflows within Life Sciences
  • Deep understanding of SDLC/Agile and DevOps for CI/CD and artifact management
  • Knowledge of AWS and Databricks security best practices and compliance standards
  • Ability to maintain 100% overlap with US East Coast (ET) business hours
  • Databricks Certified Data Engineer, Databricks Spark Developer, or AWS cloud certifications
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
428 Employees
Year Founded: 2013

What We Do

A 10+ year old AI company offering cutting-edge AI products and solutions across industries. With over a decade of experience, we help companies in their AI Transformation journey with our suite of AI Products and AI Solutions supported by our global AI Talent from underserved communities. On a mission to #DemocratizeAI, we aim to bridge the gap between AI advancement and global impact, bringing the most advanced technology solutions to the world.

Similar Jobs

Datadog Logo Datadog

Director, Sales Engineering

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Hybrid
2 Locations
6500 Employees

Samsara Logo Samsara

Director, Marketing - Mexico

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
México
4000 Employees
1M-2M Annually

CrowdStrike Logo CrowdStrike

Customer Success Manager

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
2 Locations
11000 Employees

MongoDB Logo MongoDB

Solutions Architect

Big Data • Cloud • Software • Database
Easy Apply
Hybrid
Mexico City, Cuauhtémoc, Mexico City, MEX
5550 Employees

Similar Companies Hiring

Cargill Thumbnail
Food • Greentech • Logistics • Sharing Economy • Transportation • Agriculture • Industrial
Wayzata, MN
155000 Employees
Taskrabbit Thumbnail
eCommerce • Information Technology • Sharing Economy • Software
San Franscisco, CA
450 Employees
Federal Reserve Bank of Chicago Thumbnail
Agency • Fintech • Payments • Sharing Economy • Social Impact
Chicago, IL
1515 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account