Data Engineer (Databricks/AWS)

Posted 3 Hours Ago
Indianapolis, IN, USA
In-Office
Mid level
Artificial Intelligence • Software • Consulting • Automation
The Role
Designs and maintains batch and streaming data pipelines using Databricks, AWS, Python or Scala, SQL, and orchestration tools. The role covers data quality, governance, cataloging, lineage, API integration, and regulated pharma data environments. Responsibilities include translating business needs into technical specifications, partnering across pharma functions, presenting to executives, prioritizing data initiatives, supporting adoption, and building new data domains. The engineer also applies warehousing, dimensional modeling, documentation, and stakeholder coordination practices.
Summary Generated by Built In
AWS Data Engineer - Databricks
Hybrid – Indianapolis, IN 

About the Role

We are seeking a Data Engineer with 3–5 years of experience working specifically within the pharma industry to join a pharma-focused data team. This is a senior-flavored engineering role that combines hands-on pipeline and platform work with significant business-facing responsibility — including translating business needs into technical specs, presenting to executive-level stakeholders, and helping stand up new data domains from the ground up. You will design, build, and govern the data infrastructure that powers analytics and reporting across the business, while also acting as a trusted technical partner to non-technical stakeholders.

Key Responsibilities
  • Design, build, and maintain scalable ETL/ELT pipelines (batch and streaming) using Databricks, AWS, and related orchestration tools
  • Write and optimize advanced SQL, and build data transformations in Python or Scala
  • Integrate external data sources via APIs and manage pipeline orchestration (Airflow, Databricks Workflows, AWS Glue)
  • Apply data quality, governance, cataloging, and lineage practices aligned with regulated-industry standards
  • Work within GxP-regulated data environments and apply awareness of data privacy/compliance considerations (e.g., 21 CFR Part 11, GDPR where applicable)
  • Partner with business stakeholders across the pharma value chain (R&D, Manufacturing & Quality, Commercial, Drug Development) to gather and translate requirements into technical specifications
  • Present technical work and data strategy to executive-level audiences
  • Prioritize high-impact data initiatives and proactively identify and avoid duplicated data efforts
  • Support change management and adoption of new data solutions across business teams
  • Help stand up new data domains from scratch (green-field build), not just maintain existing ones


Requirements
Required Qualifications
Data Engineering & Pipelines
  • ETL/ELT development (batch and streaming)
  • Advanced SQL (joins, window functions, query optimization)
  • Python or Scala for data transformation
  • Data pipeline orchestration (Airflow, Databricks Workflows, AWS Glue)
  • API integration for external data source ingestion
Platforms & Tools
  • Databricks (Delta Lake, Unity Catalog, Genie)
  • Cloud platforms — AWS (S3, Glue, Athena) and/or Azure/GCP equivalents
  • Data warehousing concepts (dimensional modeling, star schema)
  • BI/visualization tools (Tableau, Power BI, or similar) to understand downstream consumption
Data Quality & Governance
  • Data profiling and cleansing techniques
  • Metadata management and data cataloging
  • Master data management (MDM) principles
  • Data lineage tracking
  • Data governance frameworks (especially regulated-industry standards)
Pharma / Life Sciences Domain Knowledge
  • Familiarity with GxP-regulated data environments
  • Understanding of the pharma value chain (R&D, Manufacturing & Quality, Commercial, Drug Development)
  • Awareness of data privacy/compliance considerations (21 CFR Part 11, GDPR where applicable)
  • Knowledge of common pharma data domains (clinical, manufacturing, quality, commercial)
Stakeholder Management
  • Requirements gathering and translation (business need → technical spec)
  • Cross-functional communication (Business ↔ IT)
  • Executive-level presentation skills (given EC visibility)
  • Change management / adoption support
Analytical & Strategic Thinking
  • Prioritization frameworks (identifying high-impact vs. low-value data asks)
  • Cost-avoidance mindset (spotting duplication before it happens)
  • Ability to work with ambiguity and evolving priorities
Project & Program Skills
  • Agile/Scrum familiarity
  • Documentation discipline (data dictionaries, source-to-target mappings)
  • Vendor/partner coordination (if external data sources are involved)
Nice-to-Have Differentiators
  • Prior consulting or client-facing delivery experience
  • Experience standing up new data domains from scratch (green-field vs. maintenance)
  • Familiarity with AI/GenAI-enabled analytics tools

Skills Required

  • 3-5 years of data engineering experience, specifically within the pharmaceutical industry
  • ETL/ELT development for batch and streaming pipelines
  • Advanced SQL, including joins, window functions, and query optimization
  • Python or Scala for data transformation
  • Pipeline orchestration using Airflow, Databricks Workflows, or AWS Glue
  • API integration for external data source ingestion
  • Databricks experience, including Delta Lake, Unity Catalog, and Genie
  • AWS experience with S3, Glue, Athena, or equivalent Azure/GCP platforms
  • Data warehousing concepts, dimensional modeling, and star schema
  • Familiarity with Tableau, Power BI, or similar BI and visualization tools
  • Data profiling and cleansing techniques
  • Metadata management and data cataloging
  • Master data management principles
  • Data lineage tracking
  • Data governance frameworks, especially regulated-industry standards
  • Familiarity with GxP-regulated data environments
  • Understanding of the pharma value chain, including R&D, manufacturing and quality, commercial, and drug development
  • Awareness of 21 CFR Part 11 and GDPR data privacy and compliance considerations
  • Knowledge of clinical, manufacturing, quality, and commercial pharma data domains
  • Requirements gathering and translation into technical specifications
  • Cross-functional communication between business and IT
  • Executive-level presentation skills
  • Change management and adoption support
  • Prioritization, cost avoidance, ambiguity management, and strategic analytical thinking
  • Agile/Scrum familiarity
  • Documentation including data dictionaries and source-to-target mappings
  • Vendor or partner coordination for external data sources
  • Prior consulting or client-facing delivery experience
  • Experience standing up new data domains from scratch
  • Familiarity with AI or generative AI-enabled analytics tools
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
96 Employees
Year Founded: 2015

What We Do

RADcube is a technology consulting and software development firm founded in 2015 and headquartered in Carmel, Indiana. It helps organizations turn enterprise ideas into practical innovations through digital transformation, custom software, data analytics, artificial intelligence, cloud computing, cybersecurity, and intelligent automation. The company serves healthcare, finance, government, and manufacturing clients, combining human-centric delivery with responsible, accountable technology solutions designed to produce measurable business outcomes.

Similar Jobs

CrowdStrike Logo CrowdStrike

Business Systems Analyst

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
125K-180K Annually

CrowdStrike Logo CrowdStrike

Systems Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
140K-215K Annually

CrowdStrike Logo CrowdStrike

Development Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
85K-120K Annually

Citizens Logo Citizens

Private Banker Associate

Digital Media • Fintech • Information Technology • Machine Learning • Financial Services • Cybersecurity • Automation
In-Office or Remote
2 Locations
17000 Employees
39-49 Hourly

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account