Senior Data Engineer

Posted 2 Days Ago
Hiring Remotely in Washington, DC, USA
In-Office or Remote
Senior level
Artificial Intelligence • Software
The Role
Design, implement, and maintain Azure-based ELT/ETL pipelines and data architecture. Migrate and normalize source data into ADLS, enforce quality controls, source control, documentation, and self-service analytics. Optimize storage/processing, audit pipelines, and explore AI-assisted tooling.
Summary Generated by Built In

Senior Data Engineer

Location: Herndon, VA (Remote Work)

Must have an Public Trust Clearance

KEY RESPONSIBILITIES

•      Provide authoritative expertise on data engineering methods and best practices, including code first development approaches and modern pipeline design patterns.

•      Design, implement, and maintain the data architecture that supports products and end users, with all assets managed under source control.

•      Design, implement, and maintain ELT and ETL pipelines for efficient processing of source data in Azure Synapse and Azure Machine Learning, using both SDK V1 and SDK V2.

•      Migrate source data identified by SBA OIG into Azure Data Lake Storage.

•      Normalize entity attributes such as addresses, phone numbers, and other common fields.

•      Review, maintain, and improve existing architecture and pipelines, including periodic audits addressing bottlenecks, deprecated dependencies, and architecture drift.

•      Establish quality controls across all pipelines and introduce error handling, logging mechanisms, and validation checks.

•      Incorporate source control across all pipelines and analytics codebases so code can evolve iteratively without destabilizing the architecture.

•      Optimize ingestion, processing, and storage across a wide variety of datasets and data types, including modern columnar formats such as Parquet.

•      Develop self service capabilities that let SBA OIG analysts query and export data for investigations and audits.

•      Author robust standard operating procedures governing the authoring, development, validation, publishing, execution, and monitoring of all data pipelines and assets in the Azure environment.

•      Produce detailed documentation of the data architecture, including data dictionaries, entity relationship diagrams, and pipeline process maps.

•      Maintain and expand the environment with additional datasets and services on request, following a defined intake and testing process before production deployment.

•      Stay current with emerging AI tooling relevant to data engineering and contribute to exploratory work evaluating automation and language model assisted capabilities.


Requirements

Education

Bachelor's degree in data engineering, computer science, data science, machine learning, mathematics, or a related field. Alternatively, five years of applied work experience in any of the same fields.

  • 5 years - Maintaining SQL databases and conducting advanced operations in SQL and T-SQL.
  • 5 years - Designing, implementing, and maintaining ELT and ETL processes in cloud based data analytics environments.
  • 3 years - Working in Azure Synapse and Azure Machine Learning with the modern data stack. Certifications preferred, DP-203 or equivalent.
  • 3 years -Manipulating data in Python. Pandas is required. PySpark and Polars preferred. Experience developing reusable, modular code preferred.

PREFERRED QUALIFICATIONS

  • DP-203, Microsoft Certified Azure Data Engineer Associate, or an equivalent current certification.
  • Implementing pipelines and infrastructure using code first approaches: Python SDK, CLI, REST APIs, or infrastructure as code tooling such as Terraform or Bicep.
  • Implementing source control and continuous integration and delivery workflows for data assets.
  • Demonstrated familiarity with AI coding assistants and large language model integration patterns.
  • PySpark or Polars at production scale.
  • Entity resolution and attribute normalization across records with inconsistent addresses, names, and identifiers.
  • Building self service analytic access for non engineering users.

Benefits

We are proud to offer competitive compensation and benefits packages to include

  • Medical 
  • Dental
  • Vision
  • Basic Life 
  • Health Saving Account
  • 401K matching
  • Three weeks of PTO/Sick
  • 11 Paid Holidays
  • Pre-Approved Online Training

Skills Required

  • Must have a Public Trust clearance
  • Bachelor's degree in data engineering, computer science, data science, machine learning, mathematics, or related field OR five years of applied work experience
  • 5 years maintaining SQL databases and advanced operations in SQL and T-SQL
  • 5 years designing, implementing, and maintaining ELT and ETL processes in cloud-based data analytics environments
  • 3 years working in Azure Synapse and Azure Machine Learning with the modern data stack
  • 3 years manipulating data in Python with Pandas (PySpark and Polars preferred)
  • Experience migrating data into Azure Data Lake Storage and working with columnar formats such as Parquet
  • DP-203 or equivalent Azure Data Engineer certification
  • Implementing pipelines and infrastructure using code-first approaches (Python SDK, CLI, REST APIs) or IaC (Terraform or Bicep)
  • Implementing source control and CI/CD workflows for data assets
  • Familiarity with AI coding assistants and large language model integration patterns
  • Production-scale PySpark or Polars experience
  • Entity resolution and attribute normalization experience (addresses, names, identifiers)
  • Experience building self-service analytics access for non-engineering users
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Leesburg, VA
11 Employees
Year Founded: 2017

What We Do

We Love Building Innovative Digital Solutions using Automation and A.I./ML to Solve Complex Problems for our Customer's Mission

Similar Jobs

Samsara Logo Samsara

Senior Data Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
134K-203K Annually

SentiLink Logo SentiLink

Platform Engineer

Fintech • Information Technology • Software
Remote
United States
170 Employees
170K-230K Annually

Life360 Logo Life360

Senior Data Engineer

Kids + Family • Mobile
Remote
USA
600 Employees
148K-217K Annually

World Wide Technology Logo World Wide Technology

Senior Data Engineer

Big Data • Cloud • Hardware • Software • App development
Remote
United States
9000 Employees
112K-140K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account