Databricks Engineer

Posted 2 Days Ago
Be an Early Applicant
Adelphi, MD, USA
Hybrid
Mid level
Information Technology • Professional Services • Consulting • Cybersecurity
The Role
Design, build, and operate a Databricks-centric Data & AI platform using Medallion Architecture. Implement scalable ELT/ETL pipelines with Spark and Delta Lake, ingest data from enterprise systems, enforce data quality and governance, enable ML/AI workflows with MLflow, and manage cloud data storage on ADLS or S3. Provide documentation, monitoring, training, and weekly progress reporting.
Summary Generated by Built In

ABOUT US:


CMT Services Inc. is a dynamic and small business supporting Federal, State, and Local government agencies. As an SBA-certified HUBZone, Woman Owned Small Business (WOSB), we deliver quality, professional services to support the missions and strategic business goals of our clients.


Position Title: Databricks Engineer

Location:

University of Maryland Global Campus

3501 University Blvd. East

Adelphi, MD 20783

Period of Performance:

1 year

40 hours per week (no overtime). 

Position Summary: The Databricks Engineer will design, build, and operate a Data & AI platform with a strong foundation in the Medallion Architecture (raw/bronze, curated/silver, and mart/gold layers). This platform will orchestrate complex data workflows and scalable ELT pipelines to integrate data from enterprise systems such as PeopleSoft, D2L, and Salesforce, delivering high-quality, governed data for machine learning, AI/BI, and analytics at scale.


You will play a critical role in engineering the infrastructure and workflows that enable seamless data flow across the enterprise, ensure operational excellence, and provide the backbone for strategic decision-making, predictive modeling, and innovation.


Responsibilities

Data & AI Platform Engineering (Databricks-Centric):

  • Design, implement, and optimize end-to-end data pipelines on Databricks, following the Medallion Architecture principles.
  • Build robust and scalable ETL/ELT pipelines using Apache Spark and Delta Lake to transform raw (bronze) data into trusted curated (silver) and analytics-ready (gold) data layers.
  • Operationalize Databricks Workflows for orchestration, dependency management, and pipeline automation.
  • Apply schema evolution and data versioning to support agile data development.

Platform Integration & Data Ingestion:

  • Connect and ingest data from enterprise systems such as PeopleSoft, D2L, and Salesforce using APIs, JDBC, or other integration frameworks.
  • Implement connectors and ingestion frameworks that accommodate structured, semi-structured, and unstructured data.
  • Design standardized data ingestion processes with automated error handling, retries, and alerting.

Data Quality, Monitoring, and Governance:

  • Develop data quality checks, validation rules, and anomaly detection mechanisms to ensure data integrity across all layers.
  • Integrate monitoring and observability tools (e.g., Databricks metrics, Grafana) to track ETL performance, latency, and failures.
  • Implement Unity Catalog or equivalent tools for centralized metadata management, data lineage, and governance policy enforcement.

Security, Privacy, and Compliance:

  • Enforce data security best practices including row-level security, encryption at rest/in transit, and fine-grained access control via Unity Catalog.
  • Design and implement data masking, tokenization, and anonymization for compliance with privacy regulations (e.g., GDPR, FERPA).
  • Work with security teams to audit and certify compliance controls.

AI/ML-Ready Data Foundation:

  • Enable data scientists by delivering high-quality, feature-rich data sets for model training and inference.
  • Support AIOps/MLOps lifecycle workflows using MLflow for experiment tracking, model registry, and deployment within Databricks.
  • Collaborate with AI/ML teams to create reusable feature stores and training pipelines.

Cloud Data Architecture and Storage:

  • Architect and manage data lakes on Azure Data Lake Storage (ADLS) or Amazon S3, and design ingestion pipelines to feed the bronze layer.
  • Build data marts and warehousing solutions using platforms like Databricks.
  • Optimize data storage and access patterns for performance and cost-efficiency.

Documentation & Enablement:

  • Maintain technical documentation, architecture diagrams, data dictionaries, and runbooks for all pipelines and components.
  • Provide training and enablement sessions to internal stakeholders on the Databricks platform, Medallion Architecture, and data governance practices.
  • Conduct code reviews and promote reusable patterns and frameworks across teams.

Reporting and Accountability:

  • Submit a weekly schedule of hours worked and progress reports outlining completed tasks, upcoming plans, and blockers.
  • Track deliverables against roadmap milestones and communicate risks or dependencies.

Required Qualifications:

  • Hands-on experience with Databricks, Delta Lake, and Apache Spark for large-scale data engineering.
  • Deep understanding of ELT pipeline development, orchestration, and monitoring in cloud-native environments.
  • Experience implementing Medallion Architecture (Bronze/Silver/Gold) and working with data versioning and schema enforcement in enterprise grade environments.
  • Strong proficiency in SQL, Python, or Scala for data transformations and workflow logic.
  • Proven experience integrating enterprise platforms (e.g., PeopleSoft, Salesforce, D2L) into centralized data platforms.
  • Familiarity with data governance, lineage tracking, and metadata management tools.

Preferred Qualifications:

  • Experience with Databricks Unity Catalog for metadata management and access control.
  • Experience deploying ML models at scale using MLFlow or similar MLOps tools.
  • Familiarity with cloud platforms like Azure or AWS, including storage, security, and networking aspects.
  • Knowledge of data warehouse design and star/snowflake schema modeling.

Join Our Team:


At CMT Services, we believe that extraordinary results come from empowering exceptional people. If you're ready to lead innovative projects, solve complex challenges, and contribute to meaningful infrastructure development while advancing your career in a supportive, collaborative environment, we want to hear from you.


Disclaimer: 

By submitting your resume for this job posting, you authorize CMT Services, Inc. to forward your resume to all applicable internal and external managers, agencies, and recruitment personnel for review and consideration to hire.

Skills Required

  • Hands-on experience with Databricks, Delta Lake, and Apache Spark
  • Deep understanding of ELT pipeline development, orchestration, and monitoring in cloud-native environments
  • Experience implementing Medallion Architecture and working with data versioning and schema enforcement
  • Strong proficiency in SQL, Python, or Scala for data transformations
  • Proven experience integrating enterprise platforms such as PeopleSoft, Salesforce, and D2L
  • Familiarity with data governance, lineage tracking, and metadata management tools
  • Experience with Databricks Unity Catalog for metadata management and access control
  • Experience deploying ML models at scale using MLflow or similar MLOps tools
  • Familiarity with cloud platforms such as Azure or AWS (storage, security, networking)
  • Knowledge of data warehouse design and star/snowflake schema modeling
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
125 Employees
Year Founded: 2002

What We Do

CMT Services, Inc. is an SBA-certified HUBZone and Woman-Owned Small Business that serves as a leading business services solution provider for government and commercial entities. The company specializes in delivering a broad range of high-quality solutions across Consulting, Management, and Technology, including IT consulting, cybersecurity, and program management, focusing on supporting the strategic business goals of its diverse customer base.

Similar Jobs

PwC Logo PwC

Data Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
34 Locations
370000 Employees
77K-202K Annually

PwC Logo PwC

Data Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
35 Locations
370000 Employees
99K-232K Annually

EXL Logo EXL

Data Engineer

Information Technology • Database • Consulting
Remote or Hybrid
United States
30246 Employees
65K-87K Annually
In-Office
3 Locations
35118 Employees
109K-180K Annually

Similar Companies Hiring

Amplify Platform Thumbnail
Fintech • Financial Services • Consulting • Cloud • Business Intelligence • Big Data Analytics
Scottsdale, AZ
62 Employees
Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account