Position Summary
Scicom Infrastructure Services is seeking a Databricks Data Operations Specialist to support the operational stability, reliability, performance, and security of a large-scale Databricks data platform.
This role will work as part of a multidisciplinary Databricks team that includes data engineers, architects, business analysts, governance specialists, and technical managers. The Data Operations Specialist will be responsible for monitoring production data pipelines, responding to incidents, troubleshooting platform and data-processing issues, coordinating deployments, and ensuring that data products meet established availability, quality, and service-level requirements.
The ideal candidate will have hands-on experience operating Databricks environments and supporting production data workloads using Apache Spark, Delta Lake, SQL, Python, cloud infrastructure, monitoring tools, and automated deployment practices.
Key Responsibilities
Databricks Platform Operations
- Monitor the health, availability, performance, and capacity of Databricks workspaces, jobs, pipelines, clusters, SQL warehouses, and related cloud services.
- Provide day-to-day operational support for production Databricks environments.
- Monitor scheduled and event-driven data pipelines for failures, delays, data-quality issues, and performance degradation.
- Troubleshoot Databricks jobs, Spark workloads, cluster failures, notebook errors, connectivity issues, and data-processing exceptions.
- Restart, rerun, recover, or coordinate remediation of failed workflows according to documented procedures.
- Support Databricks Workflows, Lakeflow, Delta Live Tables, Auto Loader, Structured Streaming, and batch-processing workloads.
- Review driver and executor logs, Spark UI metrics, cluster event logs, audit logs, and application telemetry to identify root causes.
- Support configuration and maintenance of Databricks clusters, serverless compute, instance pools, cluster policies, and SQL warehouses.
- Monitor resource utilization and recommend changes to cluster sizing, autoscaling, workload scheduling, and compute policies.
- Help maintain consistent configurations across development, testing, staging, and production environments.
- Support platform upgrades, runtime-version changes, library updates, and configuration changes.
- Coordinate with Databricks support and cloud-service providers when vendor assistance is required.
Production Pipeline Support
- Monitor ETL, ELT, streaming, API, and file-based data-integration processes.
- Validate that scheduled data loads complete within required processing windows.
- Investigate missing, late, duplicated, incomplete, or inconsistent data.
- Perform controlled data reprocessing, backfills, reconciliation, and recovery activities.
- Support ingestion and transformation processes involving JSON, CSV, XML, Parquet, APIs, databases, message queues, and cloud storage.
- Monitor Bronze, Silver, and Gold data layers within a medallion architecture.
- Validate source-to-target record counts, control totals, schema consistency, and business-rule compliance.
- Coordinate with data engineers to resolve recurring pipeline or transformation defects.
- Maintain operational runbooks for common failures, recovery procedures, escalation paths, and recurring support tasks.
- Ensure production interventions are documented and performed in accordance with change-control procedures.
Incident and Problem Management
- Serve as a first or second level of support for production Databricks and data-pipeline incidents.
- Triage incidents based on severity, business impact, affected systems, and service-level commitments.
- Coordinate incident response among engineering, cloud, security, governance, and business teams.
- Provide timely status updates during active incidents.
- Escalate issues to technical leads, architects, vendors, or program leadership when appropriate.
- Conduct or support root-cause analyses for production incidents.
- Document incident timelines, contributing factors, corrective actions, and preventive measures.
- Track recurring problems and recommend permanent engineering or process improvements.
- Participate in post-incident reviews and ensure assigned corrective actions are completed.
- Help develop automated remediation for common operational failures.
- Maintain incident, problem, and service-request records in the organization’s ticketing system.
Monitoring and Observability
- Develop and maintain dashboards, alerts, and operational reports for Databricks and associated data services.
- Monitor job success rates, pipeline latency, processing duration, cluster utilization, data freshness, data quality, and compute consumption.
- Configure actionable alerts that minimize unnecessary notifications while identifying material operational issues.
- Integrate Databricks monitoring data with enterprise observability platforms.
- Support tools such as Azure Monitor, Log Analytics, CloudWatch, Datadog, Splunk, Grafana, Prometheus, or comparable platforms.
- Define operational health indicators, service-level indicators, and service-level objectives.
- Produce daily, weekly, and monthly operational metrics for program leadership.
- Identify performance trends and operational risks before they result in production failures.
- Maintain dashboards showing system availability, incident volume, recovery time, pipeline status, and data freshness.
Data Quality and Data Contract Operations
- Monitor compliance with established data contracts between data producers and consumers.
- Validate schema, field, data-type, frequency, freshness, completeness, and quality requirements.
- Detect schema drift, unexpected source-system changes, and contract violations.
- Coordinate resolution of data-contract issues with business analysts, data owners, and engineering teams.
- Execute automated and manual data-quality controls.
- Monitor data-quality rules related to accuracy, completeness, uniqueness, consistency, validity, and timeliness.
- Track exceptions and ensure data-quality issues are assigned to the appropriate owner.
- Support reconciliation of contracting, procurement, financial, and operational datasets.
- Assist with validation of data structured according to the Open Contracting Data Standard when applicable.
- Maintain operational evidence supporting data quality, lineage, audit, and governance requirements.
Deployment and Change Management
- Support the release and deployment of Databricks notebooks, jobs, workflows, libraries, configurations, and infrastructure changes.
- Coordinate deployments across development, testing, staging, and production environments.
- Verify deployment readiness, approvals, testing results, dependencies, and rollback plans.
- Participate in release validation and post-deployment monitoring.
- Support CI/CD pipelines using Azure DevOps, GitHub Actions, GitLab CI, Jenkins, or comparable tools.
- Assist with infrastructure-as-code deployments using Terraform or similar technologies.
- Maintain deployment logs, implementation records, configuration documentation, and release notes.
- Ensure emergency changes follow established approval and documentation requirements.
- Support environment comparisons and investigate configuration drift.
- Coordinate production changes with technical and business stakeholders to minimize operational disruption.
Security, Access, and Governance Support
- Support Databricks identity, access, and entitlement administration.
- Assist with user onboarding, offboarding, group membership, workspace permissions, and service-principal access.
- Support Unity Catalog permissions, catalogs, schemas, tables, volumes, external locations, and storage credentials.
- Apply role-based and least-privilege access-control principles.
- Monitor access failures, unusual activity, audit events, and policy violations.
- Coordinate access requests with security, governance, and data-ownership teams.
- Support secret management through Databricks secrets, Azure Key Vault, AWS Secrets Manager, or comparable services.
- Maintain operational documentation supporting security reviews and audits.
- Assist with data retention, archival, backup, recovery, and disaster-recovery procedures.
- Ensure production-support activities comply with applicable security and privacy requirements.
Cost and Performance Management
- Monitor Databricks and cloud-resource consumption.
- Identify idle, oversized, inefficient, or improperly configured compute resources.
- Analyze job duration, cluster utilization, query performance, storage consumption, and workload patterns.
- Recommend cluster right-sizing, autoscaling, scheduling, caching, partitioning, and workload-isolation improvements.
- Support implementation of compute policies, tagging standards, budgets, and cost alerts.
- Produce usage and cost reports for technical and program leadership.
- Work with engineers and architects to reduce unnecessary compute consumption.
- Track the operational impact of optimization initiatives.
Documentation and Continuous Improvement
- Create and maintain operational procedures, troubleshooting guides, escalation matrices, and support runbooks.
- Document platform configurations, dependencies, schedules, service accounts, integrations, and recovery requirements.
- Maintain an accurate inventory of production jobs, pipelines, data products, and system interfaces.
- Identify opportunities to automate repetitive monitoring, support, recovery, and reporting activities.
- Participate in operational-readiness reviews for new data products and pipelines.
- Ensure new solutions include appropriate monitoring, alerts, logging, support procedures, and ownership assignments before production deployment.
- Share lessons learned and common troubleshooting practices across the Databricks team.
- Contribute to platform standards, operational policies, and continuous-improvement initiatives.
Required Qualifications
- Bachelor’s degree in computer science, information technology, data engineering, information systems, or a related field, or equivalent professional experience.
- At least four years of experience in data operations, production support, data engineering, cloud operations, DevOps, platform support, or a related technical role.
- At least two years of hands-on experience supporting Databricks environments or production Databricks workloads.
- Experience supporting production ETL, ELT, batch, or streaming data pipelines.
- Working knowledge of:
- Databricks
- Apache Spark
- PySpark or Python
- SQL
- Delta Lake
- Databricks Workflows or Jobs
- Cloud object storage
- Data pipeline monitoring
- Log analysis and troubleshooting
- Experience investigating failed jobs, delayed pipelines, schema issues, data discrepancies, and performance problems.
- Familiarity with medallion architecture and Bronze, Silver, and Gold data layers.
- Experience with at least one major cloud platform: Microsoft Azure, AWS, or Google Cloud.
- Experience using ticketing and service-management tools such as ServiceNow, Jira, or a comparable platform.
- Experience with monitoring, alerting, logging, or observability tools.
- Understanding of incident, problem, change, and release-management processes.
- Ability to write and maintain technical procedures and operational documentation.
- Strong analytical, organizational, communication, and troubleshooting skills.
- Ability to work effectively with engineers, business analysts, architects, security teams, and nontechnical stakeholders.
- Ability to support multiple production priorities in a structured and timely manner.
Preferred Qualifications
- Databricks Certified Data Engineer Associate or Databricks Certified Data Engineer Professional.
- Databricks Certified Associate Developer for Apache Spark.
- Microsoft Azure, AWS, or Google Cloud certification.
- Experience with Unity Catalog, Delta Live Tables, Lakeflow, Auto Loader, Structured Streaming, or Databricks SQL.
- Experience with Azure Data Factory, Azure Data Lake Storage, Azure Monitor, Log Analytics, AWS Glue, Amazon S3, CloudWatch, or comparable cloud services.
- Experience with Terraform, Git, Azure DevOps, GitHub Actions, Jenkins, or other CI/CD tools.
- Experience with Splunk, Datadog, Grafana, Prometheus, or similar observability platforms.
- Experience supporting APIs, Kafka, Event Hubs, Kinesis, Airflow, dbt, Snowflake, or Power BI.
- Knowledge of data contracts, schema registries, metadata management, data lineage, and data-quality frameworks.
- Familiarity with OCDS, procurement data, contracting data, financial data, or public-sector data standards.
- Experience supporting government, regulated, or high-security environments.
- Experience operating systems with formal service-level agreements and 24-hour support requirements.
- Experience with disaster recovery, backup, archival, and operational-resilience testing.
Core Competencies
- Strong operational ownership and attention to detail.
- Ability to remain calm and organized during production incidents.
- Clear written and verbal communication.
- Methodical troubleshooting and root-cause analysis.
- Ability to distinguish symptoms from underlying technical causes.
- Strong documentation and process-discipline skills.
- Ability to prioritize incidents based on business and technical impact.
- Collaborative approach to working with engineering and business teams.
- Commitment to automation, reliability, security, and continuous improvement.
- Willingness to take ownership of issues through resolution.
Skills Required
- Bachelor's degree in computer science, information technology, data engineering, information systems, or related field, or equivalent experience.
- At least four years of experience in data operations, production support, data engineering, cloud operations, DevOps, platform support, or related technical role.
- At least two years of hands-on experience supporting Databricks environments or production Databricks workloads.
- Experience supporting production ETL, ELT, batch, or streaming data pipelines.
- Working knowledge of Databricks.
- Working knowledge of Apache Spark.
- Working knowledge of PySpark or Python.
- Working knowledge of SQL.
- Working knowledge of Delta Lake.
- Working knowledge of Databricks Workflows or Jobs.
- Working knowledge of cloud object storage.
- Experience with data pipeline monitoring, log analysis, and troubleshooting.
- Experience investigating failed jobs, delayed pipelines, schema issues, data discrepancies, and performance problems.
- Familiarity with medallion architecture and Bronze, Silver, and Gold data layers.
- Experience with at least one major cloud platform: Microsoft Azure, AWS, or Google Cloud.
- Experience using ticketing and service-management tools such as ServiceNow, Jira, or comparable platforms.
- Experience with monitoring, alerting, logging, or observability tools.
- Understanding of incident, problem, change, and release-management processes.
- Ability to write and maintain technical procedures and operational documentation.
- Strong analytical, organizational, communication, and troubleshooting skills.
- Ability to work effectively with engineers, business analysts, architects, security teams, and nontechnical stakeholders.
- Ability to support multiple production priorities in a structured and timely manner.
- Databricks Certified Data Engineer Associate or Databricks Certified Data Engineer Professional.
- Databricks Certified Associate Developer for Apache Spark.
- Microsoft Azure, AWS, or Google Cloud certification.
- Experience with Unity Catalog, Delta Live Tables, Lakeflow, Auto Loader, Structured Streaming, or Databricks SQL.
- Experience with Azure Data Factory, Azure Data Lake Storage, Azure Monitor, Log Analytics, AWS Glue, Amazon S3, CloudWatch, or comparable cloud services.
- Experience with Terraform, Git, Azure DevOps, GitHub Actions, Jenkins, or other CI/CD tools.
- Experience with Splunk, Datadog, Grafana, Prometheus, or similar observability platforms.
- Experience supporting APIs, Kafka, Event Hubs, Kinesis, Airflow, dbt, Snowflake, or Power BI.
- Knowledge of data contracts, schema registries, metadata management, data lineage, and data-quality frameworks.
- Familiarity with OCDS, procurement data, contracting data, financial data, or public-sector data standards.
- Experience supporting government, regulated, or high-security environments.
- Experience operating systems with formal service-level agreements and 24-hour support requirements.
- Experience with disaster recovery, backup, archival, and operational-resilience testing.
What We Do
Scicom’s singular focus is to deliver high quality, reliable and cost effective technology solutions to support our client’s business objectives. Our clients consist of the companies from the Fortune 500 and leading government organizations – where Scicom has delivered enterprise services across key technology domains including architecture, applications, infrastructure, management consulting and enterprise software. Our ability to contend with complexity allows our clients to rapidly achieve business objectives and bring back innovation in IT


.jpg)





