Senior data engineer

Posted 2 Days Ago
Be an Early Applicant
Budapest, HUN
In-Office
Senior level
Agency • HR Tech • Information Technology • Professional Services
The Role
Designs, develops, and optimizes scalable batch and streaming data pipelines using Azure Databricks, Spark, Delta Lake, Azure Data Factory, and related Azure services. Responsibilities include data modeling, performance tuning, cluster management, monitoring, governance, security, data quality, CI/CD, documentation, and collaboration with analytics and data science teams. The role also supports machine learning integration, business intelligence, and migration to Databricks as a strategic data platform.
Summary Generated by Built In

•        Data Pipeline Design & Development (ETL/ELT)

o        Develop ETL/ELT Pipelines: Build scalable ETL/ELT pipelines using Azure Databricks, integrating multiple data sources like Azure Blob Storage, Azure Data Lake, SQL Databases, etc.

o        Automation of Data Workflows: Automate data ingestion, transformation, and loading processes through Azure Data Factory (ADF), Databricks workflows, and Azure Functions.

o        Real-time Data Processing: Implement real-time streaming data pipelines using Databricks Structured Streaming for use cases such as IoT or event-driven architectures.

•        Data Transformation & Modeling

o        Data Cleaning & Transformation: Leverage Apache Spark in Databricks to process large datasets efficiently, performing data cleansing, transformation, and enrichment.

o        Data Modeling: Design and implement dimensional data models (e.g., star schema) optimized for performance and querying in Azure Synapse Analytics or other reporting layers.

o        Delta Lake Implementation: Use Delta Lake for reliable and scalable ACID-compliant data storage and to optimize data for batch and stream processing.

•        Optimization & Performance Tuning

o        Optimize Data Processing: Tune Databricks notebooks and jobs for performance, leveraging Databricks' autoscaling features and optimizing Apache Spark configurations for specific workloads.

o        Data Partitioning & Indexing: Implement best practices for partitioning large datasets, managing table storage formats (Parquet, Delta), and indexing data for faster querying.

o        Cluster Management: Manage Databricks clusters (autoscaling, sizing, and costs), ensuring efficient resource utilization on Azure.

•        Collaboration & Integration with Azure Services

o        Azure Integration: Integrate Databricks with other Azure services like Azure Data Lake Storage, Azure SQL, Azure Synapse, Azure Key Vault (for security), and Power BI for seamless data flow and analysis.

o        Monitoring & Alerts: Set up monitoring, alerting, and logging of data pipelines using tools like Azure Monitor, Databricks Jobs, and Azure Log Analytics to ensure smooth operations.

•        Data Governance & Security

o        Data Security: Ensure security measures such as data encryption, role-based access control (RBAC), and compliance with GDPR and other regulations using Azure Active Directory (AAD) and Databricks Secrets.

o        Data Quality Management: Implement data quality checks in Databricks pipelines, ensuring consistency, accuracy, and validity of data.

o        Version Control & CI/CD: Use version control systems like Git and implement CI/CD pipelines using Azure DevOps for Databricks notebooks and data workflows.

•        Collaboration & Documentation

o        Cross-functional Collaboration: Collaborate with data scientists, analysts, and business users to develop insights and solutions that drive business objectives.

o        Documentation: Maintain thorough documentation of data architectures in DF Confluence, processes, and pipelines in Databricks for future scalability and team collaboration.

•        Advanced Analytics & Machine Learning

o        Machine Learning Integration: Collaborate with data science teams to build and deploy machine learning models using Databricks MLflow and integrate with Azure's machine learning services for operationalization.

o        Data Exploration: Support exploratory data analysis and business intelligence needs using Databricks Notebooks and integrate with Azure Power BI or other visualization tools.



Requirements
  • Hands-on experience with Databricks (mandatory) and proven expertise in designing, developing, and maintaining data solutions on the Databricks platform.
  • 6-8+ years of professional experience in Data Engineering
  • Strong SQL expertise
  • Proficiency in Python for data processing, automation, and backend data development.
  • Experience with Microsoft Azure, including cloud-based data services and data platform implementations.
  • Ability to work in modern cloud-based data environments and contribute to the migration and adoption of Databricks as a strategic data platform.
  • Strong analytical and problem-solving skills with a focus on scalable and high-quality data solutions.


  • Skills Required

    • Hands-on experience with Databricks and expertise designing, developing, and maintaining Databricks data solutions
    • 6-8+ years of professional experience in data engineering
    • Strong SQL expertise
    • Proficiency in Python for data processing, automation, and backend data development
    • Experience with Microsoft Azure, including cloud-based data services and data platform implementations
    • Ability to work in modern cloud-based data environments and contribute to Databricks migration and adoption
    • Strong analytical and problem-solving skills focused on scalable, high-quality data solutions
    Am I A Good Fit?
    beta
    Get Personalized Job Insights.
    Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

    The Company
    HQ: Caracas
    225 Employees
    Year Founded: 2015

    What We Do

    IDBC Group is a digital recruitment company and an IT & HR services firm, acting as one of Hungary's leading external IT resource providers, specializing in executive search, IT contracting, RPO, and workforce leasing services.

    Similar Jobs

    Hybrid
    Budapest, HUN
    100000 Employees
    18M-27M Annually

    Nebius Logo Nebius

    Senior Data Engineer

    Artificial Intelligence • Information Technology • Consulting
    In-Office or Remote
    29 Locations
    473 Employees

    Hiflylabs Logo Hiflylabs

    Senior Data Engineer

    Business Intelligence
    Hybrid
    Budapest, HUN
    160 Employees

    Instructure Logo Instructure

    Senior Software Engineer

    Edtech • Information Technology
    Hybrid
    Budapest, HUN
    1233 Employees

    Similar Companies Hiring

    Compa Thumbnail
    Artificial Intelligence • HR Tech • Software • Business Intelligence
    Irvine, California
    75 Employees
    NODA AI Thumbnail
    Artificial Intelligence • Information Technology • Software • Cybersecurity
    Sydney, AU
    54 Employees
    Golden Pet Brands Thumbnail
    Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
    El Segundo, California
    178 Employees

    Sign up now Access later

    Create Free Account

    Please log in or sign up to report this job.

    Create Free Account