Data Engineer

Posted 3 Days Ago
Be an Early Applicant
Pune, Mahārāshtra, IND
In-Office
Junior
Software
The Role
Build and optimize scalable real-time and batch ETL/ELT pipelines using Azure Databricks, PySpark, and Apache Spark. Design cloud data lakes, lakehouses, and warehouses; process structured and unstructured industrial data; and apply data cleaning and transformation techniques. Deploy solutions across cloud platforms, automate workflows with DevOps and CI/CD, use Docker and Kubernetes, monitor pipeline reliability, and collaborate with data scientists and engineers supporting AI/ML models.
Summary Generated by Built In
Job Title: Data Engineer
Location: Pune
Job Type: Full-Time ( WFO )
About TVARIT
TVARIT GmbH specializes in developing and delivering cutting-edge artificial intelligence (AI) solutions for the metal industry, including steel, aluminum, copper, cast iron, and more. Our software products empower customers to make intelligent, data-driven decisions, driving advancements in Predictive Quality (PsQ), Predictive Maintenance (PdM), and Energy Consumption Reduction (PsE), etc.
With a strong portfolio of renowned reference customers, state-of-the-art technology, a talented research team from prestigious universities, and recognition through esteemed awards such as the EU Horizon 2020 AI Prize, TVARIT is recognized as one of the most innovative AI companies in Germany and Europe.
We are seeking a self-motivated individual with a positive "can-do" attitude and excellent oral and written communication skills in English to join our team.
Job Description
We are looking for a Data Engineer with strong expertise in Azure Databricks, PySpark, and distributed computing to develop and optimize scalable ETL pipelines for manufacturing analytics. The role involves working with high-frequency industrial data to enable real-time and batch data processing.
Key Responsibilities
  • Build scalable real-time and batch processing workflows using Azure Databricks, PySpark, and Apache Spark.
  • Perform data pre-processing, including cleaning, transformation, deduplication, normalization, encoding, and scaling to ensure high-quality input for downstream analytics.
  • Design and maintain cloud-based data architectures, including data lakes, lakehouses, and warehouses, following Medallion Architecture.
  • Deploy and optimize data solutions on Azure (preferred), AWS, or GCP, with a focus on performance, security, and scalability.
  • Develop and optimize ETL/ELT pipelines for structured and unstructured data from IoT, MES, SCADA, LIMS, and ERP systems.
  • Automate data workflows using CI/CD and DevOps best practices, ensuring security and compliance with industry standards.
  • Monitor, troubleshoot, and enhance data pipelines for high availability and reliability.
  • Utilize Docker and Kubernetes for scalable data processing.
  • Collaborate with the automation team, data scientists, and engineers to provide clean, structured data for AI/ML models.
Desired Skills and Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Information Technology, or a related field.
  • Minimum 2 years of experience in data engineering, with a strong focus on cloud platforms such as Azure (preferred), AWS, or GCP.
  • Proficiency in PySpark, Azure Databricks, Python, and Apache Spark.
  • Expertise in relational databases (e.g., SQL Server, PostgreSQL), time-series databases (e.g., InfluxDB), and NoSQL databases (e.g., MongoDB, Cassandra).
  • Experience in containerization (Docker, Kubernetes).
  • Strong analytical and problem-solving skills with attention to detail.
  • Good to have knowledge of MLOps, DevOps, and model lifecycle management.
  • Excellent communication and collaboration skills, with a proven ability to work effectively as a team player.
  • Comfortable working in a dynamic, fast-paced startup environment, adapting quickly to changing priorities and responsibilities.

Skills Required

  • Bachelor's or Master's degree in Computer Science, Information Technology, or a related field
  • Minimum 2 years of data engineering experience
  • Experience with cloud platforms such as Azure, AWS, or GCP
  • Proficiency in PySpark, Azure Databricks, Python, and Apache Spark
  • Expertise with relational databases such as SQL Server and PostgreSQL
  • Experience with time-series databases such as InfluxDB
  • Experience with NoSQL databases such as MongoDB and Cassandra
  • Experience with Docker and Kubernetes
  • Strong analytical and problem-solving skills with attention to detail
  • Knowledge of MLOps, DevOps, and model lifecycle management
  • Excellent communication and collaboration skills
  • Ability to work in a dynamic, fast-paced startup environment
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Hesse
49 Employees
Year Founded: 2019

What We Do

We are a specialized Deep Tech company based in Frankfurt, Germany, and a provider of cutting-edge TVARIT Industrial AI (TiA) Technology with a focus on the manufacturing industry, especially foundries and metalworking companies. Our sole mission is to enable a sustainable, zero-waste manufacturing using our technology by almost eliminating energy losses, waste and maximizing machine and equipment availability. Through our unique patented technologies like “Hybrid AI” & “Transfer learning,” we guarantee first results (on average -30% less scrap and -20% less energy consumption) within 2-3 months and thus an ROI under 6 months. With our technology TiA, we combat the knowledge drain at foundries and metal companies. The expert knowledge is "conserved" and continuously increased making the platform smarter and more accurate through longer use. We are creating "Digitally Empowered Operators": by using TiA less experienced operators can be "empowered" to make optimal recipe adjustments in real-time, ensuring virtually defect-free and non-disruptive production. Our customers include world-renowned manufacturing companies in the metal industry such as Aditya Birla, Kamax, Schunk Group, and Maxion Wheels. In addition to proven results from more than 55 industrial projects in reducing scrap, increasing machine availability & productivity as well as reducing CO2 emissions, we have been recognized as the best AI company in Europe in the track “Best Smart Factory Startup in Europe" (out of 8 tracks with 495 participants) by the European Data Incubator (EDI) in 2020. Check out our website www.tvarit.com for more details on what we do

Similar Jobs

CrowdStrike Logo CrowdStrike

Data Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
India
11000 Employees

Zscaler Logo Zscaler

Integration Engineer

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
India
8697 Employees
In-Office or Remote
2 Locations
5000 Employees

EXL Logo EXL

Data Engineer

Information Technology • Database • Consulting
Remote or Hybrid
India
30246 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account