- Build scalable real-time and batch processing workflows using Azure Databricks, PySpark, and Apache Spark.
- Perform data pre-processing, including cleaning, transformation, deduplication, normalization, encoding, and scaling to ensure high-quality input for downstream analytics.
- Design and maintain cloud-based data architectures, including data lakes, lakehouses, and warehouses, following Medallion Architecture.
- Deploy and optimize data solutions on Azure (preferred), AWS, or GCP, with a focus on performance, security, and scalability.
- Develop and optimize ETL/ELT pipelines for structured and unstructured data from IoT, MES, SCADA, LIMS, and ERP systems.
- Automate data workflows using CI/CD and DevOps best practices, ensuring security and compliance with industry standards.
- Monitor, troubleshoot, and enhance data pipelines for high availability and reliability.
- Utilize Docker and Kubernetes for scalable data processing.
- Collaborate with the automation team, data scientists, and engineers to provide clean, structured data for AI/ML models.
- Bachelor’s or Master’s degree in Computer Science, Information Technology, or a related field.
- Minimum 2 years of experience in data engineering, with a strong focus on cloud platforms such as Azure (preferred), AWS, or GCP.
- Proficiency in PySpark, Azure Databricks, Python, and Apache Spark.
- Expertise in relational databases (e.g., SQL Server, PostgreSQL), time-series databases (e.g., InfluxDB), and NoSQL databases (e.g., MongoDB, Cassandra).
- Experience in containerization (Docker, Kubernetes).
- Strong analytical and problem-solving skills with attention to detail.
- Good to have knowledge of MLOps, DevOps, and model lifecycle management.
- Excellent communication and collaboration skills, with a proven ability to work effectively as a team player.
- Comfortable working in a dynamic, fast-paced startup environment, adapting quickly to changing priorities and responsibilities.
Skills Required
- Bachelor's or Master's degree in Computer Science, Information Technology, or a related field
- Minimum 2 years of data engineering experience
- Experience with cloud platforms such as Azure, AWS, or GCP
- Proficiency in PySpark, Azure Databricks, Python, and Apache Spark
- Expertise with relational databases such as SQL Server and PostgreSQL
- Experience with time-series databases such as InfluxDB
- Experience with NoSQL databases such as MongoDB and Cassandra
- Experience with Docker and Kubernetes
- Strong analytical and problem-solving skills with attention to detail
- Knowledge of MLOps, DevOps, and model lifecycle management
- Excellent communication and collaboration skills
- Ability to work in a dynamic, fast-paced startup environment
What We Do
We are a specialized Deep Tech company based in Frankfurt, Germany, and a provider of cutting-edge TVARIT Industrial AI (TiA) Technology with a focus on the manufacturing industry, especially foundries and metalworking companies. Our sole mission is to enable a sustainable, zero-waste manufacturing using our technology by almost eliminating energy losses, waste and maximizing machine and equipment availability. Through our unique patented technologies like “Hybrid AI” & “Transfer learning,” we guarantee first results (on average -30% less scrap and -20% less energy consumption) within 2-3 months and thus an ROI under 6 months. With our technology TiA, we combat the knowledge drain at foundries and metal companies. The expert knowledge is "conserved" and continuously increased making the platform smarter and more accurate through longer use. We are creating "Digitally Empowered Operators": by using TiA less experienced operators can be "empowered" to make optimal recipe adjustments in real-time, ensuring virtually defect-free and non-disruptive production. Our customers include world-renowned manufacturing companies in the metal industry such as Aditya Birla, Kamax, Schunk Group, and Maxion Wheels. In addition to proven results from more than 55 industrial projects in reducing scrap, increasing machine availability & productivity as well as reducing CO2 emissions, we have been recognized as the best AI company in Europe in the track “Best Smart Factory Startup in Europe" (out of 8 tracks with 495 participants) by the European Data Incubator (EDI) in 2020. Check out our website www.tvarit.com for more details on what we do









