We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
The Role
As an AI Engineer specializing in LLM data, you will build and improve high-quality training data for foundation models across pre-training, mid-training, and post-training. Your work will include large-scale data curation and processing, LLM-based data synthesis, data quality evaluation, and experimentation to understand how different data choices affect model performance.
Key Responsibilities
Build, curate, and improve large-scale datasets for LLM pre-training, mid-training, and post-training.
Rapidly support time-sensitive data and model development tasks in a fast-moving research environment.
Develop and improve data processing pipelines including data extraction, cleaning, filtering, deduplication, quality scoring, transformation, and dataset composition.
Design and implement LLM-based data synthesis and augmentation pipelines, including prompt-based generation, filtering, refinement, and quality control of synthetic data.
Research and apply methods for improving training data quality, diversity, coverage, and efficiency.
Design experiments to understand the relationship between training data and model performance, and use model evaluation results to guide data improvements.
Develop scalable tools and workflows for processing and analyzing large datasets efficiently.
Analyze datasets using both statistical and model-based methods to identify quality issues, biases, duplication, distributional gaps, and opportunities for improvement.
Collaborate closely with researchers, model engineers, and other data teams to translate model development needs into effective data solutions.
Document datasets, experiments, data processing methodologies, and key findings clearly to support reproducibility and knowledge sharing.
Professional Experience - Required
- Strong programming skills in Python and experience building reliable engineering or research workflows.
- Hands-on experience with machine learning, deep learning, NLP, or large language models.
- Good understanding of modern LLM development, including how training data is used in pre-training and/or post-training.
- Experience working with large-scale datasets, including data processing, analysis, filtering, transformation, and quality control.
- Ability to independently investigate data or model quality issues, design experiments, and make data-driven technical decisions.
- Familiarity with common machine learning and LLM tools and frameworks such as PyTorch, Hugging Face, vLLM, or similar technologies.
- Strong problem-solving skills and the ability to work effectively in a fast-paced AI research and engineering environment.
- Strong communication and collaboration skills, with the ability to work closely with researchers and engineers across different technical areas.
Professional Experience - Preferred
Experience preparing data for large-scale LLM pre-training, continued/mid-training, supervised fine-tuning, preference optimization, reinforcement learning, or other post-training workflows.
Experience with synthetic data generation using LLMs, including generation, filtering, verification, or quality evaluation.
Experience designing or running LLM evaluations, benchmarks, model training, or fine-tuning experiments.
Understanding of how data quality, mixture, diversity, and scaling affect foundation model performance.
Experience with large-scale or distributed data processing and compute infrastructure.
Experience working with research teams on rapidly evolving foundation model or generative AI projects.
Contributions to open-source AI/ML projects, relevant publications, or demonstrated hands-on work with modern foundation models are a plus.
Skills Required
- Bachelor's degree in Computer Science, Data Science, Engineering, or related technical field
- Extensive experience in data engineering, data processing, and automation using Python
- Proficiency in designing and deploying web crawling solutions and automated data extraction
- Experience implementing and maintaining data processing pipelines and workflows
- Strong understanding of data structures, algorithms, databases, SQL, and performance optimization
- Experience with cloud infrastructure and distributed data processing frameworks (e.g., AWS, Spark, Kafka, Kubernetes)
- Excellent problem-solving abilities, attention to detail, and ability to address technical challenges rapidly
- Strong communication and collaboration skills with cross-functional teams
- Master's degree in Computer Science, Data Engineering, or related technical field (preferred)
- Proven track record supporting NLP or AI research teams with rapid and reliable data delivery (preferred)
- Experience refining outputs from large-scale AI models / LLM-generated data (preferred)
- Contributions to open-source projects or visibility in coding communities (preferred)
- Familiarity with latest advancements in NLP data processing and large language model technologies (preferred)
What We Do
First a passion, then an idea transformed into success – when it comes to pioneering automation and digitalisation technology, the ifm group is the ideal partner. Since its foundation in 1969, ifm has developed, produced and sold sensors, controllers, software and systems for industrial automation and for SAP-based solutions for supply chain management and shop floor integration worldwide. As one of the pioneers of Industry 4.0, ifm develops and implements consistent solutions to digitalise the entire value chain “from sensor to ERP”. Today, the second-generation family-run ifm group has more than 8,750 employees and is one of the worldwide market leaders. The group combines the internationality and innovative strength of a growing group of companies with the flexibility and close customer contact of a medium-sized company.







