We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
The RoleAs an AI Engineer Intern specializing in LLM data, you will work with our team to build and improve high-quality training data for foundation models across pre-training, mid-training, and post-training. You will gain hands-on experience with large-scale data processing, LLM-based data synthesis, data quality evaluation, and model experimentation.
Responsibilities
Build, curate, and improve datasets for LLM pre-training, mid-training, and post-training, including time-sensitive requests that may require fast turnaround.
Develop and improve data processing workflows for data extraction, cleaning, filtering, deduplication, transformation, and quality control.
Use LLMs to generate, refine, filter, and evaluate synthetic data.
Analyze data quality and identify issues such as duplication, low-quality samples, distribution gaps, and coverage limitations.
Run experiments to understand how different data sources and processing strategies impact model performance.
Support model inference, evaluation, and training experiments when needed to validate data quality.
Research new datasets, data processing techniques, and LLM data methodologies.
Collaborate closely with researchers and engineers on fast-moving foundation model projects.
Qualifications
Currently pursuing a Bachelor’s, Master’s, or PhD degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science, or a related technical field.
Strong programming skills in Python.
Good understanding of machine learning, deep learning, NLP, or large language models.
Hands-on experience with LLMs through research, coursework, projects, internships, or open-source work.
Comfortable working with datasets and performing data processing and analysis.
Strong problem-solving skills and willingness to learn new technologies quickly.
Ability to work independently and collaborate effectively in a fast-paced research environment.
Preferred Qualifications
Experience with PyTorch, Hugging Face, vLLM, or similar ML/LLM frameworks.
Experience with LLM data synthesis, fine-tuning, model evaluation, or prompt-based generation.
Familiarity with LLM training pipelines, including pre-training, supervised fine-tuning, or other post-training methods.
Experience working with large-scale datasets or distributed data processing.
Relevant research, open-source contributions, competitions, or personal projects in LLMs or generative AI.
Skills Required
- Currently pursuing a Bachelor's, Master's, or PhD in Computer Science, Artificial Intelligence, Machine Learning, Data Science, or a related technical field.
- Strong programming skills in Python.
- Good understanding of machine learning, deep learning, NLP, or large language models.
- Hands-on experience with LLMs through research, coursework, projects, internships, or open-source work.
- Comfortable working with datasets and performing data processing and analysis.
- Strong problem-solving skills and willingness to learn new technologies quickly.
- Ability to work independently and collaborate effectively in a fast-paced research environment.
- Experience with PyTorch, Hugging Face, vLLM, or similar ML/LLM frameworks.
- Experience with LLM data synthesis, fine-tuning, model evaluation, or prompt-based generation.
- Familiarity with LLM training pipelines, including pre-training, supervised fine-tuning, or other post-training methods.
- Experience working with large-scale datasets or distributed data processing.
- Relevant research, open-source contributions, competitions, or personal projects in LLMs or generative AI.
What We Do
First a passion, then an idea transformed into success – when it comes to pioneering automation and digitalisation technology, the ifm group is the ideal partner. Since its foundation in 1969, ifm has developed, produced and sold sensors, controllers, software and systems for industrial automation and for SAP-based solutions for supply chain management and shop floor integration worldwide. As one of the pioneers of Industry 4.0, ifm develops and implements consistent solutions to digitalise the entire value chain “from sensor to ERP”. Today, the second-generation family-run ifm group has more than 8,750 employees and is one of the worldwide market leaders. The group combines the internationality and innovative strength of a growing group of companies with the flexibility and close customer contact of a medium-sized company.







