The Role
Research and develop human-centric generative video models for facial expression control, head motion, lip synchronization, and realistic talking-head animation. Responsibilities include fine-tuning video diffusion models, creating large-scale audio/video datasets and training pipelines, exploring expression editing and multi-view consistency, and collaborating with product and creative teams. The role also involves staying current with video generation, speech-driven animation, and 3D-aware neural rendering research.
Summary Generated by Built In
Job Description:
We are seeking a Machine Learning Researcher to join our team and help advance the state of the art in human-centric generative video models. Your work will focus on improving expression control, lip synchronisation, and overall realism in models such as WAN and Hunyuan. You’ll collaborate with a world-class team of researchers and engineers to build systems that can generate lifelike talking-head videos from text, audio, or motion signals—pushing the boundaries of neural rendering and avatar animation. We are hiring remotely across the EMEA region.
Key Responsibilities
- Research and develop cutting-edge generative video models, with a focus on controllable facial expression, head motion, and audio-driven lip synchronisation.
- Fine-tune and extend video diffusion models such as WAN and Hunyuan for better visual realism and audio-visual alignment.
- Design robust training pipelines and large-scale video/audio datasets tailored for talking-head synthesis.
- Explore techniques for controllable expression editing, multi-view consistency, and high-fidelity lip sync from speech or text prompts.
- Work closely with product and creative teams to ensure models meet quality and production constraints.
- Stay current with the latest research in video generation, speech-driven animation, and 3D-aware neural rendering.
Must Haves
- Strong background in machine learning and deep learning, especially in generative models for video, vision, or speech.
- Hands-on experience with video synthesis tasks such as face reenactment, lip sync, audio-to-video generation, or avatar animation.
- Proficient in Python and PyTorch; familiar with libraries like MMPose, MediaPipe, DLIB, or image/video generation frameworks.
- Experience training large models and working with high-resolution audio/video datasets.
- Deep understanding of architectures such as transformers, diffusion models, GANs and motion representation techniques.
- Proven ability to work independently and drive research from idea to implementation.
- Strong problem-solving skills, ability to work autonomously in a remote-first environment.
Nice to Have
- PhD in Computer Vision, Machine Learning, or a related field, with publications in top-tier conferences (CVPR, ICCV, ICLR, NeurIPS, etc.).
- Familiarity with or contributions to open-source projects in lip sync, video generation, or 3D face modelling.
- Experience with real-time inference, model optimisation, or deployment for production applications.
- Knowledge of adjacent areas like emotion modelling, multimodal learning, or audio-driven animation.
- Experience working with or adapting models like WAN, Hunyuan or similar.
Skills Required
- Strong background in machine learning and deep learning, particularly generative models for video, vision, or speech
- Hands-on experience with video synthesis, face reenactment, lip synchronization, audio-to-video generation, or avatar animation
- Proficiency in Python and PyTorch
- Familiarity with MMPose, MediaPipe, DLIB, or image/video generation frameworks
- Experience training large models and working with high-resolution audio/video datasets
- Deep understanding of transformers, diffusion models, GANs, and motion representation techniques
- Ability to independently drive research from idea through implementation
- Strong problem-solving skills and ability to work autonomously in a remote-first environment
- PhD in Computer Vision, Machine Learning, or a related field
- Publications in top-tier conferences such as CVPR, ICCV, ICLR, or NeurIPS
- Familiarity with or contributions to open-source lip synchronization, video generation, or 3D face modeling projects
- Experience with real-time inference, model optimization, or production deployment
- Knowledge of emotion modeling, multimodal learning, or audio-driven animation
- Experience adapting WAN, Hunyuan, or similar models
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
BRAHMA AI is an enterprise AI content platform that enables organizations to create, manage, and distribute AI-driven media with intelligence, security, and efficiency. Formed through the integration of Prime Focus Technologies and Metaphysic, it combines CLEAR®, CLEAR® AI, ATMAN digital humans, and VAANI voice localization into one ecosystem. Guided by its Mind² philosophy, BRAHMA AI helps enterprises scale content creation, protect IP, and govern assets with provenance, consent, and trust.









