Lead the development of custom models for speech recognition, transcription, audio segmentation, and speaker identification.
Build robust data pipelines: collecting, preprocessing, cleaning, and labeling large audio datasets.
Fine-tune and evaluate state-of-the-art open-source models (e.g., Whisper, wav2vec, HuBERT, Conformer) on proprietary datasets.
Design experiments and benchmark models for quality, latency, and domain adaptability.
Work closely with product teams to embed voice capabilities into real-time applications (e.g., live summarization, AI agents, call insights).
Maintain scalable training, evaluation, and inference workflows using modern ML tooling (e.g., PyTorch, Hugging Face, Weights & Biases).
Contribute to internal knowledge sharing and best practices around audio ML.
RequirementsRequirements
Strong experience with speech or audio ML: speech recognition, speaker diarization, voice activity detection, etc.
Hands-on experience in building models from scratch and fine-tuning large models.
Deep understanding of signal processing, feature extraction, and data augmentation for audio.
Proficient in Python and common ML libraries: PyTorch, NumPy, Scikit-learn, Hugging Face.
Familiarity with end-to-end ML pipelines: data cleaning, training, tuning, evaluation, and serving.
Comfort with using cloud platforms (GCP, AWS) and containerized environments.
High agency and comfort working in fast-paced, ambiguous environments.
Experience with LLM + speech integration (e.g., Whisper + GPT pipelines).
Knowledge of real-time systems or streaming inference.
Understanding of multilingual ASR challenges and dialect modeling (Arabic dialects a plus).
Experience with tools like DVC, MLflow, or W&B for experiment tracking.
Benefits
The chance to build domain-defining voice technology from scratch.
Exposure to real-world deployments and rapid iteration cycles.
Mentorship and collaboration with a team of high-agency engineers and researchers.
Flexible, remote-friendly work culture centered around ownership and outcomes.
Skills Required
- Strong experience with speech or audio ML (speech recognition, speaker diarization, VAD)
- Hands-on experience building models from scratch and fine-tuning large models
- Deep understanding of signal processing, feature extraction, and data augmentation for audio
- Proficient in Python and ML libraries: PyTorch, NumPy, Scikit-learn, Hugging Face
- Familiarity with end-to-end ML pipelines: data cleaning, training, tuning, evaluation, serving
- Comfort with cloud platforms (GCP, AWS) and containerized environments
- High agency and ability to work in fast-paced, ambiguous environments
- Experience with LLM + speech integration (e.g., Whisper + GPT pipelines)
- Knowledge of real-time systems or streaming inference
- Understanding of multilingual ASR challenges and dialect modeling (Arabic dialects a plus)
- Experience with tools like DVC, MLflow, or Weights & Biases for experiment tracking
What We Do
Clusterlab is an AI company specializing in tailored AI solutions and Arabic language models. Through its product Callab.ai, it creates smart voice agents that automate business calls like routing, bookings, outreach, and follow-ups. The company focuses on bridging the gap in the MENA region by developing voice AI that understands various Arabic dialects to enhance operational efficiency and customer experience.








