This role is for one of the Weekday's clients
Min Experience: 4+ years
Location: India
JobType: full-time
We are looking for a highly skilled Data Scientist with expertise in Vision Language Models (VLMs), Computer Vision, and Multimodal AI to join our AI team. In this role, you will work on designing, developing, and optimizing next-generation AI models that can understand and reason across multiple data modalities, including images, text, and video. You will collaborate closely with machine learning engineers, data engineers, and product teams to build intelligent AI solutions that solve complex real-world problems.
The ideal candidate has a strong background in deep learning, computer vision, multimodal foundation models, and data-centric AI development. You should be comfortable working with large datasets, experimenting with state-of-the-art architectures, and deploying scalable AI solutions.
RequirementsKey Responsibilities
- Design, build, and optimize AI models using Vision Language Models (VLMs) and multimodal learning techniques.
- Develop computer vision pipelines for image classification, object detection, segmentation, visual reasoning, OCR, and image understanding.
- Build multimodal AI systems that effectively combine text, images, videos, and structured data to solve business challenges.
- Research, evaluate, and implement state-of-the-art deep learning models and emerging AI architectures.
- Prepare, clean, annotate, and validate large-scale datasets for machine learning and computer vision applications.
- Develop efficient data preprocessing, augmentation, and feature engineering pipelines to improve model quality.
- Fine-tune pre-trained vision and multimodal foundation models for domain-specific use cases.
- Conduct experiments, evaluate model performance using appropriate metrics, and continuously improve model accuracy and efficiency.
- Collaborate with engineering teams to deploy AI models into production and monitor their performance.
- Document experiments, methodologies, and best practices while staying updated with the latest advancements in AI research.
- Strong experience with Vision Language Models (VLMs) and multimodal AI architectures.
- Hands-on expertise in Computer Vision, including image processing and deep learning-based vision tasks.
- Solid understanding of Multimodal AI, integrating text, image, and video data for intelligent applications.
- Proficiency in Python and deep learning frameworks such as PyTorch or TensorFlow.
- Experience working with transformer-based architectures and foundation models.
- Strong knowledge of machine learning algorithms, model evaluation, optimization, and experimentation.
- Familiarity with GPU-based training and distributed deep learning workflows.
- Good understanding of data structures, algorithms, and software engineering best practices.
- Experience in Dataset Preparation, including data collection, cleaning, annotation, and quality validation.
- Expertise in Dataset Engineering for large-scale AI training pipelines.
- Hands-on experience in Fine-tuning foundation models, vision transformers, and multimodal models using modern techniques such as LoRA, PEFT, or transfer learning.
- Familiarity with MLOps tools, model versioning, and experiment tracking platforms.
- Exposure to cloud platforms such as AWS, Azure, or Google Cloud.
- Knowledge of containerization technologies like Docker and orchestration platforms such as Kubernetes is a plus.
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Data Science, Machine Learning, or a related field.
- 4–6 years of professional experience in Data Science, Computer Vision, or Applied AI.
- Strong analytical, problem-solving, and communication skills.
- Passion for researching and implementing cutting-edge AI technologies while delivering production-ready solutions.
Skills Required
- 4+ years professional experience in Data Science, Computer Vision, or Applied AI
- Experience with Vision Language Models (VLMs) and multimodal AI architectures
- Hands-on expertise in Computer Vision, image processing, and deep learning vision tasks
- Proficiency in Python
- Proficiency with deep learning frameworks such as PyTorch or TensorFlow
- Experience with transformer-based architectures and foundation models
- Familiarity with GPU-based training and distributed deep learning workflows
- Strong knowledge of machine learning algorithms, model evaluation, optimization, and experimentation
- Bachelor's or Master's degree in Computer Science, AI, Data Science, Machine Learning, or related field
- Good understanding of data structures, algorithms, and software engineering best practices
- Experience in dataset preparation, cleaning, annotation, and quality validation
- Dataset engineering for large-scale AI training pipelines
- Experience fine-tuning foundation models and vision transformers using techniques like LoRA, PEFT, or transfer learning
- Familiarity with MLOps tools, model versioning, and experiment tracking
- Exposure to cloud platforms (AWS, Azure, Google Cloud)
- Knowledge of Docker and Kubernetes
What We Do
Weekday is an AI-powered recruitment platform that helps startups hire top-tier engineering and product talent. By leveraging a massive database of white-collar professionals and advanced outreach tools, the company streamlines the hiring process through automated sourcing, AI-driven resume screening, and white-glove contingency services. Their mission is to modernize recruitment by enabling companies to discover and engage passive candidates efficiently, ensuring high-quality hires for critical roles.

.jpeg)





