- Design, train, and optimize large-scale vision foundation models across image and video modalities
- Develop multimodal AI systems using architectures such as Vision Transformers (ViT), SAM, DINOv3, CLIP, and VLMs
- Apply self-supervised learning, transfer learning, and fine-tuning approaches for downstream tasks
- Build and enhance Vision-Language Models for visual reasoning and multimodal understanding
- Develop Retrieval-Augmented Generation (RAG) pipelines and multimodal knowledge retrieval systems
- Work with embeddings, vector databases, and semantic search frameworks
- Build scalable pipelines for training, evaluation, and deployment
- Manage large-scale image, video, and multimodal datasets
- Optimize distributed training workflows and model performance
- Translate research into production-ready solutions and explore emerging approaches in multimodal AI and generative AI
- Evaluate model quality, robustness, and retrieval effectiveness
Who are you?You bring strong expertise in computer vision, foundation models, and multimodal AI systems, along with the ability to deliver scalable solutions from research to production.
- Master’s or PhD in Computer Science, Artificial Intelligence, Machine Learning, or a related field
- Extensive experience in deep learning, computer vision, or multimodal AI
- Strong programming skills in Python and experience with PyTorch
- Deep understanding of computer vision, Vision Transformers, self-supervised learning, Vision-Language Models, and multimodal systems
- Hands-on experience with foundation models such as SAM, DINOv3, CLIP, BLIP/BLIP-2, LLaVA, or diffusion-based vision models
- Experience building RAG pipelines, semantic retrieval systems, and working with embeddings and vector databases such as FAISS, Milvus, Pinecone, or Weaviate
- Experience working with large-scale image and video datasets and distributed training environments
- Familiarity with GPU acceleration and scalable ML infrastructure
- Exposure to generative AI, multimodal reasoning systems, or large-scale perception systems
- Contributions to research, publications, or open-source projects are valued
- Opportunity to work on cutting-edge AI and multimodal technologies
- A collaborative, inclusive, and innovation-driven work environment
- Opportunities to learn, grow, and advance your career
- Exposure to large-scale, real-world AI challenges and global impact
- Competitive compensation and performance-based bonus
- Flexible and hybrid working options
- Employee wellness programs and professional development support
Who are we?
HERE Technologies is a location data and technology platform company. We empower our customers to achieve better outcomes – from helping a city manage its infrastructure or a business optimize its assets to guiding drivers to their destination safely.
At HERE we take it upon ourselves to be the change we wish to see. We create solutions that fuel innovation, provide opportunity and foster inclusion to improve people’s lives. If you are inspired by an open world and driven to create positive change, join us. Learn more about us on our YouTube Channel.
About the TeamYou will be part of a highly collaborative AI/ML team focused on developing next-generation Vision Foundation Models (VFMs), Vision-Language Models (VLMs), and multimodal AI systems. The team works at the intersection of research and scalable production systems, driving innovation in large-scale image and video understanding.
Skills Required
- Master's or PhD in Computer Science, AI, ML, or related field
- Extensive experience in deep learning, computer vision, or multimodal AI
- Strong programming skills in Python and experience with PyTorch
- Deep understanding of Vision Transformers, self-supervised learning, Vision-Language Models, and multimodal systems
- Hands-on experience with foundation models such as SAM, DINOv3, CLIP, BLIP/BLIP-2, LLaVA, or diffusion-based vision models
- Experience building RAG pipelines, semantic retrieval systems, and working with embeddings and vector databases (FAISS, Milvus, Pinecone, Weaviate)
- Experience with large-scale image and video datasets and distributed training environments
- Familiarity with GPU acceleration and scalable ML infrastructure
- Exposure to generative AI, multimodal reasoning systems, or large-scale perception systems
- Contributions to research, publications, or open-source projects
HERE Technologies Compensation & Benefits Highlights
-
Healthcare Strength — Medical, dental, and vision coverage is paired with life and disability insurance and an Employee Assistance Program. Core health coverage in the U.S. is often described as solid and high quality.
-
Leave & Time Off Breadth — Vacation/PTO, sick leave, paid holidays, paid volunteer time, and a formal sabbatical policy are offered. These programs provide notable breadth beyond standard leave.
-
Parental & Family Support — Parental leave is offered and often described as generous, with maternity and paternity options referenced. Family-oriented policies are visible, though specific durations differ by location.
HERE Technologies Insights
What We Do
HERE Technologies is a location data and technology company that created the first digital map over 35 years ago. Today we are the world's leading location platform company with a global footprint across 52 countries. Although our strongest presence is in the automotive industry, we also work with leading companies across a wide range of industries, including transport and logistics, mobility, manufacturing and retail and the public sector.
Why Work With Us
At HERE, we're always excited about discovering people who share our passion for building innovative solutions that make the world easier to navigate. We believe our success is powered by our team's diversity, creativity and collaboration and we're always looking for opportunities to grow it further.
Gallery
HERE Technologies Offices
Hybrid Workspace
Employees engage in a combination of remote and on-site work.

