Applied Scientist III - VLM R&D

Posted 2 Days Ago
Be an Early Applicant
Kirkland, WA, USA
Hybrid
134K-181K Annually
Senior level
Computer Vision • eCommerce • Healthtech • Internet of Things • Security
Making great technology accessible!
The Role
Develop and deploy in-house vision-language models for smart home video understanding: track research, train and evaluate multimodal VLMs on large-scale video data, design evaluation pipelines, build prototypes, inform product direction, and publish or open-source research.
Summary Generated by Built In
At Wyze, we make smart home technology accessible to everyone. We're known for disrupting markets with high-quality, affordable products - from cameras to lighting to sensors and more. We believe technology should simplify life, not complicate it. We’re a fast-moving, customer-obsessed team driven by curiosity and powered by data. 
The Opportunity
We are looking for an Applied Scientist to drive the development of our in-house vision-language models (VLMs) for smart home video understanding. In this role, you will track breakthrough research from academia and the broader AI community, and rapidly translate it into our production VLM development. You will help build the next generation of smart home physical AI — foundation models that understand the physical world of the home.

You will work with tens of millions of authorized videos to deeply investigate user event patterns and build a physical smart home foundation model that impacts over 10 million Wyze households. We believe in advancing the field, not just our product: we encourage publishing your research and releasing open-source models and datasets to benefit the broader community. This is a rare opportunity to shape a category-defining product at the intersection of frontier multimodal research and real-world deployment at massive scale.

What You'll Do
  • Follow the latest breakthroughs in multimodal and vision-language research from academia and industry, evaluate their relevance, and apply them to our in-house VLM development for smart home video understanding
  • Train, fine-tune, and evaluate multimodal vision-language models on large-scale, real-world home video data
  • Design and run rigorous evaluation pipelines to measure model quality on video understanding tasks such as event detection, activity recognition, and temporal reasoning
  • Investigate user event patterns across tens of millions of authorized videos to inform model design and product direction
  • Contribute to the architecture and training of a physical smart home foundation model, drawing on advances in visual transformers, physical world foundation models, and embodied AI
  • Build rapid proofs of concept using AI-assisted research and development workflows, and carry promising directions from idea to validated prototype
  • Publish research at top venues and contribute open-source models and datasets that help advance the community
  • Help define research problems, set technical direction, and anticipate where academic research and industry solutions are heading

What We're Looking For
  • PhD in Computer Vision, Machine Learning, or a related field; or a Master's degree with a strong track record of research or applied impact (publications, open-source contributions, or shipped ML systems)
  • Hands-on experience training and evaluating multimodal vision-language models
  • Experience in one or more of: visual transformer algorithm innovation, physical world foundation models, or embodied AI
  • Strong research sense: the ability to define the right problems, choose promising directions, and predict how research trends will translate into industry solutions
  • Proficiency with AI-assisted research and fast POC development — you use modern AI tools to multiply your own research velocity
  • Solid engineering skills in Python and deep learning frameworks (e.g., PyTorch), with the ability to work with large-scale video data pipelines

Nice to Have
  • Publications at top venues (CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, or similar)
  • Experience with video understanding, long-context temporal modeling, or efficient inference for edge/cloud deployment
  • Experience deploying ML models in consumer products at scale



Compensation
The base pay range for this role is $134,000 – $181,000 per year.

Skills Required

  • PhD in Computer Vision, Machine Learning, or related field; or Master's with strong research/applied impact (publications, open-source, or shipped ML systems)
  • Hands-on experience training and evaluating multimodal vision-language models
  • Experience in visual transformer algorithm innovation, physical world foundation models, or embodied AI
  • Strong research sense: define problems, choose directions, and translate research to industry solutions
  • Proficiency with AI-assisted research and rapid proof-of-concept development
  • Solid engineering skills in Python and deep learning frameworks (e.g., PyTorch), and ability to work with large-scale video data pipelines
  • Publications at top venues (CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR)
  • Experience with video understanding, long-context temporal modeling, or efficient inference for edge/cloud deployment
  • Experience deploying ML models in consumer products at scale
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Kirkland, WA
350 Employees
Year Founded: 2017

What We Do

It’s our goal to become the most user-centric smart home technology company. We’re passionate about providing users access to high-quality products at great prices, we relentlessly keep costs low by partnering with the world’s most efficient manufacturers, we cut out “channel fat” by selling directly from our own website, and, unlike our competitors, we don’t seek a high-profit margin over our cost base, passing on all of these savings to our users. As we grow, we will continue to launch high-quality, affordable smart home products that enrich people’s lives and make great technology accessible to everyone!

Why Work With Us

We’re passionate about providing customers access to high-quality products at great prices. We relentlessly keep costs low by partnering with the world’s most efficient manufacturers. We cut out “channel fat” by selling directly from our own website.

Gallery

Gallery

Similar Jobs

Boeing Logo Boeing

Senior General Analyst

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
Tukwila, WA, USA
170000 Employees
127K-171K Annually

Boeing Logo Boeing

Electromagnetic Effects Design and Analysis Engineers (Associate, Experienced and Expert)ic Compatibility)

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
Tukwila, WA, USA
170000 Employees
99K-198K Annually

Boeing Logo Boeing

Operations Manager

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
Everett, WA, USA
170000 Employees
127K-155K Annually

Boeing Logo Boeing

Operations Analyst

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
Renton, WA, USA
170000 Employees
119K-161K Annually

Similar Companies Hiring

Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees
OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account