Applied Scientist III - VLM R&D in Kirkland, Washington at wyzecam.com
NewJob Function: Science
wyzecam.com
Kirkland, Washington, 98033, United States
Posted on
New job! Apply early to increase your chances of getting hired.
Explore Related Opportunities
Models, Demonstrators, and Product Promoters jobs near me in WashingtonJobs near me in WashingtonModels, Demonstrators, and Product Promoters jobs
Job Description
At Wyze, we make smart home technology accessible to everyone. We're known for disrupting markets with high-quality, affordable products - from cameras to lighting to sensors and more. We believe technology should simplify life, not complicate it. We’re a fast-moving, customer-obsessed team driven by curiosity and powered by data.
The OpportunityWe are looking for an Applied Scientist to drive the development of our in-house vision-language models (VLMs) for smart home video understanding. In this role, you will track breakthrough research from academia and the broader AI community, and rapidly translate it into our production VLM development. You will help build the next generation of smart home physical AI — foundation models that understand the physical world of the home.
You will work with tens of millions of authorized videos to deeply investigate user event patterns and build a physical smart home foundation model that impacts over 10 million Wyze households. We believe in advancing the field, not just our product: we encourage publishing your research and releasing open-source models and datasets to benefit the broader community. This is a rare opportunity to shape a category-defining product at the intersection of frontier multimodal research and real-world deployment at massive scale.
What You'll Do
What We're Looking For
Nice to Have
CompensationThe base pay range for this role is $134,000 – $181,000 per year.
The OpportunityWe are looking for an Applied Scientist to drive the development of our in-house vision-language models (VLMs) for smart home video understanding. In this role, you will track breakthrough research from academia and the broader AI community, and rapidly translate it into our production VLM development. You will help build the next generation of smart home physical AI — foundation models that understand the physical world of the home.
You will work with tens of millions of authorized videos to deeply investigate user event patterns and build a physical smart home foundation model that impacts over 10 million Wyze households. We believe in advancing the field, not just our product: we encourage publishing your research and releasing open-source models and datasets to benefit the broader community. This is a rare opportunity to shape a category-defining product at the intersection of frontier multimodal research and real-world deployment at massive scale.
What You'll Do
- Follow the latest breakthroughs in multimodal and vision-language research from academia and industry, evaluate their relevance, and apply them to our in-house VLM development for smart home video understanding
- Train, fine-tune, and evaluate multimodal vision-language models on large-scale, real-world home video data
- Design and run rigorous evaluation pipelines to measure model quality on video understanding tasks such as event detection, activity recognition, and temporal reasoning
- Investigate user event patterns across tens of millions of authorized videos to inform model design and product direction
- Contribute to the architecture and training of a physical smart home foundation model, drawing on advances in visual transformers, physical world foundation models, and embodied AI
- Build rapid proofs of concept using AI-assisted research and development workflows, and carry promising directions from idea to validated prototype
- Publish research at top venues and contribute open-source models and datasets that help advance the community
- Help define research problems, set technical direction, and anticipate where academic research and industry solutions are heading
What We're Looking For
- PhD in Computer Vision, Machine Learning, or a related field; or a Master's degree with a strong track record of research or applied impact (publications, open-source contributions, or shipped ML systems)
- Hands-on experience training and evaluating multimodal vision-language models
- Experience in one or more of: visual transformer algorithm innovation, physical world foundation models, or embodied AI
- Strong research sense: the ability to define the right problems, choose promising directions, and predict how research trends will translate into industry solutions
- Proficiency with AI-assisted research and fast POC development — you use modern AI tools to multiply your own research velocity
- Solid engineering skills in Python and deep learning frameworks (e.g., PyTorch), with the ability to work with large-scale video data pipelines
Nice to Have
- Publications at top venues (CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, or similar)
- Experience with video understanding, long-context temporal modeling, or efficient inference for edge/cloud deployment
- Experience deploying ML models in consumer products at scale
CompensationThe base pay range for this role is $134,000 – $181,000 per year.
Scan to Apply
Just scan this QR code to apply from your phone.
Job Location
Kirkland, Washington, 98033, United States
Frequently asked questions about this position
Continue to apply
Enter your email to continue. You’ll be redirected to the employer’s application.By clicking Continue, you understand and agree to JobTarget's Terms of Use and Privacy Policy.