- Job type
- Full-time
- Work mode
- On-site
- Level
- Not listed
- Department
- Research and Development (R&D)
- Experience
- Not listed
- Posted
- Aug 14, 2026
About the role
The Role
We are seeking a high-caliber, deeply innovative Vision-Language Model (VLM) expert to lead our efforts in teaching foundation models how to evaluate consumer products like an expert appraiser.
This is not a standard implementation role. You will be expected to invent new technologies, design novel architectures, and author proprietary training paradigms when existing open-source or commercial models fall short. You will go beyond simple object detection, engineering systems that can estimate highly abstract and valuable aspects of product images—such as generating rich, context-aware captions, predicting precise market price points, and evaluating subjective aesthetic quality or "beauty" scores.
If you are a pioneer who thrives on solving unsolved multimodal problems and wants your inventions to power a groundbreaking production platform, we want to hear from you.
Responsibilities
- Innovation & Invention: Pioneer novel neural architectures, loss functions, and multimodal integration techniques. We expect you to invent new AI technologies and methodologies to solve unprecedented challenges in visual product perception.
- VLM Fine-Tuning & Adaptation: Lead the deep fine-tuning, adaptation, and structural optimization of state-of-the-art Vision-Language Models (e.g., Qwen-VLM, PaliGemma2 etc.) for targeted, highly complex computer vision tasks.
- Complex Attribute Estimation: Develop the mathematical and architectural foundations to extract highly nuanced signals from product images, translating subjective concepts (like aesthetics and market value) into rigorous, learnable objectives.
- Dataset Strategy: Guide the creation, curation, and algorithmic augmentation of specialized multimodal datasets required to teach foundation models novel, domain-specific concepts.
- Evaluation & Metrics: Invent rigorous evaluation frameworks and custom metrics to accurately measure model performance on abstract tasks where standard academic benchmarks do not exist.
- Cross-Functional Collaboration: Work closely with engineering teams to ensure your proprietary research and new technologies translate seamlessly into scalable, production-ready systems.
Qualifications
- Education: Ph.D. in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a strictly related field.
- Publication Record: A strong track record of advancing the state-of-the-art, evidenced by first-author publications in top-tier AI, CV, or NLP venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR, ACL).
- Domain Expertise: Deep theoretical and practical understanding of Vision-Language Models, multimodal architectures, and modern transformer-based computer vision. You must understand the math and mechanics under the hood.
- Technical Stack: Expert-level proficiency in Python and deep learning frameworks (specifically PyTorch), with the ability to write optimized training loops when necessary.
- Problem Solving: A proven track record of formulating highly ambiguous, real-world visual problems into rigorous, solvable machine learning tasks.
- Industry Experience (bonus): Prior industrial experience as a Research Scientist or Machine Learning Engineer, specifically involving the deployment of deep learning models to large-scale production environments.
- E-commerce/Product ML (bonus): Previous experience applying machine learning to product imagery, retail technology, or computational aesthetics.
Additional
Competitive compensation
Daily catered lunch prepared by our chef
Company events