Arcade logo
Arcade

Research Scientist

USAOn-sitePosted 1 month ago

Apply opens Arcade's site. When you're back, we'll ask whether you applied.

Job type
Full-time
Work mode
On-site
Level
Not listed
Department
Research and Development (R&D)
Experience
Not listed
Posted
Aug 14, 2026

About the role

The Role

We are seeking a high-caliber, deeply innovative Vision-Language Model (VLM) expert to lead our efforts in teaching foundation models how to evaluate consumer products like an expert appraiser.

This is not a standard implementation role. You will be expected to invent new technologies, design novel architectures, and author proprietary training paradigms when existing open-source or commercial models fall short. You will go beyond simple object detection, engineering systems that can estimate highly abstract and valuable aspects of product images—such as generating rich, context-aware captions, predicting precise market price points, and evaluating subjective aesthetic quality or "beauty" scores.

If you are a pioneer who thrives on solving unsolved multimodal problems and wants your inventions to power a groundbreaking production platform, we want to hear from you.

Responsibilities

  • Innovation & Invention: Pioneer novel neural architectures, loss functions, and multimodal integration techniques. We expect you to invent new AI technologies and methodologies to solve unprecedented challenges in visual product perception.
  • VLM Fine-Tuning & Adaptation: Lead the deep fine-tuning, adaptation, and structural optimization of state-of-the-art Vision-Language Models (e.g., Qwen-VLM, PaliGemma2 etc.) for targeted, highly complex computer vision tasks.
  • Complex Attribute Estimation: Develop the mathematical and architectural foundations to extract highly nuanced signals from product images, translating subjective concepts (like aesthetics and market value) into rigorous, learnable objectives.
  • Dataset Strategy: Guide the creation, curation, and algorithmic augmentation of specialized multimodal datasets required to teach foundation models novel, domain-specific concepts.
  • Evaluation & Metrics: Invent rigorous evaluation frameworks and custom metrics to accurately measure model performance on abstract tasks where standard academic benchmarks do not exist.
  • Cross-Functional Collaboration: Work closely with engineering teams to ensure your proprietary research and new technologies translate seamlessly into scalable, production-ready systems.

Qualifications

  • Education: Ph.D. in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a strictly related field.
  • Publication Record: A strong track record of advancing the state-of-the-art, evidenced by first-author publications in top-tier AI, CV, or NLP venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR, ACL).
  • Domain Expertise: Deep theoretical and practical understanding of Vision-Language Models, multimodal architectures, and modern transformer-based computer vision. You must understand the math and mechanics under the hood.
  • Technical Stack: Expert-level proficiency in Python and deep learning frameworks (specifically PyTorch), with the ability to write optimized training loops when necessary.
  • Problem Solving: A proven track record of formulating highly ambiguous, real-world visual problems into rigorous, solvable machine learning tasks.
  • Industry Experience (bonus): Prior industrial experience as a Research Scientist or Machine Learning Engineer, specifically involving the deployment of deep learning models to large-scale production environments.
  • E-commerce/Product ML (bonus): Previous experience applying machine learning to product imagery, retail technology, or computational aesthetics.

Additional

Competitive compensation

Daily catered lunch prepared by our chef

Company events