OpusClip logo
OpusClip

Lead AI Researcher

Mountain View, USAOn-sitePosted 1 month ago

Apply opens OpusClip's site. When you're back, we'll ask whether you applied.

Job type
Full-time
Work mode
On-site
Level
Lead
Department
Engineering
Experience
Not listed
Posted
Sep 2, 2026

About the role

We want to hire a hands-on Lead AI Researcher with strong technical depth and people management is only a PLUS!

What we are looking for

  • Proven experience shipping AI/ML models from prototype to production at scale.
  • Deep expertise in model post-training, including fine-tuning, preference optimization, evaluation, and learning from user feedback. Experience building data flywheels that turn user behavior and feedback into training data and continuous model improvements.
  • Strong inference optimization experience for self-hosted models, including latency, throughput, GPU utilization, quantization, and serving cost.
  • Strong model evaluation skills, including offline benchmarks, human evaluation, online experiments, and product-quality metrics.
  • Practical experience applying LLMs or multimodal models to video applications, such as highlight detection, content curation, ranking, personalization, and editing decisions.
  • Familiarity with video-related models and systems, including but not limited to vision-language models, ASR/transcription, video understanding, and enhancement/upscaling, etc..
  • Able to make build-vs-buy decisions across proprietary APIs, open-source models, and internally trained models.
  • Capable of setting the technical roadmap, reviewing architecture, mentoring engineers, and managing a small high-performing AI team.
  • Product-oriented and pragmatic: understands how to balance model quality, latency, reliability, and infrastructure cost.

Nice to have

  • Experience with consumer video, creator tools, recommendation/content systems, or other high-scale multimodal products is plus.
  • Distributed training infrastructure and inference optimization experience for self-hosted models: including latency, throughput, GPU utilization, quantization, and serving cost.
  • Familiarity with video-related models and systems: vision-language models, ASR/transcription, video understanding/reasoning/generation, and enhancement/upscaling