AI Researcher (Multimodal Audio/Video Generation) at Series B multimodal AI lab
Jack & Jill · San Francisco, CA · 3 days ago
On-siteEngineeringFull-time
About the role
Lead research on audio-visual avatar generation at a cutting-edge lab building real-time conversational humans. You will design diffusion-based models for high-fidelity talking heads and neural avatars, bridging the gap between verbal and non-verbal communication. Partner with engineering teams to ship groundbreaking research into production-ready applications across healthcare and education.
Why this role is remarkable
- Shape the future of human-computer interaction by building AI avatars that see, hear, and respond with human-like emotional intelligence.
- Join a well-funded Series B startup backed by top-tier global venture capital firms with a culture of shipping research directly to production.
- Work directly with founders and a high-caliber research team to define the architecture for real-time, multimodal conversational timing and rendering.
What You Will Do
- Design and develop state-of-the-art diffusion models for long-video generation and audio-visual modeling to create life-like talking heads.
- Drive innovation in multimodal perception and neural rendering, capturing intricate verbal and non-verbal signals in perfectly synchronized flows.
- Mentor junior researchers and set strategic research directions while publishing impactful work at top-tier venues like CVPR and NeurIPS.
Qualifications
- Holds a PhD in Computer Science or a related field with 2-3+ years of experience applying large-scale generative models in industry.
- Possesses deep expertise in diffusion models, PyTorch, and GPU-optimized workflows with a proven track record of publication at premier AI conferences.
- Demonstrates mastery of multimodal generation spanning video and audio, ideally with experience in 3D graphics or Gaussian splatting techniques.