Multimodal AI Research Engineer at Series B-backed AI research lab
Jack & Jill · San Francisco, CA · 1 mo ago
RemoteRemoteInformation TechnologyFull-time
Series B-backed AI research lab.
About the role
As a Research Engineer, you will develop state-of-the-art multimodal foundation models to power next-generation interactive digital interfaces. You'll focus on the intersection of audio, vision, and language to create low-latency, expressive systems. This role offers the unique opportunity to conduct high-level research while shipping impactful technology to a global user base at scale.
Why this role is remarkable
- Work on cutting-edge foundation models that bridge generative video and real-time interaction at the frontier of AI.
- Join a well-funded Series B startup backed by premier global venture capital firms with a high-growth trajectory.
- Transition research directly into production systems, seeing your work reach millions of users rather than staying in papers.
Responsibilities
- Research and develop multimodal conversational models to control expressive digital behaviors in real-time with sub-second latency.
- Fine-tune audio-visual foundation models to improve responsiveness, controllable expressions, and task-specific performance.
- Partner with applied machine learning teams to migrate research prototypes into high-performance, scalable production pipelines.
Requirements
- PhD or equivalent research experience in Deep Learning, specializing in Large Multimodal Models or Generative AI.
- Expert-level proficiency in PyTorch and experience building complex deep learning pipelines for audio or video understanding.
- Proven track record of high-impact research, ideally with a history of publications at top-tier venues like NeurIPS, CVPR, or ICLR.
Location
San Francisco, USA; London, UK; or Remote.