AI Research Intern
OpusClip · California, United States · 1 mo ago
RemoteRemoteEngineering$50/hrFull-time
About the role
We're looking for an AI Research Intern to join our AI team and explore cutting-edge research across multimodal AI, LLMs, computer vision, speech, and agent systems. You'll work across OpusClip, AgentOpus, and our next-generation AI products, collaborating closely with AI researchers and engineers to investigate emerging technologies, build research prototypes, and ship features used by millions of creators worldwide.
Responsibilities
- Research and develop deep learning models in one or more of the following areas, depending on product priorities:
- Computer vision (e.g. video enhancement, super-resolution, restoration)
- Speech & audio (e.g. speech enhancement, voice cloning, voice generation)
- Multimodal understanding and generation
- LLM post-training (e.g. SFT, RLHF, DPO)
- Build AI-powered product features by integrating frontier foundation models into production systems through prompt and context engineering strategies and Agent workflows (e.g., using LangChain, RAG frameworks).
- Collaborate with product and engineering teams to rapidly prototype and ship new AI capabilities across OpusClip and AgentOpus.
- Design scalable evaluation pipelines for multimodal AI systems.
- Develop domain-specific benchmarks using automated evaluation methods (e.g. LLM-as-a-Judge) together with task-specific visual, audio, and language quality metrics.
- Keep up with the latest AI research and open-source developments.
- Reproduce state-of-the-art research and translate new advances into production-ready systems.
Requirements
- Currently pursuing or recently completed a Master's degree in Computer Science, Artificial Intelligence, Mathematics, or a related field.
- Strong understanding of Transformer architecture and Attention mechanisms; familiarity with mainstream generative model families (GANs, diffusion models, autoregressive models).
- Familiarity with media processing fundamentals (video and/or audio — e.g., ffmpeg, codecs, signal processing basics).
- Strong programming skills in Python. Familiarity with Linux development environments, Git, and data structures.
- Fluent in English with strong technical reading and writing skills, including the ability to read research papers and write technical documentation.
- Hands-on experience in one or more of the following areas:
- Computer Vision (especially low-level vision): e.g., Real-ESRGAN, SwinIR, BasicVSR++, or diffusion-based SR; NTIRE / AIM challenge participation.
- Voice / speech: voice cleaning (speech enhancement / denoising / separation), voice cloning (TTS / voice conversion), or voice generation.
- LLM fine-tuning: SFT, RLHF / DPO, LoRA / PEFT, or post-training of open-source models.
Qualifications
- Basic Qualifications
- Education: Currently pursuing or recently completed a Master's degree in Computer Science, Artificial Intelligence, Mathematics, or a related field.
- Deep Learning Foundation: Solid understanding of Transformer architecture and Attention mechanisms; familiarity with mainstream generative model families (GANs, diffusion models, autoregressive models).
- Media Processing: Familiarity with media processing fundamentals (video and/or audio — e.g., ffmpeg, codecs, signal processing basics).
- Coding Skills: Strong programming skills in Python. Familiarity with Linux development environments, Git, and data structures.
- Fluent in English with strong technical reading and writing skills, including the ability to read research papers and write technical documentation.
- Hands-on Experience In One Or More Of The Following Computer Vision (especially low-level vision): e.g., Real-ESRGAN, SwinIR, BasicVSR++, or diffusion-based SR; NTIRE / AIM challenge participation.
- Voice / speech: voice cleaning (speech enhancement / denoising / separation), voice cloning (TTS / voice conversion), or voice generation.
- LLM fine-tuning: SFT, RLHF / DPO, LoRA / PEFT, or post-training of open-source models.
Preferred Qualifications
- Experience building Agent Systems or LLM-powered product features with frontier-model APIs (e.g., ChatGPT, Claude, Gemini).
- This role contributes to both OpusClip and AgentOpus products.
- Familiarity with TypeScript is a bonus, helpful for shipping product features.
- Ownership & execution: Involvement in projects from inception to completion, with strong coding fundamentals; open-source contributions are a plus.
- Research breadth: Academic background or interest in adjacent areas — video understanding and generation, multimodal systems, agents, and model evaluation / benchmarking.
- Publications: Involvement or interest in academic research, with a focus on top-tier venues like CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, ICASSP, Interspeech, AAAI, MM, TIP, TPAMI, ACL, EMNLP etc.