Data Engineer / ML Ops at VC-backed multimodal AI startup
Jack & Jill · United States · 1 wk ago
RemoteRemoteInformation TechnologyFull-time
VC-backed multimodal AI startup is seeking its first dedicated data hire to architect the end-to-end data strategy for a platform creating authentic, real-time AI humans.
About the role
You will build massive pipelines for multimodal datasets—including video, audio, and text—powering proprietary models at the frontier of AI research. This is a foundational role defining best practices for data engineering and ML operations.
Location: Remote
Why this role is remarkable
- Own the entire data strategy and infrastructure from scratch for a pioneering multimodal AI platform.
- Join a well-funded, Series B-backed startup supported by some of the world's leading venture firms.
- Work directly with founders on complex problems with no established playbook at the intersection of video and AI.
Responsibilities
- Source, curate, and structure large-scale multimodal datasets (video, audio, and text) for proprietary model training.
- Architect and optimize high-performance data pipelines and automated labeling workflows to scale ML production.
- Define engineering best practices and proactively build the data foundation that powers next-generation conversational AI.
Requirements
- Expert proficiency in Python, SQL, and building large-scale data processing systems in a production environment.
- Proven experience working with LLMs and managing complex multimodal data challenges like video and audio.
- A startup builder mentality with extreme ownership and the ability to define standards in new problem spaces.