Principal Machine Learning Engineer
About the Role
We're hiring a Principal Machine Learning Engineer to lead the data-science function powering our content platform—the curation, quality control, and enrichment of a petabyte-scale library using computer vision and AI.
This role sits at the intersection of data science, data engineering, and computer vision. The ideal candidate possesses deep algorithmic expertise and the proficiency to manage massive data infrastructures. You will define how content is understood across our library, building upon a robust foundation and leading a high-performing team.
You'll work closely with our CTO and existing data team and play a central role in building out the data science function in San Francisco.
Responsibilities
- Multimodal Curation & QC: Manage curation and quality control for one of the market's largest content collections, ensuring library integrity at petabyte scale.
- Automated Enrichment: Deploy CV and LLM models for classification, object detection, and metadata enrichment to enhance content discoverability and value.
- Standards & Criteria Design: Define and automate grading standards tailored to various content types, building consistent and scalable evaluation models.
- Big Data Infrastructure: Execute complex algorithms across AWS and on-site lakehouse environments, focusing on video understanding and multi-modal classifiers.
- Team Building & Mentorship: Recruit, train, and supervise a growing ML team, fostering professional development and maximizing productivity through performance data.
- Strategic Alignment: Partner with Research, Product, and Engineering teams to refine Trust & Safety strategies and ensure the success of project SLAs.
- Executive Reporting: Report directly to the CTO, providing effective communication on risks, mitigation, and the evaluation of scalable tools and processes.
Requirements
- 6+ years of experience in ML/Data Science with a core specialty in computer vision (classification, object detection, content enrichment).
- Proven track record of operating on large-scale data, demonstrating fluency in both AI algorithm depth and big-data engineering.
- The rare hybrid of CV/AI algorithm depth and big-data engineering ability.
- Experience in leading and mentoring data science teams within a fast-paced environment.
- Video/multimedia expertise and a background in content marketplaces or moderation platforms are significant advantages.
What You're Building Towards
As our data lab scales, the ambition grows beyond curation and quality control. There's a genuine opportunity ahead to use our enriched, petabyte-scale library to train and publish proof-of-concept models that demonstrate to the market what better data—and better content intelligence—can do. You'll lead the data-science function at the foundation of that vision, and grow the team that brings it to life.
What You Won't Own
Production model training at scale from scratch. Your intellectual energy goes into applying computer vision and AI to curate, classify, and enrich content at volume—and into leading the data-science team that does it. That said, understanding how models are built and running small training experiments to validate the dataset and the enrichment quality are genuine advantages here, not disqualifiers.
Our Stack
- Petabyte-scale image and video content
- Computer vision + LLM models
- Amazon Nova Multimodal Embeddings
- MongoDB as a Vector DB
- dbt
- Spark
- Trino
- Apache Iceberg
- Argo Workflows and Argo Events
- Ness
We value the rare combination of CV/AI algorithm depth and big-data engineering fluency over familiarity with any single tool.