Technical Product Manager, Data Ingestion & Quality
Protege · United States · Yesterday
RemoteRemoteMarketingFull-time
About the role
The Product Manager will own the supply side of Protege's data platform, focusing on the pipeline that transforms raw data from partners into catalog-ready, trustworthy, and usable data.
Responsibilities
- Define the stages, validation gates, and quality checks for data moving from partner arrival to catalog-ready.
- Own the platform requirements that ensure the pipeline is repeatable across different modalities and verticals.
- Own metadata generation, defining what metadata gets extracted or generated at ingestion and how it gets stored and surfaced.
- Define what "catalog-ready" means and build the tooling that enforces these standards, validating data directly to ensure standards are met.
- Work with vertical stakeholders to translate their requirements into consistent platform-level standards that don't require custom engineering per deal.
Requirements
- 4–7 years of PM experience owning a data pipeline, data quality system, or data ingestion platform.
- Hands-on technical depth, able to write SQL, review pipeline logs, and understand data validation architectures.
- Experience with external data, having worked on a product that ingested messy, inconsistently formatted data from third-party partners and made it trustworthy.
- Build-versus-partner judgment, making vendor decisions in a fast-moving technical domain.
- Cross-functional credibility, holding technical conversations and product conversations in the same meeting.
Qualifications
- Nice to have: Experience with data quality frameworks, metadata standards, or catalog tooling, such as dbt, Great Expectations, data contracts, or similar.
- Familiarity with de-identification approaches for sensitive data, such as PHI, PII, or confidential enterprise data.
- Background in healthcare data operations, financial data infrastructure, or any domain where data quality has real downstream consequences.
- Experience with ML training pipelines or AI data workflows, where data fitness affects model outcomes.
- Exposure to data governance strategies.
What You'll Work On
- Ingestion pipeline product: Define stages, validation gates, and quality checks; own platform requirements.
- Metadata generation: Own product decisions around metadata extraction/generation, including transcripts, tags, confidence scores, schema inference, at what threshold, and storage/surfacing.
- QA standards and tooling: Define "catalog-ready," build tooling to enforce standards, validate data directly.
- Cross-vertical consistency: Work with vertical stakeholders to translate their requirements into consistent platform-level standards.
Success Metrics
- Ramp: Understand current data ingestion workflow, build context with relevant teams, identify quality risks.
- Take Ownership: Own first clear version of "catalog-ready," translate gaps into requirements engineering can build against.
- Operate Independently: Own roadmap for improvement, create repeatable operating rhythm for reviewing pipeline outputs, quality signals, and risks.