ML Encoder Lead (Senior Role)
The Fountain Group · South San Francisco, CA · 1 wk ago
HybridOTHR$100–$200/hrContract
This position will report to a client located in South San Francisco, CA. Local candidates able to have an onsite presence as needed are preferred. Well-qualified remote candidates able to overlap with PST working hours are also considered.
Pay
$100–$200/hour W2, based on qualifications.
Schedule
6-month contract to start; full-time hours overlapping PST business hours.
Responsibilities
- Design and develop customer representation-learning models using longitudinal transaction, sales, and interaction data.
- Define pretraining objectives and build dense customer embeddings for downstream GenAI and analytics applications.
- Develop, train, and evaluate encoder models using large-scale sparse event data.
- Design rigorous model evaluation using time-based splits, leakage detection, cold-start analysis, held-out populations, uncertainty, and baseline comparisons.
- Assess embedding quality through downstream signal, calibration, stability, drift, and subgroup performance.
- Build production-ready training and serving pipelines with reproducibility, model versioning, monitoring, and data contracts.
- Lead the transition of representation-learning research into scalable production systems.
- Present modeling results and evaluation evidence to senior stakeholders and make recommendations on model continuation or termination.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, Statistics, or a related technical field.
- Hands-on experience personally training encoder or embedding models and designing pretraining objectives.
- Deep expertise in representation learning, including self-supervised or contrastive learning.
- Strong experience with sequence/temporal modeling, Transformers, Graph Neural Networks, or recommender-system embeddings.
- Experience modeling large-scale, sparse, longitudinal event data such as transactions, claims, clickstream, customer journeys, or engagement histories.
- Experience developing inductive representations for cold-start entities using entity features rather than lookup tables.
- Advanced Python with PyTorch or JAX.
- Strong SQL and distributed data processing experience.
- Experience with cloud-based model training at scale.
- Experience building production ML pipelines, model serving, versioning, monitoring, and reproducible training workflows.
Preferred Skills
- Customer-360 representations and behavioral embeddings.
- Recommender systems or foundation models applied to event data.
- Experience addressing privacy, fairness, and re-identification risks in learned representations.
- Publications, patents, or public applied work in representation learning.
- Experience with large-scale behavioral data in consumer technology, marketplaces, streaming, financial services, payments, or advertising technology.