Member of Technical Staff - Engineering Lead, Data Ingestion
Reflection · New York, NY · 1 mo ago
On-siteEngineeringFull-time
About the role
The Data team at Reflection builds the training corpora our frontier models learn from. The Data Ingestion Lead provides front-line leadership of the team that builds the ingestion layer, spanning web crawl, data ingestion pipelines, and data lakes.
Responsibilities
- Build, mentor, and grow a high-performing team of data ingestion engineers.
- Coach and support your reports in understanding, and pursuing, their professional growth.
- Provide front-line leadership across the full ingestion stack — web crawling and acquisition, extraction and normalization pipelines, and the data lakes that version and deliver training corpora and the campaigns that run across all three.
- Stay hands-on: become familiar with the team's technical stack enough to make targeted contributions as an individual contributor.
- Manage day-to-day execution: prioritize the team's work and run data acquisition and ingestion campaigns in a highly dynamic, fast-paced environment.
- Guide technical and architectural decisions, emphasizing scalability, reliability, auditability, and cost.
- Closer with research: partner with pre-training and data quality teams to tie ingestion decisions to measurable downstream model impact, and run experiments to evaluate crawling strategies, extraction methods, and ingestion tradeoffs.
- Work across the wider data effort, data partnerships, digitization operations, and external vendors to onboard new sources within legal, licensing, and robots.txt constraints.
- Raise the bar for technical judgment, prioritization, communication, and execution in a fast-moving environment.
Requirements
- Experience building, mentoring, and growing data or infrastructure engineering teams while staying technically hands-on.
- (Comfortable growing into leading a larger team quickly if you haven't managed at that scale before).
- Deep experience building web-scale data acquisition or ingestion systems, with real ownership of production-grade pipelines at multi-TB to PB scale.
- Strong coding ability and the credibility to earn the technical trust of a strong team.
- Deep expertise in at least one of: web crawling & acquisition, large-scale extraction & ingestion pipelines, or data lakes / corpus storage & delivery with working knowledge across the others, and the ability to learn the rest.
- Familiarity with the modern large-scale data toolkit: distributed compute (Ray, Beam, Spark), orchestration (Airflow, Prefect), formats (Parquet, JSONL, WARC), and object-store / data-lake architectures.
- Familiarity with how LLMs are trained and evaluated, and an intuition for what makes data useful for training; comfortable designing experiments and using proxy quality signals to guide system improvements.
- Ability to guide strategy and drive execution across concurrent data campaigns, and to partner effectively with research teams, external data vendors, and partnerships/legal on acquisition.
Qualifications
- Fluency with the modern large-scale data toolkit: distributed compute (Ray, Beam, Spark), orchestration (Airflow, Prefect), formats (Parquet, JSONL, WARC), and object-store / data-lake architectures.
- Experience with web crawling & acquisition, large-scale extraction & ingestion pipelines, or data lakes / corpus storage & delivery with working knowledge across the others, and the ability to learn the rest.
- Experience with distributed compute (Ray, Beam, Spark), orchestration (Airflow, Prefect), formats (Parquet, JSONL, WARC), and object-store / data-lake architectures.
- Experience with web crawling & acquisition, large-scale extraction & ingestion pipelines, or data lakes / corpus storage & delivery with working knowledge across the others, and the ability to learn the rest.
Skills
- Experience with web crawling & acquisition, large-scale extraction & ingestion pipelines, or data lakes / corpus storage & delivery with working knowledge across the others, and the ability to learn the rest.
- Experience with distributed compute (Ray, Beam, Spark), orchestration (Airflow, Prefect), formats (Parquet, JSONL, WARC), and object-store / data-lake architectures.
- Experience with web crawling & acquisition, large-scale extraction & ingestion pipelines, or data lakes / corpus storage & delivery with working knowledge across the others, and the ability to learn the rest.
Benefits
- Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.
- Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options.
- Health & wellness: Comprehensive medical, dental, vision, and life, with an annual wellness allowance.
- Meals: Lunch and dinner are provided in the office daily.
- Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys.
- Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K.
- Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable.
- Team building: We have regular off-sites, happy hours, and team celebrations.
Pay
Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.
Schedule
Unlimited paid time off in the U.S. and 30 days in the U.K.