Jobs · Engineering · New York

Member of Technical Staff - Engineering Lead, Data Ingestion

Reflection · New York, NY · 1 mo ago
On-siteEngineeringFull-time

About the role

The Data team at Reflection builds the training corpora our frontier models learn from. The Data Ingestion Lead provides front-line leadership of the team that builds the ingestion layer, spanning web crawl, data ingestion pipelines, and data lakes.

Responsibilities

  • Build, mentor, and grow a high-performing team of data ingestion engineers.
  • Coach and support your reports in understanding, and pursuing, their professional growth.
  • Provide front-line leadership across the full ingestion stack — web crawling and acquisition, extraction and normalization pipelines, and the data lakes that version and deliver training corpora and the campaigns that run across all three.
  • Stay hands-on: become familiar with the team's technical stack enough to make targeted contributions as an individual contributor.
  • Manage day-to-day execution: prioritize the team's work and run data acquisition and ingestion campaigns in a highly dynamic, fast-paced environment.
  • Guide technical and architectural decisions, emphasizing scalability, reliability, auditability, and cost.
  • Closer with research: partner with pre-training and data quality teams to tie ingestion decisions to measurable downstream model impact, and run experiments to evaluate crawling strategies, extraction methods, and ingestion tradeoffs.
  • Work across the wider data effort, data partnerships, digitization operations, and external vendors to onboard new sources within legal, licensing, and robots.txt constraints.
  • Raise the bar for technical judgment, prioritization, communication, and execution in a fast-moving environment.

Requirements

  • Experience building, mentoring, and growing data or infrastructure engineering teams while staying technically hands-on.
  • (Comfortable growing into leading a larger team quickly if you haven't managed at that scale before).
  • Deep experience building web-scale data acquisition or ingestion systems, with real ownership of production-grade pipelines at multi-TB to PB scale.
  • Strong coding ability and the credibility to earn the technical trust of a strong team.
  • Deep expertise in at least one of: web crawling & acquisition, large-scale extraction & ingestion pipelines, or data lakes / corpus storage & delivery with working knowledge across the others, and the ability to learn the rest.
  • Familiarity with the modern large-scale data toolkit: distributed compute (Ray, Beam, Spark), orchestration (Airflow, Prefect), formats (Parquet, JSONL, WARC), and object-store / data-lake architectures.
  • Familiarity with how LLMs are trained and evaluated, and an intuition for what makes data useful for training; comfortable designing experiments and using proxy quality signals to guide system improvements.
  • Ability to guide strategy and drive execution across concurrent data campaigns, and to partner effectively with research teams, external data vendors, and partnerships/legal on acquisition.

Qualifications

  • Fluency with the modern large-scale data toolkit: distributed compute (Ray, Beam, Spark), orchestration (Airflow, Prefect), formats (Parquet, JSONL, WARC), and object-store / data-lake architectures.
  • Experience with web crawling & acquisition, large-scale extraction & ingestion pipelines, or data lakes / corpus storage & delivery with working knowledge across the others, and the ability to learn the rest.
  • Experience with distributed compute (Ray, Beam, Spark), orchestration (Airflow, Prefect), formats (Parquet, JSONL, WARC), and object-store / data-lake architectures.
  • Experience with web crawling & acquisition, large-scale extraction & ingestion pipelines, or data lakes / corpus storage & delivery with working knowledge across the others, and the ability to learn the rest.

Skills

  • Experience with web crawling & acquisition, large-scale extraction & ingestion pipelines, or data lakes / corpus storage & delivery with working knowledge across the others, and the ability to learn the rest.
  • Experience with distributed compute (Ray, Beam, Spark), orchestration (Airflow, Prefect), formats (Parquet, JSONL, WARC), and object-store / data-lake architectures.
  • Experience with web crawling & acquisition, large-scale extraction & ingestion pipelines, or data lakes / corpus storage & delivery with working knowledge across the others, and the ability to learn the rest.

Benefits

  • Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.
  • Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options.
  • Health & wellness: Comprehensive medical, dental, vision, and life, with an annual wellness allowance.
  • Meals: Lunch and dinner are provided in the office daily.
  • Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys.
  • Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K.
  • Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable.
  • Team building: We have regular off-sites, happy hours, and team celebrations.

Pay

Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.

Schedule

Unlimited paid time off in the U.S. and 30 days in the U.K.

Similar jobs