Jobs · Information Technology · New York

Senior ML/Data Engineer

Catapult · New York, United States · 1 wk ago
Information Technology$158k–$259k/yrFull-time

Catapult is building the future of sports performance technology, with a mission to Unleash the Potential of every athlete and team on earth. We work with over 5,000+ teams around the world, empowering coaches, managers, and trainers in premier teams in the NFL, NBA, NHL, MLS, EPL, AFL, NRL, NCAA, and more. We provide the information they need to optimize athletes’ health, game-day readiness, and performance, as well as in-game tactics.

About the role

This role is the data foundation of our AI platform. You will own the infrastructure every AI agent reasons over: the data architecture, the feature store, the sport knowledge graph, and the evaluation framework that makes every recommendation trustworthy before it reaches a practitioner. This is the highest-leverage engineering position in Phase 1 of a platform that will define what performance intelligence means in professional sport.

Responsibilities

  • Design and build systems that ingest, store, and serve athlete performance data at scale—from real-time sensor streams to longitudinal historical records—ensuring every layer is governed, isolated per client, and trustworthy at the moment of a high-stakes decision.
  • Build the feature store that makes derived metrics available to agents in milliseconds.
  • Design the sport knowledge graph that encodes relationships between athletes, loads, injuries, and outcomes.
  • Create the evaluation framework that validates every agent’s recommendations against the full distribution of real-world cases before any output reaches a practitioner.
  • Work directly with domain scientists and AI engineers, translating deep sport science requirements into durable, production-grade infrastructure that compounds in value with every season of data added to the platform.

Requirements

Non-negotiable:

  • 5+ years of production data engineering at scale, including time-series databases, data lakes, and feature stores in a real-time or near-real-time environment.
  • Experience building probabilistic evaluation frameworks or model calibration infrastructure; understanding the difference between a model that works and one that is trustworthy.
  • Strong Python, Golang, and SQL; experience with streaming ingestion (not batch-only) for live sensor or IoT data.
  • Graph database experience: designing schemas for complex relationship networks, not just querying existing ones.
  • Production experience with tenant-level data isolation at the infrastructure level, not just access control.
  • Comfort working directly with domain scientists and AI engineers; translating data requirements into durable infrastructure, not just pipelines.

Strongly Preferred:

  • Experience with causal inference or counterfactual modeling over graph structures.
  • Background in sports technology, wearable sensor data, or biomechanics data; understanding of what makes athlete time-series data structurally different from standard telemetry.
  • Experience building knowledge graphs with custom ontologies for a specific domain.
  • Familiarity with LLM evaluation frameworks and their limitations in probabilistic sport science contexts.
  • Experience working with AWS (ECS, EC2, Lambda, SNS, SQS, etc.), GraphQL, REST, gRPC, Postgres, Mongo.

The Stack

The platform requires:

  • Real-time streaming ingestion.
  • Time-series optimized storage at lake scale.
  • Millisecond feature serving.
  • Graph database with custom sport ontology.
  • Causal and simulation modeling layer.
  • Full provenance and audit logging.
  • Per-club isolation at the infrastructure level.

Why Catapult?

Catapult has spent twenty years collecting ground-truth athlete data from hardware on the body and on the field, across 40+ sports and 100+ countries. We are now building the intelligence layer that turns this data into actionable insights for coaches in high-stakes decision moments. The data engineer who builds this foundation will create the infrastructure layer underneath a platform that compounds in value with every decision made on it.

Pay

The target total compensation range for this position is $157,945 - $259,480 per year. This range includes base salary and a target incentive plan (which may include equity, commission, or other bonus structures). Your specific compensation within this range will be determined by factors such as geographic location, relevant experience, and job-related skills.

Benefits

  • Generous paid leave and recognized company holidays.
  • Comprehensive benefits package, including Health, Dental, and Vision insurance.
  • 401(k) retirement plan with company match.

Similar jobs

Senior ML / Data Engineer

CatapultNew York, NY· 1 wk ago
Information Technology$158k–$259k/yrapply on job-boards.greenhouse.io

Senior AI/ML & Data Engineer

Accenture Federal ServicesChantilly, VA· 3 wk ago
Information Technology$100k–$203k/yrapply on boards.greenhouse.io