Senior ML / Data Engineer
About the role
Catapult is building the future of sports performance technology, with a mission to Unleash the Potential of every athlete and team on earth. We are looking for a Senior ML / Data Engineer to build the data infrastructure that powers the next generation of performance intelligence. This is a senior production engineering role where you will own significant parts of the architecture that ingest, store, transform, serve, and evaluate athlete performance data. The systems you build will support real-time and machine learning use cases across a global, multi-tenant platform.
You will work closely with data scientists, ML engineers, software engineers, and sports scientists to turn complex performance requirements into reliable, scalable production infrastructure. This role is suited to an engineer who has spent several years building and operating production systems and is comfortable taking ownership of architecture and technical decisions.
Responsibilities
- Design and build production data infrastructure for high-volume athlete performance and sensor data.
- Build and operate real-time and near-real-time ingestion systems for streaming data.
- Design storage and data architectures for high-volume time-series data and longitudinal athlete records.
- Build infrastructure that makes production features and derived metrics available to machine learning systems and AI agents with low latency.
- Design and implement graph data models and schemas representing relationships between athletes, training loads, injuries, performance, and outcomes.
- Build data and ML evaluation infrastructure that helps measure model reliability, calibration, and performance across real-world cases.
- Design systems that maintain strong tenant-level data isolation across clubs and customers.
- Establish appropriate data provenance, lineage, auditability, and observability across the platform.
- Work with ML and AI engineers to provide reliable data foundations for model training, inference, and evaluation.
- Work with sport scientists and domain experts to translate complex requirements into durable production systems.
- Make pragmatic technology and architecture decisions as the platform evolves.
Requirements
- 5+ years of full-time professional software or data engineering experience, excluding internships, university placements, coursework, and academic projects.
- Proven experience designing, building, and operating production data infrastructure at scale.
- Strong experience working with time-series data or time-series databases, such as InfluxDB, TimescaleDB, Prometheus, ClickHouse, or equivalent technologies.
- Significant experience with real-time or streaming data ingestion, using technologies such as Kafka, Kinesis, Flink, Spark Streaming, Pulsar, or equivalent.
- Experience designing graph data models or graph database schemas, not simply querying or consuming an existing graph database.
- Experience designing or operating multi-tenant systems with tenant-level data isolation.
- Strong Python and SQL skills. Professional experience with Go is highly desirable.
- Experience working with production systems where reliability, scalability, observability, and data correctness matter.
- Ability to take ownership of ambiguous technical problems and turn them into practical production architectures.
- Experience working directly with data scientists, ML engineers, or other technical domain specialists.
Preferred Qualifications
- Experience building probabilistic evaluation, model calibration, or model monitoring infrastructure.
- Experience with causal inference, counterfactual modelling, or simulation.
- Experience working with wearable sensors, IoT data, biomechanics, sports technology, or other high-frequency telemetry.
- Experience building knowledge graphs or domain-specific ontologies.
- Experience with LLM or AI evaluation frameworks and an understanding of their limitations.
- Experience with AWS including ECS, EC2, Lambda, SNS, SQS, or related services.
- Experience with GraphQL, REST, gRPC, Postgres, MongoDB, or similar technologies.
About the Platform
The platform requires capabilities including:
- Real-time streaming ingestion
- High-volume time-series storage
- Data lake and analytical infrastructure
- Low-latency feature serving
- Graph databases and domain-specific ontologies
- ML evaluation and calibration
- Causal and simulation modelling
- Data provenance and audit logging
- Strong tenant-level isolation
- Production observability and reliability
No individual vendor or technology is locked in. We value engineers who understand the underlying architectural trade-offs and can choose the right technology for the problem.
Why Catapult?
Catapult has spent more than twenty years collecting ground-truth athlete data from hardware on the body and on the field, across more than 40 sports and 100 countries. That data represents a significant opportunity to build new forms of performance intelligence. The challenge is turning that data into systems that are reliable, explainable, and useful at the point where coaches and performance staff need to make decisions.
You will have the opportunity to work on a technically challenging combination of real-time data, machine learning, time-series infrastructure, graph data, and AI evaluation, with direct impact on products used by elite sporting organisations around the world.
Pay
The target Total Compensation range for this position is $157,945 to $259,480 per year. This range is inclusive of base salary and a target incentive plan, which may include equity, commission, or other bonus structures. Your specific compensation will be determined by factors including geographic location, relevant experience, and job-related skills.
Benefits
- Paid leave and recognised company holidays
- Comprehensive benefits package including Health, Dental, Vision
- 401(k) retirement plan with company match