Jobs · Information Technology · California

ML Infra Engineer (Data Systems)

Physical Intelligence · San Francisco, CA · 1 mo ago
On-siteInformation TechnologyFull-time

About the Company

Physical Intelligence is bringing general-purpose AI into the physical world. We are a group of engineers, scientists, roboticists, and company builders developing foundation models and learning algorithms to power the robots of today and the physically-actuated devices of the future.

About the Team

The Infrastructure organization builds the foundations that make large-scale learning possible at PI. This includes training systems, data platforms, evaluation pipelines, and the tooling that allows researchers and roboticists to work with massive datasets safely and efficiently.

Responsibilities

  • Data Ingestion & Processing: Design and build high-throughput pipelines that validate, transform, and featurize raw multimodal data.
  • Batch & Streaming Systems: Operate large-scale batch and streaming workflows over massive datasets.
  • Storage Systems: Design object storage layouts, metadata systems, and efficient access patterns; choose file formats with performance and scalability in mind.
  • Data Lifecycle Management: Build systems for backfills, dataset rebuilds, garbage collection, and large-scale transformations.
  • Training-Time Performance: Optimize dataloaders, sharding, prefetching, caching, and throughput to reduce time from data arrival → model training.
  • Metadata & Indexing: Build scalable metadata stores for datasets, annotations, and training artifacts.
  • Data Movement: Move petabytes efficiently across clusters and environments.
  • Operational Correctness: Implement observability, validation, and guardrails to prevent silent data regressions.
  • Cross-Functional Collaboration: Work closely with cross-functional teams of researchers, engineers, and roboticists to translate evolving data needs into robust systems.

Requirements

  • Strong software engineering fundamentals.
  • Experience building distributed systems or large-scale data pipelines.
  • Comfort reasoning about performance, memory, I/O, and storage efficiency.
  • Familiarity with batch and/or streaming processing systems.
  • Experience with object storage systems and data format tradeoffs.
  • Ownership mindset: design, build, operate, and iterate on systems end-to-end.
  • Enjoy working closely with researchers and unblocking fast-moving projects.

Nice to Have

  • Experience with large ML training pipelines or dataloading systems.
  • Knowledge of columnar or custom data formats.
  • Experience with systems like ClickHouse, Ray, Flink, Spark, or similar.
  • Hands-on experience operating petabyte-scale datasets.
  • Debugging and fixing performance bottlenecks in data-heavy systems.

Similar jobs

Data Engineer

Wolfe, LLCPittsburgh, PA· Yesterday
Information Technologyapply on apply.wolfe.com

Data Engineer

Ascot GroupDallas, TX· Yesterday
Information Technology$110k–$125k/yrapply on fa-emkq-saasfaprod1.fa.ocs.oraclecloud.com

Data Engineer

ProlaioChicago, IL· 2 days ago
Analyst$134k/yrapply on job-boards.greenhouse.io

Data Engineer

Evlo AINew York, NY· 2 days ago
RemoteEngineeringapply on evlo.ai

Data Engineer

Chenega MIOS SBUHuntsville, AL· 1 mo ago
apply on chenega.jibeapply.com

Data Engineer

HaystackWashington, DC· 1 mo ago
RemoteInformation Technology$62k–$141k/yrapply on haystack.cv