Jobs · Engineering · California

Research Member of Technical Staff- Robot Learning Systems & Reliability

Rhoda AI · Mountain View, CA · 4 wk ago
On-siteEngineering$200k–$300k/yrFull-time

About the role

We're looking for a senior or staff-level Research Engineer or ML Systems Engineer to make our robot-learning pipeline reliable, reproducible, and measurable from end to end.

Responsibilities

  • Own the robot-learning pipeline end to end.
  • Establish and maintain a trusted workflow spanning robot data collection, data ingestion and compilation, post-training, checkpoint generation, inference, and real-robot evaluation.
  • Define the supported golden path.
  • Maintain known-good combinations of code, datasets, configurations, checkpoints, robot software, hardware settings, task stations, and evaluation procedures.
  • Build automated validation at every interface.
  • Develop checks for timestamp synchronization, sensor and action integrity, episode completeness, schema compatibility, dataset migrations, dataloader outputs, preprocessing behavior, model inputs, and configuration correctness.
  • Create end-to-end regression tests.
  • Build representative smoke tests that exercise data compilation, training, checkpoint loading, inference, replay or simulation, and real-robot execution.
  • Create small-scale overfit and canary experiments that catch correctness regressions before expensive training runs begin.
  • Ensure training and inference consistency.
  • Identify and prevent discrepancies in image processing, sensor normalization, temporal context, action representation, model configuration, and other transformations used across training and deployment.
  • Make failures observable and diagnosable.
  • Build instrumentation and debugging tools that help determine whether a performance regression originated in data, model code, infrastructure, inference, robot software, hardware configuration, the physical environment, or evaluation execution.
  • Improve real-robot evaluation reliability.
  • Partner with researchers and robot operations to establish stable benchmark stations, reference baselines, clear rubrics, repeatable trial protocols, operator procedures, and tracking of environmental variables that affect performance.
  • Lead cross-functional root-cause investigations.
  • Drive ambiguous failures to resolution across Research, Data Infrastructure, Model Infrastructure, Software, and Robot Operations.
  • Turn incidents and regressions into durable tests, monitors, documentation, and interface contracts.
  • Establish release and compatibility standards.
  • Define the validation required before changes to robot software, data systems, training code, or inference systems become part of the supported research workflow.
  • Measure and improve research reliability.
  • Track pipeline success rates, reproducibility, regression frequency, time to root cause, benchmark stability, and other metrics that reflect the health of the research workflow.

Qualifications

  • A strong track record owning, integrating, or debugging complex ML, robotics, autonomy, or other sensor-rich systems across multiple layers.
  • Excellent software engineering skills, including strong Python, maintainable system design, automated testing, CI, code review, and production-quality debugging practices.
  • Hands-on experience with PyTorch or a comparable ML framework.
  • You should be able to train or fine-tune models, inspect datasets and batches, interpret losses and outputs, load checkpoints, and debug inference behavior.
  • Experience building or operating ML workflows that span data processing, training, evaluation, and deployment - not only one isolated part of the stack.
  • Strong systems-debugging ability.
  • You can take an ambiguous report such as "robot performance dropped" and turn it into a structured investigation with controlled experiments, measurements, and a defensible root cause.
  • Familiarity with data-pipeline correctness, including schemas, versioning, temporal data, migrations, provenance, validation, and reproducibility.
  • Experience with Linux, containers, and modern compute environments such as Kubernetes, Slurm, or distributed GPU clusters.
  • Strong attention to detail and a willingness to investigate silent failures that do not produce obvious exceptions or crashes.
  • Ability to work effectively across research, infrastructure, software, and operations teams and to drive decisions without relying on formal managerial authority.
  • Clear written communication, including the ability to document interfaces, test plans, root-cause analyses, release criteria, and operational procedures.
  • Comfort working directly with physical systems and dealing with the variability and ambiguity inherent in real-world robotics.
  • A degree in Computer Science, Robotics, Electrical Engineering, or a related discipline - or equivalent practical experience.
  • A PhD and publication record are not required.
  • Nice To Have, But Not Required:
  • Experience in robotics, autonomous driving, drones, industrial automation, warehouse automation, or another domain that combines learned models with physical systems.
  • Familiarity with robot data collection, teleoperation, camera and sensor systems, calibration, proprioception, force/torque sensing, or action synchronization.
  • Experience with ROS or ROS2, real-time robot systems, simulation, replay infrastructure, or hardware-in-the-loop testing.
  • Experience with imitation learning, behavior cloning, robot post-training, reinforcement learning, video models, multimodal models, or learned control policies.
  • Experience building golden datasets, model-quality regression suites, data contracts, or automated dataset validation.
  • Familiarity with columnar datasets and data infrastructure, including formats such as Parquet and systems for indexing, querying, or migrating large datasets.
  • Experience debugging distributed training, GPU environments, checkpointing, networking, storage, Kubernetes, Slurm, or InfiniBand-related failures.
  • Experience designing statistically sound evaluation protocols for systems with noisy or stochastic performance.
  • Prior experience in ML platform engineering, research infrastructure, autonomy validation, release engineering, SRE, or systems integration.

Compensation Range

$200K - $300K

Similar jobs