Research Member of Technical Staff- Robot Learning Systems & Reliability
Rhoda AI · Mountain View, CA · 4 wk ago
On-siteEngineering$200k–$300k/yrFull-time
About the role
We're looking for a senior or staff-level Research Engineer or ML Systems Engineer to make our robot-learning pipeline reliable, reproducible, and measurable from end to end.
Responsibilities
- Own the robot-learning pipeline end to end.
- Establish and maintain a trusted workflow spanning robot data collection, data ingestion and compilation, post-training, checkpoint generation, inference, and real-robot evaluation.
- Define the supported golden path.
- Maintain known-good combinations of code, datasets, configurations, checkpoints, robot software, hardware settings, task stations, and evaluation procedures.
- Build automated validation at every interface.
- Develop checks for timestamp synchronization, sensor and action integrity, episode completeness, schema compatibility, dataset migrations, dataloader outputs, preprocessing behavior, model inputs, and configuration correctness.
- Create end-to-end regression tests.
- Build representative smoke tests that exercise data compilation, training, checkpoint loading, inference, replay or simulation, and real-robot execution.
- Create small-scale overfit and canary experiments that catch correctness regressions before expensive training runs begin.
- Ensure training and inference consistency.
- Identify and prevent discrepancies in image processing, sensor normalization, temporal context, action representation, model configuration, and other transformations used across training and deployment.
- Make failures observable and diagnosable.
- Build instrumentation and debugging tools that help determine whether a performance regression originated in data, model code, infrastructure, inference, robot software, hardware configuration, the physical environment, or evaluation execution.
- Improve real-robot evaluation reliability.
- Partner with researchers and robot operations to establish stable benchmark stations, reference baselines, clear rubrics, repeatable trial protocols, operator procedures, and tracking of environmental variables that affect performance.
- Lead cross-functional root-cause investigations.
- Drive ambiguous failures to resolution across Research, Data Infrastructure, Model Infrastructure, Software, and Robot Operations.
- Turn incidents and regressions into durable tests, monitors, documentation, and interface contracts.
- Establish release and compatibility standards.
- Define the validation required before changes to robot software, data systems, training code, or inference systems become part of the supported research workflow.
- Measure and improve research reliability.
- Track pipeline success rates, reproducibility, regression frequency, time to root cause, benchmark stability, and other metrics that reflect the health of the research workflow.
Qualifications
- A strong track record owning, integrating, or debugging complex ML, robotics, autonomy, or other sensor-rich systems across multiple layers.
- Excellent software engineering skills, including strong Python, maintainable system design, automated testing, CI, code review, and production-quality debugging practices.
- Hands-on experience with PyTorch or a comparable ML framework.
- You should be able to train or fine-tune models, inspect datasets and batches, interpret losses and outputs, load checkpoints, and debug inference behavior.
- Experience building or operating ML workflows that span data processing, training, evaluation, and deployment - not only one isolated part of the stack.
- Strong systems-debugging ability.
- You can take an ambiguous report such as "robot performance dropped" and turn it into a structured investigation with controlled experiments, measurements, and a defensible root cause.
- Familiarity with data-pipeline correctness, including schemas, versioning, temporal data, migrations, provenance, validation, and reproducibility.
- Experience with Linux, containers, and modern compute environments such as Kubernetes, Slurm, or distributed GPU clusters.
- Strong attention to detail and a willingness to investigate silent failures that do not produce obvious exceptions or crashes.
- Ability to work effectively across research, infrastructure, software, and operations teams and to drive decisions without relying on formal managerial authority.
- Clear written communication, including the ability to document interfaces, test plans, root-cause analyses, release criteria, and operational procedures.
- Comfort working directly with physical systems and dealing with the variability and ambiguity inherent in real-world robotics.
- A degree in Computer Science, Robotics, Electrical Engineering, or a related discipline - or equivalent practical experience.
- A PhD and publication record are not required.
- Nice To Have, But Not Required:
- Experience in robotics, autonomous driving, drones, industrial automation, warehouse automation, or another domain that combines learned models with physical systems.
- Familiarity with robot data collection, teleoperation, camera and sensor systems, calibration, proprioception, force/torque sensing, or action synchronization.
- Experience with ROS or ROS2, real-time robot systems, simulation, replay infrastructure, or hardware-in-the-loop testing.
- Experience with imitation learning, behavior cloning, robot post-training, reinforcement learning, video models, multimodal models, or learned control policies.
- Experience building golden datasets, model-quality regression suites, data contracts, or automated dataset validation.
- Familiarity with columnar datasets and data infrastructure, including formats such as Parquet and systems for indexing, querying, or migrating large datasets.
- Experience debugging distributed training, GPU environments, checkpointing, networking, storage, Kubernetes, Slurm, or InfiniBand-related failures.
- Experience designing statistically sound evaluation protocols for systems with noisy or stochastic performance.
- Prior experience in ML platform engineering, research infrastructure, autonomy validation, release engineering, SRE, or systems integration.
Compensation Range
$200K - $300K