Research Member of Technical Staff- Robot Learning Systems & Reliability
Rhoda AI
At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality. About the Role We're looking for a senior or staff-level Research Engineer or ML Systems Engineer to make our robot- learning pipeline reliable, reproducible, and measurable from end to end. You will own the supported path from robot data collection through dataset generation, model training, inference, and real-robot evaluation. You will build the validation systems, regression tests, observability, and operating practices that allow researchers to trust their experiments and quickly distinguish model limitations from data, software, infrastructure, hardware, or evaluation failures. This is a core research systems role - not generic DevOps, infrastructure support, or traditional QA. You will need to understand the semantics of robot data and model behavior, work directly with training and inference code, investigate failures on physical robots, and drive fixes across organizational boundaries. We are hiring across senior and staff levels. Because this will be the first dedicated owner of the end-to-end workflow, we are primarily looking for someone with staff-level ownership and systems judgment. What You'll Do • Own the robot-learning pipeline end to end. Establish and maintain a trusted workflow spanning robot data collection, data ingestion and compilation, post-training, checkpoint generation, inference, and real-robot evaluation. • Define the supported golden path. Maintain known-good combinations of code, datasets, configurations, checkpoints, robot software, hardware settings, task stations, and evaluation procedures. • Build automated validation at every interface. Develop checks for timestamp synchronization, sensor and action integrity, episode completeness, schema compatibility, dataset migrations, dataloader outputs, preprocessing behavior, model inputs, and configuration correctness. • Create end-to-end regression tests. Build representative smoke tests that exercise data compilation, training, checkpoint loading, inference, replay or simulation, and real-robot execution. Develop small-scale overfit and canary experiments that catch correctness regressions before expensive training runs begin. • Ensure training and inference consistency. Identify and prevent discrepancies in image processing, sensor normalization, temporal context, action representation, model configuration, and other transformations used across training and deployment. • Make failures observable and diagnosable. Build instrumentation and debugging tools that help determine whether a performance regression originated in data, model code, infrastructure, inference, robot software, hardware configuration, the physical environment, or evaluation execution. • Improve real-robot evaluation reliability. Partner with researchers and robot operations to establish stable benchmark stations, reference baselines, clear rubrics, repeatable trial protocols, operator procedures, and tracking of environmental variables that affect performance. • Lead cross-functional root-cause investigations. Drive ambiguous failures to resolution across Research, Data Infrastructure, Model Infrastructure, Software, and Robot Operations. Turn incidents and regressions into durable tests, monitors, documentation, and interface contracts. • Establish release and compatibility standards. Define the validation required before changes to robot software, d