RoboDojo: New Benchmark Reveals Gaps in Robot Manipulation Between Simulation and Reality
July 20, 2026
Early findings show that performance gains in simulation do not reliably translate to real-world success, underscoring the need for robust generalization, memory use, long-horizon planning, and reliable open-ended instruction handling in policies.
RoboDojo is a unified benchmark developed by a consortium of top universities to evaluate general robot manipulation policies across simulation and real-world tasks, addressing limitations of prior short-horizon or domain-limited benchmarks.
The benchmark tests five core manipulation capabilities—generalization, memory, long-horizon execution, fine-grained control, and open instruction understanding—across 42 tasks to diagnose policy weaknesses.
Real-world evaluation highlights safety, stability, and failure recovery, showing that partial progress can still leave complete task success precarious due to jitter, unstable contact, and execution latency.
RoboDojo demonstrates strong evaluation efficiency and reproducibility, with high-throughput parallel simulation and standardized real-world testing to rapidly diagnose weaknesses and inform future directions.
Open results reveal substantial limitations across capability dimensions: generalization and long-horizon execution are weak, fine-grained control and memory are difficult, and open semantic manipulation remains a major hurdle.
In the real world, the leading policy achieves only about a 12.8% success rate across 18 tasks, far from human teleoperation, underscoring reliability gaps not evident in simulation.
In simulation, top policies reach only modest average success rates (best around 8.8%), with humans at roughly 76%, signaling a large cross-domain transfer gap.
RoboDojo comprises a Simulation Platform, a Real-World Evaluation Platform (RoboDojo-RealEval), and XPolicyLab to standardize task construction, training, deployment, and evaluation through unified interfaces and data formats.
The initiative introduces dual leaderboards for simulation and real-world performance and standardizes resets, deployment, and evaluation protocols to improve reproducibility and enable meaningful cross-domain comparisons.
Summary based on 1 source
