RoboDojo: New Benchmark Reveals Gaps in Robot Manipulation Between Simulation and Reality

July 20, 2026
RoboDojo: New Benchmark Reveals Gaps in Robot Manipulation Between Simulation and Reality
  • Early findings show that performance gains in simulation do not reliably translate to real-world success, underscoring the need for robust generalization, memory use, long-horizon planning, and reliable open-ended instruction handling in policies.

  • RoboDojo is a unified benchmark developed by a consortium of top universities to evaluate general robot manipulation policies across simulation and real-world tasks, addressing limitations of prior short-horizon or domain-limited benchmarks.

  • The benchmark tests five core manipulation capabilities—generalization, memory, long-horizon execution, fine-grained control, and open instruction understanding—across 42 tasks to diagnose policy weaknesses.

  • Real-world evaluation highlights safety, stability, and failure recovery, showing that partial progress can still leave complete task success precarious due to jitter, unstable contact, and execution latency.

  • RoboDojo demonstrates strong evaluation efficiency and reproducibility, with high-throughput parallel simulation and standardized real-world testing to rapidly diagnose weaknesses and inform future directions.

  • Open results reveal substantial limitations across capability dimensions: generalization and long-horizon execution are weak, fine-grained control and memory are difficult, and open semantic manipulation remains a major hurdle.

  • In the real world, the leading policy achieves only about a 12.8% success rate across 18 tasks, far from human teleoperation, underscoring reliability gaps not evident in simulation.

  • In simulation, top policies reach only modest average success rates (best around 8.8%), with humans at roughly 76%, signaling a large cross-domain transfer gap.

  • RoboDojo comprises a Simulation Platform, a Real-World Evaluation Platform (RoboDojo-RealEval), and XPolicyLab to standardize task construction, training, deployment, and evaluation through unified interfaces and data formats.

  • The initiative introduces dual leaderboards for simulation and real-world performance and standardizes resets, deployment, and evaluation protocols to improve reproducibility and enable meaningful cross-domain comparisons.

Summary based on 1 source


Get a daily email with more AI stories

More Stories