Reality Check Leaderboard
PAW-GEN-10
Share of episodes rated Ok or Excellent.
Share of the task's steps completed successfully.
Excellent-vs-Ok share of successful episodes.
Mean active duration of successful episodes. Lower is better.
Task steps completed per active minute, pooled over the successful episodes.
Spectral arc length (SPARC) of the arm's motion, a unitless measure ≤ -1: the closer to -1, the smoother. It is measured over each whole episode and averaged over episodes.
How abruptly the arm's acceleration changed: the end effector's jerk, filtered at 10 Hz and averaged over each episode, then over episodes. It grows with speed, so a faster motion reads as rougher. Lower is better.
Share of failed episodes that ended with both arms ok, from the arms' own status when the recording stopped.
The largest force either arm exerted on its surroundings during an episode, averaged over episodes. It is based on the FR3's joint torque sensors. Lower is better.
| Rank and model | Overall score and number of evaluations | Mid-training DataM0 has no mid-training data and post-trains the model directly on the environment demonstrations. M100 first mid-trains it on 100 hours of embodiment data (coming soon). | Data ScalingScores after training on ~10, ~100 or ~300 demonstrations per environment (D10, D100, D300). The sets are nested. | Spatial GeneralizationNominal uses the training placements for evaluation. Interpolation uses unseen placements between them. Extrapolation uses unseen placements outside their convex hull. | Per EnvironmentThe model's score on every environment of the benchmark, e01 onwards. Hover an environment's column for its name. | Model details | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| # | Model | # Evals | Open model | |||||||||||||||||||
| 1 | MolmoAct 2* | 28% | 3600 | 28% | Coming soon | 4% | 37% | 44% | 31% | 25% | Coming soon | 10% | 44% | 52% | 15% | 21% | 4% | 42% | 52% | 18% | 24% | |
| 2 | Pi 0.5* | 21% | 3600 | 21% | Coming soon | 5% | 26% | 33% | 23% | 19% | Coming soon | 4% | 39% | 66% | 2% | 9% | 3% | 13% | 59% | 13% | 2% | |
| 3 | DiT-Flow | 19% | 3600 | 19% | Coming soon | 2% | 24% | 29% | 21% | 16% | Coming soon | 2% | 17% | 48% | 2% | 0% | 1% | 32% | 63% | 16% | 4% | |
| 4 | GR00T N1.7* | 12% | 3600 | 12% | Coming soon | 4% | 14% | 19% | 13% | 12% | Coming soon | 3% | 14% | 25% | 1% | 1% | 1% | 14% | 50% | 13% | 0% | |
* All VLAs were started from their official checkpoints, but fine-tuned by us on the target tasks and data regimes.
Data Scaling
D10, D100, and D300 indicate how many demonstration episodes per environment were used to train the model. Whiskers are 95% confidence intervals over the evaluation episodes.
Spatial Generalization
Nominal uses training placements; interpolation uses unseen placements between them; extrapolation uses unseen placements outside their convex hull. Whiskers are 95% confidence intervals over the evaluation episodes.
Per Environment (SuccessProgressExecution qualityAvg execution timeExecution speedSmoothnessAvg jerkSafe failure rateMax contact force)
e01–e10 represent the environments in PAW-GEN-10, each value in the selected metric averaged across every data tier and placement condition.
