Who imitates best?
Nine models, fifteen variants, scored in 40 evaluation tasks in two ways: an automated metric in simulation, and blind human comparisons in the Arena for both simulation and real world.
This board — and every other model comparison on the site — is scored on a fixed protocol: 5 seen and 5 unseen tasks, each at all four imitation levels, in both simulation and the real world. IG-10K's full task suite is larger; see the 50-task reference below for how far that goes.
Full task-suite reference (50 tasks)
Every comparison on this site — the board above, the Arena, and the paper's headline numbers — is scored on the same fixed protocol: 5 seen and 5 unseen tasks, each evaluated at all four imitation levels (L0-L3), in both simulation and the real world. IG-10K itself, however, ships with 50 manipulation tasks, and we encourage the community to go further: train and evaluate on the complete task suite for a fuller picture of imitation ability, not just the ten tasks used for ranking.
As a starting point, we trained DP and ACT (DINOv2 encoder) on the entire IG-10K simulation corpus — all 50 tasks, 50 demonstrations each, 200 task-level variants — and evaluated both on every task at every level. As with the main board, object placement is perturbed by a small random amount in both the training and evaluation scenes, so success still requires robust imitation rather than fixed-layout replay.
| Task | DP | ACT | Mean | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| L0 | L1 | L2 | L3 | L0 | L1 | L2 | L3 | DP | ACT | |
Put your model on this board
One dataset format, one model interface, templated tasks and robots — any policy that takes in human video as prompt can be easily adapted to the unified training / evaluation protocol. Join us and submit your model to win the Imitator Game!