The Imitator Game
Benchmarking Robot Imitative Ability Beyond Action Prediction
Can a robot imitate what a human intends — not just what they do? We widen the gap between the demonstration and the robot's own scene until trajectory replay stops working, and measure what is left.
episodes
environments
on the board
Three minutes on the whole thing
assets/media/demo_video.mp4
Humans imitate at the level of intent: given a demonstration, we infer its goal and carry it out with whatever tools, objects, and layouts are at hand. Current robot policies instead learn observation-to-action mappings from visual inputs and language instructions, without explicitly inferring the demonstrated task. Learning from human video thus remains largely trajectory-level: models can replay motions in near-identical scenes, but still struggle to imitate what the demonstrator intends rather than merely what they do. We introduce The Imitator Game, a four-level benchmark (L0-L3) that progressively widens the gap between the human demonstration and the robot's own scene, isolating where trajectory replay ceases to suffice and task understanding becomes necessary. We pair it with IG-10K, the largest environment-aligned paired human-robot dataset to date and the only one instantiated across all four levels in both real and simulated settings (20,000+ paired episodes, 50+ tasks, 6 domains), and Imitator Arena, an open platform for blind A/B human evaluation. Across nine state-of-the-art models, performance is stable from L0 to L2 but collapses at L3, identifying functional substitution — achieving the same intent through a different object affordance — as the decisive barrier to intent-level imitation. Human-video-conditioned models outperform caption-conditioned ones, yet every model falls below 13% zero-shot success on unseen tasks; fine-tuning IG-10K-pretrained models with only 10 paired human-robot demonstrations yields large gains that grow with pretraining scale.
One task, four levels of imitation
Every level asks for the same outcome. What changes is how much of the demonstrated trajectory still applies. Pick a level below to see its reference footage and its real-world pretrain + fine-tune success rate, averaged over the four representative models.
Paired human-robot data, simulation and real world
The largest environment-aligned paired human-robot dataset we know of. Simulation and real world datasets fit in one unified format, policies will be trained and evaluated under the same interface. Each episode is multi-view and richly annotated with 3D MANO hand pose, segmentation masks, and multi-abstraction language.
Experimental findings
We evaluate 9 state-of-the-art models (15 trained variants) across simulation, real-world deployment, and human Arena judgments — along three axes. Ten manipulation tasks across L0-L3, total 40 task-level pairs was evaluated.
These ten tasks — 5 seen, 5 unseen — are the fixed protocol behind every model comparison on this site, in both simulation and the real world. IG-10K itself has 50 tasks, and we encourage the community to evaluate on the complete suite for a fuller picture of imitation ability. As a starting reference, we trained DP and ACT on the entire IG-10K simulation corpus (all 50 tasks, 200 task-level variants) and evaluated them on every task; object placement is randomized slightly in both training and evaluation scenes. See the full 50-task reference on the leaderboard.
Seen — how faithfully, not just whether
Long-horizon or dexterous behaviour the models were trained on: stirring with a thin spoon, folding a towel, hanging a mug on a rack, picking up a sheet of paper, and racking plates. Object placements are still perturbed at test time, so the question is fidelity — does the imitation hold up as the scene drifts from L0 to L3, not just at the layout it memorised.
Hover to see the five tasksThe five seen tasks
Unseen — can a new skill come from the video alone
Closer to atomic skills — single-arm and bimanual placement, pouring, and an articulated folding task — but their demonstrations, objects, and layouts never appear in pretraining. Evaluated three ways on the same task: zero-shot, from ten scratch demonstrations, and IG-10K-pretrained then fine-tuned on the same ten.
Hover to see the five tasksThe five unseen tasks
Which imitation interface is strongest?
The paradigm landscape. Each trained variant on the(seen-SR, P+FT-SR) plane, automated channel; the video families cluster upper-right while the VLA family stretches along the diagonal.
Does the visual encoder matter?
Encoder ablation. Within the video-conditioned family, DINOv2 and SigLIP2 are consistently ahead of VideoMAE, in simulation and on hardware alike. The representation is a design choice that shows up in the score.
Does a bigger corpus help?
Corpus scaling. Simulation P+FT going from 15 to 45 pretraining tasks improves few-shot success almost everywhere. Real-world zero-shot is a different story: the three video-conditioned models improve monotonically while the caption-conditioned π₀.₅ stays at the floor.
Does scale help at every level?
Level × corpus size. Under pretrain + fine-tune, success rises with corpus size at every one of L0-L3, both simulation and real world. Simulation zero-shot is around the floor at all four levels.
Where does the hierarchy become hard?
The level profile. Real-world pretrain + fine-tune SR by level, with the human imitation score on the right axis. With levels rise, the success rate drop, especially on L3, the intent-level transfer, which requires functional transfer.
Can you trust the automated metric?
Metric validity. All 15 trained variants × 4 transfer regimes — 60 points — plotted automated SR against the Arena's human-judged SR on the same rollouts (r = 0.858 for SR and r = 0.861 for Q). Resolved by level on the right, the sub-goal rate tracks the human imitation score just as closely.
The leaderboard, right now
Top five by seen-task success in simulation. The full table also carries the Arena results. New models land on a rolling basis.
These standings, like the experimental findings above, are scored on the fixed 5-seen / 5-unseen, L0-L3 protocol. For a full-coverage reference across all 50 IG-10K tasks, see the full task-suite table on the leaderboard page.
Four things you can use today
The benchmark keeps running: new models enter, the Arena keeps collecting judgments, and the board moves.
Leaderboard
Every variant evaluated in sim and real, will be judged with automated metrics and human. Sortable, and honest about how models imitate.
See the standingsArena
A reference clip and two anonymous rollouts will be presented. Just follow your heart, score each, and pick the real winner.
Judge a pairGallery
50+ task groups at all four levels, with the human demonstration, the real-robot rollout and the simulated one lined up side by side.
Browse the tasksSubmit
We provide a simple and unified interface for creating new simulation environments and training policies. Submit your model and win the game.
Enter your policyJoin Imitator Game Community !
The Imitator Game is an open-source community for robot imitation from human video. Tasks, models, embodiments, assets, demonstrations and evaluation ideas are all welcome!
BibTeX
@misc{zhou2026imitatorgamebenchmarkingrobot,
title={The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction},
author={Xunzhe Zhou and Yiyang Cai and Fengyi Wang and Ran Ju and Hanxiang Ren and Ruizhe Liu and Yu Zhang and Qian Luo and Feng Chen and Pei Zhou and Yi Ma and Yanchao Yang},
year={2026},
eprint={2608.22301},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2608.22301},
}