Intent imitation, L0-L3 20,000+ paired episodes Human evaluation Arena

The Imitator Game

Benchmarking Robot Imitative Ability Beyond Action Prediction

Xunzhe Zhou1,2,*,† · Yiyang Cai2,3,* · Fengyi Wang2,3,* · Ran Ju1,2,* · Hanxiang Ren2,4 · Ruizhe Liu1
Yu Zhang1 · Qian Luo1,2 · Feng Chen1 · Pei Zhou1,2 · Yi Ma1,2 · Yanchao Yang1,2,‡

1The University of Hong Kong · 2TranscEngram · 3Fudan University · 4Zhejiang University

* Equal contribution · † Project lead · ‡ Corresponding author
Accepted by Conference on Robot Learning (CoRL) 2026
Accepted by NeurIPS 2026 Workshop on On-Device Intelligence

Can a robot imitate what a human intends — not just what they do? We widen the gap between the demonstration and the robot's own scene until trajectory replay stops working, and measure what is left.

0 intent-based hierarchical imitation levels
0 paired human-robot
episodes
0 scalable simulation
environments
0 evaluation variants
on the board
Overview

Three minutes on the whole thing

Overview film
The 2-3 minute explainer: setup, the four levels, a real rollout, the Arena.
assets/media/demo_video.mp4
Abstract

Humans imitate at the level of intent: given a demonstration, we infer its goal and carry it out with whatever tools, objects, and layouts are at hand. Current robot policies instead learn observation-to-action mappings from visual inputs and language instructions, without explicitly inferring the demonstrated task. Learning from human video thus remains largely trajectory-level: models can replay motions in near-identical scenes, but still struggle to imitate what the demonstrator intends rather than merely what they do. We introduce The Imitator Game, a four-level benchmark (L0-L3) that progressively widens the gap between the human demonstration and the robot's own scene, isolating where trajectory replay ceases to suffice and task understanding becomes necessary. We pair it with IG-10K, the largest environment-aligned paired human-robot dataset to date and the only one instantiated across all four levels in both real and simulated settings (20,000+ paired episodes, 50+ tasks, 6 domains), and Imitator Arena, an open platform for blind A/B human evaluation. Across nine state-of-the-art models, performance is stable from L0 to L2 but collapses at L3, identifying functional substitution — achieving the same intent through a different object affordance — as the decisive barrier to intent-level imitation. Human-video-conditioned models outperform caption-conditioned ones, yet every model falls below 13% zero-shot success on unseen tasks; fine-tuning IG-10K-pretrained models with only 10 paired human-robot demonstrations yields large gains that grow with pretraining scale.

The hierarchy

One task, four levels of imitation

Every level asks for the same outcome. What changes is how much of the demonstrated trajectory still applies. Pick a level below to see its reference footage and its real-world pretrain + fine-tune success rate, averaged over the four representative models.

Hover to see the rollout sample
Human demonstration (shared reference)
Human demonstration
assets/media/levels/human_ref.mp4
Rollout sample · L0
Rollout sample
assets/media/levels/l0_sample.mp4
L0

Scene-identical execution

The imitator scene matches the demonstration. Replaying the trajectory is enough.

0.00 real-world P+FT success at L0, averaged over 4 representative models
IG-10K

Paired human-robot data, simulation and real world

The largest environment-aligned paired human-robot dataset we know of. Simulation and real world datasets fit in one unified format, policies will be trained and evaluated under the same interface. Each episode is multi-view and richly annotated with 3D MANO hand pose, segmentation masks, and multi-abstraction language.

0 paired episodes
0 task suite
0 simulation environments
0 imitation levels
Evaluation

Experimental findings

We evaluate 9 state-of-the-art models (15 trained variants) across simulation, real-world deployment, and human Arena judgments — along three axes. Ten manipulation tasks across L0-L3, total 40 task-level pairs was evaluated.

These ten tasks — 5 seen, 5 unseen — are the fixed protocol behind every model comparison on this site, in both simulation and the real world. IG-10K itself has 50 tasks, and we encourage the community to evaluate on the complete suite for a fuller picture of imitation ability. As a starting reference, we trained DP and ACT on the entire IG-10K simulation corpus (all 50 tasks, 200 task-level variants) and evaluated them on every task; object placement is randomized slightly in both training and evaluation scenes. See the full 50-task reference on the leaderboard.

5 tasks · trained on

Seen — how faithfully, not just whether

Long-horizon or dexterous behaviour the models were trained on: stirring with a thin spoon, folding a towel, hanging a mug on a rack, picking up a sheet of paper, and racking plates. Object placements are still perturbed at test time, so the question is fidelity — does the imitation hold up as the scene drifts from L0 to L3, not just at the layout it memorised.

Hover to see the five tasks
5 tasks · trained on

The five seen tasks

5 tasks · never trained on

Unseen — can a new skill come from the video alone

Closer to atomic skills — single-arm and bimanual placement, pouring, and an articulated folding task — but their demonstrations, objects, and layouts never appear in pretraining. Evaluated three ways on the same task: zero-shot, from ten scratch demonstrations, and IG-10K-pretrained then fine-tuned on the same ten.

Hover to see the five tasks
5 tasks · never trained on

The five unseen tasks

All ten tasks, all four levels, continuously — hover a clip to pause
Q1 · Fig. C1

Which imitation interface is strongest?

The paradigm landscape. Each trained variant on the(seen-SR, P+FT-SR) plane, automated channel; the video families cluster upper-right while the VLA family stretches along the diagonal.

Q1 · Fig. C2

Does the visual encoder matter?

Encoder ablation. Within the video-conditioned family, DINOv2 and SigLIP2 are consistently ahead of VideoMAE, in simulation and on hardware alike. The representation is a design choice that shows up in the score.

Q2 · Fig. C3

Does a bigger corpus help?

Corpus scaling. Simulation P+FT going from 15 to 45 pretraining tasks improves few-shot success almost everywhere. Real-world zero-shot is a different story: the three video-conditioned models improve monotonically while the caption-conditioned π₀.₅ stays at the floor.

Q2 · Fig. C4

Does scale help at every level?

Level × corpus size. Under pretrain + fine-tune, success rises with corpus size at every one of L0-L3, both simulation and real world. Simulation zero-shot is around the floor at all four levels.

Q3 · Fig. C5

Where does the hierarchy become hard?

The level profile. Real-world pretrain + fine-tune SR by level, with the human imitation score on the right axis. With levels rise, the success rate drop, especially on L3, the intent-level transfer, which requires functional transfer.

Fig. C7

Can you trust the automated metric?

Metric validity. All 15 trained variants × 4 transfer regimes — 60 points — plotted automated SR against the Arena's human-judged SR on the same rollouts (r = 0.858 for SR and r = 0.861 for Q). Resolved by level on the right, the sub-goal rate tracks the human imitation score just as closely.

Standings

The leaderboard, right now

Top five by seen-task success in simulation. The full table also carries the Arena results. New models land on a rolling basis.

These standings, like the experimental findings above, are scored on the fixed 5-seen / 5-unseen, L0-L3 protocol. For a full-coverage reference across all 50 IG-10K tasks, see the full task-suite table on the leaderboard page.

Join Imitator Game Community !

The Imitator Game is an open-source community for robot imitation from human video. Tasks, models, embodiments, assets, demonstrations and evaluation ideas are all welcome!

Citation

BibTeX

@misc{zhou2026imitatorgamebenchmarkingrobot,
      title={The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction}, 
      author={Xunzhe Zhou and Yiyang Cai and Fengyi Wang and Ran Ju and Hanxiang Ren and Ruizhe Liu and Yu Zhang and Qian Luo and Feng Chen and Pei Zhou and Yi Ma and Yanchao Yang},
      year={2026},
      eprint={2608.22301},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2608.22301}, 
}