Open submission

Apply for evaluation

Tell us your team and how to reach you — we will look you up on your preferred channel.

Why it scales

Simple and unified interface

Every one of the fifteen variants already on the board went through the same three contracts your submission will. Follow them and your policy, task, or robot slots in cleanly.

One dataset schema

Your rollouts and any paired demonstrations you contribute use the same manifest and episode fields as IG-10K — L0-L3 level, task ID, domain — so they compose with the existing corpus instead of living beside it.

One model interface

Every policy on the board, whatever its paradigm, implements the same inference API: a human video in, an action chunk out. That is what makes a like-for-like leaderboard possible at all.

Template-driven tasks & robots

Proposing a new task or embodiment, not just a policy? It follows the same _template/ contracts and stays in sync with upstream ManiSkill — the benchmark grows outward without forking the simulator.

What counts

Rules of the board — evaluation integrity & anti-gaming

The held-out tasks stay held out

No unseen-task demonstrations in pretraining. Few-shot regimes use the same demonstration budget per unseen task at every level.

One human video in, one rollout out

The policy sees the demonstration video for the episode it is executing. No privileged state, no per-episode hand tuning, no unreported retries.

Declare the corpus and the regime

Entries are listed with their pretraining corpus size and training regime. A 45-task entry is not compared against a 15-task entry without saying so.

Rollouts are public

Approved videos go into the Gallery and the Arena. Weights stay yours; the behaviour goes on the record so anyone can see what the number means.

Application

Apply for evaluation

Tell us your team and how to reach you — we will look you up on your preferred channel. Nothing is sent from this page: submitting opens a pre-filled tracking issue, or you can copy the summary and send it however you prefer.

Simulation runs are GPU-hosted. Real-robot time is scheduled only after the package reproduces in simulation.
Choose one. This is how we confirm your identity on the channel below.
We will find you with the contact information above. When requesting to join a group, please note your team name.
This is about the checkpoint itself — what data it was trained on — and is independent of "Evaluation type" above. A sim-trained checkpoint can still be evaluated on real hardware (and vice versa); tell us which so we can place your entry correctly on the Leaderboard's Sim / Real-world boards.
Please be specific enough that a maintainer can tell, without asking, whether your reported numbers are comparable to the sim-trained or real-trained baselines already on the board.
Optional at this stage — the package itself arrives with your pull request.

After applying, open a pull request in the repository to integrate your policy. Questions before you start? Ask on Discord or WeChat.