armnet eval — evaluate a policy on real robots

armnet.dev

Pick a task, pick or paste a policy, press Run, and watch 20 rollouts on a real SO-101 cell. Success is auto-scored by the instrumented BusyBox — no human judging. Anyone here can watch live evals; if more than one cell is running, pick which feed to follow.

Task

Push the green button. Push the green button on the BusyBox panel. The button is momentary and springs back, so the scene resets itself.

How to use this Space

  1. Train a LeRobot-compatible policy on either the task-specific armnet/busybox_push_green_button dataset or the language-conditioned armnet/busybox_multitask dataset. Make sure the trained policy is uploaded to Hugging Face and publicly available.
  2. Copy your repo ID into the text box on the Live tab and submit it. Armnet will automatically run and score the policy on a real SO-101 cell. If every cell is busy, the job waits in the shared queue. If more than one eval is running, pick which cell's live feed to watch.

You can run a policy more than once. The more rollouts a policy accumulates, the closer its lower confidence interval matches its average success rate — see the Leaderboard tab.

Every task has its own dataset and its own board, so pick the one you want at the top of the page before you start.