BEYOND MOTION EXECUTION Anonymous supplement

Beyond Motion Execution: Benchmarking Embodied Agents on Wet-Lab Operations and Procedures

From liquid transfer to multi-step experiments, explore policy execution and dataset demonstrations across twenty simulated laboratory tasks.

20Laboratory tasks
680Evaluation episodes
14Dataset videos

Policy execution

π0.5 Policy Evaluation

Success and failure rollouts from the 100,000-step checkpoint. Select a task to compare its overlook videos.

Explore dataset demos

Results by task

All evaluation episodes
Success rate and mean score for all 20 tasks. Select a task name to view its policy rollouts.
TaskSuccess
rate
Mean
score

Success rate is the fraction of successful episodes. Mean score averages all episodes on a 0–1 scale.

Statistics use every evaluated episode, including failures. Tasks 01–14 each have 40 episodes across instruction, background, foreground, and mixed settings. Zero-shot tasks 15–20 each have 20 mixed-setting episodes.

How the examples were selected

One success and one failure are shown when both exist. The successful rollout is closest to the median successful episode length. The failure uses the same setting where possible and is closest to that setting’s median failure score. Tasks with no successful episode show a failure only. Videos retain their original resolution, frame rate, and full duration.

How to read the collection

Instructions, abilities
& scoring.

Each task pairs a visual example with its original language instruction and evaluation rubric.

Policy success rates & scores

The policy section reports π0.5 at 100,000 training steps. Success follows the evaluator’s recorded flag: the episode score equals 1 within a tolerance of 0.000001. A task’s success rate is successful episodes divided by all evaluated episodes; its mean score includes both successes and failures. The score under a video belongs to that individual rollout.

Language & task splits

Training tasks (01–14) use the shortest instruction in each task’s training metadata. Zero-shot tasks (15–20) use the shortest instruction in the corresponding zero-shot evaluation metadata; these tasks have no training split.

Capability labels
Horizon
Number of atomic skill steps: Short, 1–5; Medium, 6–10; Long, more than 10.
Precision
Required positional accuracy: Low, >2 cm; Medium, >5 mm and ≤2 cm; High, ≤5 mm.
Reaction
Whether the task requires a timely response to a changing state.
Reading
Whether the task requires reading text or numerical values.
Metrics, scores & skills

Expand “Metric, Score & Skills” on any task to inspect its rubric. Scores are criterion weights and sum to 1.00 per task; they are not model success rates. S1, S2, and subsequent labels identify evaluation stages. “×N” indicates N independently scored instances. A dash means no skill was specified.

Zero-shot task