BEYOND MOTION EXECUTION Anonymous supplement

Beyond Motion Execution: Benchmarking Embodied Agents on Wet-Lab Operations and Procedures

From liquid transfer to multi-step experiments, explore robotic manipulation across twenty simulated laboratory tasks.

20Laboratory tasks
14Video demonstrations
06Zero-shot tasks

How to read the collection

Instructions, abilities
& scoring.

Each task pairs a visual example with its original language instruction and evaluation rubric.

Language & task splits

Training tasks (01–14) use the shortest instruction in each task’s training metadata. Zero-shot tasks (15–20) use the shortest instruction in the corresponding zero-shot evaluation metadata; these tasks have no training split.

Capability labels
Horizon
Number of atomic skill steps: Short, 1–5; Medium, 6–10; Long, more than 10.
Precision
Required positional accuracy: Low, >2 cm; Medium, >5 mm and ≤2 cm; High, ≤5 mm.
Reaction
Whether the task requires a timely response to a changing state.
Reading
Whether the task requires reading text or numerical values.
Metrics, scores & skills

Expand “Metric, Score & Skills” on any task to inspect its rubric. Scores are criterion weights and sum to 1.00 per task; they are not model success rates. S1, S2, and subsequent labels identify evaluation stages. “×N” indicates N independently scored instances. A dash means no skill was specified.

Zero-shot task