BEYOND MOTION EXECUTION Anonymous supplement

Beyond Motion Execution: Benchmarking Embodied Agents on Wet-Lab Operations and Procedures

From liquid transfer to multi-step experiments, explore policy execution and dataset demonstrations across twenty simulated laboratory tasks.

20Laboratory tasks
680Evaluation episodes
14Dataset videos

Benchmark organization

Table 1. Overview of the 20-Task Benchmark

Tasks combine laboratory operations from four families. Scope distinguishes a basic operation from an experimental procedure.

HP Handling & PlacementTM Transfer & MixingIO Instrument OperationMD Measurement & Dosing
Paper Table 1: task names, operation families, scope, four capabilities and distinctive skills.
IDFamilyTaskScopeHorizonPrecisionQRDistinctive skills

Scope: Op., basic operation; Proc., experimental procedure. Q: Quantitative-target requirement. R: Reaction-observation requirement. Distinctive skills are representative atomic skills. Task numbers 01–20 follow the paper’s Table 1 order.

Policy execution

π0.5 Policy Evaluation

Success and failure rollouts from the 100,000-step checkpoint. Select a task to compare its overlook videos.

Explore dataset demos

Results by task

All evaluation episodes
Success rate and mean score for all 20 tasks. Select a task name to view its policy rollouts.
TaskSuccess
rate
Mean
score

Success rate is the fraction of successful episodes. Mean score averages all episodes on a 0–1 scale.

Statistics use every evaluated episode, including failures. Tasks 01–14 each have 40 episodes across C-only, B-only, I-only, and Mixed (C+B+I) settings, with 10 episodes per setting. C/B/I denote task-configuration, background, and instruction variations. Zero-shot tasks 15–20 each have 20 Mixed (C+B+I) episodes.

How the examples were selected

One success and one failure are shown when both exist. The successful rollout is closest to the median successful episode length. The failure uses the same setting where possible and is closest to that setting’s median failure score. Tasks with no successful episode show a failure only. Videos retain their original resolution, frame rate, and full duration.

How to read the collection

Instructions, abilities
& scoring.

Each task pairs a visual example with its original language instruction and evaluation rubric.

Policy success rates & scores

The policy section reports π0.5 at 100,000 training steps. Success follows the evaluator’s recorded flag: the episode score equals 1 within a tolerance of 0.000001. A task’s success rate is successful episodes divided by all evaluated episodes; its mean score includes both successes and failures. The score under a video belongs to that individual rollout.

Language & task splits

Task names, operation families, scope, and capability labels follow Table 1 of the paper. Demonstrated tasks (01–14) use the shortest instruction in each task’s training metadata. Zero-shot tasks (15–20) use the shortest instruction in the corresponding zero-shot evaluation metadata; these tasks have no training split.

Generalization settings

Table 3 defines C-only, B-only, I-only, and Mixed (C+B+I). Task-configuration (C) changes task-relevant physical or experimental parameters, such as object positions, liquid amount or color, and duration for time-dependent tasks. Background (B) changes tabletop appearance, room background, and illumination. Instruction (I) changes descriptive detail while preserving task semantics. Mixed varies C, B, and I jointly.

Capacity profile: four dimensions
Horizon
Number of atomic skill steps: Short, 1–5; Medium, 6–10; Long, more than 10.
Precision
Required positional accuracy: Low, >2 cm; Medium, >5 mm and ≤2 cm; High, ≤5 mm.
Quantitative-target requirement
Whether task success requires reaching a specified quantitative target, rather than merely obtaining a numerical readout.
Reaction-observation requirement
Whether task success requires observing a reaction-induced visual state, such as a titration endpoint, precipitate formation, or color change.
Metrics, scores & skills

Expand “Metric, Score & Skills” on any task to inspect its rubric. Scores are criterion weights and sum to 1.00 per task; they are not model success rates. S1, S2, and subsequent labels identify evaluation stages. “×N” indicates N independently scored instances. A dash means no skill was specified.

Zero-shot task