Beyond Motion Execution: Benchmarking Embodied Agents on Wet-Lab Operations and Procedures
From liquid transfer to multi-step experiments, explore policy execution and dataset demonstrations across twenty simulated laboratory tasks.
Benchmark organization
Table 1. Overview of the 20-Task Benchmark
Tasks combine laboratory operations from four families. Scope distinguishes a basic operation from an experimental procedure.
| ID | Family | Task | Scope | Horizon | Precision | Q | R | Distinctive skills |
|---|
Scope: Op., basic operation; Proc., experimental procedure. Q: Quantitative-target requirement. R: Reaction-observation requirement. Distinctive skills are representative atomic skills. Task numbers 01–20 follow the paper’s Table 1 order.
Policy execution
π0.5 Policy Evaluation
Success and failure rollouts from the 100,000-step checkpoint. Select a task to compare its overlook videos.
Results by task
All evaluation episodes| Task | Success rate | Mean score |
|---|
Success rate is the fraction of successful episodes. Mean score averages all episodes on a 0–1 scale.
Statistics use every evaluated episode, including failures. Tasks 01–14 each have 40 episodes across C-only, B-only, I-only, and Mixed (C+B+I) settings, with 10 episodes per setting. C/B/I denote task-configuration, background, and instruction variations. Zero-shot tasks 15–20 each have 20 Mixed (C+B+I) episodes.
How the examples were selected
One success and one failure are shown when both exist. The successful rollout is closest to the median successful episode length. The failure uses the same setting where possible and is closest to that setting’s median failure score. Tasks with no successful episode show a failure only. Videos retain their original resolution, frame rate, and full duration.
Dataset Demonstrations
Training demonstrations and Zero-shot task scenes.
20 matching tasks
Demonstrated Tasks
Zero-shot Tasks
Six held-out tasks for zero-shot evaluation. Each image shows the task setup or a representative scene.
No matching tasks
Try a different task name, skill, or capability.
How to read the collection
Instructions, abilities
& scoring.
Each task pairs a visual example with its original language instruction and evaluation rubric.
Policy success rates & scores
The policy section reports π0.5 at 100,000 training steps. Success follows the evaluator’s recorded flag: the episode score equals 1 within a tolerance of 0.000001. A task’s success rate is successful episodes divided by all evaluated episodes; its mean score includes both successes and failures. The score under a video belongs to that individual rollout.
Language & task splits
Task names, operation families, scope, and capability labels follow Table 1 of the paper. Demonstrated tasks (01–14) use the shortest instruction in each task’s training metadata. Zero-shot tasks (15–20) use the shortest instruction in the corresponding zero-shot evaluation metadata; these tasks have no training split.
Generalization settings
Table 3 defines C-only, B-only, I-only, and Mixed (C+B+I). Task-configuration (C) changes task-relevant physical or experimental parameters, such as object positions, liquid amount or color, and duration for time-dependent tasks. Background (B) changes tabletop appearance, room background, and illumination. Instruction (I) changes descriptive detail while preserving task semantics. Mixed varies C, B, and I jointly.
Capacity profile: four dimensions
- Horizon
- Number of atomic skill steps: Short, 1–5; Medium, 6–10; Long, more than 10.
- Precision
- Required positional accuracy: Low, >2 cm; Medium, >5 mm and ≤2 cm; High, ≤5 mm.
- Quantitative-target requirement
- Whether task success requires reaching a specified quantitative target, rather than merely obtaining a numerical readout.
- Reaction-observation requirement
- Whether task success requires observing a reaction-induced visual state, such as a titration endpoint, precipitate formation, or color change.
Metrics, scores & skills
Expand “Metric, Score & Skills” on any task to inspect its rubric. Scores are criterion weights and sum to 1.00 per task; they are not model success rates. S1, S2, and subsequent labels identify evaluation stages. “×N” indicates N independently scored instances. A dash means no skill was specified.