Beyond Motion Execution: Benchmarking Embodied Agents on Wet-Lab Operations and Procedures
From liquid transfer to multi-step experiments, explore robotic manipulation across twenty simulated laboratory tasks.
Task gallery
20 matching tasks
Training Tasks
Zero-shot Tasks
Six held-out tasks for zero-shot evaluation. Each image shows the task setup or a representative scene.
No matching tasks
Try a different task name, skill, or capability.
How to read the collection
Instructions, abilities
& scoring.
Each task pairs a visual example with its original language instruction and evaluation rubric.
Language & task splits
Training tasks (01–14) use the shortest instruction in each task’s training metadata. Zero-shot tasks (15–20) use the shortest instruction in the corresponding zero-shot evaluation metadata; these tasks have no training split.
Capability labels
- Horizon
- Number of atomic skill steps: Short, 1–5; Medium, 6–10; Long, more than 10.
- Precision
- Required positional accuracy: Low, >2 cm; Medium, >5 mm and ≤2 cm; High, ≤5 mm.
- Reaction
- Whether the task requires a timely response to a changing state.
- Reading
- Whether the task requires reading text or numerical values.
Metrics, scores & skills
Expand “Metric, Score & Skills” on any task to inspect its rubric. Scores are criterion weights and sum to 1.00 per task; they are not model success rates. S1, S2, and subsequent labels identify evaluation stages. “×N” indicates N independently scored instances. A dash means no skill was specified.