Definition
Plain language
A simulated science-lab environment for testing whether AI agents can carry out experimental procedures.
As stated in the literature
A text-based interactive environment for evaluating LLM agents on multi-step science tasks like growing plants or measuring temperatures, built on a PDDL-style simulator.
Also called: SciWorld
Why it matters: Multi-step procedural tasks expose failure modes that single-question benchmarks miss, like agents that get partway through a procedure and then forget what they were doing.
For example, the agent might be told to grow a plant, and has to find seeds, fill a pot with soil, add water, and place it near a light source — each as a discrete action in the simulator.
Heard on the show
“ALFWorld, ScienceWorld, TextCraft — these are PDDL-style simulators where the system internally knows exactly what objects exist, what rooms exist, what actions are valid.”Episode 052 — An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents