Recursive-Play: AI-Generated Interactive Challenges for Improving AI Agents
Summary
AI-assisted task construction can supply executable challenges that evaluated AI agents do not yet complete, creating a concrete resource for subsequent improvement. Recursive-Play implements this idea with 150 interactive puzzle environments, ten mechanic families and six sequential levels per task. Author-only constructive witnesses and deterministic state verifiers establish reachability, while players receive pixels and neutral controls without the rules or solutions. Across eleven configurations and 1,650 independently eligible outcomes, 15 tasks remain unfinished by every configuration under the fixed 512-action / 720-second protocol, and two yield no completed level for any configuration.
Details
- Independent replay verifies all 900 canonical seed–level witnesses within the action budget.
- Each fresh episode uses the original model identifier, canonical seed and blind prompt, with 512 charged actions and 720 seconds across all six levels. Partial progress is credited and reported separately from complete six-level wins.
- The strongest observed configuration completes 135 of 150 tasks; the remaining fifteen have no full win from any of the eleven configurations under the recorded budgets and setups.
- The paper specifies a subsequent improvement loop: select verified failures, acquire corrected trajectories, train a candidate, and test it on newly frozen held-out tasks with no-regression checks.
Scope and limits
The study establishes the challenge supply and evaluation substrate; training gains remain to be tested. The release does not demonstrate trained improvement, persistent adaptation or recursive self-improvement, and the exposed task bank is not a sealed held-out test set.
Citation
@misc{envloop_recursive_play,
title = {Recursive-Play: AI-Generated Interactive Challenges for Improving AI Agents},
author = {{EnvLoop Research}},
url = {https://huggingface.co/datasets/EnvLoop/Recursive-Play-Bench}
}