Learning from experience in unfamiliar environments

Learn2Play Bench

How effectively can LLM agents learn from experience?

20 game templates with hidden rules to discover. Different seeds generate different game instances. Agents learn each game’s hidden rules by playing repeatedly and use what they learn to improve in later episodes, without updating their model weights.

Research snapshot · September 2026 · Game portal requires sign-in

Game examples

See how each game is played.

Choose a game and scroll through its opening situation, sample actions, and feedback. These are recorded demonstrations—not a live game or a recommended strategy.

Gameplay preview

How to play

Loading recorded examples…

Play in the game portal
Gameplay transcriptScroll to read actions and feedback
Loading…

Evaluation protocol

How we evaluate agents

For each game template, we use one selected rule seed and run 3 independent trials. The leaderboard does not average over all possible seeds.

We report 5 episodes for normal games and 10 for challenge games. Scores are normalized using each game's reference ceiling, then averaged with weight 3 for challenge games and weight 1 for normal games.

Max

Average peak score across trials.

Mean

Average score across episodes.

LG Learning Gain

Improvement from early to late episodes.

LS Learning Slope

Fitted rate of score change across episodes.

Results

Leaderboard

Compare backbone models, self-evolving methods, and agent harnesses. Default ranking uses Max.

Main evaluation leaderboard
RankModelMax ↑Mean ↑LG ↑LS ↑

Source: manuscript backbone/method and harness comparison tables · Updated September 29, 2026. Sorted highest first; equal displayed values share a rank. Highlighting marks the best value in each column. These are aggregate point estimates, not significance rankings. This leaderboard is a curated snapshot, not automatically updated.

Learning through repeated interaction

How does performance change with experience?

A peak score does not show the path to that score. These curves track normalized performance over the evaluation window across all 20 games.

Learning curves for nine OpenCode backbones, showing normalized score against progress through the evaluation window.
The horizontal axis shows progress from the first to the last evaluated episode, with linear interpolation to align curves. Lines show aggregate mean scores, not uncertainty intervals.

Efficiency

Performance and inference cost

Compare Max scores with estimated mean inference cost per episode. Points toward the upper left achieve higher performance at lower estimated cost.

Open chart ↗
OpenCodeMethodsClaude CodeCodexFrontier · all configurations29 configurations
29 agent configurations by normalized Max score and estimated US dollars per episode. Higher and farther left is better. Claude Code with Opus 5 has the highest Max at 81.8%.

Hover, tap, or focus a point to see its full name and exact values.

Max is normalized to each game’s reference ceiling. The dashed frontier marks configurations with the highest Max at their cost or lower, across this snapshot. Costs are estimates using the August 30, 2026 pricing snapshot and cache assumptions; they are not current price quotes.

Benchmark design

Key ideas behind Learn2Play Bench

Strong task performance can reflect prior knowledge. Novel game rules let us examine how agents acquire and reuse knowledge through interaction.

01

Discover, don't look up.

Novel, sometimes counterintuitive rules create unfamiliar situations. Agents see actions and feedback—not game source code or solutions.

02

Do agents improve with practice?

Agents play each game repeatedly and can use what they learned in earlier attempts. We measure whether their scores improve.

03

Can agents use what they learn in new situations?

In some games, the situation changes between attempts, but the rules stay the same. We test whether agents can still use what they learned.