Discover, don't look up.
Novel, sometimes counterintuitive rules create unfamiliar situations. Agents see actions and feedback—not game source code or solutions.
Learning from experience in unfamiliar environments
How effectively can LLM agents learn from experience?
20 game templates with hidden rules to discover. Different seeds generate different game instances. Agents learn each game’s hidden rules by playing repeatedly and use what they learn to improve in later episodes, without updating their model weights.
Research snapshot · September 2026 · Game portal requires sign-in
Game examples
Choose a game and scroll through its opening situation, sample actions, and feedback. These are recorded demonstrations—not a live game or a recommended strategy.
Evaluation protocol
For each game template, we use one selected rule seed and run 3 independent trials. The leaderboard does not average over all possible seeds.
We report 5 episodes for normal games and 10 for challenge games. Scores are normalized using each game's reference ceiling, then averaged with weight 3 for challenge games and weight 1 for normal games.
Average peak score across trials.
Average score across episodes.
Improvement from early to late episodes.
Fitted rate of score change across episodes.
Results
Compare backbone models, self-evolving methods, and agent harnesses. Default ranking uses Max.
| Rank | Model | Max ↑ | Mean ↑ | LG ↑ | LS ↑ |
|---|
Source: manuscript backbone/method and harness comparison tables · Updated September 29, 2026. Sorted highest first; equal displayed values share a rank. Highlighting marks the best value in each column. These are aggregate point estimates, not significance rankings. This leaderboard is a curated snapshot, not automatically updated.
Learning through repeated interaction
A peak score does not show the path to that score. These curves track normalized performance over the evaluation window across all 20 games.
Efficiency
Compare Max scores with estimated mean inference cost per episode. Points toward the upper left achieve higher performance at lower estimated cost.
Hover, tap, or focus a point to see its full name and exact values.
Benchmark design
Strong task performance can reflect prior knowledge. Novel game rules let us examine how agents acquire and reuse knowledge through interaction.
Novel, sometimes counterintuitive rules create unfamiliar situations. Agents see actions and feedback—not game source code or solutions.
Agents play each game repeatedly and can use what they learned in earlier attempts. We measure whether their scores improve.
In some games, the situation changes between attempts, but the rules stay the same. We test whether agents can still use what they learned.