Interactive rule-discovery benchmark
WitnessBench
Unseen Witness-style line-puzzle games with novel rules. Agents are not told the rules: they must discover them by playing, using only experiments and verbal reasoning (reflecting, adding and deleting rule hypotheses, and planning). Each solved level is scored against the average human's action count (RHAE) and the optimal action count (OAE).
The official WitnessBench site: model evaluation results, replays of recorded rollouts on two demo games, and the demo games playable in the browser.
- Models
- 18
- Unseen games
- 17
- Levels / game
- 20
- Eval seeds
- 5
- Top RHAE-L5
- 36.4Fable-5
Replay Opus-5.5 · Bastion Leagues · seed 1
Replay of a recorded rollout: Opus-5.5 plays Bastion Leagues (seed 1) and beats levels 1 to 4. Use 'open full replay' to see every step.
Recorded demo rollout, levels 1–4 of 20 (Opus-5.5 is not in the table; see Demo Games). Real frames, reasoning and rules; pauses at each LLM call are added for reading, and the glowing line is drawn over the agent's actual path.
Model Evals
Model Evals using our canonical harness (unseen games w/ novel rules, 20 lvls/game, 5 eval seeds).
RHAE-L5 / RHAE-L20: RHAE computed over the first 5 levels (the benchmark number) and over all 20 levels. The first 5 levels are four rule-teaching levels with curriculum design plus a non-teaching testing level; the remaining levels are more complex levels of larger grid or harder.
RHAE-uncap: the human-baseline score without the two caps.
OAE (objective action efficiency): the same level-weighted efficiency score against the optimal action count.
Deepest level: the most levels completed in any single rollout.
Evaluated in max-effort where the endpoint exposed effort tiers.
| Fable-5 | 36.4 ± 3.7 | 414/1700(24%) | 4.87/20 | 20.9 / 5.4 | 156.3 / 34.3 | 20/20 | 2/85 | 8.3 ± 2.7 |
| Opus-5 | 30.6 ± 2.0 | 405/1700(24%) | 4.76/20 | 21.2 / 6.6 | 210.9 / 38.6 | 20/20 | 2/85 | 9.4 ± 2.2 |
| GPT-6-Astra-Pro | 29.8 ± 2.0 | 399/1700(23%) | 4.69/20 | 19.5 / 9.1 | 224.1 / 59.5 | 20/20 | 10/85 | 12.8 ± 1.1 |
| GPT-6-Astra | 27.6 ± 0.9 | 389/1700(23%) | 4.58/20 | 17.6 / 7.7 | 182.7 / 50.0 | 20/20 | 10/85 | 11.3 ± 0.9 |
| Kimi-K3 | 24.6 ± 4.0 | 249/1700(15%) | 2.93/20 | 14.9 / 1.8 | 119.2 / 15.2 | 20/20 | 1/85 | 3.1 ± 1.6 |
| GPT-5.6-Sol | 23.0 ± 3.1 | 262/1700(15%) | 3.08/20 | 13.8 / 2.1 | 92.7 / 10.8 | 14/20 | 0/85 | 3.4 ± 1.0 |
| Opus-4.8 | 22.8 ± 4.3 | 203/1700(12%) | 2.39/20 | 14.4 / 1.1 | 106.2 / 7.8 | 9/20 | 0/85 | 1.8 ± 0.3 |
| Muse-Spark-1.2 | 14.9 ± 0.4 | 194/1700(11%) | 2.28/20 | 8.2 / 0.7 | 28.2 / 2.3 | 11/20 | 0/85 | 1.2 ± 0.1 |
| Grok-4.6 | 12.9 ± 2.6 | 240/1700(14%) | 2.82/20 | 5.8 / 0.7 | 43.2 / 4.8 | 20/20 | 1/85 | 1.6 ± 1.1 |
| DeepSeek V4 Pro | 7.1 ± 2.3 | 132/1700(8%) | 1.55/20 | 4.0 / 0.3 | 47.8 / 3.4 | 5/20 | 0/85 | 0.5 ± 0.2 |
| Qwen-3.8-Max | 6.9 ± 1.0 | 132/1700(8%) | 1.55/20 | 4.0 / 0.3 | 26.8 / 1.9 | 5/20 | 0/85 | 0.5 ± 0.1 |
| Qwen3.8-27B | 6.4 ± 1.2 | 112/1700(7%) | 1.32/20 | 3.9 / 0.3 | 16.8 / 1.2 | 5/20 | 0/85 | 0.5 ± 0.1 |
| GLM-5.2-Max | 5.6 ± 0.8 | 109/1700(6%) | 1.28/20 | 2.8 / 0.2 | 20.8 / 1.5 | 5/20 | 0/85 | 0.4 ± 0.1 |
| GLM-5.3 | 4.7 ± 2.6 | 90/1700(5%) | 1.06/20 | 2.7 / 0.2 | 9.1 / 0.7 | 5/20 | 0/85 | 0.3 ± 0.2 |
| Qwen3.5-35B-A3B | 2.8 ± 1.3 | 58/1700(3%) | 0.68/20 | 1.7 / 0.1 | 3.6 / 0.3 | 4/20 | 0/85 | 0.2 ± 0.1 |
| Glim-30 | 2.7 ± 0.7 | 57/1700(3%) | 0.67/20 | 1.9 / 0.1 | 5.6 / 0.4 | 4/20 | 0/85 | 0.2 ± 0.1 |
| Dots-3-Note-Preview | 2.6 ± 1.7 | 54/1700(3%) | 0.64/20 | 0.9 / 0.1 | 10.6 / 0.8 | 5/20 | 0/85 | 0.2 ± 0.1 |
| Qwen3.5-9B | 1.0 ± 0.8 | 44/1700(3%) | 0.52/20 | 0.6 / 0.0 | 1.4 / 0.1 | 3/20 | 0/85 | 0.1 ± 0.1 |
RHAE (Relative Human Action Efficiency): per level, scoreL = min(115, (baselineL / actionsL)² × 100) if the level is completed, else 0, with baselineL = the average human action count on that level; a game's score is the completion-capped, level-weighted mean of scoreL over the first 5 (RHAE-L5) or 20 (RHAE-L20) levels.
Demo Games
Recorded rollouts on two demo games ; click a model name (or watch ▶) to see its replay: the board, the model's reasoning, rule hypotheses, plan and actions.
Opus-5.5 was released too recently to be included in the table above; its demo rollouts are shown for illustration.
| model | seed | levels beaten | actions | LLM calls | |
|---|---|---|---|---|---|
| Opus-5.5 | 0 | 4/20 | 906 | 109 | watch ▶ |
| Opus-5.5 | 1 | 2/20 | 694 | 94 | watch ▶ |
| Opus-5.5 | 2 | 6/20 | 1491 | 173 | watch ▶ |
| Qwen3.8-27B | 0 | 0/20 | 300 | 82 | watch ▶ |
| Qwen3.8-27B | 1 | 0/20 | 300 | 64 | watch ▶ |
| Qwen3.8-27B | 2 | 1/20 | 466 | 92 | watch ▶ |
| model | seed | levels beaten | actions | LLM calls | |
|---|---|---|---|---|---|
| Opus-5.5 | 0 | 2/20 | 330 | 38 | watch ▶ |
| Opus-5.5 | 1 | 11/20 | 802 | 80 | watch ▶ |
| Opus-5.5 | 2 | 2/20 | 334 | 37 | watch ▶ |
| Qwen3.8-27B | 0 | 2/20 | 417 | 84 | watch ▶ |
| Qwen3.8-27B | 1 | 3/20 | 778 | 174 | watch ▶ |
| Qwen3.8-27B | 2 | 2/20 | 491 | 100 | watch ▶ |
Play
Play the demo games yourself, in the browser ; arrows / WASD move, Space / Enter confirms, R resets the level ; 20 levels per game, any level selectable.

