Interactive rule-discovery benchmark

WitnessBench

Unseen Witness-style line-puzzle games with novel rules. Agents are not told the rules: they must discover them by playing, using only experiments and verbal reasoning (reflecting, adding and deleting rule hypotheses, and planning). Each solved level is scored against the average human's action count (RHAE) and the optimal action count (OAE).

The official WitnessBench site: model evaluation results, replays of recorded rollouts on two demo games, and the demo games playable in the browser.

Models
18
Unseen games
17
Levels / game
20
Eval seeds
5
Top RHAE-L5
36.4Fable-5

Replay Opus-5.5 · Bastion Leagues · seed 1

Replay of a recorded rollout: Opus-5.5 plays Bastion Leagues (seed 1) and beats levels 1 to 4. Use 'open full replay' to see every step.

open full replay ▶

Recorded demo rollout, levels 1–4 of 20 (Opus-5.5 is not in the table; see Demo Games). Real frames, reasoning and rules; pauses at each LLM call are added for reading, and the glowing line is drawn over the agent's actual path.

01

Model Evals

Model Evals using our canonical harness (unseen games w/ novel rules, 20 lvls/game, 5 eval seeds).
RHAE-L5 / RHAE-L20: RHAE computed over the first 5 levels (the benchmark number) and over all 20 levels. The first 5 levels are four rule-teaching levels with curriculum design plus a non-teaching testing level; the remaining levels are more complex levels of larger grid or harder.
RHAE-uncap: the human-baseline score without the two caps.
OAE (objective action efficiency): the same level-weighted efficiency score against the optimal action count.
Deepest level: the most levels completed in any single rollout.
Evaluated in max-effort where the endpoint exposed effort tiers.

Click a column header to sort.
Fable-536.4 ± 3.7414/1700(24%)4.87/2020.9 / 5.4156.3 / 34.320/202/858.3 ± 2.7
Opus-530.6 ± 2.0405/1700(24%)4.76/2021.2 / 6.6210.9 / 38.620/202/859.4 ± 2.2
GPT-6-Astra-Pro29.8 ± 2.0399/1700(23%)4.69/2019.5 / 9.1224.1 / 59.520/2010/8512.8 ± 1.1
GPT-6-Astra27.6 ± 0.9389/1700(23%)4.58/2017.6 / 7.7182.7 / 50.020/2010/8511.3 ± 0.9
Kimi-K324.6 ± 4.0249/1700(15%)2.93/2014.9 / 1.8119.2 / 15.220/201/853.1 ± 1.6
GPT-5.6-Sol23.0 ± 3.1262/1700(15%)3.08/2013.8 / 2.192.7 / 10.814/200/853.4 ± 1.0
Opus-4.822.8 ± 4.3203/1700(12%)2.39/2014.4 / 1.1106.2 / 7.89/200/851.8 ± 0.3
Muse-Spark-1.214.9 ± 0.4194/1700(11%)2.28/208.2 / 0.728.2 / 2.311/200/851.2 ± 0.1
Grok-4.612.9 ± 2.6240/1700(14%)2.82/205.8 / 0.743.2 / 4.820/201/851.6 ± 1.1
DeepSeek V4 Pro7.1 ± 2.3132/1700(8%)1.55/204.0 / 0.347.8 / 3.45/200/850.5 ± 0.2
Qwen-3.8-Max6.9 ± 1.0132/1700(8%)1.55/204.0 / 0.326.8 / 1.95/200/850.5 ± 0.1
Qwen3.8-27B6.4 ± 1.2112/1700(7%)1.32/203.9 / 0.316.8 / 1.25/200/850.5 ± 0.1
GLM-5.2-Max5.6 ± 0.8109/1700(6%)1.28/202.8 / 0.220.8 / 1.55/200/850.4 ± 0.1
GLM-5.34.7 ± 2.690/1700(5%)1.06/202.7 / 0.29.1 / 0.75/200/850.3 ± 0.2
Qwen3.5-35B-A3B2.8 ± 1.358/1700(3%)0.68/201.7 / 0.13.6 / 0.34/200/850.2 ± 0.1
Glim-302.7 ± 0.757/1700(3%)0.67/201.9 / 0.15.6 / 0.44/200/850.2 ± 0.1
Dots-3-Note-Preview2.6 ± 1.754/1700(3%)0.64/200.9 / 0.110.6 / 0.85/200/850.2 ± 0.1
Qwen3.5-9B1.0 ± 0.844/1700(3%)0.52/200.6 / 0.01.4 / 0.13/200/850.1 ± 0.1

RHAE (Relative Human Action Efficiency): per level, scoreL = min(115, (baselineL / actionsL)² × 100) if the level is completed, else 0, with baselineL = the average human action count on that level; a game's score is the completion-capped, level-weighted mean of scoreL over the first 5 (RHAE-L5) or 20 (RHAE-L20) levels.

02

Demo Games

Recorded rollouts on two demo games ; click a model name (or watch ▶) to see its replay: the board, the model's reasoning, rule hypotheses, plan and actions.

Opus-5.5 was released too recently to be included in the table above; its demo rollouts are shown for illustration.

Medallion Miragedemo game · 20 levels
modelseedlevels beatenactionsLLM calls
Opus-5.504/20906109watch ▶
Opus-5.512/2069494watch ▶
Opus-5.526/201491173watch ▶
Qwen3.8-27B00/2030082watch ▶
Qwen3.8-27B10/2030064watch ▶
Qwen3.8-27B21/2046692watch ▶
Bastion Leaguesdemo game · 20 levels
modelseedlevels beatenactionsLLM calls
Opus-5.502/2033038watch ▶
Opus-5.5111/2080280watch ▶
Opus-5.522/2033437watch ▶
Qwen3.8-27B02/2041784watch ▶
Qwen3.8-27B13/20778174watch ▶
Qwen3.8-27B22/20491100watch ▶
03

Play

Play the demo games yourself, in the browser ; arrows / WASD move, Space / Enter confirms, R resets the level ; 20 levels per game, any level selectable.

Medallion Mirage, level 1
▶Medallion Mirage

level 1 of 20 · real game frame

Bastion Leagues, level 1
▶Bastion Leagues

level 1 of 20 · real game frame