WitnessBench ← back — / —
0 / 0
Levels: filter action history:

Game

↑ ← ↓ →
⟲ ✓
faded path = the path just submitted (display overlay; the game resets the path to the start on a failed submit)
red rings (if any) = cells the game flagged as violations

Now

Action history ◆ LLM = LLM call ; dim = plan exec

    Plan

    Memory

    Active rules

    running state at this step
    ● verified ● proposed ● transferred

      Compacted (last LLM-fed snapshot)

      
            

      Last change

        Memory growth

        Reasoning trace

        What the LLM saw (semantic ASCII board)

        
            

        Rules fed to the reflection (## Current Knowledge)

        
            

        Observations (recent action effects)

        
            

        LLM response (meta = thinking, add/del = rule churn, plan = next actions)

        What is the agent doing right now?

        The agent alternates between LLM reflection calls and plan execution. The current stage shows in the Now panel and is color-banded in the timeline above:

        All reflections