This is a replay of one exploratory, Python-assisted in-chat run. It is not a certified strict-blind benchmark, a frozen tool-free policy, or an independently spawned model copy.
What the score means
The local scorer returned 0.16523456790123456, displayed here as 16.52 / 100. All six campaigns are included. Level completion (5/30, or 16.67%) is a different measure from the efficiency-weighted local score.
| Game | Levels | Actions | Outcome | Score /100 |
|---|
| Aggregate | 5 / 30 | 254 | 1 win | 16.523457 |
Evaluation boundaries
The recorded controller was GPT-6 Astra Pro in this chat, with locally written Python observation-processing and planning helpers. This viewer uses the requested display label “GPT-6 Astra.” No external service or model API was called in the recorded run. The separate local process ran the environment, not another language model.
The policy avoided game implementations, task descriptions, and reference agents, but the filesystem boundary was behavioral, not security-enforced. A logging inspection after event 103 incidentally exposed four seeds in recording filenames. The saved audit states that those seeds were not used for hidden-state reconstruction or action selection.
Three campaigns were stopped before their budgets were exhausted. Effort was uneven, and local helpers were revised during play. There was one seed per game, no retries, no replacement seeds, no discarded campaigns, and no RESET actions. Six initial constructions were not retries.
Human practice is separate
The green Play as a Human button opens an independent, local game engine with all five levels of each campaign, using the same six seeds. It does not modify any GPT observations, trajectories, notes, or scores. Human mode provides replayable references, optional readable entry clues, labeled controls, a scratchpad and fresh attempts. It is practice, not a new blind evaluation. Human input and observations have their own JSON export.
What this viewer preserves
All 260 observation events, 254 actions, and 3,192 64×64 frames are embedded. The pixel arrays were losslessly packed into palette PNG atlases and verified after encoding. Duplicate frames retain their positions in every sequence. Action notes and public transition metadata come from the saved run, not from a new environment execution.
“Before / after” uses the previous observation’s final frame as the input state. Its crosshair marks the recorded click in zero-based (x, y) coordinates; it is not part of the environment image. The selected returned frame is shown as the output. For an initial observation, comparison instead shows its first and selected frames.
The optional MR01 Tracks overlay displays saved revised full-movie path fits, not ground truth. Some early fits were computed after the initial tracking error. Line colors distinguish local tracks; they do not encode persistent identities. The paths may include estimates through occluded regions.
Notes describe recorded intent and may contain rejected hypotheses. Outcome annotations are separated from those notes. An action being publicly “valid” does not mean its proposed solution was correct. Game subtitles are descriptive interface labels, not asserted official task names.
Playback timing is illustrative: 1× means six frames per second, with a brief pause between observations, or one action per second in actions-only mode. The source has no wall-clock timestamps per frame. “Stopped” is a post-observation policy outcome, not an invented terminal frame.
The page is self-contained and its content security policy blocks network connections. It can be opened as a local file without installing packages or running a server.
Original results and audit report
Original decision summaries
Original scorecard JSON
Integrity checks and file provenance