The question
An answer in a chat window tells you only so much about an agent. ChronoBench puts language models inside Chrono Trigger, where they must interpret the screen, choose actions, and make progress through a connected world.
The benchmark covers a defined opening segment, with eight primary checkpoints. Reaching the final checkpoint means completing that segment—not the entire game.
What I built
I built the evaluation and presentation around a vision-based game agent: a versioned record of runs, checkpoint evidence, resource use, and the places where models struggle. The public site makes those outcomes inspectable rather than reducing them to a single score.
The current harness adds evidence validation and a secondary exploration track. That separates progress through the story from curiosity about the world around it.
Making the evidence useful
- Progress: primary checkpoints establish how far each run got.
- Behavior: secondary checkpoints, stuck events, and diagnostics show a different side of performance.
- Resources: cycle counts, token use, and recorded inference costs give results context.
- Evidence: available screenshots connect a checkpoint to something a reader can actually inspect.
Runs remain grouped by harness version. Earlier versions are archived because changes in the agent environment matter when interpreting results. Inferred checkpoint timings and estimated costs are labeled where they appear.
Explore the results
The current standings give a quick overview. The explorer lets you compare individual runs and inspect their checkpoint evidence. Methodology explains the constraints behind the numbers.
