← Selected work

ChronoBench

What can an AI agent accomplish when it has to see the screen, make a plan, and find its own way?

Independent projectBenchmark design · Agent systems · Engineering
An agent’s recorded view of Chrono Trigger
CHRONOBENCH / FIELD STUDY 038 CHECKPOINTS

The question

An answer in a chat window tells you only so much about an agent. ChronoBench puts language models inside Chrono Trigger, where they must interpret the screen, choose actions, and make progress through a connected world.

The benchmark covers a defined opening segment, with eight primary checkpoints. Reaching the final checkpoint means completing that segment—not the entire game.

What I built

I built the evaluation and presentation around a vision-based game agent: a versioned record of runs, checkpoint evidence, resource use, and the places where models struggle. The public site makes those outcomes inspectable rather than reducing them to a single score.

The current harness adds evidence validation and a secondary exploration track. That separates progress through the story from curiosity about the world around it.

Making the evidence useful

  • Progress: primary checkpoints establish how far each run got.
  • Behavior: secondary checkpoints, stuck events, and diagnostics show a different side of performance.
  • Resources: cycle counts, token use, and recorded inference costs give results context.
  • Evidence: available screenshots connect a checkpoint to something a reader can actually inspect.

Runs remain grouped by harness version. Earlier versions are archived because changes in the agent environment matter when interpreting results. Inferred checkpoint timings and estimated costs are labeled where they appear.

Explore the results

The current standings give a quick overview. The explorer lets you compare individual runs and inspect their checkpoint evidence. Methodology explains the constraints behind the numbers.

A visible record of what autonomous agents can actually do.

More work →