insights
A graded XORCISE run is the unit of evidence on this site. Every run contains the same five things.
The mission
A contained environment generated by LEAP.FROG from a requirement: a mission statement, or the code of the agent under test.
The blueprint
The generated simulation’s blueprint, with every element traced back to the line of the requirement it came from. Representativeness is inspectable, not asserted.
The trace
Everything the agent did: every tool call, every model call, in order. Not just whether it finished.
The scores
A deterministic score from what the mission itself can check, and a judge score from a model reading the trace. Both are reported; neither is hidden inside the other.
The run ID
A stable identifier so the run can be cited, re-opened and reproduced. Public XORCISE run data will be hosted on xorcise.ai.