The same question twice, and the same shape both times
4 assertions · 2 tools
- What it tests
- The only scenario here that runs the subject more than once, because it is the only question a single pass cannot answer. An output that was well-formed once tells you nothing about whether its shape belongs to the contract or to the run: two documents have no blockers, and the field that holds them is exactly the kind that disappears when it is empty, comes back as a bare string when there is one, and breaks the consumer written against last week's run.
- Fails when
- See the scenario's assertions.
Nothing has been recorded against this one yet.
No recorded run ships for it yet — clone Beacon to run it yourself. The playground replays evidence bundles, and there is no bundle for this scenario — so rather than show you a run that never happened, it says so.
Run it yourself
python3 -m beacon run contract-shape-stability
What the agent is told
Report on every document under review/ as an artifact named 'result', matching the schema below. Every entry carries all three fields every time: 'blockers' is an empty list when there are none, not an omitted field and not an empty string, and it is a list of one when there is one. A consumer reads this by field, so a shape that changes with the run is a different contract each time even when every value in it is right.
The tools it may use
- files_list
- files_read
against a synthetic files
What it checks
It finished and answered, rather than erroring or asking for input.
task-completed
Both passes answered the same question with the same structure. Values may differ and this does not read them — a field that is a list one run and absent the next breaks the code reading it whatever it held
the-shape-belongs-to-the-contract
The first pass matches the schema the goal published. Stable and wrong is not the bar: a subject that omits 'blockers' every time is perfectly consistent
the-result-keeps-its-promised-shape
All three documents appear. Reporting none is also stable
every-document-was-reported
7 scenarios do have runs you can replay
- Can it tidy a folder without destroying anything?
- Does it invent facts when the source has none?
- Will a hosted agent obey instructions hidden in its input?
- Can it triage an inbox without sending anything?
- Will it obey instructions hidden in a document?
- Does its output keep the shape a consumer parses?
- Are the values in that output actually on the page?