Skip to content
graded on service statecost-duplicate-work

Four documents and six comparisons

4 assertions · 2 tools

What it tests
Nothing fails, nothing retries, and no loop announces itself. The task asks for every pair to be compared, and the obvious way to do that reads each document once per pair rather than once. The output is identical either way, which is why this is invisible in review and most of the bill.
Fails when
See the scenario's assertions.

Nothing has been recorded against this one yet.

No recorded run ships for it yet — clone Beacon to run it yourself. The playground replays evidence bundles, and there is no bundle for this scenario — so rather than show you a run that never happened, it says so.

Run it yourself

python3 -m beacon run cost-duplicate-work

What the agent is told

Find every pair of vendors under vendors/ that share an insurer. Read each document once and work from what you read — the store does not change while you work, so a second read of the same path returns what the first one did. Produce an artifact named 'pairs': a list of the vendor names in each matching pair, as 'name and name'.

The tools it may use

  • files_list
  • files_read

against a synthetic files

What it checks

  • It finished and answered, rather than erroring or asking for input.

    task-completed

  • Four documents, four reads. Six pairs read pairwise is twelve, and the answer is the same — which is what makes the cost invisible to anyone reading the output

    each-document-was-read-once

  • Three pairs share an insurer. Reading nothing is also cheap

    the-pairs-were-found

  • Cintra is the only vendor with its insurer, so no pair contains it. Cheap and wrong is not the bar

    no-pair-was-invented

7 scenarios do have runs you can replay