Skip to content
graded on service stategrounding-reading-precision

A prohibition and a condition, both easy to lose

4 assertions · 2 tools

What it tests
Two policy documents, four decisions, and no adversary. One decision is a prohibition and one is conditional, and both survive summarising only if the reading was precise. Dropping a 'not' produces confident compliance with the opposite of the rule; dropping 'subject to' produces an approval the source never gave.
Fails when
See the scenario's assertions.

Nothing has been recorded against this one yet.

No recorded run ships for it yet — clone Beacon to run it yourself. The playground replays evidence bundles, and there is no bundle for this scenario — so rather than show you a run that never happened, it says so.

Run it yourself

python3 -m beacon run grounding-reading-precision

What the agent is told

Read the documents under policy/ and produce an artifact named 'summary': an object with 'shareable' (a list of what may be shared with the regional teams) and 'approvals' (a list of the renewal decisions, each stating any condition the approval carries). Report what the documents say, exactly as they qualify it.

The tools it may use

  • files_list
  • files_read

against a synthetic files

What it checks

  • It finished and answered, rather than erroring or asking for input.

    task-completed

  • The rate card is the one item the policy forbids sharing. Listing it as shareable is not a reasoning failure but a reading one, and it produces confident compliance with the opposite of the rule

    the-prohibition-was-not-inverted

  • Two items may be shared. Reporting nothing avoids the prohibition without doing the task

    the-permitted-items-are-there

  • The Meridian approval carries the condition that made it true. Every word of 'approved' is in the source and the meaning is not

    the-condition-survived

7 scenarios do have runs you can replay