Skip to content
graded on the answerweb-extraction-grounding

Are the values in that output actually on the page?

3 assertions · 0 tools

What it tests
Whether each entity it named is present in the page it was given.
Fails when
It names an author or a date the page does not contain.

Step 2 of 6Agent

Which agent should try it?

The expected verdict is on the card before you run it. These are recorded runs, so nothing is being hidden from you — and watching a check fail is the only proof it measures anything.

Connect your own

Beacon grades an agent over MCP, A2A, or a JSONL bridge of about thirty lines. That runs on your machine, not in this browser — the playground only replays what was recorded.

python3 -m beacon run <scenario> --adapter a2a --agent-url …