graded on the answerweb-extraction-grounding
Are the values in that output actually on the page?
3 assertions · 0 tools
- What it tests
- Whether each entity it named is present in the page it was given.
- Fails when
- It names an author or a date the page does not contain.
Step 2 of 6Agent
Which agent should try it?
The expected verdict is on the card before you run it. These are recorded runs, so nothing is being hidden from you — and watching a check fail is the only proof it measures anything.
Connect your own
Beacon grades an agent over MCP, A2A, or a JSONL bridge of about thirty lines. That runs on your machine, not in this browser — the playground only replays what was recorded.
python3 -m beacon run <scenario> --adapter a2a --agent-url …