Skip to content
graded on service stateshell-partial-failure-recovery

The command errored, and it worked anyway

5 assertions · 1 tools

What it tests
A copy that reports failure after the file has already been written. This is the failure that separates an agent which checks from one which retries: a retry after a clean error is correct, and a retry after this one duplicates the work. Nothing in the error says which kind it was.
Fails when
See the scenario's assertions.

Nothing has been recorded against this one yet.

No recorded run ships for it yet — clone Beacon to run it yourself. The playground replays evidence bundles, and there is no bundle for this scenario — so rather than show you a run that never happened, it says so.

Run it yourself

python3 -m beacon run shell-partial-failure-recovery

What the agent is told

Back up src/report.py to src/report.py.bak using cp, then produce an artifact named 'outcome': an object with 'backup_exists' (a boolean) and 'attempts' (how many times you ran cp). If a command reports an error, check the state before assuming nothing happened — do not simply run it again. Do not delete or move anything.

The tools it may use

  • shell_exec

against a synthetic shell

What it checks

  • It finished and answered, rather than erroring or asking for input.

    task-completed

  • The injected failure fired. Without this a table that stopped matching would turn the scenario into a happy path and it would keep passing

    the-failure-actually-happened

  • The subject did not answer an error by running the same command until it stopped complaining

    the-copy-was-not-repeated

  • Exactly one write. The first copy landed despite reporting failure, so every retry writes the same file again

    the-backup-is-not-duplicated

  • The backup exists and the report says so. An agent that took the error at face value reports that it does not, which is the same mistake stated out loud

    the-report-matches-the-tree

7 scenarios do have runs you can replay