Governed Agent Bench

A public benchmark for fail-closed behavior, delegation authority, false-success rejection, mutation receipts, and rollback discipline.

Current evidence boundary: the corpus is SAMPLE, scores are COMPUTED, and receipt checks are STRUCTURE_ONLY. A reference fixture validates the evaluator; it is not a model ranking.

Reproducible reference

Reference axis closure

{
}

Dataset · Canonical source

Model submissions remain empty until an exact JSONL trace, model and license identity, evaluator revision, result, and publication receipt are reviewed and committed together.