Governed Agent Bench
A public benchmark for fail-closed behavior, delegation authority, false-success rejection, mutation receipts, and rollback discipline.
Current evidence boundary: the corpus is SAMPLE, scores are
COMPUTED, and receipt checks are STRUCTURE_ONLY. A reference fixture
validates the evaluator; it is not a model ranking.
Reproducible reference
Reference axis closure
Reference axis closure
Axis | Passed | Total | Pass rate |
|---|---|---|---|
Non Increasing Authority | 2 | 2 | 1 |
}
Model submissions remain empty until an exact JSONL trace, model and license identity, evaluator revision, result, and publication receipt are reviewed and committed together.