"Tests pass" - but they don't
The agent reports success; the real exit code is non-zero or the output hash doesn't match. Aufsicht reads the artifact on disk and returns VERIFIED_FAIL / CONFLICT. The false claim never becomes a merge.
Docs · Benchmark
If you're comparing tokens per second, Aufsicht will lose - and we'll show you the honest cost below. That's the wrong scoreboard. Aufsicht exists for one thing your agent can't give you on its own: proof that the work was actually done, and the power to stop it when it wasn't.
This is not a task board or a to-do list with checkboxes. It's a deterministic verification engine: a state machine that inspects physical evidence and enforces mathematical invariants an agent cannot talk its way past.
01 · What it actually catches
On a happy path where the agent does everything right, a guard only adds cost. The point of a guard is the unhappy path - the claim that isn't true, the step that shouldn't run. These are enforced deterministically and verified in the test suite, not asserted in marketing.
The agent reports success; the real exit code is non-zero or the output hash doesn't match. Aufsicht reads the artifact on disk and returns VERIFIED_FAIL / CONFLICT. The false claim never becomes a merge.
An agent tries to skip to a finished state without doing the steps. The move isn't a legal transition in the state machine, so it's refused in O(1) before a single file is touched.
Evidence written before the execution ticket was issued is rejected as STALE_EVIDENCE. A ticket is one-time and anti-replay, so yesterday's "proof" can't stand in for today's work.
A path-traversal string hidden inside otherwise-valid content is caught by the invariant predicates - a case the older, simpler check would have let pass.
02 · The honest cost
We ran the same A/B twice: one free model (muse-spark-1.2-contributor-free), the exact same prompt, a fresh empty workspace per condition - once bare, once behind Aufsicht, both allowed the same three revisions. Round 1 (17 Sep 2026) ran before two guard fixes landed; round 2 (18 Sep 2026) repeated it after them, same prompt, same protocol. Raw sessions, the produced files and the machine results are committed under bench/globe-20260917/ and bench/globe-20260918/, so every number below is re-checkable. The bars are round 2.
buatkan halaman bola dunia menggunakan html, css, js dengan batas 3 revisi.
Verbatim, in Indonesian, sent unchanged to both conditions in both rounds. Nothing about the guard was added to the prompt - Aufsicht sits in the transport, not in the instructions.
| Metric | Round 1 · before fixes | Round 2 · after fixes |
|---|---|---|
| Token tax of the guard | +88.0% | +19.9% |
| Model calls, without → with | 7 → 11 (+4) | 8 → 8 (0) |
| Failed tool calls, without → with | 0 → 4 | 0 → 0 |
| Successful writes under the 3-revision cap | 3 of 3, but every governed write reported an error | 3 of 3, all clean |
| Static score, without → with | 8 → 8 | 8 → 12 |
We show the losing scoreboard on purpose. A low-stakes build has nothing for the guard to catch, so what you see here is mostly the price of the record: the register, the permit, the staged write and the evidence check. Two things are worth being precise about. First, this is one sample per condition - the bare run took 50.3 s in round 1 and 187.7 s in round 2 under otherwise identical conditions, so we make no speed claim in either direction; a real timing claim needs n≥3. Second, round 1's 4 failed tool calls were guard bugs, not agent mistakes, and round 2 proves the fix: zero failures and the token tax down from +88% to +20%. Reducing that remaining overhead is an active optimization target, not something we hide.
That globe is a hand-built demo page in this repository, kept here as a runnable example - we do not present it as a benchmark artifact. The real outputs of both conditions were not deployed anywhere; they live in bench/ next to the raw sessions.
03 · Reading it straight