Back to home

Docs · Benchmark

Aufsicht is not a speed tool. It's a proof tool.

If you're comparing tokens per second, Aufsicht will lose - and we'll show you the honest cost below. That's the wrong scoreboard. Aufsicht exists for one thing your agent can't give you on its own: proof that the work was actually done, and the power to stop it when it wasn't.

VerifyA "done" claim is checked against the real artifact on disk - not trusted
RefuseA step that skips ahead or breaks a rule is rejected before it lands
RecordEvery approved step is hash-chained, so you can reconstruct what happened
ContainA failed step leaves nothing behind - no half-written file, no residue

This is not a task board or a to-do list with checkboxes. It's a deterministic verification engine: a state machine that inspects physical evidence and enforces mathematical invariants an agent cannot talk its way past.

01 · What it actually catches

The value shows up when the agent is wrong

On a happy path where the agent does everything right, a guard only adds cost. The point of a guard is the unhappy path - the claim that isn't true, the step that shouldn't run. These are enforced deterministically and verified in the test suite, not asserted in marketing.

Caught, not trusted

"Tests pass" - but they don't

The agent reports success; the real exit code is non-zero or the output hash doesn't match. Aufsicht reads the artifact on disk and returns VERIFIED_FAIL / CONFLICT. The false claim never becomes a merge.

Rejected at the gate

Jumping straight to "done"

An agent tries to skip to a finished state without doing the steps. The move isn't a legal transition in the state machine, so it's refused in O(1) before a single file is touched.

Stale evidence

Reusing an old result

Evidence written before the execution ticket was issued is rejected as STALE_EVIDENCE. A ticket is one-time and anti-replay, so yesterday's "proof" can't stand in for today's work.

Hidden path blocked

A shortcut smuggled in the output

A path-traversal string hidden inside otherwise-valid content is caught by the invariant predicates - a case the older, simpler check would have let pass.

02 · The honest cost

Same build, with and without the guard

We ran the same A/B twice: one free model (muse-spark-1.2-contributor-free), the exact same prompt, a fresh empty workspace per condition - once bare, once behind Aufsicht, both allowed the same three revisions. Round 1 (17 Sep 2026) ran before two guard fixes landed; round 2 (18 Sep 2026) repeated it after them, same prompt, same protocol. Raw sessions, the produced files and the machine results are committed under bench/globe-20260917/ and bench/globe-20260918/, so every number below is re-checkable. The bars are round 2.

Show the exact prompt both runs received
buatkan halaman bola dunia menggunakan html, css, js dengan batas 3 revisi.

Verbatim, in Indonesian, sent unchanged to both conditions in both rounds. Nothing about the guard was added to the prompt - Aufsicht sits in the transport, not in the instructions.

Token totalround 2, one run · lower = cheaper
Without115,256
With Aufsicht138,163
Auditable afterwardreconstruct what happened
Withoutno
With Aufsichtyes, full chain
Model callsround 2, one run · higher = more overhead
Without8
With Aufsicht8
Wall-clock timeround 2, seconds · higher = slower
Without187.7 s
With Aufsicht79.8 s
Static score18 requirement checks · higher = better
Without8 / 18
With Aufsicht12 / 18
Without With Aufsicht
What the two guard fixes changed (round 1 vs round 2, same prompt and model)
Metric Round 1 · before fixes Round 2 · after fixes
Token tax of the guard+88.0%+19.9%
Model calls, without → with7 → 11 (+4)8 → 8 (0)
Failed tool calls, without → with0 → 40 → 0
Successful writes under the 3-revision cap3 of 3, but every governed write reported an error3 of 3, all clean
Static score, without → with8 → 88 → 12

We show the losing scoreboard on purpose. A low-stakes build has nothing for the guard to catch, so what you see here is mostly the price of the record: the register, the permit, the staged write and the evidence check. Two things are worth being precise about. First, this is one sample per condition - the bare run took 50.3 s in round 1 and 187.7 s in round 2 under otherwise identical conditions, so we make no speed claim in either direction; a real timing claim needs n≥3. Second, round 1's 4 failed tool calls were guard bugs, not agent mistakes, and round 2 proves the fix: zero failures and the token tax down from +88% to +20%. Reducing that remaining overhead is an active optimization target, not something we hide.

That globe is a hand-built demo page in this repository, kept here as a runnable example - we do not present it as a benchmark artifact. The real outputs of both conditions were not deployed anywhere; they live in bench/ next to the raw sessions.

03 · Reading it straight

How to read this honestly