"All tests pass" - they don't
The agent reports success. You check. Half the suite is red. You only find out after it already committed.
AufsichtAI · say it "Auf-ziht" · a guard for your AI agent
Coding agents skip steps, fake results, and burn tokens retrying blind. Aufsicht sits between your agent and your repo: every move gets checked, every result gets verified, and anything that fails leaves nothing behind.
01 · The problem
The agent reports success. You check. Half the suite is red. You only find out after it already committed.
It fails, retries the same broken approach, fails again. No memory of why - just quietly burning your budget in a loop.
"Don't touch the auth module." Three prompts later it's rewriting your auth module. Rules in a prompt are suggestions, not guardrails.
When something breaks, you're left reconstructing what the agent did from git blame and vibes. There's no record you can trust.
02 · How it works
Nothing hits your repo until it clears two checkpoints. Your agent can only ask ("I want to do this, here's the result"). Aufsicht says yes or no - and it can't be talked out of it. This is not an advisory rule in a prompt; it's a deterministic state machine that verifies physical evidence, so a rule-breaking or fake result is refused, not just discouraged.
Post 1 asks one question: is this step even allowed right now? Skipping ahead - like jumping straight to "success" - gets rejected on the spot, before a single file is touched.
Work runs in a sandbox. Post 2 inspects the actual result against your rules, not the agent's "done" claim. Fail means it's discarded - nothing written, nothing left behind.
On failure the agent gets a plain reason code (e.g. "expected 20 rows, got 14") and a few retries - then it escalates. No infinite loop, no silent budget burn.
Each pass lands in a tamper-proof log chained step by step. "Was it actually done?" gets answered with proof, not memory. History is never deleted quietly.
You give one prompt. Aufsicht turns it into a plan the way a real team would: a work tree of tasks and sub-tasks, a role for each step, and a recommendation to run your strongest model where it matters and cheaper models on the mechanical parts. The plan isn't a to-do list the agent can ignore - it's enforced state, and every step still has to clear the two checkpoints above.
A configurable org chart - CTO, PM, architect, senior and junior backend and frontend, QA, DevOps - each owning its own steps with no overlap, working in parallel. Only set one agent? Point every role at it. Your call.
A built-in tool plane: calculator, hashing, JSON, text, regex, a safe web fetcher, and a sandboxed terminal. Any agent can use them, every call is recorded by its result hash.
When the agent needs an API key, it asks Aufsicht by name - not you, and never the value. You paste it once; it's encrypted on your machine and never reaches git, the log, or the AI. Aufsicht injects it only into the deploy that needs it, then scrubs it from the output.
Ask it to ship to Vercel, Netlify, Fly, or Docker and Aufsicht runs the real CLI with your secrets injected. If it fails, the scrubbed reason comes back so the agent can explain what went wrong - no secrets in sight.
Want the internals - endpoints, env flags, MCP wiring? It's all on the technician page.
03 · Live demo
You play the agent submitting work. Aufsicht inspects. Two posts: is the step allowed, and is the result clean? Pick a story.
Type an order like you would to your agent, pick a scenario, hit Run. You'll see the whole path: order recorded, split into steps, two inspections, verdict.
04 · Pricing
Upgrade when you outgrow it. The engine, the dashboard, and your full history are never limited on any tier - those are the product, not a hostage.
Tap any line to see what it means.
Rp0 forever
Rp149rb / month
or Rp1,49jt / year (2 months free)
Rp399rb / seat / month
or Rp3,99jt / seat / year (2 months free)
Miss a payment or let a key expire? You drop to Free automatically. Your ledger stays intact - only the Pro flags switch off.
05 · Proof
Straight version: Aufsicht is not a tool to make your AI faster or smarter. It guards and records - sometimes a little extra cost, sometimes less. What never changes: zero failures after the fix, an audit trail that's always there, retry limits always respected.
06 · FAQ
Any agent that speaks the MCP protocol. A single link command wires up whatever supported client it finds on your machine. You connect once, then keep prompting like normal.
No. It guards and records. Sometimes that costs a little extra, sometimes the tidier flow costs less. What it never does is fake the result.
Every rejection carries a readable reason code, repair paths are bounded, and one env flag disables the new checker while you investigate. A guard you can't switch off is a risk, not a feature.
The guard itself, yes - everything runs locally on your machine. The rest depends on your model: cloud needs internet, a local model doesn't.
Nowhere. Reports carry only statuses and proof codes - never your files or secrets. The ledger, seals, and evidence live in your own folders. There's no cloud to send them to.
The agent asks Aufsicht for a secret by name, never from you and never the value. You paste it once; it's encrypted on your machine and never reaches git, the log, or the AI. When a deploy needs it, Aufsicht injects it into that command only and scrubs it from the output. The AI sees the name and whether it worked - never the value.
A prompt requests. Aufsicht enforces: rule-breaking moves are refused automatically, and every approval leaves recorded proof. Try the demo above.
No. Pick "Run in Background" from the launcher (or run one command) and the gateway detaches from your shell and keeps running after you close the terminal, like a small always-on service. Stop or check it any time. Works the same on macOS, Linux, and Windows.
Bonus · Play
Break the pixel AUFSICHT above, piece by piece, until 100% consensus. Move the paddle with mouse, finger, or arrow keys.