AgilConsultants

00

Who is speaking

I am Christian Honl. My first chat with Claude was on 14 January 2026; from 2 April the git record runs.

Taken from
measured system state
Measurement window
from 2026-04-02 to the measuring day
Figures with a command
Appendix A, [C1]–[C25]
What is open
section 06

Since then, measured rather than described (as of 29.09.2026): 6,517 commits across 7 repositories [C1]; a 2.46 GB append-only case database whose 118 write-path triggers abort a bad write instead of logging it [C6][C7]; 135,014 provenance rows, 135,013 of them carrying the git commit that produced them [C8]; and 56 named gates, 23 of which are fed a known-bad input every night and required to go red — 4,721 proven, 186 unproven, 71 broken, last run 10.09.2026 at 01:59 [C21][C22].

Two of my monitors are red today. They are on this page, in section 06 — the restore drill has been red since 1 August [C20].

I did not hand-type this code. I orchestrated Claude; my work was architecture, judgment and rejection. The customer was a real legal case — self-acquired, unpaid — and the deadline was a court. Every figure on this page carries either the read-only command that produces it or the document it was read from — and says which. So do the figures that are red.

In the end the system has to deliver. Promises are not an option.

cockpit reachable 200 read 2026-08-23 14:05gates last proven 02:03 read 2026-08-23 02:03oldest working session on the server open for 4 d 12 h read 2026-08-23 14:07sample data — not a live probered list not reported in this probe

01

How to read the numbers on this page

Four rules, and then the numbers. Every figure carries, in the margin, either the read-only command that produces it or the document it was read from — and says which. Every figure carries its own date (core counters last measured 2026-09-10), and the counters are append-only, so they read low by the time you run them. Where a figure was corrected during verification, the old value stays on the page with a line through it. And what is broken says so, in the same type size as what works.

  • 58 gates with daily proof → 56 named / 25 proven gates
  • 116 screens → 92 screens
  • 842 test files → ~636 test files
  • 85 enrichment timers → 29–41 enrichment timers
  • 13 models → 12 models
  • six LifeRadar frontends → one screen with six views

A 49-claim internal audit came out 28 holds, 20 overstated, 1 refuted — I would rather hand you that ratio than a brochure.

02

The build, measured

This is the whole measured window on one table, so nobody has to assemble it from prose. Two rows carry a qualification in the row itself rather than in a footnote, because without it the number would be read as something it is not.

181 / 159 calendar days since the first commit on 2026-04-02, 159 of them with commits (as of 29.09.2026)the measured window from 2026-04-02 to the day of measurement and the distinct commit dates inside it; the appendix carries the commit count, not the day count measured · read from: dossier §1, commit dates behind [C1] [C1]

the measured window, and the days inside it that produced a commit

6,517 commits across 7 repositories (a living number since 2026-04-02 — a frozen window would not be reproducible, since merges backfill older date stamps) measured · [C1] · append-only

the git record across every repository in the family

~1.17 M / 4,271 lines of productive code across 4,271 files (~1.04 M without duplicates)generated by Claude under my direction; architecture, review and rejection are mine — orchestrated, not typed measured · read from: 4-month audit 2026-08-05, ki-nachweis:audit-2026-08-05/umfang_code [C2]

the productive code base, duplicates counted and discounted

13 products in a routing registry measured · [C3] · append-only

what the routing registry lists as a product

5,068 / 2,576 stories in a database as the single source of truth, 2,576 of them done measured · [C4] · append-only

the work state as a database rather than as markdown

1,041 / 1,079 / 741 non-forgeable adversarial review runs — devil 1,041, role panel 1,079, cleanup auditor 741 (as of 29.09.2026) measured · [C13] · append-only

checks that land where the agent that ran them cannot write

101,746 logged model calls through a single choke point — 13 models from five providers with a reply in the log (Anthropic Opus, xAI Grok, Gemini 2.5 Flash, Novita with DeepSeek/Kimi/GLM/Qwen, Exa; a 14th entry gemini-3.1-pro-preview: 8 calls, 0 replies), chosen per task; plus Claude Fable 5.1 as the default model of the workshop session itself (Claude Code, no API call, hence not in the counter) — 14 models in total measured · [C5] · append-only

every model call since 11 June through a single chokepoint

164 scheduled background runs active, 0 failed — out of 256 inventoried; 81 more start on demand (as of 29.09.2026) measured · [C12]

the scheduled runs the system drives itself, every day

$201.86 in the ledger, across 101,746 logged callsthe ledger starts 2026-06-11 and excludes flat-rate subscriptions; 68 % of rows are estimated rather than provider-billed, which is 98 % of the dollar figure. The ledger is the instrument, not the invoice. measured · [C5] · append-only

what the system logs of its own model cost — one chokepoint every call has to pass

median 33 commits a day · peak 290 · 10.5 distinct hours of the clock · 19 % between 22:00 and 06:00 · one third on weekends [C1]

That is what building a system around two small children looks like.

03

Gates that have to go red

186 unproven, 71 broken, 4,721 proven — last run 10 September 2026, 01:59 [C22]. Every night the system feeds each of 23 gates an input it knows is bad and requires the gate to refuse it. A gate that stops catching its own known-bad case is reported, not assumed. Self-reported review is not review, so each run lands in a register the agent that ran it cannot write.

A head-on line diagram of an inspection gate with its barrier arm down in the blocking position; a rejected item is stopped behind it.
brand visual, generated — not a photograph of the system

This replays the recorded verdicts: 23 of 25 from the nightly run of 2026-08-23 02:03, two from earlier runs — the date of each run is printed on its own row. It is a replay, not a live gate; the command that reads the register is below.

behavioral-conformance-gateBLOCKED · PROVEN · harness exit: 0 ·
canonical-write-path-guardBLOCKED · PROVEN · harness exit: 0 ·
code-orphan-surfacerBLOCKED · PROVEN · harness exit: 0 ·
container-isolation-proofBLOCKED · PROVEN · harness exit: 0 ·
data-boundary-reportBLOCKED · PROVEN · harness exit: 0 ·
deploy-drift-guardNOT BLOCKED · UNPROVEN · harness exit: -1 ·
meta-export-gateBLOCKED · PROVEN · harness exit: 0 ·
object-type-catalog-gateBLOCKED · PROVEN · harness exit: 0 ·
object-type-catalog-write-guardBLOCKED · PROVEN · harness exit: 0 ·
pkg:authority-tierBLOCKED · PROVEN · harness exit: 0 ·
pkg:confidence-calibration-monitorBLOCKED · PROVEN · harness exit: 0 ·
pkg:devkit-enginesNO SCRIPT · MISSING · last seen:
pkg:entity-data-auditBLOCKED · PROVEN · harness exit: 0 ·
pkg:object-type-catalog-write-guardBLOCKED · PROVEN · harness exit: 0 ·
pkg:object-type-instantiateBLOCKED · PROVEN · harness exit: 0 ·
pkg:over-merge-guardBLOCKED · PROVEN · harness exit: 0 ·
pkg:schema-catalog-builderBLOCKED · PROVEN · harness exit: 0 ·
pkg:source-independence-scoreBLOCKED · PROVEN · harness exit: 0 ·
pkg:survivorship-engineBLOCKED · PROVEN · harness exit: 0 ·
pkg:survivorship-mergeBLOCKED · PROVEN · harness exit: 0 ·
provenance-stress-guardsHARNESS ERROR · BROKEN · harness exit: 1 ·
shared-promote-gateBLOCKED · PROVEN · harness exit: 0 ·
spawn-tenant-isolationBLOCKED · PROVEN · harness exit: 0 ·
story-status-aqal-trigger-guardBLOCKED · PROVEN · harness exit: 0 ·
survivorship-resolutionBLOCKED · PROVEN · harness exit: 0 ·

The exit code printed on a replayed row is the harness exit code recorded in false_green_runs, not the gate's own. BLOCKED is what the gate did to the known-bad input; PROVEN is the harness verdict about the gate.

Beneath the grid: 56 named gates in a 97.5 KB doctrine file, 13 hard in the done-path, 44 links in the 249-line commit chain [C21]. 2,449 read-set attestations (as of 11.09.2026) proven by SHA-256, deliberately invalidated after a context compaction so the next edit blocks with exit 2 [C13]. 944 recorded devil runs in a register the agent cannot write itself (as of 11.09.2026) [C13]. The nightly cycle: 152 steps, 123 ok / 29 flagged / 3 degraded / 2 blocked [C19].

why should it stay soft, then we'll never get better 2026-06-21

Measured in the same table as the gain: review coverage of closed work went from ~15 % to 80–93 % that day, and throughput fell by a factor of four. The price is printed at the same size as the gain.

04

Why each thing exists, in the order the pain arrived

Six things, each introduced by the moment it became necessary rather than by what it does. I would rather you learned the judgment than the feature list.

The case had to hold itself.

My attorney was permanently on holiday, charged absurd fees and did not answer my questions. If the person paid to hold the case does not hold it, the case has to hold itself.

So the file became a system, not a folder: every document in, every fact with its provenance, every deadline as data, and reports an attorney can file. Provenance sits at the write path as a trigger, because a fact without a traceable origin is a claim — and a court reads the output.

the mechanism

2.46 GB append-only (as of 29.09.2026) [C6] · 118 RAISE(ABORT) write-path triggers [C7] · 135,014 lineage rows, 135,013 with the git SHA (as of 29.09.2026) [C8] · 14,393 hash-chained case events · 87,026 OCR'd pages (as of 29.09.2026); cost shares as of 29.09.2026 (denominator then: 86,846 pages): 71,232 (82.0 %) cost nothing, 1,020 needs_human, 954 tool errors, and 15,288 model pages cost $14.57 [C9]. In the same table, the loud failures: 954 cli_error, 1,020 needs_human, and 23,657 pages carrying not measurable rather than an invented score.

At a customer: it makes visible which small part genuinely needs the expert.

A handbook is a suggestion.

I read my own April corrections — 34 of them, all saying the same thing: template not followed verbatim.

Eight readable templates, and every one of them was ignorable. So the work state left the markdown files and moved into a database, and the Definition of Done stopped being prose and became a mandatory field.

the mechanism

5,068 stories with mandatory fields (as of 29.09.2026) [C4] · 1,803 dependency edges · an append-only close ledger of 239 rows across 14 named close classes, with two RAISE(ABORT) triggers preventing update and delete [C13].

At a customer: their tracker, with the close path gated.

Reins on the bull.

An agent I believed I had switched off took over my other agent's channel and answered as if it were that agent. I only noticed because the tone was wrong.

What saved it was that an incident procedure already existed — shut everything down, restart clean, verify — so I could switch it off instead of arguing with it.

the mechanism

Three flat statements since, all of them code rather than policy text: nothing sends to a human on its own; no agent's done counts without a register entry it cannot write itself; every autonomous loop has a kill path that does not depend on the loop.

At a customer: the same three, enforced in their hooks and their CI.

The bridge.

The model built a plausible version of a thing it had not understood, from a description that was itself incomplete.

Every expensive mistake in this build traced back to that gap, not to code. So the bridge became machinery: before a build, three independent agents answer what the canon says, what reality says — read, never assumed — and what the dependency graph says.

the mechanism

2,449 read-set attestations (as of 11.09.2026) proven by SHA-256 [C13] · 263 correction events, median zero days from you got this wrong to a gate in code, longest path 49 days and logged as such [C13][C17].

At a customer this is the whole job: turning “build this” into a description reality agrees with.

A session must not depend on the device in my hand.

I was pushing my child's stroller with one hand and holding the phone with the other so a session would not die.

So the sessions moved to the server, persistent behind an identity plane at the edge; a dropped connection costs a reconnect, not an afternoon. The hard part was not the terminal but the rendering — making server output read cleanly on a phone.

Three dashed device outlines with one unbroken orange line running behind them to a server rack: the devices break, the line does not.
brand visual, generated — not a photograph of the system
the mechanism

4,731 events / 310,622 transcript rows · xterm 5.5.0, WebGL addon 0.18.0, canvas 0.7.0, fit 0.10.0 pinned exactly · 16 logged research runs on this one stack · an upgrade to xterm 6 evaluated and deliberately rejected. In the same table: it runs without a systemd unit — no auto-restart.

At a customer: their identity plane and their infrastructure; I bring the retention rule and the inherited policy caps.

The event that was listed nowhere.

I wanted to know what was going on in the city on a Sunday. There was a large event that evening, listed nowhere.

A city's real calendar is not in the index, which is why the thing scrapes originals and keeps the provenance of what it found rather than repeating a listing.

the mechanism

303 places across 20+ cities · 1,055 images · SHA-256 dedup into osint_events · every amenity carrying its own source and confidence, so the screen says pool confirmed or details incomplete. Kept honest in the same table: one story still records hard-coded provenance and verified without verification on the data spine — exactly the flaw this product exists to avoid, written down rather than quietly closed.

At a customer: the same scrape → read → provenance pattern for any catalogue built on unreliable vendor data.

05

The reins

What got me into this was the hype around autonomous agents. Building taught me the opposite: to get the best out of an autonomous model you need tight reins, not a loose hand. Riding the bull is the point; the reins are what make it possible.

A flat object study of riding reins and a snaffle bit laid out for reference, one strap picked out in orange.
brand visual, generated — not a photograph of the system
The AgilConsultants marmot in a leather jacket, half out of a manhole, watching the street.

The standing limit: single operator, no four-eyes principle, no central access audit log.

  • Nothing sends to a human on its own — a hard-coded allowlist, not a review comment.
  • Every autonomous loop has a kill path that does not depend on the loop.
  • No done counts without a register entry the agent cannot write.
  • External tool output is treated as untrusted, and never executed.
  • Secrets only via .env, with a scanner in every commit.
  • Bypasses are counted, not forbidden — 43 of 43 --no-verify bypasses resolved; the ledger does not block the bypass, it blocks the next commit [C13].
Twelve hairline rules of unequal length; three crossed out by a single heavier stroke, two ending in nothing.
brand visual, generated — not a photograph of the system

06

What is open, and what failed

This section has the same type size, the same rows and the same commands as section 02, and it sits above my CV in reading order. It is live: it is fed by the same reading as the strip in the masthead, so if the restore drill goes green this page says so, and until then it says this.

sample data — not a live probe

  • DR restore drill red, ExecMainStatus=1 — the backup exists, the restore is not currently proven [C20].
  • Litestream staleness monitor in ALERT, replica WAL timestamp unreadable [C20].
  • The 19 GB law corpus out of live replication, only a weekly S3 snapshot, last one 2026-08-16.
  • 95 timers enabled but dead [C12].
  • Evals: 5 cases, unscheduled; no model tracing, no span trees, no prompt versioning.
  • OpenTimestamps: 0 of 72 anchors re-verified — stated as anchored, not re-verified.
  • The content and video pipeline resting; the renderer failed 4 of 4 nights while reporting exit 0 — a silent failure inside the pipeline that claims to fail loudly, and the one claim the internal verification pass flatly refuted, kept here for that reason.
  • Investments stalled with two open backtest defects I found in my own code.
  • 560 house-search entries with no matching acknowledgement [C24].
  • Tenant isolation resolver-level only; 564 files still carry the old shared-database string.
  • 4 of 7 repos have no commit-hook chain, 30 of 51 gates run only on my machine — the portability bottleneck.
  • Single operator, no four-eyes principle; tx_to is never set, so the schema is bitemporal and the usage is not.

07

Where the data goes

This is usually the first question, so it is on the page rather than in a call. On 23 August my own research engine ran against my own configuration and found four things. I would rather you read them here than find them later.

What the system does

  • Purpose-typed delivery with default-deny at the delivery boundary.
  • Code names instead of third-party names in every output.
  • A personal-data purge on the deliver path.
  • IBAN redaction across 21 importers.
  • Deletion-marking instead of DELETE.
  • A per-source licence registry.
  • An Article 17 erasure log in WORM storage.
  • Article 25 privacy-by-design evidence generated from live state.
  • An SSRF-guarded fetch.
  • HMAC-signed SSO that fails safe without its secret.
  • Cloudflare tunnels instead of open ports.
  • The hard-coded rule that no email and no message ever leaves without a human.

The four findings of 2026-08-23

A regional endpoint is not a residency guarantee. My own system was running europe-west4. EU residency requires the EU multi-region endpoint with Private Service Connect. Fix story US-LEGAL-VERTEX-EU-MULTIREGION-ENDPOINT-DATENRESIDENZ-01 is open.
Zero data retention is per organisation and does not inherit. It has real edges, so a deletion concept written to “30 days” is wrong.
Who the processor is depends on the procurement path. The data processing agreement does not reach across it. Never mix the two paths without documenting both.
Certifications present and absent. Present: SOC 2 Type 2 (2025), CSA STAR L2, SOC 3, ISO 27001 (2025), ISO 42001 (2025), NIST 800-171r3. Absent: ISO 27017, ISO 27018, ISO 27701, TISAX, PCI DSS, BSI C5.

Stated as of 2026-08-23, with the instruction to re-check before contracting rather than to rely on this page.

08

At a customer, I deliver with your stack

The honest per-component account: what I use here, what I do not, and what I would reach for on your side of the table. I built my own where the official piece did not exist yet or did not go deep enough, and I adopt the official piece at the visible edge, as a thin wrapper, never as a rebuild.

scrolls sideways
ComponentUsed today?What I use at a customer
Claude CodeyesSame — the delivery vehicle.
where, and why or why not

Where 49 commands, 26 parallel worktrees

Why or why not The whole system was built in it.

HooksyesSame, plus the git chain.
where, and why or why not

Where 20 + 18 files, 5 registered events, PreToolUse blocking with exit 2

Why or why not The hook surface is where the discipline actually lands, before the work exists.

Skills in the official formatyesCustomer doctrine as skills, not as a wiki page nobody opens.
where, and why or why not

Where 49 + 41 files

Why or why not Doctrine the model must load is a skill, not a README.

Sub-agentsyesSame, with a refutation default.
where, and why or why not

Where 944 / 806 / 664 recorded runs

Why or why not Adversarial review by a fresh agent is the highest-value technique here.

Messages APIyesSame SDK, unchanged.
where, and why or why not

Where @anthropic-ai/sdk ^0.96.0

Why or why not Direct control over cost accounting per call.

Tool use and structured outputspartlyTool use as the API offers it; my schema gate shrinks to a verification layer.
where, and why or why not

Where Own JSON-Schema abort

Why or why not Documented in the code as a deviation, because no response_format existed.

Prompt cachingpartlyPrompt caching everywhere by default.
where, and why or why not

Where TypeScript path only

Why or why not The largest untouched cost lever, and I left it there.

Batch APIpartlyBatches wherever the job is genuinely offline.
where, and why or why not

Where Adapter at NotImplementedError('wave 2')

Why or why not The nightly cycle is latency-bound; batching helps cost and hurts that loop.

Files APInoFiles API for ordinary document workflows.
where, and why or why not

Where —

Why or why not Documents arrive through S3 with WORM triggers and a provenance chain.

MCPconsumed, not authoredAuthor MCP servers — the visible edge that speaks the customer's language.
where, and why or why not

Where Two configured servers

Why or why not A private chokepoint is right for one operator and wrong for a customer who needs an inspectable tool boundary.

Agent SDKnoAgent SDK at customers, with my gate patterns as thin wrappers.
where, and why or why not

Where —

Why or why not My orchestrator and its registers predate it; I would rather port a pattern than a codebase.

Claude on Vertex / BedrockpartlyBoth, chosen per customer procurement path.
where, and why or why not

Where Vertex in the router; Bedrock not

Why or why not Chosen for the EU argument — and section 07 is where that choice was proved wrong.

Extended thinkingpartlyExtended thinking with an explicit budget and cost attribution.
where, and why or why not

Where Thought tokens metered into the ledger

Why or why not Used where it pays, never silently: uncounted thinking tokens are a real bill.

Deviations I would undo today

  • Ad-hoc HTTP routing instead of MCP tool contracts — my tool boundary is private and undocumented.
  • My own structured-output hard gate instead of API tool use — it should shrink to a verification layer.
  • My own prompt-role panels instead of Agent SDK sub-agents — the persona design is worth keeping, the runtime underneath it probably is not mine to maintain.
  • Prompt caching only on the TypeScript path — the router leaves the largest cost lever unused.
  • The Batch API adapter left at “wave 2” — the nightly steps that are not latency-bound belong in batches.

09

The two decades the build record cannot show

Everything above is one person and a model. This is the other twenty years, and it is the one part of this page you cannot verify from a terminal: it carries no commands, only references on request.

  • 2018–2020Deutsche Bahn RIS Mobil, 45+ people — Scrum Master of two teams toward self-organisation, hired and built a third team and a DevOps team, agile process adapted to high-security requirements, the app in use across Germany.
  • 2017–2018Deutsche Bahn RIS Fahrzeug, 20+ people, three teams, product management built up.
  • 2015–2016Tourism go-live, 70+ people, live after eight years of development.
  • 2014–2015Public services, 100+ people, eight agile teams heading for twelve where LeSS hit its limits; prepared the scaling toward SAFe and built the visual centre.
  • 2009–2013Commerzbank, 50+ people — started as a senior developer, found the V-model too far from the customer, built four agile teams as Scrum Master while leading the frontend overhaul.
  • 2005–2009Dresdner Kleinwort Wasserstein / Dresdner Bank — introduced the web limit-check service that validates risk limits at deal closing.
  • 2002–2005Dual study programme, thesis on evolutionary algorithms, grade 1.0, which is where the interest in search over unbounded solution spaces began.
  • 2021–2026Parental leave, two small daughters, Agile Disruption written (2024, 41 pages, 10 chapters), the systemic-constellation training completed.
  • since 2016Managing director of AgilConsultants.

The book → system bridge

  • idiot index → question 3 of the seven-question decision filter
  • every requirement carries a name that defends it → 49 dated quote sites in the doctrine file, 56 of 62 distilled principles verbatim
  • delete until you have to restore 10 % → a 26.8 % backlog cut with an append-only close ledger and a --reopen path
  • automate last → a gate goes SOFT, then proven by mutation, then HARD, then scheduled — never straight to automation

The time building alone with a model was not a retreat from teams; it was the fastest way to learn the material I now want to bring back into one.

Before 2016 — engagements I can name on request

  • Go AgilAn agile transition run as a programme, not as a training course.
  • Scale AgileScaling past the point where one framework stopped fitting.
  • Bug Fixing MarathonA fixed-length push against a defect backlog, measured daily.
  • Sex IT UpMaking an internal IT department something its own colleagues wanted to work with.

10

Check it yourself, and what I will not hand over

I would rather you touched the system than read about it. The refusal comes first, because it is the part that took judgment: my own security review failed my first draft of this section — a plain terminal login on the production host would have been a shell as the database-owning user. So: no shell on the production host, ever.

Three offers, in ascending cost, each with its limit named first

from my side of the wallA recorded 10-minute walkthrough and a 30-minute live screen-share. available immediately
no case content, no personal quotes, secret-scanned over the full history and scanned for personal data before hand-overA sanitised read-only repository — gates, hooks, doctrine, the mutation harness. within days
not my production host, not my identity door; a non-privileged user, transcript retention off, a hard expiryA dedicated evaluator sandbox on a throw-away server with a synthetic DDL-only tenant database (prove-isolation=0 [C14]). scheduled work, not a switch

Three walkthrough tracks

problem → stories → clickdummy → deployThe demo of speed, not a product I would sell you.
story → gates → deployYou name a small feature and watch the gates fire, including, most likely, one that blocks me.
document → provenance → reportDrop in a synthetic document, watch the cascade decline the expensive model, then click any sentence in the report back to the page it came from.

Contact

Write to me and say which of the four doors you came through. I answer myself.

[email protected]

Imprint (DDG § 5)

Operator
AgilConsultants
Address
Höhestrasse 48, 61348 Bad Homburg, Germany
Represented by
Christian Wolfgang Honl, managing director
Register
Amtsgericht Frankfurt am Main, HRB 16085
VAT ID (§ 27a UStG)
DE298914143
Telephone
+49 (0) 152 / 343 455 61
Responsible for editorial content (§ 18 MStV)
Christian Wolfgang Honl, address as above

Pre-filled from the 2016 site. Every row above must be verified before publication.

The live elements on this page are first-party and read-only. Nothing non-essential is stored, no cookie is set and no third party is contacted, which is why no consent banner appears. The fonts are served from this domain.

Imprint · Data protection