By Yianna Kokalas & Claude
How this paper was written. This is a working paper from a production side project: software built for the Magic: The Gathering community, with real users, built and operated by one engineer (an enterprise engineer by day) together with Claude, Anthropic’s AI model. Every mechanism claim was fact-checked against the running code before publication. The audit found nine places where our own documentation had drifted, and two claims we had repeated for months were tested and killed by measurement before this went out; all of it is published in sections 7 through 9 rather than quietly fixed. Both names on the byline mean it.
Position: the working unit of serious agentic engineering is not one strong agent with a long leash. It is a supervisor delegating to role-scoped subagents, each with its own context window, its own permissions, and a model chosen for the risk of its specific job. The supervisor holds the plan and the verdicts; the subagents hold the work. This is the same trust arc the previous paper in this series walked (The Drain): autonomy is extended only as far as verification allows, and role boundaries are what make verification tractable. A reviewer that shares no context with the implementer cannot inherit its blind spots. A fix agent that cannot merge cannot bury its own mistake.
This paper is also an audit, and the audit question comes first:
Do role-scoped subagents on deliberately different models beat one strong agent, and is each of our model assignments still evidence-justified, or a habit nobody re-measured?
The honest answer, argued in section 7: the role-scoping is incident-justified and we would rebuild it; the model assignments mostly are not. Of 33 model-assignment slots in the system, 18 carry no recorded evidence at all, 4 are backed by real measurements (all of them the local models, not the frontier ones), and the rest are recorded decisions or incident-backed pins that were reasoned once and never measured. One flagship topology has direct measured counter-evidence against itself, and the skill now says so in its own body. We print all of it. The paper’s response to the unevidenced rows is not backfilled justification: section 7.1 restates each contested assignment as a falsifiable claim with a named kill condition and the live alternatives beside it. Two of those claims were tested before publication. One killed a number from the previous paper in this series. The other killed the system’s most-repeated security rationale at 0 refusals in 12, for thirteen dollars.
Every mechanism claim below was verified against the code and skill definitions at HEAD on 2026-08-30. The file:line citations throughout reference our private orchestrator code and skill definitions, so a reader cannot follow them; they are retained anyway, as evidence that each claim was checked against a specific line of running code rather than recalled from memory.
1. The shape of the unit
Some vocabulary, so the paper reads cold. The system is a planning vault: a git repository of markdown ticket specs, plus skills (versioned prompt programs a session can invoke) and a headless dispatcher called the drain that turns approved tickets into merged pull requests with no human in the loop. A “supervisor” here is the one session that owns a run: it reads the ticket, spawns every worker agent itself, scores their results, and is the only writer of consequential state. A “subagent” is a worker spawned for exactly one role, with a fresh context containing only what the role needs.
Three properties define the unit, and none of them is stylistic:
- The supervisor spawns everything. The harness blocks subagents from spawning
subagents, so delegation is one level deep by construction (
implement/SKILL.md:11). This is a constraint we chose to lean into rather than fight: a flat topology means there is exactly one place where verdicts are aggregated and exactly one agent whose claims must be re-verified against ground truth. - Roles are scoped by permission, not by politeness. Investigator agents are
read-only by convention and by tool restriction (
agents/index.md:3). The implementation agent may not merge its own pull request or advance ticket state (implement/SKILL.md:218). The fix agent fixes; it does not re-litigate the review (“it is a fixer, not a critic”,implement/SKILL.md:281). When a role needs isolation from the supervisor’s own working tree, it builds a separate git worktree and works there. - Models are assigned per role, per risk. The strongest models go where judgment is scarcest (adversarial review), cheaper models go where the work is bounded (mechanical synthesis, batch classification), and one entire task family runs on local open-weight models where the economics demand it. Whether each assignment is still justified is the audit of section 6.
The trust-ladder frame carries over from the previous paper: none of this delegation was
trusted on day one. Skills ran interactively with a human watching before any supervisor
was allowed to run them over subagents, and subagent output is treated the way the drain
treats dispatch output: as a claim, verified against ground truth (git log, GitHub PR
state, the test suite) before anything advances. The specific rule “confirm via git, not
via the agent’s say-so” was earned: named background subagents were observed going idle
and never returning result blocks, so the flagship pipeline spawns its agents unnamed
and synchronous and verifies their work products directly (implement/SKILL.md:203,275).
2. The flagship: the /implement pipeline
/implement executes one ticket end to end: implementation, two independent reviews, a
bounded fix loop, then merge eligibility. It is the pipeline the drain dispatches
headlessly, which means every role boundary in it must hold with nobody watching.
supervisor (thin dispatcher)
reads ticket, spawns all agents, scores verdicts,
only writer of ticket state
|
+------------------+
|
v
IMPLEMENTATION AGENT
fresh context, builds its own
git worktree, implements the
ticket, runs the tests
model: inherits the dispatcher
may NOT merge or advance state
|
| work product verified via git, not via self-report
v
+---------------------------+---------------------------+
| |
v v
CRITICAL REVIEWER SECURITY REVIEWER
correctness, design, tests auth, injection, secrets
model: Fable interactively, model: Opus in ALL modes,
Opus under headless drain deliberately never Fable
(cost); on refusal or death, (false-refusal risk on the
re-spawned on Opus; a missing exact auth/JWT/credential
review is NEVER scored PASS code it must read)
| |
+---------------------------+---------------------------+
|
both PASS? | findings?
v
FIX AGENT (model: sonnet)
applies review findings only;
merge-state gate runs first so a fix
never lands on an already-merged PR
|
v
BOTH REVIEWERS RE-RUN, FRESH CONTEXTS
(max 2 fix cycles, then verdicts decide)
|
v
merge eligibility: human opt-in flag AND
validation PASS AND tree-integrity PASS
(missing = fail closed) AND both reviews PASS
|
v
gh pr merge --auto --squash
(the pipeline never writes "done";
the drain does, after deploy verification)
The texture is in the gates, and each gate has a scar:
- Fresh re-reviews. After a fix cycle, both reviewers re-run with fresh contexts
rather than being asked “is it fixed now?” (
implement/SKILL.md:401). A reviewer that carries the prior conversation has already committed to a framing; a fresh one has not. - The merge-state gate. Before every fix cycle the supervisor checks actual PR state,
because in June 2026 a fix was committed to a branch whose PR had already merged
(
implement/SKILL.md:339-348). Findings against a merged PR become a follow-up ticket, never a commit. - PASS-with-warnings is PASS. An operator ruling dated 2026-08-19, made after a green
PR sat unmerged for a day over advisory nits (
implement/SKILL.md:402). Only an explicit FAIL blocks. - A missing review is never a pass. If the critical reviewer’s model refuses or dies,
the same prompt re-runs on Opus; the absence of a verdict is treated as the most
suspicious outcome, not the cleanest (
implement/SKILL.md:279). - The self-collision ban. Once preflight passes, re-deriving “is another implement
running” from process listings is forbidden: on 2026-08-18 three headless dispatches
found their own process with
ps, concluded a collision, and self-aborted with zero work done (implement/SKILL.md:126).
Why two reviewers instead of one stronger one? Because the two roles fail differently.
The critical reviewer’s failure mode is missing a subtle bug, so it gets the strongest
available judgment. The security reviewer’s failure mode is refusing to engage at all:
the model family that is strongest at general bug-finding carries safety classifiers
that can false-refuse on legitimate auth, JWT, and credential code, which is precisely
the code a security reviewer must read. So the security seat is pinned to Opus in all
modes, and the strongest model is deliberately kept out of it
(implement/SKILL.md:269,280). That is the rationale as the code records it. Honesty
requires the caveat up front: the false-refusal claim was recorded in three places with
no dated incident artifact behind it, and when this audit finally tested it (section 7.1),
Fable reviewed 12 of 12 auth-adjacent diffs without a single refusal. The exclusion
survives on cost alone. The diagram above shows what the code says; section 7.1 shows
what the measurement says.
3. The team shape: /plan-create
/plan-create turns a discovery report into reviewed, implementation-ready ticket
specs. Where /implement is a serial pipeline with parallel review, /plan-create is a
team: creators, reviewers, fixers, and re-checkers, phased, with workers shut down
between phases so no context bleeds across roles.
- Creators (model: opus) write ticket specs in dependency waves: backend tickets
before the frontend tickets that consume their APIs (
plan-create/SKILL.md:218-224). A creator that discovers a missing sibling ticket may surface up to three evidence-gated discovery blocks, but only the supervising lead mints files: one writer per file, always (plan-create/SKILL.md:246-267). - A deterministic lint gate runs between creation and review: structural problems
(missing sections, invalid frontmatter) are caught by a script, not by a model, so
reviewer attention is spent only on judgment (
plan-create/SKILL.md:273-293). - Reviewers (model: opus) run in parallel, one per ticket, thirteen dimensions each.
- Fixers (model: sonnet) run only on tickets scored NEEDS_FIXES, and get exactly one
cycle; a ticket still failing after fix and re-review is reported for manual
attention, not looped (
plan-create/SKILL.md:337-345,564). - Every wave arms a stall alarm (600 seconds, one per spawn batch), because the
documented failure mode of background agents is not crashing loudly but going quiet
(
stalled-agent-protocol.md:12).
The halt boundary is the same one the drain honors: created tickets land at
status: todo, never execution-ready. Promotion to the auto-merge queue is a human act,
in both systems, always (plan-create/SKILL.md:265,434).
4. Read-only investigators, and the topology that failed
The cheapest role in the system is the read-only investigator: an agent that can search and read but not write, spawned when a question spans more files than the supervisor wants in its own context. The scoping does real work: an investigator cannot “helpfully” fix what it finds, so its findings arrive as claims with citations, and the supervisor decides what becomes a ticket.
The fan-out pattern peaks in two skills:
- /plan-explore spawns up to four parallel investigators (code search, docs, user
scenarios) on sonnet, then hands the draft to a two-agent synthesis pair: an Analyst
on the default model and a Critic escalated to the strongest model, because
adversarial validation of a discovery draft is the highest-judgment step in the flow.
The debate is capped at two rounds and the lead decides (
plan-explore/SKILL.md:85,283). - /plan-review-swarm runs five adversarial reviewer personas in parallel (security
and auth, test coverage, codebase fact-checking, race conditions and state, operator
validation), each capped at three findings so the synthesizer dedups signal rather
than noise. The personas run on sonnet, with a recorded rationale: “Haiku is too thin
for adversarial-with-grep; Opus is overkill per persona”
(
plan-review-swarm/SKILL.md:63,99). It has shipped as a spike since 2026-05-26 and is documented as augmenting, never replacing, the sequential review.
And then there is /debug-swarm, which this audit is obligated to report as the honest
counter-example. The design is four blind investigators on opus (reproduction,
telemetry, git history, hypothesis), each forbidden to see the others’ work, with a hard
evidence rule: a conclusion with an empty evidence array is demoted to hypothesis, and
“not reproduced” requires the exact commands and output of the attempt
(debug-swarm/SKILL.md:12,121,244). The rule is good. The topology, measured, was not:
on 2026-07-20 a four-investigator swarm returned zero of four result blocks, all four
going idle, and the entire root-cause analysis was produced by working the angles
inline. A second bug investigated inline from the start reached a sharper answer in
about 15 minutes (debug-swarm/SKILL.md:56-58). The skill’s own body now steers bounded
bugs inline and reserves the swarm for genuinely parallel evidence-gathering. That
correction is the audit method working: the measurement beat the architecture diagram,
and the diagram yielded.
5. Models chosen per risk: how assignment actually works
Model assignment in this system is not one decision; it is 33 small ones, made in skill definitions, agent frontmatter, and dispatcher code. The mechanisms worth naming:
- The dispatch pin. Every headless dispatch passes an explicit
--model(default sonnet). This exists because on 2026-08-19 a bare dispatch inherited the operator’s interactive model selection and silently ran the drain on the most expensive model available, at roughly twice Opus cost, until the round errored out (bin/drain:1257-1269). The planning pipeline pins its phase dispatches identically (pipeline_runner.rb:235-239). The pin is evidence-backed. The choice of sonnet specifically is precedent (“every prior healthy dispatch had run on sonnet”), not measurement. - The mode gate. The critical reviewer runs the strongest model interactively but
drops to Opus under headless dispatch, trading peak judgment for cost on the path
that runs unattended many times a day (
implement/SKILL.md:264-265). The always headless fixer pipeline resolves the same gate to Opus permanently, keeping the three-attempt self-heal loop off the 2x-cost model (fixer-agent/SKILL.md:93-99). - The security exception. The strongest model is excluded from the security seat in all modes, as described in section 2.
- The local family. The backlog-hygiene classifiers (freshness, dedup, dispatch
triage) run on local open-weight models via Ollama, pinned in a code registry with
the benchmark table reproduced beside the pin (
hygiene_models.rb:18-73). This is the only corner of the system where model choice was measured properly: a 2026-06-01 benchmark of six models on 20 hand-labeled tickets, and a 2026-08-21 re-benchmark on a 70-entry hand-verified corpus. It is also where measurement delivered the most uncomfortable result, reported in section 7.
6. The model-assignment inventory
The audit’s core deliverable. Every model-bearing role in the system, with the evidence for its assignment or the honest entry “habit, no recorded evidence”. “Reasoned, no measurement” means a rationale is written down but nothing was ever measured against an alternative. Full citations per row are in the internal fact sheet.
| Role | Model | Evidence for the choice |
|---|---|---|
| Drain-dispatched /implement session | sonnet (pinned) | The pin: incident 2026-08-19 (silent drift to the 2x-cost model). The choice of sonnet: precedent, not measurement (bin/drain:1257-1269) |
| Planning-pipeline phase dispatches | sonnet (pinned) | Same incident, cited at the pin (pipeline_runner.rb:235-239) |
| /implement implementation agent | unpinned (inherits dispatcher) | Habit, no recorded evidence for leaving it unpinned |
| /implement critical reviewer | fable interactive / opus under drain | Decision 2026-07-01: Fable is 2x Opus cost; “Opus already SOTA at bug-finding” is asserted, not measured |
| /implement security reviewer | opus, all modes | Recorded false-refusal rationale, three places; no dated refusal incident artifact found; tested and unsupported (section 7.1) |
| /implement fix agent | sonnet | “A fixer, not a critic”: stated rationale only; habit |
| /fixer-agent both reviewers | opus | Derived from the mode gate + cost rationale; no measurement |
| /plan-create creators | opus | Habit, no recorded evidence |
| /plan-create reviewers | opus | Habit, no recorded evidence |
| /plan-create fixers | sonnet | Habit (mirrors the fixer-lighter pattern) |
| /plan-explore investigators | sonnet | Stated rationale, no measurement |
| /plan-explore Critic | fable (opus fallback) | Decision 2026-07-01: highest-judgment step; reasoned, no measurement |
| /plan-explore Analyst | default (unpinned) | Habit, no recorded evidence |
| /debug-swarm investigators x4 | opus | Habit for opus specifically; the topology has measured counter-evidence (0 of 4, 2026-07-20) |
| /plan-review-swarm personas x5 | sonnet | Reasoned (“Haiku too thin, Opus overkill”), no measurement |
| /plan-review-swarm synthesizer | sonnet | Reasoned (“mechanical”), no measurement |
| /plan-stale-check workers | haiku | Reasoned (“bounded, deterministic”), no measurement |
| /plan-triage workers | sonnet | Habit, no recorded evidence |
| /fable-review | fable | Decision 2026-07-01: human-invoked, one artifact, bounded cost |
| /drain-select | fable | Decision 2026-07-01: one call per session, no security code in scope |
| /copy-review editor | fable | Habit, no recorded evidence (no decision artifact found) |
| /support-triage classifier | fable (opus fallback) | Habit; follows the copy-review pattern by its own admission |
| /support-reply drafter | fable (opus fallback) | Habit, no recorded evidence |
| /doc-gen main agent | “Opus 4.6” (stale string, unenforced) | Reasoned (“model quality IS doc quality”); no measurement; version string stale |
| /doc-gen article reviewers | sonnet | Habit, no recorded evidence |
| Nine skills’ worker subagents (decide, plan-discuss, plan-fix, plan-finalize, plan-prioritize, plan-review, plan-batch-review, scenario-draft, scenario-explore) | sonnet | Habit: the uniform house default |
| user-scenario-researcher agent | opus in frontmatter, sonnet at its main call site | Habit, and self-contradictory (section 8) |
| /summary | opus | Habit, no recorded evidence |
| Hygiene freshness classifier | gemma4:e4b (local) | Measured twice: 2026-06-01 (6 models, 20 tickets) and 2026-08-21 (70-entry corpus) |
| Hygiene freshness escalation | qwen3:30b-a3b (local) | Measured: 100% tool-call reliability; sibling model rejected for 2 HIGH-confidence false-cancels |
| Hygiene dedup + drain triage classifier | gemma4:e4b (local) | Same benchmark lineage; an entire model family hard-banned in code on a measured 0% tool-call reliability |
Tally: 33 slots. 18 habit with no recorded evidence. 4 measurement-backed, and all four are the local models. The remainder are recorded decisions or incident-backed pins whose reasoning was written down once and never tested against an alternative.
The pattern is worth stating plainly: the cheaper the model, the better its evidence. The local classifiers were benchmarked twice on hand-labeled corpora because they were suspected of being weak. The frontier-model assignments were never benchmarked because they were assumed to be strong. Suspicion produced measurement; confidence produced habit. One caveat on the word: “habit” throughout this inventory describes the state of the record, not necessarily the state of the reasoning. Section 7 returns to the distinction.
7. The audit answer, honestly
Do role-scoped subagents beat one strong agent? The role-scoping itself is the best-evidenced part of the architecture, but the evidence is incident-shaped, not comparison-shaped. Fresh-context re-reviews, the merge-state gate, the one-writer rule, read-only investigators, the never-score-a-missing-review-as-PASS rule: each exists because a specific failure occurred and the role boundary is the repair. We would rebuild all of it. What does not exist is a controlled comparison against the alternative: no experiment ever ran the same ticket population through a single strong agent with a long context to see what the pipeline’s overhead buys. Nor was the pattern adopted blind: the operator attests it was weighed critically against what contemporary open-source orchestration frameworks were doing when it formed. That comparison lives outside this record, so this paper does not cite it as evidence; it is noted so the absence of an artifact is not misread as an absence of thought. And where a topology was measured against inline work, the topology lost: the debug swarm’s zero-of-four (section 4) is the only direct topology measurement on file, and it favored the single agent for bounded bugs. The honest position: role separation earns its keep through verifiability (each gate is checkable in isolation) rather than through any measured quality delta, and fan-out parallelism specifically is overclaimed relative to its evidence.
Are the model assignments evidence-justified? Mostly no, per the inventory. Three specific embarrassments deserve print:
- The tiering story we tell is not what runs. The recorded pattern is “the fix
agent runs one tier lighter than the implementation agent.” But the implementation
agent is unpinned and inherits the dispatcher, which under the drain means sonnet:
the same tier as the fix agent. The tiering only exists in interactive sessions.
Nobody decided that; it fell out of an unpinned spawn call
(
implement/SKILL.md:194-201). - Reviewer value went unmeasured for three months despite being instrumented, and
the first measurement (run on review of this draft, 2026-08-30) rewrote the
question. The measurement found: a reviewer produced an actionable finding on
26.8% of 347 dispatches lifetime and 77.3% since August; the critical reviewer was
the sole source of 28 of the 39 hard FAIL verdicts, the security reviewer of 2; and
the obvious query is a trap, because the telemetry field
implement_review_verdictrecords the verdict AFTER fix cycles and therefore reads PASS on 12 of 12 completed dispatches while the same dispatch logs carry 160 hard reviewer FAILs. The field measures survival, not catching, and its own commit message claims the opposite of what it does. The previous paper’s “fix-cycle counter read 0 on 55 of 55 rows” was itself an artifact twice over: a field-identity mistake (that counter belongs to the post-deploy self-heal fixer, and /implement’s own counter shipped only on 2026-08-28) and a telemetry partition that is 99.3% content-free noise rows, now ticketed. Two gaps survive the measurement: the Fable interactive critical path has zero telemetry coverage, and nothing can yet attribute a catch to a seat by query (the prerequisite instrumentation is ticketed). - The one properly measured classifier is a no-op. The 2026-08-21 hygiene
re-benchmark recorded the production freshness pin at 42.9% accuracy with 0 of 40
non-fresh tickets recalled: it stamps everything fresh. The decision record calls
itself “a DISPROOF, not a selection” (
hygiene_models.rb:28-41). The gate machinery around it is real and fail-closed; the classifier inside it currently discriminates nothing. The docs that present the sweep as a working filter are ahead of the evidence (section 8).
Claims this paper deliberately does not make: that Fable false-refused on our security code on a specific date (asserted in three places, no incident artifact; tested on review of this draft and unsupported, see the next section), that Opus is “SOTA at bug-finding” (uncited capability claim in a decision record), and any per-seat reviewer catch rate (the field that could carry it does not exist yet).
7.1 From habit to falsifiable claims
Forcing retroactive justification onto the unevidenced rows would repeat the failure this audit exists to catch: a plausible rationale backfilled after the fact is indistinguishable from drift, and it forecloses the question instead of answering it. The operator’s direction on review of this draft (2026-08-30) is the alternative this section adopts: restate each assignment as a claim stated precisely enough to be killed, name the cheapest test that could kill it, and put the live alternatives beside it with their tradeoffs. A claim nobody has tested is not evidence. But it is honest in a way a backfilled rationale is not, because it names exactly what would change our mind. The hygiene benchmark set the precedent that a disproof is a first-class outcome; these are written to die the same way.
| Seat cluster | The claim, stated so it can die | Cheapest falsifier | Alternatives, with tradeoffs |
|---|---|---|---|
| /plan-create creators + reviewers (opus) | Opus-authored specs fail review less, and survive implementation with fewer fix cycles, than sonnet-authored ones would | Run one discovery’s ticket batch with sonnet creators; compare NEEDS_FIXES rates from the review scoring that already exists | Sonnet at a fraction of the token cost; the risk is spec gaps that surface downstream, where they cost more than the savings |
| /implement critical reviewer (fable, interactive) | Fable finds confirmed bugs Opus misses often enough to cover its 2x cost | Both models review the same N merged diffs; count unique confirmed findings per side. The 08-30 measurement covered the opus drain seat (sole source of 28 of 39 hard FAILs: the seat itself earns its keep) but the fable path has zero telemetry coverage, so the 2x question is fully open | Opus everywhere (already the drain-mode behavior); half the cost, and if the delta is near zero the interactive escalation is theater |
| Security-reviewer exclusion (never fable) | Fable false-refuses on legitimate auth, JWT, and credential diffs | RUN 2026-08-30, claim killed. 12 merged auth-adjacent diffs (JWT decode, secret comparison, HMAC webhooks, OAuth consent, CSRF, OTP), the real security-reviewer prompt, $12.97: 0 refusals, 0 hedges, 0 errors. Fable returned FAIL with critical findings on the four diffs with the most authorization surface, engaging hardest exactly where the claim predicted a flinch. Today’s rate cannot retro-test July, when the claim was recorded; it kills the claim’s current basis | The exclusion’s stated basis is dead; cost alone (2x Opus) still argues for Opus in the seat, and is now the only stated basis. Whether to keep the exclusion is the fable-vs-opus reviewer comparison above, not a refusal question |
| Security-reviewer tier (opus, upgraded from sonnet 2026-07-01) | The opus upgrade catches findings sonnet would miss on auth-adjacent diffs | Parallel sonnet and opus security reviews on the same diffs for one drain week; compare findings directly. Measured context from 08-30: sole-source FAIL on 2 of 347 dispatches, but this is a tail-risk seat, so hit rate is close to the wrong metric and the experiment must compare findings, not counts | Sonnet at lower cost; the risk is exactly the miss that does not happen, which no hit-rate number prices |
| /implement fix agent (sonnet) | Fix quality is bounded by the review findings, not the fixer’s tier, so re-review pass rates match across tiers | Sonnet vs opus fix cycles compared on re-review pass rate; the counter wiring is now verified sound (14 of 14 cross-validation), so this is runnable once per-seat verdicts ship | An opus fixer pays critic prices for directed work if the claim holds. Cost note, stated carefully: 77.3% of August dispatches run at least one fix cycle, and cycle-bearing dispatches average ~$13-14 against ~$5 without (n = 14, directional at best). That delta buys the whole loop, not the fixer: each cycle re-runs BOTH reviewers on fresh contexts, opus twice under drain, so the re-reviews are plausibly the larger share of it. Cost is only measured per dispatch; no per-seat attribution exists |
| Uniform sonnet worker default (investigators, swarm personas, the nine-skill cluster) | Grep-anchored bounded investigation saturates at sonnet: opus adds no findings that survive synthesis, haiku loses some | Downward: haiku on the same persona prompts, count findings surviving the synthesizer. Upward: opus personas on the same tickets | Haiku at a fraction of the cost (the stale-check precedent) vs opus (“overkill per persona” is reasoned, untested) |
| /plan-stale-check workers (haiku) | Signal extraction plus grep is bounded enough that model tier does not change verdicts | Sonnet on the same five-ticket chunks, diff the verdicts: our existing fixture discipline, applied to a model | Sonnet is a pure cost increase if the claim holds |
| /implement implementation agent (unpinned) | None. This row is an accident, not a claim; the first step is making it a decision at all | Pin the tier explicitly, then compare opus vs sonnet implementation on review verdicts and fix-cycle counts | Inheriting the dispatcher means the tier silently changes with the operator’s interactive model selection, which is how the 2026-08-19 incident started |
| Fable in the copy and support seats | Operator-voice drafting is where the strongest model’s judgment actually shows | Blind operator read of paired Fable and Opus drafts; n is small, but the operator is the only judge that matters for these seats | Opus at half cost; these seats are low-volume, so the spend difference may be noise either way |
Run order matters more than coverage. The security-refusal corpus test is nearly free and settles a claim this audit found asserted three times with zero artifacts. The reviewer comparison attacks the two most expensive seats and rides on the telemetry work already ticketed. The rest can wait for a reason to exist: an untested falsifiable claim is still a better artifact than a backfilled rationale, because it carries its own kill condition.
8. What we believed vs what the code does
The audit rule from the previous paper applies: every mechanism claim gets checked against code before it gets printed, and the drift found becomes corrections or tickets before the paper does. Nine findings this round:
- Drain telemetry lies about the dispatch model by default. Dispatch runs sonnet
(
bin/drain:1267); telemetry records an Opus-class label from a second constant with a different default (bin/drain:2282). Same env var, different fallbacks: any per-model cost slicing with the env unset is wrong. Already ticketed by the previous paper’s audit, still open. - “Fable is reserved for two skills only” is stale. A memory note and a comment in
bin/drain:1261both say it; the code pins Fable in six additional seats (implement’s interactive critical review, plan-explore’s Critic, copy-review, support-triage, support-reply). The “never for security review” half remains true everywhere. Memory corrected with this paper. - One agent definition contradicts its own call site.
user-scenario-researcherpins opus in frontmatter; its primary caller spawns it with sonnet under a blanket rule (agents/user-scenario-researcher.md:4vsplan-explore/SKILL.md:85,232). One of the two is not what runs. Ticketed. - The agents index lists two agents that do not exist on disk, both artifacts of retired integrations. Ticketed with the doc sweep.
- /plan-explore still prompts for a wiki that was deprecated and tells its
synthesizer to cross-reference a tracker the same skill elsewhere says to skip
(
plan-explore/SKILL.md:198,207,308). Same sweep ticket. - /doc-gen mandates “Opus 4.6”, a version not in the current lineup, and enforces
the mandate nowhere (
doc-gen/SKILL.md:3,14). Same sweep ticket. - The hygiene sweep’s docs are ahead of its evidence. Project notes present the sweep as a working freshness filter; the model registry itself records the classifier as effectively a no-op (finding 3 of section 7). The docs need the disproof stated where the confidence is. Same sweep ticket.
- /debug-swarm’s headline oversells its own default. The description advertises four investigators; the body, corrected after the 0-of-4 measurement, steers most bugs inline. Not a contradiction, but the summary a model reads first describes the path the skill itself no longer recommends. Same sweep ticket.
- The “fix agent one tier lighter” belief (finding 1 of section 7): true only interactively, and by accident rather than decision.
The drift pattern matches the previous paper’s diagnosis exactly: code corrects at incident speed, prose corrects at documentation speed, and memory hooks correct slowest of all. It is why this series verifies against code, not against our own notes, and why the inventory’s “habit” entries are labeled as habit instead of being backfilled with plausible rationales.
9. Open problems
- The claims table got its first result the day it was written, and the result was humbling in the right direction. The reviewer measurement ran on review of this draft. It killed one published number (the previous paper’s 55-of-55, corrected there), validated the critical-reviewer seat, exposed a 99.3%-noise telemetry partition, and proved the per-seat question needs one instrumentation change before any model decision is defensible. The security-refusal corpus test ran the same day and killed the claim outright (0 of 12). Next kills, in order: the first-pass verdict field, then the fable-vs-opus and opus-vs-sonnet reviewer comparisons the new field enables. A claims table nobody ever tests degrades back into habit with better formatting.
- A rationale can outlive its evidence by two months and be repeated in three places without anyone checking. The false-refusal claim was recorded on 2026-07-01, copied into two skill definitions and a memory note, and cited as settled in the fact-gather for this paper. Thirteen dollars disproved its current basis. The open problem is not that one claim was wrong; it is that nothing in the system flags a rationale that has never been tested, and the claims table is a manual substitute for that.
- Something moved the fix-cycle rate five-fold at the end of July, and we do not know what. 14.9% of dispatches ran a fix cycle before 2026-08-01; 77.3% after, with the 2-cycle cap now the modal outcome. Reviewer models were constant across the boundary. The step localizes to roughly 07-29 through 08-01, a window holding both a pipeline hardening batch (including the stalled-reviewer enforcement gate) and the start of a harder feature batch, and the four dispatches that could separate those explanations cannot carry an attribution. Until this is understood, no fix-cycle trend is interpretable as reviewer behavior alone.
- The freshness classifier’s future is an open decision: move the task to a frontier model, retire automated stamping, or accept the no-op as a cheap TTL refresher. The 2026-08-21 decision records the question; nothing has answered it.
- Fan-out ergonomics remain the weakest harness surface. The 0-of-4 idle failure, the stall alarms, the unnamed-synchronous workaround: all compensate for background agents that fail quiet. The system’s topology choices are partly downstream of that reliability gap, and would deserve re-audit if it closes. This paper’s own fact-gathering and refusal-corpus agents both went idle after finishing their work and before writing it up; the supervisor re-derived the results from their artifacts. The pattern the paper describes happened to the paper.
10. References
- The companion fact sheet (internal): every mechanism claim above with its full citation, compiled by reading the skill definitions, agent definitions, and dispatcher source end to end before this paper was written.
- Internal measurement notes: the reviewer catch-rate measurement (2026-08-30, 347 dispatches from forensic logs, fixture-validated) and the Fable security-refusal corpus test (2026-08-30, 12 diffs, every output verified from raw JSON), both quoted with their sample sizes and limits in section 7.
- Internal decision records: the two local-model benchmarks (2026-06-01 and 2026-08-21) behind the only measured rows in the inventory.
- The Drain: the dispatch layer under this paper, and the trust-ladder frame it borrows.
- Files as Prose, Store as Truth: the substrate both run on.
- The machine itself: the implement, plan-create, plan-explore, debug-swarm, and plan-review-swarm skills, the agent definitions, the dispatcher, the pipeline runner, and the model registry, private, at HEAD 2026-08-30.
This is a paper in the Yianna & Claude Whitepapers series: working papers on running a production side project at the output of a team, with a fleet of AI agents, alongside a full-time engineering job, written from the system’s own measurements and incident reports. Evidence citations reference our internal decision records and postmortems; where a claim could not be traced to a measurement, the text says so.