
# Files as Prose, Store as Truth

<!-- publication-record:start -->
*Published 2026-08-26 · Revised 2026-09-13 · [Revision history](#revision-history)*

*Technical verification is documented with the individual claims.*
<!-- publication-record:end -->

**By Yianna Kokalas & Claude**

> **How this paper was written.** This is a working paper from a production side
> project: software built for the Magic: The Gathering community, with real users,
> built and operated by one engineer (an enterprise engineer by day) together
> with Claude, Anthropic's AI model.
> The design and this paper grew from our working sessions, incident reports, and
> measurements.

<!-- document-note:editorial:start -->
**Authorship and editing**

These papers began as a collaboration between me and Claude. I now review and revise them with Astra, an OpenAI model, as my editor. I make the final calls on what they say and what gets published.
<!-- document-note:editorial:end -->

**Position**: ticket prose stays as Markdown files synced by Git; state, audit, triggers,
and search belong behind one small store service. The instructions need to be easy to
read and edit. The status needs to remain trustworthy when several agents are working,
and its history needs to survive mistakes. Search needs both the instructions and the
status. These are different jobs, so they do not all need the same home. The August
design keeps SQLite for state, adds sqlite-vec for search, and puts an API in front of
both. The service and search index were proposed work, not a completed deployment.

<!-- claim:design-status -->
<a id="claim-design-status"></a>

<!-- change-note:design-status:claim:start -->
> **Corrected · 2026-09-13** — The earlier opening sounded like an operating service, while the paper's build list described proposed work. The article now identifies the August design consistently; it does not certify a deployment. [View revision →](#revision-change-design-status)
<!-- change-note:design-status:claim:end -->



<!-- document-note:implementation:start -->
**Implementation note**

This paper records the design chosen in August 2026. At that point, ticket state and an
event log existed in a local SQLite library; the network service, ticket-vector index,
consumer cursors, and service-managed leases were proposed. Sections 3–8 describe the
intended architecture and its requirements, not a verified deployment. Historical
measurements retain their original dates; this editorial revision is not a fresh audit
of the whole system.
<!-- document-note:implementation:end -->

---

<a id="0-the-setting-context-for-a-cold-reader"></a>

## 0. The setting

One engineer runs a production side project, alongside a full-time enterprise engineering
job, by operating a fleet of coding agents.
The constraint that shaped everything in this paper is hours, not headcount: the system
has to make real progress while its operator is at work or asleep, which is why so much
of it is built to run headless behind explicit approval gates.

The planning side lives in a git repository of markdown files (the "vault"): tickets,
decision records, discovery reports, design papers like this one. A ticket is a written
specification: what to change, what belongs in scope, and how to tell whether it worked.
We call that the ticket's prose. It is intended to give an implementer enough context
to proceed, while leaving status and approval to the store.
[See a short ticket example](#example-prose-spec).

<!-- example:prose-spec:start -->
<a id="example-prose-spec"></a>

### Example: A prose ticket

Simplified ticket body for illustration. Live status and approval are tracked separately in the store.

```markdown
# Keep collection search after editing a card

## Objective
Keep the current search when a collector saves a card edit.

## In Scope
Preserve the search text and matching card list after save.

## Out of Scope
Changing search behavior or saving filters between visits.

## Implementation Shape
Refresh the edited card without resetting the search state.

## Acceptance Criteria
- [ ] Saving a card edit keeps the current search text.
- [ ] The matching list shows the updated card details.
- [ ] Clearing the search still shows the full collection.

## Test Plan
Search, edit a matching card, save, and verify the search remains.
```
<!-- example:prose-spec:end -->

Three mechanisms matter for this paper:

- **The drain**: a headless orchestrator that takes explicitly approved tickets
  ("promoted" by the human) and runs each one end-to-end: implementation, code review,
  fix cycles, tests, merged pull request, and, when the ticket calls for it, deployment
  to production. It proceeds without routine human intervention, but stops for human
  review when a gate or failure requires it.
- **The store**: the single source of truth for ticket *state* (status, priority,
  blocking relationships), a SQLite database at the August design point. State moved out of the markdown
  files' YAML frontmatter (the metadata block at the top) and into the store in an
  earlier migration; the files kept the prose.
- **The event log**: ordinary status transitions append an event with actor, reason, and
  timestamp. This already exists in the store and has paid for itself (section 3).

The question is how to keep local ticket files useful while giving several workers
reliable shared state. Reading and searching local files fits the agents' existing
tools. Reserving work, recording changes, and searching by meaning need different
capabilities.

## 1. Abstract

This paper separates four jobs carried by a ticket: prose, state, audit, and retrieval.
Files keep instructions easy to read and edit; a shared service coordinates state
changes, records their history, and searches by meaning and status. The sections
below explain that choice, walk the proposed data flows, and identify the recovery
and coordination requirements still to be proved.

## 2. The four workloads hiding inside "a ticket"

| Workload | What it is | Where it lives | Why |
|----------|-----------|------------------|-----|
| Prose | Objective, scope, acceptance criteria; the spec a human or implementer reads | Markdown files in git | Agents use existing file tools; humans can use Obsidian; reading and drafting work offline |
| State | status, priority, claims, gates; small records that workers may update concurrently | Store service rows with atomic reservations that expire unless renewed | Reserving a ticket must check availability and record its owner as one operation |
| Audit | who changed what, when, why | Append-only event log in the same store | In the July incident, events preserved the state history needed for repair (section 3) |
| Retrieval | "similar tickets, backend only, open only" | sqlite-vec chunks beside the state columns | Filters use state from the query's database snapshot, avoiding a separate metadata mirror |

Compare-and-swap (CAS) changes a value only if it still matches the expected value. For
a claim, the service must atomically check that a ticket is available and record its
owner. A lease makes that ownership expire unless renewed. Retrieval divides prose
into sections, or chunks, and represents their meaning as numeric vectors called
embeddings; a search finds nearby vectors, then restricts results by ticket state.

<a id="3-evidence"></a>

<a id="4-design-principles"></a>

## 3. Design principles

1. **Per-workload ownership, no dual source of truth.** Each field has exactly one
   authoritative substrate (section 2). Cross-references are by value (path + content
   hash), never by shared authority. The vector chunks live beside the state columns they
   filter on, so filters read the query's state snapshot rather than a separate mirror.
2. **State changes need a history independent of prose edits.** In July 2026, both
   Markdown and the store said `todo` for 12 finished tickets. The event log preserved
   the history used to restore their statuses. That supports keeping an audit record;
   it does not prove that every ticket field can be reconstructed. The design goes
   further: current state would be a projection built from events. That requires
   complete creation data, every subsequent change, and a tested replay procedure.
3. **Retire the old write path when ownership moves.** Our migration dragged on while
   we kept shipping features. Stale file metadata could still overwrite the store
   until we retired the old reconciliation path. A cutover must remove the old
   path's authority, not just introduce a new owner.
4. **Keep ordinary prose reads local.** Prose reads stay grep/read on local files.
   Retrieval that files are bad at (semantic, filtered) becomes one API call that is one
   SQL query inside. State and audit reads use the store, including the done-date
   lookup moving out of ticket-file Git history.
5. **Save the change and its event together; let each consumer track its progress.** An event that should wake a
   listener is written in the same transaction as the state change (outbox), so a missed
   notification does not erase the pending work. Each consumer owns a durable cursor row (its last processed
   event id) and sweeps forward from it; a shared processed-flag would break the moment
   there is a second consumer, so cursors are a day-one requirement, not an optimization.
   A consumer advances only after successful processing. A crash can cause redelivery,
   so repeating an event must have the same effect as processing it once. Retention and
   backup must preserve events that consumers still need.
6. **Sync is not backup.** Sync propagates deletions faithfully; backup must survive them.
   Two mechanisms, deliberately.

<!-- document-note:source:start -->
**Source**

[Corrected store incident report](#source-store-incident), July 19, 2026. Twelve finished-ticket regressions were part of a larger repair; recovery depended on the surviving event log.
<!-- document-note:source:end -->

<!-- source-details:store-incident:start -->
<a id="source-store-incident"></a>

### Source details: The July ticket-state incident

Summary of the corrected internal incident report dated July 19, 2026.

State was moving from file metadata into the store, but the old reconciliation
path remained active while product work delayed the cutover. Both representations
could still determine state.

Six of 44 test suites could resolve the production database. Fixture rows established
that at least one had written to it; the count of exposed suites was not a count of
proven polluters. The tests could invoke a reconciliation path that copied stale file
metadata into the store without logging the overwritten statuses.

The report records 18 rows restored from events overall. Twelve finished tickets
formed the regression cohort discussed in this paper: both their files and current
store rows said `todo`, so comparing those copies would not reveal their prior status.
The event log retained that history. This was a status repair with a surviving event
table, not a test of rebuilding a lost database in full.

The report describes retiring the reconciliation path and adding a test-process
guard. On September 13, 2026, we checked the current local implementation and reran
the two regression suites for reconciliation and ticket registration: all 17 tests
passed. These checks show that the tested file-import paths preserve existing ticket
state. They did not rerun the historical data repair or verify the running deployment
and its test-isolation configuration.
<!-- source-details:store-incident:end -->

<a id="5-architecture"></a>

## 4. Architecture

The proposed components are grouped by what they own. Host labels below describe
roles, not the private deployment topology.

- **Ticket files**: `tickets/*.md` (and decision records, discovery reports, ...) in the
  vault git repo, cloned on every host that runs agents. A hosted git remote (GitHub) is
  the sync hub and offsite copy. Optionally a peer-to-peer file sync tool (Syncthing) for
  non-git corpora and phone-side files; explicitly not for the vault working tree, which
  git already syncs.
- **The store service**: a Ruby (Rack) API around the
  existing store library, on an always-on service host. Inside: one SQLite file holding
  the state projection, append-only event log,
  claims with leases, outbox with per-consumer cursor rows, and sqlite-vec chunks keyed
  `(slug, section, section_hash)`. The service process is the only thing that touches the
  file; local and remote workers use the same API, and store callers become clients.
  The DB file never enters git.

  SQLite supports multiple local connections but one write transaction at a time.
  Sharing the database file across machines introduces network-filesystem locking
  and durability risks. The proposed API keeps database access on the service host
  while giving every worker the same state operations. See SQLite's
  [network guidance](https://www.sqlite.org/useovernet.html) and
  [transaction model](https://www.sqlite.org/lang_transaction.html).
- **Listeners**: the drain (wakes on ready-events, claims via CAS + lease), the indexer
  (pulls its vault clone on wake, re-embeds changed sections into sqlite-vec in the same
  file), an analytics shipper (mirrors events for
  dashboards). All idempotent, all resumable from their cursors, all clients of the
  service.
- **Operator access**: a remote terminal lets the operator watch or steer the agents. It
  operates a worker host rather than maintaining another copy of ticket state.
- **Backup**: database snapshots and copies of any
  non-git corpora. The git remote is the offsite copy for ticket prose. The vector
  chunks are derived data, rebuildable by re-embedding, but they ride along in the store
  dump for free.
- **Other corpora**: material with no ticket-state join can use a separate retrieval
  system. It is not in the ticket path.

<a id="51-topology-objective"></a>

### 4.1 Topology (objective)

This diagram is the objective, not a deployment inventory. At the August design point,
the store was still a local library. The **planning host** is where the human plans:
interactive human-and-agent sessions author
prose locally and talk to the store service for state (register at creation, promote,
status, search); it can run a drain of its own for tickets labeled as needing a human,
which the headless drain's queue refuses by design. The **service host** runs the store
service; workers can request claims through its API. Cross-host execution also needs
the recovery protections described in section 5. **Supporting services** handle
observability, backups, and non-ticket retrieval outside the ticket request path.

Shared access does not establish round-the-clock operation. The August 14–16, 2026
feasibility study found only five tickets passing all drain gates in its store sample.
It modeled dispatch durations with quota assumptions that still needed confirmation;
it was not a 24-hour runtime trial.

<!-- claim:runtime-evidence -->
<a id="claim-runtime-evidence"></a>

<!-- change-note:runtime-evidence:claim:start -->
> **Corrected · 2026-09-13** — The feasibility study modeled continuous operation using quota assumptions that still needed confirmation. It was not a 24-hour runtime trial, and host placement alone does not establish round-the-clock operation. [View revision →](#revision-change-runtime-evidence)
<!-- change-note:runtime-evidence:claim:end -->



```
  Planning host / workers              Store service
  +-------------------------+         +---------------------------+
  | local prose + git clone |-- API -> | one SQLite file           |
  | approve / implement     |         | state | events | claims   |
  | read specs locally      |         | outbox | ticket vectors   |
  +------------+------------+         +-------------+-------------+
               |                                   |
          git push / pull                  events / cursor sweeps
               |                                   |
               v                                   v
  +-------------------------+         +---------------------------+
  | git remote: prose sync  |<-- pull--| workers / indexer clones  |
  | and offsite copy        |         | analytics consumer        |
  +-------------------------+         +---------------------------+

  Database snapshots -> independent backup storage
  Non-ticket retrieval -> separate system, outside this request path
```

The API carries state commands and prose references; git carries the ticket files.
The indexer writes derived vectors through the service. Those roles do not make git
the authority for state or the store the author of prose.

The proposed handoff pushes committed prose before sending a command that depends on
it. The write appends the event (with the
prose ref: path + content hash) in the same transaction; the drain's sweep wakes on it,
`git pull`s, and verifies the named hash is present before running (section 4.4). The
signal channel is the store, the content channel is git; neither needs a webhook.

<a id="52-write-and-event-flow"></a>

### 4.2 Write and event flow

```
        PROSE WRITE                              STATE WRITE
        ===========                              ===========

  planner edits tickets/<slug>.md      client calls the store API:
            |                          transition(slug, from:, to:)
            v                                        |
     git commit + push                               v
            |                          service: CAS row update
            v                             (lease honored)
     git remote (offsite)                 [same transaction]
            |                                        |
            v                        event appended: actor, reason, ts,
     other hosts git pull                    prose ref = path + hash
                                                     |
                                                     v
                                        outbox row (same tx)
                                                     |
                     +-------------------------------+------------------------+
                     |                               |                        |
                     v                               v                        v
              drain (cursor sweep)            indexer (cursor sweep)  analytics shipper
              wake on ready-event             git pull; read .md      mirror events
              claim via CAS + lease           at event's hash         for dashboards
              run the implement               re-embed changed
              pipeline                        sections into
                     |                        sqlite-vec (same file)
                     v
              transition(...)
              [loops back to STATE WRITE]
```

Each of the three consumers sweeps forward from its own cursor row on wake and on a
periodic tick; an in-process nudge after each write is a latency optimization, the outbox
plus cursor is the delivery contract.

The proposed workload has one producer, the store service, and three consumers:
the drain, indexer, and analytics shipper. An outbox and separate consumer cursors
are the chosen starting point; the design does not require a separate message broker.

<a id="53-read-paths"></a>

### 4.3 Read paths

```
  agent needs the spec                  agent needs "similar open backend tickets"
            |                                        |
            v                                        v
  grep / read local tickets/*.md        one store-API call, one SQL query inside:
  no network hop, batchable,            vector match JOIN state
  works offline                           WHERE repo='BE' AND status='todo'
                                        filters use the query snapshot;
                                        text relevance can still lag
                                        behind prose edits
```

An August 25, 2026 sweep of Git-history calls in the vault's scripts, skills, and
hooks found one use of ticket-file history: a support utility taking a file's last
commit date as its default done-date. Other calls targeted code repositories or pull
requests. That lookup is planned to move to the state event log (section 8). This
finding describes those tools, not whether people consult prose history.

For retrieval, the July 19, 2026 decision recorded 1,585 tickets and about 20 MB of
prose—not database size. Its index sizing was an estimate: roughly 10,000 chunks
and 150 MB indexed. The August 26 decision chose brute-force search, which compares
the query with every vector, expecting millisecond responses. No benchmark on this
corpus was found in the records inspected; latency and filtered recall still need
testing. A dedicated vector database remains the fallback if the corpus outgrows
this approach, while index freshness needs checking either way (section 6).

<!-- claim:evidence-scope -->
<a id="claim-evidence-scope"></a>

<!-- change-note:evidence-scope:claim:start -->
> **Corrected · 2026-09-13** — Corpus and index sizes are dated estimates, not benchmark results. The expected search latency was not established by a located test on this corpus; latency and filtered recall remain to be measured. [View revision →](#revision-change-evidence-scope)
<!-- change-note:evidence-scope:claim:end -->



<a id="54-isnt-pull-based-sync-a-burden"></a>

### 4.4 Isn't pull-based sync a burden?

Git is pull-based, so it is fair to ask who pulls, and when. In this design the store
event tells consumers that work is pending. The drain and indexer sweep their cursors,
pull committed prose, and check its content hash before acting. A hash identifies the
expected content; it does not itself retrieve a missing file revision. The handoff
must preserve a way to fetch the corresponding committed content. Prose-only edits
with no state event are caught by a periodic reconcile sweep. This avoids requiring
a separate inbound webhook from the git host, at the cost of waiting for the next sweep.

The drain must stop on a mismatch. The indexer has a deliberately weaker fallback:
after bounded retries it can index current content and record the divergence (section
5). That preserves progress through the queue, but gives up exact historical content
for that event. Neither pulling alone nor the presence of a hash proves freshness.

The continuous-sync alternative (a peer-to-peer file-sync tool like Syncthing on the
working tree) loses on a sharper point than convenience: the reconciliation mechanism
requires retained, identifiable revisions. Git provides committed versions that can be
named and inspected; continuous file sync alone does not bind a version to a store event.
Conflict copies such as `.sync-conflict-*` can also be overlooked by an agent expecting
one ticket path. Git rejects a non-fast-forward push, and merging or rebasing may then
require conflict resolution; it does not flag every pair of edits as a conflict. The
cost accepted in exchange is explicit synchronization: a host may be behind between
sweeps, and a worker must verify the prose it is about to use.

<!-- claim:prose-freshness -->
<a id="claim-prose-freshness"></a>

<!-- change-note:prose-freshness:claim:start -->
> **Corrected · 2026-09-13** — Pulling and checking a hash require retained, retrievable committed content. The indexer's bounded fallback can deliberately index different content, so the earlier universal freshness guarantee was too strong. [View revision →](#revision-change-prose-freshness)
<!-- change-note:prose-freshness:claim:end -->



<a id="6-flows-worth-walking"></a>

## 5. Flows worth walking

1. **Ticket authored.** The planner (human or agent) writes `tickets/<slug>.md` locally,
   commits and pushes the prose, then registers it through the store API with a
   reference to the retrievable committed content. A new row is initialized from authoring
   metadata. This is a one-time initialization, not permission to overwrite existing
   store state from a later file edit. The registration event carries the prose
   ref, and the indexer embeds on that event. Creation itself is prose-only and works
   offline; registration waits until the committed prose is available to consumers
   and the service is reachable. A reconcile sweep backstops missed registrations
   under the same prerequisites, without overwriting rows that already exist.
2. **Prose edited after registration.** Push updates the file. The indexer picks the
   change up on its next wake: any event naming the slug, or its periodic reconcile
   sweep, which compares working-tree section hashes against the chunk keys. If the
   indexer's clone cannot yet retrieve content matching the event's hash, it leaves its
   cursor in place and retries on the next sweep. The wait is bounded: after N sweeps
   without the named content (the writer amended or
   never pushed that exact content), the indexer indexes the current working-tree state,
   records the divergence, and advances its cursor so a missing revision does not
   hold up every later event indefinitely. The retry limit is a design choice still
   to be set. A dedicated `prose_changed` post-push event could reduce the delay.
3. **Drain wake.** Promote emits an event; drain workers sweep the outbox and race to
   claim. The proposed claim operation must atomically select one owner, whichever
   hosts the workers run on. A crashed worker's lease can expire, making the ticket
   reclaimable, but expiry alone does not stop an old worker that resumes later.
   Reclaim is not a bare re-run: the reclaiming worker first establishes ground truth
   about the crashed run's leftovers (branch, worktree, possibly an open PR) using the
   same verification checks the implement pipeline already runs, and recovers context
   from the dead session's transcript rather than starting blind. The service must
   reject writes from an expired owner, and recovery must handle external effects such
   as an already-created pull request. These are requirements before replacing local
   locking; a lease is not a guarantee that a ticket's effects happen only once.
4. **Recovery.** The store restores from backup snapshots. The event log can support
   repair of recorded state changes; a full rebuild remains to be proved. Vector chunks are derived and rebuild by
   re-embedding. Prose restores from retained Git history. The restored data is checked
   against the prose references carried by events; those checks also need testing.
5. **Phone session.** A remote terminal attaches to a worker host to watch or steer
   the agents. The phone does not introduce another state store or prose writer.

<!-- claim:claim-ownership -->
<a id="claim-claim-ownership"></a>

<!-- change-note:claim-ownership:claim:start -->
> **Corrected · 2026-09-13** — A lease can expire without stopping an old worker. Safe cross-host recovery needs ownership enforcement and handling of prior external effects before local locking can be replaced; the article no longer treats leases alone as proof of safe execution. [View revision →](#revision-change-claim-ownership)
<!-- change-note:claim-ownership:claim:end -->



<!-- document-note:implementation:start -->
**Implementation note**

The August code, and the copy inspected on September 13, 2026, write a ticket's initial fields but record only its initial
status in the creation event. That event is insufficient to recreate the whole row.
Database snapshots remain necessary; complete event replay is an unverified design
requirement, not an additional backup already proved by the July repair.
<!-- document-note:implementation:end -->

<!-- claim:event-replay -->
<a id="claim-event-replay"></a>

<!-- change-note:event-replay:claim:start -->
> **Corrected · 2026-09-13** — Repairing statuses from surviving events did not prove full database reconstruction. Creation events omit initial fields needed for replay, so snapshots remain necessary and complete replay remains an unverified requirement. [View revision →](#revision-change-event-replay)
<!-- change-note:event-replay:claim:end -->



<a id="7-failure-modes-and-mitigations"></a>

## 6. Failure modes and mitigations

| Failure | What stops or goes wrong | Response |
|---------|-------------|------------|
| Store service down | No state transitions or semantic search; workers cannot safely claim new work | Prose stays readable/editable on clones; restore service access and check the existing state before resuming. Restore a backup when data is lost or damaged, not for an ordinary outage |
| Supporting service down | Analytics, backup delivery, or non-ticket retrieval unavailable | Ticket requests need not depend on these services; resume delivery on return and monitor backup age |
| Client cannot reach the service | No state commands or search from that client | Prose authoring remains local; retry commands when reachability returns; registration reconcile backstops missed creation |
| Git remote down | No prose sync | Local reading and drafting continue; a worker must wait if it cannot obtain the required committed content |
| Indexer lag | Stale embeddings and missed or less relevant matches | State filters use the query snapshot, but do not repair stale prose representations; a re-index can replace outdated chunks |
| Indexer bug | Wrong or incomplete retrieval | Hash keys help identify versions, but cannot prove that chunking, embedding, or replacement is correct; validate those paths separately |
| Indexer's clone behind the event's hash | Embedding delayed | Cursor stays put and retries; after N sweeps it indexes current content, records the divergence, and advances (no head-of-line blocking) |
| Two hosts edit the same .md | Divergent commits, possibly a merge conflict | A rejected push requires integration; overlapping edits may need manual resolution, while semantic conflicts still need review |
| Consumer misses a nudge | Late wake-up | Outbox rows are durable and each consumer sweeps forward from its own cursor row on a periodic tick; the nudge is an optimization, outbox + cursor is the contract |
| Sync propagates a bad deletion | File disappears from current checkouts | Restore from retained git history; independent snapshots cover data outside git. A synced remote alone does not guarantee retention |
| Store and file disagree on prose | Ambiguity about what was implemented | The proposed handoff must retrieve committed content and verify path + hash before execution; the indexer's bounded fallback is explicitly weaker |

<!-- claim:retrieval-guarantees -->
<a id="claim-retrieval-guarantees"></a>

<!-- change-note:retrieval-guarantees:claim:start -->
> **Corrected · 2026-09-13** — State filters use the query's database snapshot. They do not guarantee fresh or correct embeddings, and hash keys do not prove correct indexing. Lag and indexer bugs now have separate limits. [View revision →](#revision-change-retrieval-guarantees)
<!-- change-note:retrieval-guarantees:claim:end -->



<a id="8-rejected-alternatives"></a>

## 7. Rejected alternatives

Recorded so they are not re-litigated. Fuller arguments live in internal decision records;
the shape of each rejection is reproduced here.

- **Postgres + pgvector as the substrate** (v1 of this paper): an engine that could
  hold both workloads, but it put a full store migration in front of everything
  else. The store service gets the same consolidation (state + vectors in one
  transactional file) without an engine migration: the existing SQLite file and library
  carry over.
- **A dedicated vector database for the ticket corpus** (an interim revision of this
  paper): the service is a cost of the state plane and gets built regardless, so a
  separate vector store for tickets means a second stateful system in the ticket path
  plus the machinery to keep two systems honest (a payload mirror, a staleness class,
  verify-on-read), which exists only because of the split. Cancelled unbuilt; the vector
  database keeps the corpora with no state join, and stays the named escalation if the
  ticket corpus ever outgrows brute-force search.
- **ssh-wrapped CLI + registration deferred to git pulls** (another interim revision,
  briefly): separate ways to let another machine reach the store. We chose one API
  instead of maintaining a different remote-access path for each operation.
- **Kafka (or any broker) for events**: unused value at one producer and ~3 consumers;
  operational weight without a workload. Outbox + cursor sweeps cover the need; NATS
  JetStream is the named escalation if consumers ever multiply beyond a cursor table's
  comfort.
- **Everything into the store (no local files)**: forfeits grep/read economics, offline
  operation, Obsidian, and our fixture-testing discipline, to gain nothing the split does
  not already provide. The network hop was never the main cost; primitive loss is.
- **Everything in files, state included**: plain Markdown edits do not supply the
  atomic claims or leases this design requires; adding those would require a
  coordination layer. The store is the chosen place for that coordination.
- **A vector database as the PRIMARY store**: the candidate considered in the July
  research did not provide the claim and state-plus-event transaction contract needed
  here. Search filters alone do not establish that contract. This is a requirement to
  check for a particular engine, not a claim about every vector database.
- **An analytics/columnar database as primary store**: the engine evaluated in July
  lacked the transactional claim operation this workload needed. It kept the
  analytics-mirror role; this is not a universal limitation of columnar storage.
- **Peer-to-peer file sync for the vault repo**: using two sync mechanisms for a working
  tree adds competing ways to reconcile edits, and continuous file sync cannot pin content to
  events the way hash-addressed commits do (section 4.4). Git is the vault's sync; that
  tool's lane is non-git corpora and phone files.
- **A second document service for prose**: prose already has a substrate.

<a id="9-what-actually-remains-to-be-built"></a>

<a id="9-build-areas-identified-in-august"></a>

## 8. Build areas identified in August

These are the six build areas identified in August 2026, not six equally small tasks
or a current completion checklist. A September 3 code exploration found the service
and retrieval work still unbuilt and showed that client mode must cover direct library
callers as well as the CLI. The integration work is larger than a wrapper alone.

1. **The store service**: a Rack API around the existing store library, on the
   service host. The service process is the sole application owner of the
   SQLite file.
2. **sqlite-vec inside it**: chunk schema keyed `(slug, section, section_hash)`, the
   embed pipeline, and a `search(query, filters)` endpoint that joins vectors against
   live state columns.
3. **Client mode for all store callers**: a library-level backend for the CLI, drain,
   and other callers, preserving the existing operations. Direct SQL callers need named
   API operations too; changing the CLI alone cannot complete the boundary.
4. **The indexer**: a script, not a service. Sweeps its cursor, pulls its clone, chunks
   changed files by the ticket spec's required sections, embeds, and replaces outdated
   chunks. Repeating a pass must not create duplicates; deletion and replacement need
   validation as well as keys that identify each version.
5. **Prose refs on events**: path + file hash stamped at registration and refreshed at
   claim; plus the outbox cursor rows for the three consumers.
6. **Migrate the one done-date read** from ticket-file git history to the event log,
   closing the ticket-file history lookup identified in section 4.3.

The workload split preserves local prose authoring, review, and reading. Registration
and promotion do gain an explicit committed-prose handoff. Not in scope:
any new message broker; any engine migration; any move of `tickets/*.md` out of git.

<!-- claim:client-cutover -->
<a id="claim-client-cutover"></a>

<!-- change-note:client-cutover:claim:start -->
> **Corrected · 2026-09-13** — Client mode must cover direct library and SQL callers as well as the CLI. The six build areas were not a complete small-task implementation plan, and registration gains a committed-prose handoff. [View revision →](#revision-change-client-cutover)
<!-- change-note:client-cutover:claim:end -->



<a id="10-open-questions"></a>

## 9. Open questions

- **Embedding model and refresh policy** for the ticket corpus (a hosted embedding API vs
  a local model is undecided; the engine seam is decided, the model is not).
- **The board view is probably dead**: a kanban view regenerates on every store write,
  but the operator stopped using it: at 1,600+ tickets a board is a wall, not a view. The
  service cutover is the natural moment to retire the regeneration rather than port it.
  (An honest lesson for agent-native tooling: views built for human project management
  are the first thing to fall out of use once agents do the tracking.)
- **sqlite-vec maturity watch**: the August decision selected a pre-1.0 engine for
  brute-force search. Recheck the chosen version before implementation. Risk is bounded by the service seam
  (a swap to a dedicated vector database is one contained rewrite), but worth a periodic
  check.
- **Multi-writer prose**: if concurrent drains need to edit the same ticket file
  outside PR flow, the coordination rules need revisiting. The August design assumed
  they would not.
- **Phone-side prose access**: a read-only vault on the phone (Obsidian plus a sync path)
  is attractive but unscoped; it must not become a third writer.

## Discrepancies found while checking the paper

These findings keep the unfinished requirements visible alongside the design. The
September 13, 2026 editorial review checked historical records and selected code;
it did not establish current completion of the service or its safeguards. The
follow-ups below are work to revisit, not fixes this revision has delivered.

1. **The event log did not contain enough information for full replay.** Creation
   events recorded the initial status but omitted other initial fields. July's repair
   used surviving events to recover statuses. On September 13, 2026, we added complete
   ticket-state reconstruction to the store-service plan, with a starting snapshot
   for existing tickets and Ruby tests to prove replay. That work remains planned;
   database snapshots are still essential to recovery.
2. **A lease does not stop an expired worker.** The design permits work to be reclaimed
   after expiry, but an old worker may resume. Before replacing local locking, prove
   that expired owners cannot change state and that recovery handles existing branches
   and pull requests without repeating completed effects.
3. **A content hash does not fetch the promised prose.** Define how the handoff retains
   and retrieves the committed revision. The execution path must stop on a mismatch.
   The indexer's fallback may use newer content; set its retry limit and keep that
   divergence observable rather than describing both paths as equally strict.
4. **Search estimates were not a benchmark, and hash keys were not an indexing test.**
   Measure latency and relevant-result coverage on the intended corpus. Test changed
   and deleted sections as well as initial indexing: current state filters cannot
   repair stale or incorrect embeddings.
5. **The client cutover reaches beyond the CLI.** The September 3 exploration found
   direct library and SQL callers. Account for those paths before calling the service
   the sole application owner of the database; switching the command-line client
   alone would leave the boundary incomplete.

---

*This is a paper in the Yianna & Claude Whitepapers series: working papers on running a
production side project at the output of a team, with a fleet of AI agents, alongside a
full-time engineering job, written from the system's own measurements and incident
reports. Historical sources are internal decision records and postmortems; proposed
mechanisms and unverified requirements are identified in the text.*

## Revision history

### 2026-09-13 · Editorial

Clarified the original collaboration with Claude and ongoing editing with Astra. Yianna retains final editorial and publishing decisions.

### 2026-09-13 · Editorial, Correction, New evidence

Clarified the distinction between the August design and implemented behavior, moved evidence beside its claims, and shortened the migration lesson. Corrected recovery, coordination, and retrieval guarantees; preserved visible follow-ups and recorded a scoped local regression check.

<a id="revision-change-design-status"></a>

- **Correction · Important:** The earlier opening sounded like an operating service, while the paper's build list described proposed work. The article now identifies the August design consistently; it does not certify a deployment. [Affected claim](#claim-design-status)

<a id="revision-change-evidence-scope"></a>

- **Correction · Important:** Corpus and index sizes are dated estimates, not benchmark results. The expected search latency was not established by a located test on this corpus; latency and filtered recall remain to be measured. [Affected claim](#claim-evidence-scope)

<a id="revision-change-event-replay"></a>

- **Correction · Important:** Repairing statuses from surviving events did not prove full database reconstruction. Creation events omit initial fields needed for replay, so snapshots remain necessary and complete replay remains an unverified requirement. [Affected claim](#claim-event-replay)

<a id="revision-change-claim-ownership"></a>

- **Correction · Important:** A lease can expire without stopping an old worker. Safe cross-host recovery needs ownership enforcement and handling of prior external effects before local locking can be replaced; the article no longer treats leases alone as proof of safe execution. [Affected claim](#claim-claim-ownership)

<a id="revision-change-prose-freshness"></a>

- **Correction · Important:** Pulling and checking a hash require retained, retrievable committed content. The indexer's bounded fallback can deliberately index different content, so the earlier universal freshness guarantee was too strong. [Affected claim](#claim-prose-freshness)

<a id="revision-change-retrieval-guarantees"></a>

- **Correction · Important:** State filters use the query's database snapshot. They do not guarantee fresh or correct embeddings, and hash keys do not prove correct indexing. Lag and indexer bugs now have separate limits. [Affected claim](#claim-retrieval-guarantees)

<a id="revision-change-client-cutover"></a>

- **Correction · Important:** Client mode must cover direct library and SQL callers as well as the CLI. The six build areas were not a complete small-task implementation plan, and registration gains a committed-prose handoff. [Affected claim](#claim-client-cutover)

<a id="revision-change-bounded-mechanism-language"></a>

- **Correction · Routine:** Distinguish exposed test suites from demonstrated pollution, narrow claims about files and database categories, clarify Git integration, and state the processing and retention conditions behind durable cursor delivery. A scan of tools does not establish that people never consult prose history; only ordinary prose reads are promised to remain local.

<a id="revision-change-outage-recovery"></a>

- **Correction · Routine:** An ordinary service outage calls for recovering access and checking existing state. Restore a database backup when data is lost or damaged; the earlier response could imply replacing healthy state with an older snapshot.

<a id="revision-change-runtime-evidence"></a>

- **Correction · Important:** The feasibility study modeled continuous operation using quota assumptions that still needed confirmation. It was not a 24-hour runtime trial, and host placement alone does not establish round-the-clock operation. [Affected claim](#claim-runtime-evidence)

<a id="revision-change-local-overwrite-regressions"></a>

- **New evidence · Routine:** On September 13, 2026, both local regression suites for stale-metadata reconciliation and ticket registration passed: 17 tests. This checks the local overwrite protections, not current deployment health or full database recovery.
