Cheaper judgments with Typesafe.ai did not lower workflow costs

Field Note 002 · September 20, 2026

I wanted cheaper claim checks to reduce the cost of my agent-assisted writing workflow. Adding TypeSafe raised measured drafting and review costs from $2.84 to $4.33. Its own calls were inexpensive. The surrounding Codex work cost more.

TypeSafe checked whether candidate article claims had source support and helped the argument. Ruby checked exact quotations and routed unresolved claims for investigation or review. Codex still wrote and reviewed the article.

The first comparison did not isolate TypeSafe

Fresh agents worked in separate worktrees from the same interview, outline, and evidence. But they selected different claims: 20 checks without TypeSafe and 18 with it. We had compared different drafting and correction work, not identical judgments. That was a mistake in the experiment design.

Measured worker cost Without TypeSafe With TypeSafe
Codex $2.84134 $4.32776
TypeSafe $0 $0.00264
Total $2.84 $4.33

The 52% increase describes these runs, not a general effect of TypeSafe. The cheap calls had not demonstrated savings in the surrounding work.

The same claims, checked three ways

We then froze 30 claims and 60 questions: support and relevance for supported statements and deliberately unsupported variants. Two independent reviewers prepared reference answers. Disagreements were reconciled and the answers locked before testing.

Fresh Astra and Luna agents, both at medium reasoning effort, and TypeSafe’s Jev 1.13 received identical claims, sources, and criteria. This test generated judgments, not article prose.

Fixed-batch result Astra Luna Jev
Support agreement 29/30 28/30 22/30
Relevance agreement 28/30 25/30 22/30
Unsupported claims marked supported 0/12 0/12 2/12
Estimated cost $0.6904 $0.0184 $0.00124

Agreement uses the locked reference’s acceptable labels, including five ambiguous questions. The references were model-assisted, not human-certified, and may favor related OpenAI models. This was one diagnostic batch, not a representative benchmark.

Jev accepted that the schema required evidence references when it allowed an empty array. It also accepted that I had no outcome measurements. Their absence from the interview did not establish that I had none.

Three Jev probability distributions summed to 0.99, failing our strict validator. The table scores their labels separately; the original responses remain unchanged.

Costs use published standard API rates for OpenAI and TypeSafe and cover measured workers, excluding setup and coordinator usage. They are not subscription charges. Reference preparation cost another $1.60. Codex’s agent/tool execution and Jev’s direct API call also make timing comparisons unequal.

I wanted lower workflow costs. This integration did not deliver them. Luna came close to Astra on support judgments at about two cents per batch. The next workflow test needs to show that a cheaper checker removes expensive work while still catching unsupported claims.