Three separate marks represent supported evidence, an interpretive judgment, and a condition for proceeding. There is no mark indicating that the gate has passed.
Evidence / Judgment / Gate

How I evaluate work in Claude Code

I evaluate my Claude Code workflow by keeping the evidence for each task: what it was supposed to accomplish, which checks ran, what reviewers found, and how far the change actually got. Those records also let me compare the cost and behavior of the workflow as I change it.

The process lives in a custom /implement skill and the programs around it. The skill coordinates implementation, review, and fixes. Scripts run checks and save results. An outer runner captures the execution and collects usage. Claude Code supplies the agent environment; the evaluation process is something I have built around it.

Let’s follow one real implementation from September 29, 2026: adding a required consent checkbox to a signup form. Its tests passed, both reviewers returned PASS, and the pull request merged. The delivery check still ended red. Following those records explains what my eval process can tell me, and where it needs more evidence.

1. Decide what would count as success

Before running the agent, write down the behavior the task must demonstrate and how you will check it. The task is the starting point for the evaluation.

For the signup change, these were some of the requirements:

Required behavior Evidence needed
Consent starts unchecked A check of the initial form state
Submitting without consent sends no request A submit-path check that observes network activity
Completing signup records acceptance A real browser signup followed by a database check

That last requirement matters. A frontend test can use a mock response containing acceptance data. It cannot, by itself, establish that the backend saved anything. This ticket explicitly required the browser action and the resulting database rows before the work could be considered done.

For your own task, make this mapping before implementation. If a criterion needs a browser, a database, or a human judgment, name that evidence now. Otherwise, the easiest available check can quietly become the definition of success.

2. Capture a run you can inspect later

My /implement instructions live in .claude/skills/implement/SKILL.md. A project skill is a Markdown instruction file with metadata that Claude Code can load. The specific implementation procedure and helper commands are my own. Claude Code skill documentation

The outer runner starts the skill in Claude Code’s non-interactive mode and saves the emitted event stream. It writes each line before trying to parse it, so an unexpected event remains available for investigation. A separate JSON run record connects the task, session, code location, phase checkpoints, and usage.

You can try the capture mechanism without asking Claude to change code. With an authenticated Claude Code installation, Bash, and jq, run this from a trusted temporary directory:

run_dir=$(mktemp -d)
claude -p 'Reply with exactly OK.' \
  --tools "" --strict-mcp-config --mcp-config '{"mcpServers":{}}' \
  --no-session-persistence --output-format stream-json --verbose \
  > "$run_dir/events.jsonl" 2> "$run_dir/stderr.log"
printf '%s\n' "$?" > "$run_dir/exit-code"
printf '%s\n' "$run_dir"

This exercise disables built-in tools and supplies an empty MCP configuration. stream-json produces one JSON event per line; the redirections retain those events and stderr separately. The final two lines save the process exit code and print the directory containing the files. This example was executed with Claude Code 2.1.284 while preparing the article. Programmatic execution documentation

In the same shell, inspect the final result:

jq -s '
  [.[] | select(.type == "result")] | last
  | if . == null then error("No final result") else
      {is_error, num_turns, total_cost_usd}
    end
' "$run_dir/events.jsonl"

The reader deliberately errors if the final event is missing. Also inspect is_error and the saved exit code: a file containing JSON does not establish a successful run. This command only checks capture and parsing; it has not evaluated an implementation.

For a real task, retain a task identifier and a unique run identifier alongside the stream. Keep retries distinguishable. Preserve reviewer results explicitly, too: the parent’s event stream does not necessarily include every subagent’s prose. Subagent streaming details

This is where hooks fit into the picture. Claude Code can run hooks at lifecycle events, but my main evaluation records come from the runner and explicit helper calls. An instruction to record a review still depends on the agent making that call; a hook runs when its configured event fires. Check which mechanism actually produces each record in your setup. Hook documentation

3. Save what the checks actually ran against

During implementation, the agent edits code and runs targeted checks. Before handing the change onward, my impl-validate command selects the repository’s validation steps, runs them, and saves their output and exit status. Its record includes digests of the working tree so later steps can check that the validation still describes the code being handed over.

Here is a selected excerpt from the signup run’s full-suite step:

{
  "name": "test_full",
  "exit": 0,
  "summary": {
    "tests_passed": 5048,
    "tests_failed": 0,
    "tests_skipped": 9
  }
}

This gives me an inspectable suite result. The full record points to the log, records the command, and carries the tree-integrity result. After fixes change the code, the workflow requires fresh validation; the merge helper checks for stale evidence.

The same record counts three changed test files against a configured limit of two, leaving its file-budget check false even though the overall validation flag is green. The ticket separately permits two new test files and one modified test file. Those details need assessment together; the overall flag alone does not explain the discrepancy.

For a smaller setup, have your test wrapper save the command, exit code, output path, and tested revision. If tests run before a commit, include a snapshot or digest of the uncommitted work: the commit ID alone won’t identify it. Keep acceptance checks that have not run visibly outstanding.

4. Review the implementation and preserve the first verdict

The Claude session following /implement acts as the dispatcher: it sends the change to separate correctness and security reviewers. Each gets the task and code to inspect in its own agent context. That separation gives the workflow another examination of the work, though the reviewers can still share blind spots. Claude Code subagents

The correctness review in this run began with these fields:

verdict: PASS
reviewer: correctness
finding_count: 3

All three findings were warnings under that review’s contract. One questioned the tests: the reviewer reported that a test replaced the expected API response with a reduced response and used a stubbed authentication provider. That is useful feedback about what a passing test may fail to exercise. It is also a reviewer assertion that needs assessment, rather than an independently proven defect.

The security reviewer returned PASS with no findings. I retain that as the review outcome; it does not prove the change has no security problems.

When adapting this step, give the reviewer the acceptance criteria, the diff, and the relevant surrounding code. Require a verdict, concrete findings, locations, and severity. Keep a missing or failed review distinct from a completed review with no findings.

My skill then tells the dispatcher to write a checkpoint with the reviewer, fix-cycle number, verdict, and finding count. It preserves the first-pass verdict before any fixes can replace the current one. The final report carries both initial and final verdicts.

That distinction makes later comparisons useful. A final PASS alone loses how much correction was needed to get there. In this particular run, both reviews passed on their first pass and the recorded review-fix count was zero; the warning findings remained part of the record.

5. Verify delivery after the implementation finishes

Now return to the signup task’s browser-and-database requirement.

The implementation record says the pull request merged. The dispatcher’s final report also names the browser/database check as still outstanding. Then the deployment pipeline runs its own checks, and its test stage fails on an invoice-detail test. The retained record does not establish why that test failed or whether the signup change caused it.

The separate verification record contains:

{
  "gate": "fe-preview",
  "outcome": "red",
  "attempts": 0
}

There are no successful browser-scenario results in that record. The evaluation of this attempt is therefore:

Stage Recorded outcome
Implementation suite Passed, with nine tests skipped
Correctness review PASS with three warnings
Security review PASS with zero findings
Merge Completed
Deployment pipeline Stopped at a failing test stage
Requested browser/database proof Still outstanding in the final report

This is an implementation that merged with delivery unverified in the recorded run. It is not evidence that the deployed signup feature was broken. It is also insufficient evidence to count the task as successfully delivered.

My dispatcher uses SUCCESS for the implementation stage. Reading that label alongside the delivery record is essential. If your evaluation currently has a single success field, spell out the event it measures and retain the later checks separately.

6. Compare the cost of reaching an outcome

Once the records exist, I can examine the workflow over many attempts. My local usage report extracts cost, tokens, messages, tool calls, duration, and review-fix counts. It can compare time windows grouped by repository, including implementation success and timeout rates.

For the signup implementation session, the record reports about $2.20 in model cost, about six minutes and fourteen seconds, and 32 tool calls. The model-cost records specify list rates. That figure describes this implementation session; it is neither an invoice nor a measured cost for successfully delivering the whole task. It is one example, not a typical-performance benchmark.

Central telemetry makes those records easier to query. My setup sends dispatch summaries and phase markers through Vector into ClickHouse, alongside Claude Code’s native OpenTelemetry events. Grafana and a command-line scoreboard read them. The native export is configured to omit prompt and tool content; local logs retain the more detailed execution evidence.

The coverage has limits. Shipping is best effort. Interactive implementation runs do not get the same phase telemetry when they lack the runner’s session identifier. The scoreboard’s main-session join also leaves out separately launched fixer and backlog-checking sessions. A displayed cost therefore needs an explicit scope before I use it in a comparison.

You can begin locally with one record per attempt and links to the evidence:

Keep Why it matters
Task criteria, run ID, workflow version, model settings Establish what was attempted and under which conditions
Tested revision and validation logs Identify the code and checks behind the result
Initial findings, fixes, and final review Retain correction effort as well as the ending verdict
Delivery outcome and outstanding criteria Distinguish finished implementation from demonstrated behavior
Tokens, reported cost, duration, and accounting scope Describe the resources included in the measurement

For the signup attempt, a compact summary could look like this. This is a suggested summary format assembled from the retained records; it is not the runner’s native schema. The task label is shortened, and private identifiers and paths are omitted.

{
  "task": "Signup consent",
  "tests": {"passed": 5048, "failed": 0, "skipped": 9},
  "correctness": {"initial": "PASS", "final": "PASS", "warnings": 3},
  "security": {"initial": "PASS", "final": "PASS", "findings": 0},
  "review_fix_cycles": 0,
  "merged": true,
  "browser_database_check": "outstanding",
  "delivery": "unverified",
  "cost": {"usd": 2.1971172, "basis": "list", "scope": "implementation session"}
}

In your own record, include the actual task and run identifiers, tested revision or tree digest, and paths to the underlying evidence. Keep the workflow version and model settings too, so a later comparison can identify what changed. This summary makes the unresolved outcome visible alongside the cost; the linked records let you inspect why.

For a first comparison, choose one workflow change, such as moving repeated command selection from the agent into a script. Keep the acceptance checks stable and use the same task cases and starting revisions where practical. Record failed attempts and repairs as well as clean runs. If you use production tasks from two different weeks, describe the task mix; a cheaper week alone cannot establish what caused the difference.

Report completion rate alongside cost. Cost per completed task should include the failed attempts and repair work within the declared scope, divided by the tasks whose required outcomes were actually verified. If none completed, leave that measure undefined. If part of the cost was not captured, disclose the missing part.

The signup attempt would remain in the costs and in the unresolved-outcome count for this evaluation window. It would enter the delivered-task count only when the required evidence arrived. That rule is worth deciding before looking at which version of the workflow appears cheaper.


This walkthrough uses retained task, execution, validation, review, and delivery records from September 29, 2026. Excerpts omit private identifiers and unrelated fields. It describes that recorded attempt, not the feature’s current production status. The CLI exercise was executed separately while preparing the article; it does not rerun the signup implementation. The central pipeline description was checked against configuration and source, not a fresh successful ingestion query.

Prepared with Codex from Yianna’s workflow and retained records. Reviewed and approved by Yianna.