Can Jev make Qwen better at discovery?

Field Note 003 · September 29, 2026

Adding Jev to a local Qwen workflow helped repair a mistake that Qwen’s own review left unresolved. It did not make the workflow reliably better overall. On the final comparison, the Jev pipeline produced more fully acceptable answers, but also more failed responses and a lower average score than Qwen reviewing its own work.

In an earlier field note, I tested Jev as a claim checker in a writing workflow. This time I wanted to know where a local model could benefit from another model’s decisions. We built the experiment around a small source-reading task: explain what a piece of Ruby code would do, and cite the supplied evidence. This tested part of an exploring workflow, not an agent completing a software project.

Qwen wrote an initial answer. Either Qwen or TypeSafe’s Jev then decided which parts to keep, repair, supplement with citations, or flag as unresolved. Qwen performed the requested repairs in both versions. Jev chose the work; it did not write the replacement answer.

Each comparison shared the same draft, source passages, questions and available repair actions. Source files came from saved Git commits. Questions, reference answers and scoring rules were fixed before the relevant comparison, so neither pipeline could benefit from a more convenient source revision.

Earlier development tests showed why the distinction between deciding and repairing mattered. Among 15 completed pairs, Jev flagged all 17 source-invalid draft items, while Qwen flagged 14. Both pipelines repaired 10. Finding more problems had not produced more fully acceptable answers. Some repairs still contained false statements; others omitted citations needed to support otherwise correct explanations.

Answer structure mattered too. In a separate development comparison, changing only the JSON field order so the explanation preceded the result increased fully acceptable answers from 4/16 to 10/16. Two outputs failed citation validation, so that intervention also missed its full acceptance rule. We kept the same explanation-first repair format for both decision models in the final comparison.

The final test used eight new questions from four previously unused Ruby files, with two repetitions per question. The baseline kept Qwen’s draft unchanged. The other two versions applied their selected repairs.

Final result Unchanged Qwen review Jev review
Fully acceptable answers 4/16 6/16 7/16
Questions acceptable in both repetitions 2/8 3/8 3/8
Average quality score /100 61.98 73.44 66.15
Failed outputs 2/16 2/16 4/16

“Fully acceptable” required every finding to be correct and supported by its citations, every requested outcome to be covered, and any open questions to represent real gaps. The numerical score split credit equally between supported findings and covered requirements. Failed outputs scored zero. Eight questions repeated twice remain eight questions, not sixteen independent examples.

One question asked how a warning parser groups these lines:

File: app/a.rb
Confidence: Low
Category: SQL

The parser starts a record when it sees File. A new Confidence line saves that record and starts another. The result should therefore contain two records: one with the file, and another with confidence and category.

Qwen’s initial answer combined them into one. Its explanation described saving the first record, but its stated result still contradicted that explanation. Qwen’s own decision-and-repair path left the answer unacceptable. With Jev choosing repairs, Qwen produced an acceptable answer to the whole question in both repetitions. That is a concrete place where the combination helped.

In the earlier development comparison, both review paths corrected an answer but left its evidence incomplete. A failed lint check stopped execution before a later budget check. The repaired answer described that correctly, but its citations omitted the definition needed to establish why lint took that branch. The correction improved the answer without making it fully acceptable under our scoring rules.

The failed responses matter just as much. In both repetitions of a different question, Qwen invented citation identifiers instead of using the supplied identifiers. Those shared drafts failed validation before either decision model could act, costing all three versions an answer each time.

The Jev branch had two additional failures. Its returned probabilities summed to 0.99 for one decision in each response. Our adapter required a sum within 0.001 of one and rejected both. One rejection discarded a path whose draft was already fully acceptable. That was a failure at the interface between Jev and our validator; it does not by itself establish a wrong judgment about the source.

The earlier claim-checking study had also encountered distributions totaling 0.99. This experiment retained the strict validator and counted rejected responses as failures. Relaxing it after seeing the results would test a different pipeline. Looking only at the 12 pairs where both pipelines returned valid outputs, Jev scored higher, but that view leaves out failures the workflow actually encountered.

We set the acceptance rule before the final comparison: Jev needed more questions answered correctly in both repetitions, no loss of previously acceptable answers, no increase in failed outputs, and a meaningful average-score gain. It missed that rule. The complete experimental record covers 128 output slots across the three comparisons, including failures; none of those counts establishes broad model superiority.

Source judgments were prepared with Codex and remain provisional, without independent certification. The tests ran on a shared machine with memory-pressure pauses, so their timings do not establish isolated hardware performance. No stronger model was used for repair or escalation.

I would keep Qwen as the baseline and consider Jev for specific decisions where it has earned its place. The next investigation would need to address citation handling, the probability-response contract, and whether selected repairs actually remove the error. This round is finished. It gave us a useful repair example and a clearer account of what still prevents that example from becoming a dependable workflow.

Experiment record: September 24, 2026. Qwen 3.6 35B-A3B, Q4_K_M quantization with MTP enabled; Jev 1.13.0. Prepared with Codex from the preserved protocols, source snapshots, responses and scorecards. No workflow change was promoted from this round.