---
title: "The Judge Is Part of the Instrument"
subtitle: "Two judges read the same corpus on the same day — one saw four confirmed inferences, the other none. When your measuring device is a language model, its error profile is part of every number it produces."
date: 2026-08-19
events: "17.8.2026"
tags: ["somnus"]
summary: "A 17× speed anomaly exposed an unrecorded judge swap. A pre-registered test then killed our favourite hypothesis about the judge's errors — and field data revealed the real one: the fast judge misses genuine inferences the careful judge confirms unanimously. Structure is not a property of a corpus. It is a property of a corpus–judge pair."
canonical: https://sf3d.fi/blog/the-judge-is-part-of-the-instrument
html: https://sf3d.fi/blog/the-judge-is-part-of-the-instrument
author: Petri Korhonen, SF3D AI Lab (Sunrise Software Oy, Tampere, Finland)
site: https://sf3d.fi (index of all articles in Markdown: https://sf3d.fi/llms.txt)
numbers: measured on the lab's own hardware unless the text marks them as held for the AMD R9700 whitepaper; the Evidence block at the end of an experimental article names the run, the date, n and the pre-registration each claim rests on
---
<p class="standfirst">Part 1 introduced a machine that refutes its own ideas; Part 2 showed its value function being rebuilt to respect structure. This is the story of the component that <em>reads</em> that structure — the entailment judge — and of the week we learned that every graph it draws carries the judge's fingerprint. It ends with a division of labour, a rule about locking decisions before seeing data, and one number that should worry anyone using an LLM as a measuring device: <code>0 of 4</code> versus <code>4 of 4</code>.</p>

<aside class="primer">

**PRIMER**

Somnus is a research loop on one GPU box: a local model proposes candidate research questions from gaps between existing findings, a second local model screens them, and the best are escalated to cloud models that try to refute them. Claims that survive go into a store, and a local judge decides whether one claim follows from another; those edges are the structure most parts of this series measure.

</aside>

## The instrument in question

Somnus accumulates verified claims — crystals — and the interesting question is never a single claim. It is the *edges*: does claim A entail claim B? Contradict it? Or are they merely neighbours in embedding space with nothing to say to each other? A local judge model reads claim pairs and issues one of three verdicts — *follows*, *contradicts*, *neutral*. This is the classic natural-language-inference task (Bowman et al., 2015), with the model cast in the role the evaluation literature now calls LLM-as-judge (Zheng et al., 2023) — and out of thousands of these verdicts the pipeline assembles an entailment graph. Components of that graph are what the synthesis stage feeds on. If the judge is wrong, everything downstream is wrong, quietly.

<aside class="plain">

**IN PLAIN TERMS**

Ask two radiologists to read the same X-ray. One says "clear", the other points at a shadow. The scan did not change — the reader did. This article is about discovering that our system had two radiologists on staff, that nobody had written down which one read which scan, and that they disagree far more than anyone assumed.

</aside>

For weeks there had been two judges on the box without the pipeline ever framing them as alternatives: **Mistral Small 24B** with thinking off — the workhorse screen, roughly three seconds a pair — and a **Qwen thinking model** — the deliberate one, nearly a minute a pair, seventeen times slower. Different services, different ports, one at a time in VRAM. The assumption, never written down and therefore never examined, was that they were interchangeable enough.

## The speedup that wasn't

The assumption fell to a latency number. A screening run finished at three seconds per pair when the baseline recorded before the run said fifty-three. A `17×` speedup that nobody built is not a speedup — it is a different instrument. The logs settled it: one experiment's forced-pair matrix had run on the thinking judge, a later experiment's screen on the fast one, and the two sets of readings had been compared as if they came from the same device.

Re-judging the same fifteen pairs with both models measured the gap directly: the two judges disagreed on `40%` of verdicts, and of the pairs where each judge found *something*, only three were found by both. Neither judge is the truth. But readings from different judges are not comparable, and nothing in the data format had been forcing anyone to remember which judge produced which number.

Two permanent rules came out of that afternoon. Every verdict row now carries a judge tag. And a baseline recorded before a run is not bureaucracy — it is the tripwire that turns "nice, it got faster" into "wait, *why* is it faster?"

<aside class="nerd">

**NERD BOX**

#### How a config file lies and a log doesn't

The judge endpoint is a config key: set, it points at the thinking judge; unset, calls fall through to the fast screen model. Reading the *current* config tells you which judge runs *now* — not which judge ran two days ago. The proof came from service logs: in the exact two-hour window of the earlier run whose judge was in question, the thinking judge's service wrote `22,837` log lines and the screen wrote `5`. The lesson generalises: **measure the topology from logs, never infer it from configuration.** This failure class has now been caught three times in this project, in three different subsystems.

</aside>

## The first ground truth

Suspicion needs a benchmark. The physics-pair experiment (part of the question-form arc — pre-registered questions whose edge structure is known in advance) provided the first one: pairs *planted* in the corpus with known answers. A general mechanism and its parameter-free algebraic consequence — that edge should exist. Two systems that merely instantiate the same mechanism with different parameters — no edge, analogy is not entailment.

On the four measurable planted pairs: the thinking judge scored `4/4`; the fast judge `3/4`, and its one error was a false positive — it read an entailment where there was only siblinghood, two claims descending from a common premise. Four pairs is not a measurement, but it was enough to name a hypothesis: the fast judge has a *sibling error* — show it two children of the same parent and it will call one the parent of the other.

## The cheapest experiment of the month

The sibling-error hypothesis got the full treatment, at a total compute cost of zero euros: twenty-five hand-assembled pairs — ten sibling pairs from the corpus, ten built fresh across unrelated domains, and five genuine positive controls so that a judge who answers *neutral* to everything cannot win. Both judges, same pairs. And crucially, **the decision rule was locked before the run**: only if the fast judge's false-positive rate on siblings reached `25%` — and the careful judge stayed clean, and found the planted positives — would edge-critical screening move to the slow judge. Otherwise its `17×` throughput advantage stands.

The hypothesis died cleanly. The fast judge produced **zero false entailments in twenty sibling pairs** — against a predicted `30%`. The error we had extrapolated from a sample of four simply is not a tendency; re-run, the original offending pair judged *neutral*. This is what a pre-registered decision rule is for — a discipline borrowed wholesale from the replication-crisis literature (Nosek et al., 2018). On intuition alone we would have switched judges that afternoon, citing an error profile the data had just refuted.

But the test revealed the real profiles, and they were nothing like the suspicion. The fast judge's actual failure modes: it declares **false contradictions** — handed the ideal gas law and the statement that Boyle's law follows from it at constant temperature, it called the pair *contradictory* — and it **misses genuine positives**, finding only three of the five planted inferences. The careful judge found **five of five**, at the price of one sibling false positive (`5%` — landing exactly on its predicted point). One judge is conservative to a fault; the other is sensitive and slightly trigger-happy. Neither profile is "better". They are different instruments.

<figure>
<img src="/images/blog4-profiles.svg" alt="Two error-profile columns, fast judge versus careful judge: sibling false entailments 0 of 20 versus 1 of 20; sibling-pair false contradictions two each; planted positives found 3 of 5 versus 5 of 5." width="880" height="520" />
<figcaption>Two instruments, two failure modes. Sibling-pair false contradictions: fast judge 2 (one bidirectional), careful judge 2 (one-directional each) — and of the fast judge's two planted-positive misses, one was an active <em>contradicts</em> verdict.</figcaption>
</figure>

<aside class="math">

**MATH BOX**

#### Two axes, one theorem

Signal detection theory has priced this trade-off since the 1960s (Green & Swets, 1966): a detector has a *sensitivity* — the fraction of real signals it catches — and a *specificity* — the fraction of non-signals it correctly rejects, and no threshold setting improves one without spending the other. From this article's own numbers: the fast judge ran at sensitivity `3/5 = 0.60` with sibling-pair specificity `20/20 = 1.00`; the careful judge at sensitivity `5/5 = 1.00` with specificity `19/20 = 0.95`. Neither dominates — they occupy different corners of the same ROC space, which is exactly why the answer is a division of labour rather than a winner. The three-repetition majority vote rests on an older result still: Condorcet's jury theorem (1785). If a single verdict is correct with probability `p > 0.5`, the majority of three is right with probability `p³ + 3p²(1−p)` — at `p = 0.7`, the majority reaches `0.78`; at `p = 0.8`, `0.90`. Repetition buys reliability — agreement with itself — which is necessary but not sufficient: a judge can be reliably, repeatably wrong, which is psychometrics' distinction between reliability and validity (and the reason planted ground truth, not re-asking, settles what a verdict is worth). Chance-corrected agreement between two judges is its own measured quantity (Cohen's κ, 1960) — the raw `60%` agreement here overstates how aligned the judges are, since three-verdict tasks agree by luck alone more often than intuition suggests.

</aside>

## The field decides

The same day supplied the verdict that mattered, from live data rather than a benchmark. Two fresh ignite runs had produced twelve new crystals, and their internal edge structure was measured with both judges — every pair, forced, nothing left to sampling.

The fast judge found **zero edges in either run**. The careful judge found **four — every one of them confirmed by a three-repetition majority vote**: a component of three crystals in each run, real inferential structure, unanimous on re-ask. Set beside the planted-positive misses, the day's ledger reads: fast judge `0/4` field edges, careful judge `4/4`.

That number resolved two open questions at once. An efficiency experiment earlier that day had appeared to show that a cheaper model configuration produced crystals with *no structure* — alarming, if true. It was the screen's blindness, not the configuration's sterility: the structure was there, the fast judge could not see it. And a chain-form experiment carried a pre-registered kill branch — *if fewer than two chain edges form, the chain explanation is wrong* — that would have **fired falsely** had its edges been counted by the fast judge alone.

There is no way to say this gently: **structure is not a property of a corpus. It is a property of a corpus–judge pair.** A falsification machine whose judge misses inferences will report an absence of structure that is not an absence — the most expensive kind of wrong, because it reads as a clean negative result.

<figure>
<img src="/images/blog4-two-graphs.svg" alt="The same twelve crystals drawn twice: on the left, judged by the fast model, twelve isolated nodes; on the right, judged by the careful model, two three-node components with directed edges, one of them mutual." width="880" height="520" />
<figcaption>Same corpus, same day, two judges — every right-hand edge unanimous over three repetitions.</figcaption>
</figure>

## A division of labour, not a winner

The decision that closed the week was deliberately not "switch to the better judge". The fast judge earned its place with the benchmark's cleanest line: zero false entailments in the false-positive test, at seventeen times the throughput — exactly what you want from a **pre-screen** that grinds thousands of candidate pairs and must, above all, not invent edges. The careful judge became the **edge-measurer**: any number that feeds a decision — component sizes, chain-edge counts, bridge checks, kill branches — is measured in the slow topology, at a cost of about twenty-five minutes per run, before anyone leans on it.

The general form of the rule travels beyond this system. When an LLM is your measuring device, sensitivity and specificity are separate axes, they trade off differently in different models, and no single judge is "the" instrument. You either measure the profile of the one you have and staff it accordingly — or you inherit its blind spots as facts about the world. The calibration behind these thresholds — and the judges' measured throughput and power on this bench — belongs to the whitepaper.

<aside class="plain">

**IN PLAIN TERMS**

Football solved this years ago. The referee on the pitch is fast and almost never awards a goal that did not happen — but sprinting alongside play, they miss real ones. The video assistant is slow, sees everything, and occasionally over-calls. So the game uses both: the referee runs the match, and no decisive call stands until the slow camera has looked at it. That is precisely the arrangement the two judges ended up in.

</aside>

## What survives this section

Four rules, all paid for:

1. **Tag every reading with the judge that produced it.** Untagged readings from mixed judges are not a dataset; they are an accident report waiting to be written.
2. **Lock decision rules before the data arrives.** The locked threshold is the only reason an intuitive-but-wrong judge swap didn't happen here.
3. **Latency is also data.** The `17×` anomaly was the entire discovery; a run that finishes suspiciously fast has usually answered a different question than the one you asked.
4. **A negative structural result is only as strong as the judge's sensitivity.** Before reporting "no structure", ask the sensitive judge — the fast one's silence proves nothing.

The machine still refutes its own ideas with executable code. This week it refuted an idea about itself — and the refutation was cheap, pre-registered, and wrong in the most instructive possible direction.

<aside class="reachout">

**FOR RESEARCHERS**

#### Working on something like this?

If you use LLMs as judges — for entailment, for RAG evaluation, for anything where the model's verdict becomes a number in a table — the judge-conditionality problem here is probably in your data too. If you have measured it, or want to compare error profiles across judge models, I'd like to hear from you: [pete@sunrisesoftware.app](mailto:pete@sunrisesoftware.app).

</aside>

## Sources & further reading

- S.R. Bowman et al., *A large annotated corpus for learning natural language inference* (EMNLP, 2015) — the NLI task this judge performs.
- L. Zheng et al., *Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena* (NeurIPS, 2023) — LLM judges and their biases as a research topic.
- D.M. Green & J.A. Swets, *Signal Detection Theory and Psychophysics* (1966) — sensitivity and specificity as the two axes of any detector.
- Marquis de Condorcet, *Essai sur l'application de l'analyse à la probabilité des décisions* (1785) — why a majority of imperfect voters beats one.
- J. Cohen, *A coefficient of agreement for nominal scales* (1960) — chance-corrected inter-rater agreement.
- B.A. Nosek et al., *The preregistration revolution* (PNAS, 2018) — locking decision rules before the data arrives.

<aside class="evidence">

**EVIDENCE**

#### The sibling-pair test (T), the chain question (K3) and the effort run · 17.8.2026 · `$0` for T, cloud runs for K3

- pre-registrations: `SOMNUS-T-KOE-ENNAKKOREKISTEROINTI` v1, locked 2026-08-17T15:31Z before the run (commit `f29e1ff` in the Somnus repo as recorded; data `experiments/entailment/t-sibling-*.jsonl`); the chain question `SOMNUS-KETJUKYSYMYS` v1, 2026-08-16T18:05Z; results `SOMNUS-T-KOE-TULOKSET` v1 (2026-08-17T16:07Z) and `SOMNUS-K3-TULOKSET` v1 (2026-08-17T17:55Z); Somnus `LN-20260817-kolme-koetta-ja-tuomariehdollisuus`.
- claim: the sibling-error hypothesis for the fast judge fell — false "follows" on `20` sibling pairs: `0/20` (predicted `30%` [`10–60`], below the lower bound); the careful judge `1/20` (predicted `5%`, a point hit); true positives `3/5` vs `5/5` — status: measured. The decision rule locked before the run ("move the screen only if T1 ≥ 25 % and T2 ≤ 10 % and T3 ≥ 3/5") did not fire: the fast judge stays as the screen; its measured error profile is false *contradictions* and missed positives, not false entailments.
- claim: structure is judge-conditional — on the chain run the fast judge found `0/4` field edges and the careful judge `4/4`, in three independent datasets; the chain's answer key `4/4`; chain edges `2` against a predicted `3` [`2–4`]; the Titchmarsh condition named without being asked — source `SOMNUS-K3-TULOKSET` v1 — status: measured
- correction on the record: the throughput ratio quoted in the text (`3 s` vs `53 s` per pair, `17×`) rests on the earlier estimate; the T run measured the careful judge at about `45 min` for `20` pairs (`~105 s` per pair) against about `2 min` for the fast one — the direction holds, the ratio in the text is stale — status: overturned (the number), measured (the direction)
- claim: the day's cloud runs cost `$8.99` for `12` crystals — the one headline number of the day; the per-run series is held with the AMD R9700 whitepaper — status: measured
- not shown: n is `20` sibling pairs and `5` positives, one corpus, two judges — the sensitivity/specificity figures are indications, not a calibration; the "caught three times" in the text names no incidents in the record.

</aside>

---

<p class="sig">— SF3D</p>
