Part 3 showed that the judge reading the corpus is part of the instrument. This part goes one stage deeper, to the component the whole funnel exists to feed: abduction — the step that is supposed to look at a cluster of verified claims and produce the one mechanism that explains them. In production it has never produced one that survived validation. Not once. The story of why is the most instructive zero this project has generated, and it ends somewhere unexpected: with a map.
The gate that never opens
The word is Charles Sanders Peirce’s: abduction, the third mode of inference beside deduction and induction — the leap from observations to the hypothesis that would best explain them (Peirce, 1903). Somnus operationalises the leap and then validates it the hard way. A candidate mechanism, abduced from a cluster of verified claims, must retrodict: predict held-out claims it never saw. A gate scores that — a lift comparison against a genericised twin of the hypothesis, and a specificity floor — and only a hypothesis that clears both enters the knowledge base. The thresholds themselves, and their calibration, belong to the whitepaper. The design assumption was that the gate would be a filter. In practice it was a wall: run after run, zero passes.
The definitive measurement: fifteen clusters, six hours and forty minutes, two batches, zero errors — and zero passes. Twelve of the fifteen hypotheses were so generic that the specificity floor stopped them before a single prediction was ever scored; the three that reached scoring produced nothing above judge noise. Numbers like these do not say the threshold is slightly too strict. They say the phenomenon the gate was calibrated to detect is not occurring.
The proximate cause was visible in the output. Asked for one minimal mechanism, with trivial generalisation explicitly forbidden in the prompt, the model produced — word for word, in translation — “the system’s internal mechanisms and parameters are tightly optimised”. Karl Popper put his finger on why that sentence is worthless as an explanation: the empirical content of a statement lives in what it forbids (Popper, 1959). A generality forbids nothing — it is compatible with every held-out claim and forced by none, so it can never score. The prohibition was in the prompt, the example of what not to write was in the prompt, and the model wrote it anyway. Worse: three of the four failures echoed the forbidden example’s own vocabulary — the ban had been working as a template. Elsewhere in this codebase the same pattern has hardened into a rule: the prompt requests, the code enforces. A ban that cannot be checked mechanically is not a ban.
Ruling things out, one paid run at a time
What makes this zero valuable rather than merely disappointing is the list of explanations that were killed before the diagnosis settled — each by a measurement, not an opinion.
The held-out set really was broken: random sampling had handed a geometry cluster audio-programming claims to predict. It was fixed — nearest non-members instead of random draws — and the prediction, registered before the re-run, was that accuracy would rise. It fell, from 10.0% to 6.7%. The fix was correct and it was not the bottleneck; a registered prediction is the only reason that distinction survived (the practice is preregistration, imported wholesale from the replication-crisis literature — Nosek et al., 2018). The threshold could not be recalibrated: twelve of fifteen hypotheses never even reached the stage the threshold governs, and skips plus noise offer no distribution to calibrate against. Finer clustering raised specificity and lowered predictive accuracy. The judge had been separately validated. Contradiction handling was deliberately switched off for the decisive run, so the zero could not hide behind a second variable.
And the run’s one apparent survivor did not survive inspection. The single cluster that produced any lift also carried the run’s highest specificity score — for an afternoon, that looked like a direction: make the hypothesis precise enough and the chain works. A deep-dive dissolved both halves. The lift sat at the level of ordinary judge noise — every scored cluster threw the odd stray hit across repeated passes — and the specificity metric turned out to measure how much rare vocabulary a hypothesis carries, not whether it names a mechanism; the run’s “most specific” hypothesis was, read plainly, a generality in ornate dress. Even the flattering detail had to be killed by hand.
The planted control
At that point there were three suspects left: the model is too small for abduction, the prompt cannot carry the requirement, or the corpus cannot support a mechanism. One run separated them — at a compute cost of zero euros, by the same move the judge story used: plant material with a known answer.
Two invented clusters were written around known mechanisms — one in which every stage buffers exactly two blocks, one in which pores only function in pairs — each with eighteen consequence claims and ten held-out. A third cluster was the control’s control: fluent, same-register claims with no mechanism at all underneath. All three went through the untouched production pipeline.
On the coherent clusters, the same prompt, gate and judge that produce nothing in production produced lifts of 0.7 to 1.0, on both models — the smaller one found the mechanisms, the larger one found them perfectly, and its hypotheses even refined the planted mechanism into something sharper than what was hidden. The machinery works. The prompt carries. The bottleneck is the soil, not the plough.
The fabrication leak
The planted run also surfaced the result that matters far beyond this system. On the mechanism-free control cluster, the small model’s vague guesses scored 0.0–0.1 and the gate held, all three times. The larger model wrote a chemically fluent pseudo-mechanism — plausible vocabulary, confident structure, no truth underneath — and it beat its genericised twin on the held-out claims. The gate passed fabrication, twice out of twice (the third repetition produced no hypothesis at all).
Read that carefully, because it inverts the usual upgrade logic. The gate’s comparison measures specificity, not truth — a specific fabrication outscores a vague one, and fluency at fabrication grows with model capability. A validation gate that holds with a weaker model can leak with a stronger one. Swapping in the better model without a fabrication defence would have made the pipeline worse while looking like an upgrade — the same lesson Part 3 taught about judges, one stage deeper: every measuring device in the loop has an error profile, and capability shifts it.
The archipelago
If the soil is the problem, map the soil. The whole deduplicated corpus — 233 verified claims — was screened for entailment — 779 candidate pairs drawn from its embedding neighbourhoods — with every non-neutral verdict confirmed by a three-repetition majority (87% of the non-neutral verdicts survived confirmation). The result is the picture this arc had been missing.
The corpus is seventeen components. The genuine mechanism islands are clean and small — an audio-filter inference chain of three claims at full density, a self-referential cluster about affective profiles — and two-to-three-claim islands are exactly too small to abduce from: nothing left to derive, nothing left to hold out. (One of those islands would later turn out to be wrong in an instructive way — a later part’s story.) The one large component turned out to be five variants of the same fact, not a structure. And the conflict half of the graph told its own story: most conflict edges radiate from a single falsification claim mechanically demolishing a family of numerology, one edge at a time — the refutation machinery working exactly as designed, in the same graph where consolidation has nothing to consolidate.
The corpus is not empty and not broken. It is fragmented — an archipelago where abduction needs a continent. That single sentence replaced weeks of suspicion aimed at prompts, thresholds and models, and it re-aimed the effort: the next lever is growing the corpus connectedly, not tuning the machinery. A first targeting experiment made the point immediately — an expensive run aimed straight at the strongest island produced content, and zero structure: 135 forced pairs between the island and the new claims, no edges at all. Aiming at the island was not enough; something about the form of the question decides whether new material connects. That question got its own experiment series — and its own part in this diary.
What survives this section
- A null result becomes a direction when you kill its explanations one at a time. Fifteen zeros plus a ruled-out list is a measured research problem; fifteen zeros alone is a bad mood.
- Plant a control whose answer you know. A zero-euro fixture separated broken machinery from thin soil — and it is reusable.
- Fabrication risk grows with capability. A specificity gate cannot tell explanation from fluent invention, and a gate that holds with a weak model can leak with a strong one. Test the gate against your best model before upgrading to it.
- Register the prediction before the fix. “The fix was correct AND it was not the bottleneck” is a distinction only a pre-registered prediction can make.
The machine set out to explain its corpus and discovered, with some precision, that its corpus was not yet explainable — and that is a result: mapped, priced, and pointing at the next experiment. The diary’s next part is about what happened when the question itself went under the microscope.
Sources & further reading
- C.S. Peirce, Lectures on Pragmatism (1903) — abduction as the logic of explanatory hypothesis.
- K. Popper, The Logic of Scientific Discovery (1959) — empirical content as prohibition; falsifiability.
- J. Rissanen, Modeling by shortest data description (1978) — MDL, the series’ pricing frame (Part 2).
- R. Solomonoff (1964) and A.N. Kolmogorov (1965) — algorithmic probability and complexity, the roots of “explanation = compression”; modern treatment in P. Grünwald, The Minimum Description Length Principle (2007).
- B.A. Nosek et al., The preregistration revolution (PNAS, 2018) — the register-before-you-run discipline used throughout this project.
— SF3D