---
title: "Abduction Produces Generalities"
subtitle: "In production, the synthesis gate has never passed a single hypothesis — and it was right every time. What a machine's most stubborn zero taught us about explanation, fabrication, and why the corpus turned out to be an archipelago."
date: 2026-08-19
events: "10.–13.8.2026"
tags: ["somnus"]
summary: "In production, the abduction stage has never produced an explanation that survived validation — fifteen clusters, zero passes. Killing every alternative explanation, one paid run at a time, led to a planted control, a fabrication leak that grows with model capability, and a map: the corpus is an archipelago, and abduction needs a continent."
canonical: https://sf3d.fi/blog/abduction-produces-generalities
html: https://sf3d.fi/blog/abduction-produces-generalities
author: Petri Korhonen, SF3D AI Lab (Sunrise Software Oy, Tampere, Finland)
site: https://sf3d.fi (index of all articles in Markdown: https://sf3d.fi/llms.txt)
numbers: measured on the lab's own hardware unless the text marks them as held for the AMD R9700 whitepaper; the Evidence block at the end of an experimental article names the run, the date, n and the pre-registration each claim rests on
---
<p class="standfirst">Part 3 showed that the judge reading the corpus is part of the instrument. This part goes one stage deeper, to the component the whole funnel exists to feed: <em>abduction</em> — the step that is supposed to look at a cluster of verified claims and produce the one mechanism that explains them. In production it has never produced one that survived validation. Not once. The story of why is the most instructive zero this project has generated, and it ends somewhere unexpected: with a map.</p>

<aside class="primer">

**PRIMER**

Somnus is a research loop on one GPU box: a local model proposes candidate research questions from gaps between existing findings, a second local model screens them, and the best are escalated to cloud models that try to refute them. Claims that survive go into a store, and a local judge decides whether one claim follows from another; those edges are the structure most parts of this series measure.

</aside>

## The gate that never opens

The word is Charles Sanders Peirce's: *abduction*, the third mode of inference beside deduction and induction — the leap from observations to the hypothesis that would best explain them (Peirce, 1903). Somnus operationalises the leap and then validates it the hard way. A candidate mechanism, abduced from a cluster of verified claims, must *retrodict*: predict held-out claims it never saw. A gate scores that — a lift comparison against a genericised twin of the hypothesis, and a specificity floor — and only a hypothesis that clears both enters the knowledge base. The thresholds themselves, and their calibration, belong to the whitepaper. The design assumption was that the gate would be a filter. In practice it was a wall: **run after run, zero passes.**

The definitive measurement: fifteen clusters, six hours and forty minutes, two batches, zero errors — and zero passes. Twelve of the fifteen hypotheses were so generic that the specificity floor stopped them before a single prediction was ever scored; the three that reached scoring produced nothing above judge noise. Numbers like these do not say *the threshold is slightly too strict*. They say the phenomenon the gate was calibrated to detect is not occurring.

The proximate cause was visible in the output. Asked for one minimal mechanism, with trivial generalisation explicitly forbidden in the prompt, the model produced — word for word, in translation — *"the system's internal mechanisms and parameters are tightly optimised"*. Karl Popper put his finger on why that sentence is worthless as an explanation: the empirical content of a statement lives in what it **forbids** (Popper, 1959). A generality forbids nothing — it is compatible with every held-out claim and forced by none, so it can never score. The prohibition was in the prompt, the example of what not to write was in the prompt, and the model wrote it anyway. Worse: three of the four failures echoed the forbidden example's own vocabulary — the ban had been working as a template. Elsewhere in this codebase the same pattern has hardened into a rule: **the prompt requests, the code enforces.** A ban that cannot be checked mechanically is not a ban.

<aside class="plain">

**IN PLAIN TERMS**

A horoscope fits everyone — which is exactly why it tells you nothing about anyone. The machine kept writing horoscopes for its data: statements so broad they were compatible with everything and committed to nothing. The gate kept refusing them, run after run. The gate was the one part of the loop doing its job.

</aside>

<aside class="math">

**MATH BOX**

#### What a mechanism is, in bits

Part 2 priced hypotheses by Minimum Description Length: transmit the law, then the exceptions — `L(H) + L(D | H)` (Rissanen, 1978; the lineage runs back through Kolmogorov, 1965, and Solomonoff, 1964). A generality sits at the degenerate end of that trade: `L(H)` is tiny, but it compresses nothing — `L(D | H)` stays as long as `L(D)` — so total description length never drops. Popper's dictum is the same fact in logic instead of bits: content = prohibition. The gate's genericised-twin comparison is a working approximation of this test — strip the hypothesis of its specifics, and if the stripped twin retrodicts held-out claims just as well, the specifics were decoration, not mechanism. What the gate *measures* is therefore compression relative to a shadow of itself; what it *cannot* measure is whether the compression comes from truth. That gap is this article's fourth section. Full formalism and the gate's calibration: the whitepaper.

</aside>

## Ruling things out, one paid run at a time

What makes this zero valuable rather than merely disappointing is the list of explanations that were killed before the diagnosis settled — each by a measurement, not an opinion.

The held-out set really was broken: random sampling had handed a geometry cluster audio-programming claims to predict. It was fixed — nearest non-members instead of random draws — and the prediction, **registered before the re-run, was that accuracy would rise. It fell**, from `10.0%` to `6.7%`. The fix was correct and it was not the bottleneck; a registered prediction is the only reason that distinction survived (the practice is preregistration, imported wholesale from the replication-crisis literature — Nosek et al., 2018). The threshold could not be recalibrated: twelve of fifteen hypotheses never even reached the stage the threshold governs, and skips plus noise offer no distribution to calibrate against. Finer clustering raised specificity and *lowered* predictive accuracy. The judge had been separately validated. Contradiction handling was deliberately switched off for the decisive run, so the zero could not hide behind a second variable.

And the run's one apparent survivor did not survive inspection. The single cluster that produced any lift also carried the run's highest specificity score — for an afternoon, that looked like a direction: make the hypothesis precise enough and the chain works. A deep-dive dissolved both halves. The lift sat at the level of ordinary judge noise — every scored cluster threw the odd stray hit across repeated passes — and the specificity metric turned out to measure how much rare vocabulary a hypothesis carries, not whether it names a mechanism; the run's "most specific" hypothesis was, read plainly, a generality in ornate dress. Even the flattering detail had to be killed by hand.

## The planted control

At that point there were three suspects left: the model is too small for abduction, the prompt cannot carry the requirement, or the corpus cannot support a mechanism. One run separated them — at a compute cost of zero euros, by the same move the judge story used: **plant material with a known answer.**

Two invented clusters were written around known mechanisms — one in which every stage buffers exactly two blocks, one in which pores only function in pairs — each with eighteen consequence claims and ten held-out. A third cluster was the control's control: fluent, same-register claims with **no mechanism at all** underneath. All three went through the untouched production pipeline.

On the coherent clusters, the same prompt, gate and judge that produce nothing in production produced lifts of `0.7` to `1.0`, on both models — the smaller one found the mechanisms, the larger one found them perfectly, and its hypotheses even *refined* the planted mechanism into something sharper than what was hidden. The machinery works. The prompt carries. **The bottleneck is the soil, not the plough.**

<aside class="nerd">

**NERD BOX**

#### The fixture that settles three arguments at once

Each planted cluster is a self-contained fixture: a hidden mechanism, eighteen claims that follow from it, ten held-out claims for retrodiction. The negative control shares surface register with the positives but no generating mechanism — so a pipeline that "finds" one there is fabricating. Predictions were registered before the run. Both positive-gate predictions under-shot — `5` of `6` reps passed against a predicted `~50%`, `6` of `6` against `~75%` — and the negative-control prediction was refuted outright, which turned out to be the most valuable line in the run: it is the next section. The fixture is reusable — it has since become the calibration instrument for any change to the abduction stage.

</aside>

## The fabrication leak

The planted run also surfaced the result that matters far beyond this system. On the mechanism-free control cluster, the small model's vague guesses scored `0.0–0.1` and the gate held, all three times. The larger model wrote a chemically fluent pseudo-mechanism — plausible vocabulary, confident structure, no truth underneath — and it **beat its genericised twin on the held-out claims. The gate passed fabrication, twice out of twice** (the third repetition produced no hypothesis at all).

Read that carefully, because it inverts the usual upgrade logic. The gate's comparison measures *specificity*, not *truth* — a specific fabrication outscores a vague one, and fluency at fabrication grows with model capability. **A validation gate that holds with a weaker model can leak with a stronger one.** Swapping in the better model without a fabrication defence would have made the pipeline worse while looking like an upgrade — the same lesson Part 3 taught about judges, one stage deeper: every measuring device in the loop has an error profile, and capability shifts it.

## The archipelago

If the soil is the problem, map the soil. The whole deduplicated corpus — `233` verified claims — was screened for entailment — `779` candidate pairs drawn from its embedding neighbourhoods — with every non-neutral verdict confirmed by a three-repetition majority (`87%` of the non-neutral verdicts survived confirmation). The result is the picture this arc had been missing.

The corpus is **seventeen components**. The genuine mechanism islands are clean and small — an audio-filter inference chain of three claims at full density, a self-referential cluster about affective profiles — and two-to-three-claim islands are exactly too small to abduce from: nothing left to derive, nothing left to hold out. (One of those islands would later turn out to be wrong in an instructive way — a later part's story.) The one large component turned out to be five variants of the same fact, not a structure. And the conflict half of the graph told its own story: most conflict edges radiate from a single falsification claim mechanically demolishing a family of numerology, one edge at a time — the refutation machinery working exactly as designed, in the same graph where consolidation has nothing to consolidate.

<figure>
<img src="/images/blog5-archipelago.svg" alt="The entailment graph drawn as a map of seventeen islands: a three-claim inference chain at full density, a five-claim clump that is one fact in five variants, a conflict hub radiating refutation edges, and a scatter of two-claim islets." width="880" height="560" />
<figcaption>The corpus, mapped: falsification connects, consolidation has no continent to stand on. Solid edges entail, dashed edges refute.</figcaption>
</figure>

<aside class="plain">

**IN PLAIN TERMS**

Imagine being asked to reconstruct the plot of a novel from three pages torn out of it — and the library turns out to be shelves of three-page fragments from different books. That is the corpus. The machine was not failing to find the story; in most places there was not yet enough of any one story to find.

</aside>

**The corpus is not empty and not broken. It is fragmented** — an archipelago where abduction needs a continent. That single sentence replaced weeks of suspicion aimed at prompts, thresholds and models, and it re-aimed the effort: the next lever is growing the corpus *connectedly*, not tuning the machinery. A first targeting experiment made the point immediately — an expensive run aimed straight at the strongest island produced content, and **zero structure**: `136` forced pairs between the island and the new claims, no edges at all. Aiming at the island was not enough; something about the *form of the question* decides whether new material connects. That question got its own experiment series — and its own part in this diary.

<figure>
<img src="/images/blog5-gate.svg" alt="Left panel: production — twelve hypotheses stopped at the specificity floor before scoring, three scored at judge-noise level. Right panel: planted controls scoring 0.7 to 1.0 on both models, and the fabrication pair — vague guesses held below the gate, one fluent pseudo-mechanism passing it." width="880" height="520" />
<figcaption>Same machinery, three soils. Production stops at the floor or scores at noise; planted mechanisms sail through; and one fluent fabrication slips past a gate that measures specificity, not truth.</figcaption>
</figure>

## What survives this section

1. **A null result becomes a direction when you kill its explanations one at a time.** Fifteen zeros plus a ruled-out list is a measured research problem; fifteen zeros alone is a bad mood.
2. **Plant a control whose answer you know.** A zero-euro fixture separated broken machinery from thin soil — and it is reusable.
3. **Fabrication risk grows with capability.** A specificity gate cannot tell explanation from fluent invention, and a gate that holds with a weak model can leak with a strong one. Test the gate *against* your best model before upgrading to it.
4. **Register the prediction before the fix.** "The fix was correct AND it was not the bottleneck" is a distinction only a pre-registered prediction can make.

The machine set out to explain its corpus and discovered, with some precision, that its corpus was not yet explainable — and *that* is a result: mapped, priced, and pointing at the next experiment. The diary's next part is about what happened when the question itself went under the microscope.

<aside class="reachout">

**FOR RESEARCHERS**

#### Working on something like this?

If you run abduction or hypothesis-generation over claim stores — especially with retrodiction gates, genericised-twin baselines, or LLM-scored validation — the fabrication leak here is worth checking in your own stack: an LLM-scored gate that holds with a weaker generator can leak with a stronger one, because fluency rises faster than truth. If you have measured that boundary, I'd like to compare notes: [pete@sunrisesoftware.app](mailto:pete@sunrisesoftware.app).

</aside>

## Sources & further reading

- C.S. Peirce, *Lectures on Pragmatism* (1903) — abduction as the logic of explanatory hypothesis.
- K. Popper, *The Logic of Scientific Discovery* (1959) — empirical content as prohibition; falsifiability.
- J. Rissanen, *Modeling by shortest data description* (1978) — MDL, the series' pricing frame (Part 2).
- R. Solomonoff (1964) and A.N. Kolmogorov (1965) — algorithmic probability and complexity, the roots of "explanation = compression"; modern treatment in P. Grünwald, *The Minimum Description Length Principle* (2007).
- B.A. Nosek et al., *The preregistration revolution* (PNAS, 2018) — the register-before-you-run discipline used throughout this project.

<aside class="evidence">

**EVIDENCE**

#### The abduction arc: the gate run, the planted control (A1), the entailment graph (A3), the targeted run (D1) · 10.–14.8.2026

- source records: `SOMNUS-ABDUKTIO-PULLONKAULA` v1 (2026-08-10T14:25Z) and its rebuttal `SOMNUS-ABDUKTIO-VASTINE` v2 (2026-08-10T19:34Z, which corrected three of the first record's measurement claims); `SOMNUS-A1-TULOKSET` v1 (2026-08-12T08:23Z); `SOMNUS-A3-TULOKSET` v1 (2026-08-13T14:45Z); `SOMNUS-D1-ENNAKKOREKISTEROINTI` v1, locked 2026-08-13T17:35Z; `SOMNUS-D1-IGNITE-TULOKSET` v4 (2026-08-22T16:32Z: the forced-pair count recomputed from data, `136`, not `135`); Somnus `LN-20260810`, `LN-20260812`, `LN-20260813`.
- claim: the production gate passed nothing — `15` clusters, `6 h 40 min`, `0` errors, `0` passes; specificity median `0.2481` against a gate of `0.33` — status: measured. Correction on the record: "`14` × lift `0.0`" was not fourteen measured zeros; the gate skips judging below the specificity floor, so `12/15` clusters were never judged and only `3` lifts were measured — source `SOMNUS-ABDUKTIO-VASTINE` v2 §8.2 — status: overturned (the reading), measured (the count)
- claim: the planted control (A1): known-mechanism clusters passed — fast model `5/6` (lifts `0.7–0.9`), careful model `6/6` (lift `1.0`); the no-mechanism cluster — fast model `0/3` (the gate held), careful model `2/2` passed (the gate leaked: a fluent pseudo-mechanism, specificity `0.53`, lift `0.3`); the predictions underestimated both (`~50% → 83%`, `~75% → 100%`) — source `SOMNUS-A1-TULOKSET` v1 — status: measured
- claim: the entailment graph (A3): `779` candidate pairs → `28` consequence edges + `31` conflict edges, `17` components; three-repeat agreement `87%` (`9/68` did not repeat); the prediction held on the strict reading (`n_pred_min = 5`: `1` component) and fell optimistically on the loose one (`= 3`: `6`) — source `SOMNUS-A3-TULOKSET` v1 — status: measured / partly overturned
- claim: the targeted run (D1) produced content and no structure — `0` non-neutral verdicts over `136` forced pairs — source `SOMNUS-D1-IGNITE-TULOKSET` v4 — status: measured
- the "`10.0% → 6.7%`" accuracy change in the text has no record in the channel that this block could cite — status: recorded in the article, not re-verified here
- cost: the local runs `$0`; the D1 cloud run's cost is held with the AMD R9700 whitepaper
- not shown: the "fabrication risk grows with capability" generalisation rests on `2/2` leaks on one planted cluster with one model pair — in this fixture, with these two models.

</aside>

---

<p class="sig">— SF3D</p>
