---
title: "The Value Function That Reads the Shape"
subtitle: "A weighted sum of relevances is a single cosine in disguise — a theorem, not an accident, and reading the shape of a question beats summing it"
date: 2026-07-13
tags: ["somnus", "quaesitor", "information-theory"]
summary: "To focus on one topic, the machine had to price every candidate question against seven sub-questions — and the obvious way to do that silently collapses all seven into one average direction. The theorem behind that collapse, the fix that reads a shape instead of a sum, and the discipline that catches it before a euro is spent."
canonical: https://sf3d.fi/blog/the-value-function-that-reads-the-shape
html: https://sf3d.fi/blog/the-value-function-that-reads-the-shape
author: Petri Korhonen, SF3D AI Lab (Sunrise Software Oy, Tampere, Finland)
site: https://sf3d.fi (index of all articles in Markdown: https://sf3d.fi/llms.txt)
numbers: measured on the lab's own hardware unless the text marks them as held for the AMD R9700 whitepaper; the Evidence block at the end of an experimental article names the run, the date, n and the pre-registration each claim rests on
---
<p class="standfirst">Part 1 was about a machine that refutes its own ideas with executable code. This is a shorter, sharper story from one stage of it: the moment the machine had to put a number on <em>what to think about next</em>. A value function is a measurement, and the fastest way to ship a broken measurement is to trust an average — so this one is built not to.</p>

<aside class="primer">

**PRIMER**

Somnus is a research loop on one GPU box: a local model proposes candidate research questions from gaps between existing findings, a second local model screens them, and the best are escalated to cloud models that try to refute them. Claims that survive go into a store, and a local judge decides whether one claim follows from another; those edges are the structure most parts of this series measure.

</aside>

## The system doing the thinking

A word first on what is now running the loop, because the shape of the hardware is the shape of the argument.

Somnus is **three model families in a funnel, feeding a cloud oracle**:

- **Generation** — Gemma 3 12B on an AMD Radeon RX 7900 XT proposes candidate questions from the gaps between existing findings.
- **Screening** — Mistral Small 24B on an AMD Radeon AI PRO R9700 rejects only certain noise and ranks the rest.
- **Adjudication** — Claude (Sonnet 5 researchers, an Opus 4.8 assembler) runs the expensive, decisive step in the cloud.

The second card is new, and it is worth being exact about why it earns its slot, because it is not the obvious reason. A second GPU buys throughput, and throughput was never the constraint — the generator's queue was already saturated. What the 7900 XT bought is **prior diversity, and the end of time-sharing**. Three lineages — Gemma, Mistral, Claude — hold three different sets of blind spots, so their agreements and disagreements carry signal a single family cannot manufacture; the funnel's value is that divergence, not raw compute. And with two cards, generation and screening run *concurrently* instead of one card swapping models in and out of VRAM between roles. Three families, three stages, one cloud — that is the whole instrument, and everything below is one stage of it, examined closely.

<aside class="nerd">

**NERD BOX**

#### Two cards, run for stability

`ROCR_VISIBLE_DEVICES` pins Gemma to the 7900 XT (gfx1100) and Mistral to the R9700 (gfx1201); both `llama-server` instances run `--no-mmap`, so the weights sit in VRAM instead of paging a 16 GB host into a swap storm. Single-card, time-shared was the honest baseline until recently — one card, one model at a time, swapped by a button. The dual-card step earns its place by removing that swap, not by adding speed. The measured power and throughput on this bench belong to the whitepaper.

</aside>

## A value function is a measurement

For most of its life Somnus was a grab bag: generate every plausible cross-domain question, screen it, keep the good ones. It never ran dry and it never arrived anywhere — a pile with no gradient. The deliberate fix was to **focus**: take one goal question, break it into a handful of falsifiable sub-questions, and grind that topic until it is compressed into something validated.

Focus is precisely what turns *generate more* into a pricing problem. At any moment thousands of candidate questions touch the topic, against a fixed budget: the local models are effectively free, the cloud oracle is not. So the machine must put a price on each candidate — *resolve this locally, or spend real money on it at the frontier?* — and that price is a measurement of one thing: how much the candidate would reduce the topic's remaining uncertainty. Get the measurement wrong and every downstream euro is misspent with confidence. That is the whole reason this stage gets the scrutiny it does.

The price has to respect the **structure** of the topic. Seven sub-questions are seven distinct directions to be uncertain in. A value function that flattens them into one has not found a cheaper measurement — it has found a wrong one, and thrown away the only thing focusing was for.

<figure>
<img src="/images/blog2-mdl.svg" alt="A U-shaped curve: the cost of stating a law rises with precision, the cost of its exceptions falls, and their sum bottoms out at the best explanation." width="880" height="520" />
<figcaption>The floor under all of it: the best hypothesis is the one that compresses the data to its shortest total description — the bottom of the U, neither too generic to predict nor so precise it memorises the noise.</figcaption>
</figure>

<aside class="math">

**MATH BOX**

#### What "value" means, in bits

Part 1 set the frame and it holds here: understanding a topic and compressing it are the same act. The Minimum Description Length principle (Rissanen, 1978) prices a hypothesis `H` against data `D` by the total bits to transmit both:

`L(H) + L(D | H)`

`L(H)`, the cost of the law, rises as the hypothesis sharpens; `L(D | H)`, the cost of the exceptions it still can't explain, falls. Their sum is a U-curve, and its minimum is — in the idealised, uncomputable setting — the *algorithmic sufficient statistic* — the shortest description that loses nothing. A sub-question's uncertainty is how many bits are still unpaid on it; a candidate's value is how many bits it is expected to retire. None of this is improvised vocabulary — it is MDL (Rissanen), the Information Bottleneck (Tishby, Pereira, Bialek), and compression-progress-as-curiosity (Schmidhuber), the same lineage Part 1 named. The machine obeys the theory; it does not invent it.

</aside>

## The trap the algebra guarantees

The obvious value function almost writes itself: score a candidate `c` by its relevance to each sub-question, weight by how uncertain that sub-question still is, and add them up. It is also, provably, a trap — and a trap you can see coming with one line of algebra, which means it is a trap you check for rather than discover.

A weighted sum of cosine similarities is *identically* a single cosine to one resultant vector. The seven directions collapse into their weighted average **before** the candidate is ever compared to them; the sum can express exactly one direction, no matter how many sub-questions you feed it. So the prediction is not subtle: this value function will rank candidates purely by centrality to the topic's mean, and the sub-question structure will be inert.

Measured on the live container — 1793 candidates, seven sub-questions — that is exactly what it does, at `r = 1.000` against plain cosine-to-the-centroid. Six of the seven sub-questions were doing no work at all. The point of the measurement was never surprise; it was confirmation that a known failure mode was, or was not, present. It was.

<figure>
<img src="/images/blog2-collapse.svg" alt="Two scatter plots of candidate value against centrality. Left, the linear sum: a perfect diagonal, r equals 1.00. Right, per-candidate centering: a formless cloud, r about 0.03." width="960" height="470" />
<figcaption>Same candidates, two aggregations. The linear sum (left) is a perfect proxy for one direction — centrality — exactly as the algebra says it must be. Centring each candidate first (right) breaks the tie and hands the structure back.</figcaption>
</figure>

<aside class="math">

**MATH BOX**

#### Why an average cannot tell directions apart

`Σᵢ wᵢ·cos(c, q̂ᵢ)  =  c · (Σᵢ wᵢ q̂ᵢ)  =  |V|·cos(c, V̂)`,   where `V = Σᵢ wᵢ q̂ᵢ`

The sub-question directions are summed into one resultant `V̂` before `c` meets them, and every candidate is then scored by its angle to that single vector. Re-weighting the `wᵢ` only *moves* `V̂`; it never restores seven directions from one. This is why the collapse is structural, not a tuning artefact: a mean is a lossy summary — it keeps the centre of a set and discards its shape — and a linear aggregation over directions is a mean wearing a formula. Whenever the signal you care about lives in the *differences* between directions, summing them is the one operation guaranteed to erase it.

</aside>

## Reading the shape, not the sum

The fix is two moves, and the order matters, because the obvious repair does nothing alone.

First, **read a peak, not a total.** Replace the linear sum with a log-mean-exp, which interpolates smoothly between the average (cold) and the maximum (hot) and can therefore reward a candidate that hits *one* sub-question hard rather than everything weakly. But on this corpus the relevance profiles are nearly flat, and a soft maximum of a flat profile is still its mean — so log-mean-exp alone left the collapse standing.

The move that actually mattered was **per-candidate centring**: subtract each candidate's own mean relevance before aggregating. That removes the level of the profile — which *is* the centrality that went to `r = 1.000` — and leaves only its shape. The correlation fell from `1.000` to `0.03`; three of the seven sub-questions took the lead across different candidates. Absolute relevance is real signal, so it returns — but as a **multiplier on the shape**, never an additive term, because adding the mean back is, by the algebra above, re-introducing the exact quantity that caused the collapse.

<figure>
<img src="/images/blog2-beta.svg" alt="A curve peaking sharply at a low-to-mid temperature: the number of distinguishable sub-questions is maximal at a critical beta and decays toward one at both hot and cold extremes." width="880" height="520" />
<figcaption>β as temperature. Too hot and every candidate melts back into the average; too cold and only the single sharpest sub-question ever registers. The structure is richest in a narrow critical band — and that is where the aggregator is set to sit.</figcaption>
</figure>

<aside class="math">

**MATH BOX**

#### Temperature, and the band that carries the structure

Centre each candidate's profile, aggregate the *shape*, and scale by the level as a separate factor:

In words: each candidate's relevance profile is centred on its own mean, the centred profile is aggregated with a temperature-controlled average (a log-mean-exp over the sub-questions), and the level multiplies the result as a separate factor. The exact form and the temperature are the whitepaper's.

The log-mean-exp is, up to sign, the free energy of statistical mechanics — the log of a Boltzmann partition sum, divided by the inverse temperature `β`. It is a temperature-controlled average: as `β → 0` it is the mean, as `β → ∞` it is the max. By analogy, `β` plays the part the Information Bottleneck's multiplier plays (Tishby, Pereira, Bialek, 1999) in trading compression against relevance — here, between rewarding a candidate that covers many sub-questions weakly and one that answers a single sub-question sharply. Swept across `β`, the number of sub-questions that act as distinct attractors peaks in a narrow band and collapses to one at either extreme; the aggregator is tuned to sit inside that band. The level enters as a coefficient so that a strongly-relevant sharp candidate beats a weakly-relevant sharp one, without ever summing the mean back in.

</aside>

The value function is one stage of the funnel, and the fix only matters because of where it sits.

<figure>
<img src="/images/blog2-funnel.svg" alt="A three-stage funnel: Gemma on the 7900 XT generates candidates, Mistral on the R9700 screens with a loose reject gate, the container value function ranks and flags the frontier, and Claude adjudicates only at the frontier." width="960" height="440" />
<figcaption>Where the value function lives. Generation and screening are free and run on the two local cards; the value function orders what to grind locally and flags the frontier; only the frontier — relevant, uncertain, and locally unresolvable — reaches the paid oracle.</figcaption>
</figure>

## The discipline that catches it

None of this was found by staring at outputs and getting a bad feeling. A value function is not trusted here until it is *shown* to preserve the structure it claims to use, and the showing is one measurement: aggregate the value, correlate it against the single-direction projection, and if `r → 1.0`, the structure is gone. One number, run before a euro is spent on the ranking it produces — the same *look before you build* that runs through the whole instrument, pointed at a formula instead of a database.

The honest result — `r = 1.000` — is the entire value of the check. A value function that *looks* decisive while secretly ranking on one axis is worse than an obviously broken one, because it spends real money with false confidence. And the same test was then run across the rest of the pipeline — the generator's prior, the screener's score, the essence relevances — to see whether the averaging trap had leaked anywhere else. It had not; the container's core was the only genuine collapse. That is the cheapest insurance there is: one correlation that stops a structural error from living in production for months.

## What it buys

With the structure restored, the container does the job it was built for — rank the topic's queue locally for nothing, and spend only at the frontier. The first paid cloud ignite landed the way Part 1's thesis says it should: handed a candidate, the three cloud researchers refused it, catching a **category error in the question itself** and verifying the refutation with a null model in code rather than argument. The loop fed itself two validated facts and five testable hypotheses, and the machine's honesty held under its own money.

The numbers that make this quantitative — the critical `β`, the resolvability calibration, the escalation threshold, the cost of a frontier decision in bits per euro — are the subject of their own paper. What belongs here is the shape of the lesson: **linear aggregation destroys structure whenever the structure lives in the peaks and not the average**, and the discipline is to know that in advance, forbid the average on purpose, and prove the replacement before trusting it. A machine that can hold that line inside its own value function is, in one small and specific way, doing the thing it was built to do.

<aside class="reachout">

**FOR RESEARCHERS**

#### Working on something like this?

If you research or build in this territory — value functions over structured question spaces, local-versus-oracle escalation, aggregation that has to preserve the structure it measures — I'd like to hear from you: [pete@sunrisesoftware.app](mailto:pete@sunrisesoftware.app).

</aside>

<aside class="evidence">

**EVIDENCE**

#### The value-function check · July 2026, before the labnote practice

- source records: the article record `SF3D-ARTIKKELI-ARVOFUNKTIO` v1 in the Atlas channel (2026-07-31T21:13Z) and the publication package `somnus/docs/blog/publish/blog-02` (13.7.2026). No pre-registration: this was a structural check, not an experiment, and it predates the pre-registration practice of August 2026.
- claim: a weighted sum of cosines is identically one cosine against the resultant vector — an algebraic identity, not a measurement — status: stated (and derivable from the text)
- claim: on the live container, `1,793` candidates against `7` sub-questions, the aggregated value correlated `r = 1.000` with plain centroid distance; after per-candidate centring `0.03`, and `3` of `7` sub-questions took the lead on different candidates — source the article record — status: measured. The check could not fail given the identity; its value is that the number was taken before thresholds were set.
- claim: the same trap was looked for elsewhere in the pipeline (the generator prior, the screen scoring, the essence relevances) and not found — source the article record — status: recorded, no run data in the channel
- held with the AMD R9700 whitepaper: the exact value function and its temperature, the resolvability calibration, the escalation threshold, bits per euro (the publication package's own guardrail note)
- not shown: whether the centred function ranks better downstream — the text offers one anecdote (the first paid run); no ranking comparison was recorded.

</aside>

---

<p class="sig">— SF3D</p>
