---
title: "Three Budget Raises and the Wrong Lever"
subtitle: "The model's reasoning kept getting cut off. We raised its budget three times and the truncation stayed at 51%. Then we tried the opposite — gave it less — and the problem vanished, the quality held, and the bill halved."
date: 2026-08-27
events: "17.8.2026"
tags: ["somnus"]
summary: "A short one about a reflex. When a machine runs out of room, you give it more room. We did that three times to a model whose reasoning was being truncated, and the number never moved. The fix was a knob that pointed the other way — and the only reason we can call this an experiment rather than a lucky guess is that the answer key was sealed before the run."
canonical: https://sf3d.fi/blog/three-budget-raises-and-the-wrong-lever
html: https://sf3d.fi/blog/three-budget-raises-and-the-wrong-lever
author: Petri Korhonen, SF3D AI Lab (Sunrise Software Oy, Tampere, Finland)
site: https://sf3d.fi (index of all articles in Markdown: https://sf3d.fi/llms.txt)
numbers: measured on the lab's own hardware unless the text marks them as held for the AMD R9700 whitepaper; the Evidence block at the end of an experimental article names the run, the date, n and the pre-registration each claim rests on
---
<p class="standfirst">Here is a question with an obvious answer that is wrong. Something you are running keeps getting cut off before it finishes. What do you do? You give it more room. Anyone would. We did it three times, carefully, measuring each time — and the cut-off rate did not move by a single point. The thing that fixed it was a second knob that pointed in the opposite direction: we asked the model to think <em>less hard</em>, and it started finishing. Quality held. The bill halved. This part is about why the obvious lever was the wrong one, and about the small discipline that turned a configuration change into a result we can stand behind.</p>

<aside class="primer">

**PRIMER**

Somnus is a research loop on one GPU box: a local model proposes candidate research questions from gaps between existing findings, a second local model screens them, and the best are escalated to cloud models that try to refute them. Claims that survive go into a store, and a local judge decides whether one claim follows from another; those edges are the structure most parts of this series measure.

</aside>

## The symptom, and the reflex

The machine's job here is to reason its way through a physics question and hand back claims the rest of the pipeline can check. Roughly half the time, it did not hand back reasoning at all. The thinking block came back empty — started, cut off, gone. `51%` of them.

The reading is natural: it ran out of room. So we gave it more room. The thinking budget went up. It went up again. It went up a third time, ending more than four times where it started — from `4,500` to `20,000`.

The truncation rate stayed at `51%`.

That is the part worth sitting with. Not "it improved a little" or "we were still tuning". Three deliberate increases, each one measured, and the number stood exactly still. A budget you raise four-fold that changes nothing is not an underpowered budget. It is telling you it was never the thing in the way.

## The other knob

There were two controls on this model, not one, and they are easy to conflate. The budget says *how much room the reasoning may use*. A separate setting — effort — says *how hard to push before answering at all*. They sound like the same dial at different resolutions. They are not.

So instead of a fourth raise, we went the other way and turned the effort setting **down** one notch, from its highest level to the middle one. Nothing else changed. Same question, word for word. Same pipeline, same models throughout.

Truncation went from `51%` to **zero**. Not "improved" — zero cut-offs in the whole run, one warning total.

And the things you brace for did not happen:

| | before | after |
|---|---|---|
| truncated reasoning | `51%` | `0%` |
| answer key | `4/4` | `4/4` |

*Cost, wall clock and token count moved the same way; their series is held for the AMD R9700 whitepaper. The one number this part keeps is the cost per verified crystal, which roughly halved.*

The quality metric is the one that matters and it is the one that did not move. Four out of four on a sealed answer key before, four out of four after — and the lower-effort run named each of the corresponding physical quantities explicitly rather than leaving the reader to infer them.

So: three rounds of raising a budget bought nothing, and one notch in the other direction removed the problem, cut the wall clock by more than half and the bill with it. The right lever had been sitting in the documentation the whole time.

<aside class="plain">

**IN PLAIN TERMS**

This is the household version of a very common mistake. The tap is barely trickling, so you open it wider. Nothing. You open it wider again. Still nothing — because the restriction is not the tap, it is a kink in the hose behind the wall. Opening the tap harder is a perfectly reasonable act that cannot possibly work, and the only thing that tells you so is that you measured it and the flow did not change.

The part that generalises past our lab: we only knew the raises had failed because each one was measured against the same number. Three unmeasured raises would have felt like progress.

</aside>

<figure>
<img src="/images/blog7-two-knobs.svg" alt="Two panels. Left, the budget control: three deliberate raises from 4,500 to 20,000, and beside each the truncation rate unchanged at 51 percent. Right, the effort control: one notch down, and five measured outcomes — truncation 51 to 0 percent, answer key 4 of 4 to 4 of 4, cost $10.71 to $4.63, wall clock 65 to 25 minutes, tokens 1.34 to 0.80 million. Footer: $1.53 to $0.77 per verified crystal." width="880" height="470" />
<figcaption>Three changes on the left, one on the right. The column that moved everything is the one nobody reached for — and the quality metric is the one that stayed put.</figcaption>
</figure>

## Why this is a result and not an anecdote

Everything above would be a nice story and nothing more, except for one habit: **the predictions were written down and locked before the run.**

Seven of them, each as a point estimate *and* an interval — cost, duration, tokens, truncation rate, answer-key score, how many wrong physics claims we expected to slip through, and how many *crystals* the run would yield: verified claims that survive the pipeline's gates and are added to the archive permanently. Crystals are the unit the whole system exists to produce, which is why the cost figures below are quoted per crystal rather than per run. Along with them, the decision rule: *if the answer key holds at `3/4` or better and the errors do not increase, the lower setting stays in production.*

All seven landed inside their intervals. One — the crystal count — hit its point estimate dead on: six predicted, six delivered.

That sealed sheet is what separates this from tuning. Without it, "we changed a setting and things got better" is an anecdote, and a cheaper, faster run is exactly the kind of result a person talks themselves into. With it, the claim has a shape that could have failed: had the answer key dropped to `2/4`, the rule we had already signed said the change loses, however pretty the invoice looked.

<aside class="math">

**MATH BOX**

#### A constraint you raise and nothing happens is not the binding one

The failure mode has a precise name in optimisation. In any constrained problem, a constraint is either *binding* — the solution sits against it, and relaxing it moves the answer — or *slack*, in which case relaxing it changes nothing at all. The **shadow price** of a slack constraint is exactly zero (Dantzig, 1963; the complementary-slackness half of the Karush–Kuhn–Tucker conditions). Three raises with an unchanged outcome are not weak evidence that the budget was slack. Within measurement noise they are a *direct reading of its shadow price*, and it read zero every time.

The same fact is more familiar in its computing dress. Amdahl's law says an improvement applied to a fraction *p* of the work is capped at 1 / (1 − *p* + *p*/*s*) no matter how large the speed-up *s* — so when *p* ≈ 0, buying an enormous *s* buys you nothing (Amdahl, 1967). Both statements say one thing: **before you scale a resource, establish that it is the one you are up against.** Neither is exotic. Both are routinely skipped, because raising a limit feels like doing something and measuring whether it was the limit feels like delay.

The economics, for the record: the cost per verified crystal roughly halved (`$1.53 → $0.77`) for the same question, at unchanged answer-key quality. That is the one number this part keeps; the cost-per-unit series and the token accounting are the whitepaper's subject.

</aside>

<figure>
<img src="/images/blog7-ledger.svg" alt="Seven preregistered predictions, each on its own normalised interval, with the predicted point marked and the measured value plotted. Answer key 4 of 4. Truncation 0 percent against a bound of under 10. Wall clock 25 minutes against 40 predicted in a 25 to 65 interval. Tokens 0.80 million against 0.9 predicted. Cost $4.63 against $6.50 predicted. Crystals 6 against 6 predicted — a dead-on point hit. Wrong physics claims 0 against 1 predicted. All seven inside their intervals." width="880" height="470" />
<figcaption>The sheet that separates a result from a cheaper invoice. Note what is <em>not</em> on it: nothing was locked for structure, which is why the next section's finding is an observation.</figcaption>
</figure>

## The finding that was not on the list

The honest part of the run is the part we had not thought to predict.

The lower-effort run produced six crystals, and **not one of them connected to anything.** We forced the comparison rather than waiting for it — every combination of the new claims, fifteen pairs, put in front of the judge. Fifteen neutral verdicts. The run before it, at the higher setting, had produced a real connection and the pipeline had found it on its own.

So we had gained a halved bill and possibly lost the structure that is the entire point.

It goes in the record with a note we would rather not have had to write: **no prediction was locked for structure.** The seven sealed numbers covered cost, speed, truncation and correctness, and simply did not cover whether the output would still connect to anything. The metric was not part of the experiment, so what it did that day is an observation, not a result — and it gets labelled as one.

Two explanations fit, and this run cannot separate them. The lower setting may genuinely fail to carry a derivation far enough to crystallise. Or the content that day may have been the wrong shape: it leaned into refutation work, and refutations do not form connections at the higher setting either.

<aside class="nerd">

**NERD BOX**

#### How the ambiguity was actually settled

The discriminating test was already scheduled: a differently-shaped question — a five-step chain in which each step derives from the one before, so no step can import an outside parameter — run at the *same* lowered setting. If structure came back, the content profile was the explanation and the cheaper setting was safe. If it did not, the effort suspicion stood.

It came back: two chain edges, inside the predicted two-to-four.

With a caveat that turned out to matter more than the answer. Those two edges were visible to one judge and invisible to the other. The same corpus, the same day, the same claims: one judge found the structure, the other found none. Had we used the second judge as the meter — as we had been, without knowing there were two — the chain question's own kill condition would have fired, and we would have concluded the cheaper setting destroys structure. It does not. Our meter was blind.

That is a different story with its own part in this series, and it is the reason this one ends with a caveat rather than a clean win: the run that looked like it lost structure was, at least in part, a run we could not see properly.

</aside>

## What survives

Four things, none of them specific to this model or this setting:

1. **Raising a limit that does not move the outcome is a measurement, not a failed attempt.** It has told you the limit is slack. Three raises with a flat result is the cheapest diagnostic anyone will ever hand you — if you were counting.
2. **Look for the second knob.** Systems that expose an amount usually also expose an intensity, and reaching for the one you know is not the same as reaching for the one in the way.
3. **A configuration change becomes an experiment the moment the answer key precedes it.** The cost of writing seven numbers down beforehand is ten minutes. It is the whole difference between a result and a story about a cheaper invoice.
4. **Preregister the metrics you would be embarrassed to lose, not only the ones you expect to win.** Ours covered money, speed and correctness — everything except the property the system exists to produce. That gap is why the structure observation is written here as an open question instead of a finding.

The cheaper setting stayed in production. It is still there.

<aside class="reachout">

**FOR RESEARCHERS**

#### Working on something like this?

If you tune inference-time reasoning controls — thinking budgets, effort or verbosity levels, adaptive stopping — and you have measured a case where the intuitive control turned out to be slack while a second one carried the whole effect, I would like to compare notes. Especially: has anyone measured whether lowering effort costs *structural* output quality, separately from correctness? That is the question we could not close: [pete@sunrisesoftware.app](mailto:pete@sunrisesoftware.app).

</aside>

## Sources & further reading

- G.M. Amdahl, *Validity of the single processor approach to achieving large scale computing capabilities* (AFIPS, 1967) — the bound on improving a component that is not the bottleneck.
- G.B. Dantzig, *Linear Programming and Extensions* (Princeton, 1963) — shadow prices, and why a slack constraint's is zero.
- H.W. Kuhn & A.W. Tucker, *Nonlinear Programming* (Berkeley Symposium, 1951) — complementary slackness: the formal statement that a non-binding constraint has no price.
- E.M. Goldratt & J. Cox, *The Goal* (1984) — the same result as an operating discipline: improvements away from the constraint are not improvements.
- B.A. Nosek et al., *The preregistration revolution* (PNAS, 2018) — the register-before-you-run habit these runs are built on.
- C.D. Chambers & L. Tzavella, *The past, present and future of Registered Reports* (Nature Human Behaviour, 2022) — locking the decision rule before the data exists.

<aside class="evidence">

**EVIDENCE**

#### Effort experiment #1 · candidate `8afaaa40`, session `res_551ee1a558f9` · 17.8.2026

- pre-registration: `SOMNUS-EFFORT-ENNAKKOREKISTEROINTI` v1, locked 2026-08-17T15:45Z before the run; result `SOMNUS-EFFORT-TULOKSET` v1, 2026-08-17T16:20Z. The same question (K1) word for word, the same pipeline; the only change was the effort setting, xhigh → medium.
- claim: seven predictions, each a point and an interval, all seven inside — truncated reasoning `51% → 0%` (bound `< 10%`), answer key `4/4` (predicted `4/4` [`3–4`]), crystals `6` (predicted `6` [`3–10`], a point hit), obvious wrong physics claims `0` (predicted `1` [`0–3`]); cost, wall clock and token count inside their intervals — status: measured
- headline number: the cost per verified crystal roughly halved (`$1.53 → $0.77`); the run costs, wall-clock and token series are held with the AMD R9700 whitepaper, available on request — status: measured
- claim: the medium run's six crystals produced `0/15` non-neutral verdicts in the forced-pair matrix (same judge as the xhigh comparison); no prediction had been locked for structure, which the record states — source `SOMNUS-EFFORT-TULOKSET` v1 — status: measured, unpredicted
- claim: the chain question at medium produced `2` chain edges against a predicted `3` [`2–4`] — source `SOMNUS-K3-TULOKSET` v1, Somnus `LN-20260817-kolme-koetta-ja-tuomariehdollisuus` — status: measured
- not shown: which of two explanations accounts for the lost structure (the run's content profile, or the lower effort) — the run cannot separate them; the binding constraint behind the truncation (a token ceiling, a timeout, the thinking length) was not identified.

</aside>

---

<p class="sig">— SF3D</p>
