---
title: "The Bottleneck Wasn't the Hardware, Again"
subtitle: "Three times the machine looked like the problem: a night that produced nothing, a two-card speed-up that never arrived, and a box that froze in total silence. Three times we measured before we replaced anything — and the hardware was innocent every time."
date: 2026-08-27
events: "21.7.–2.8.2026"
tags: ["somnus"]
summary: "Hardware is the most satisfying suspect in computing: visible, blameable, replaceable. This part is three cases where it was the obvious answer and the wrong one — including a freeze we still cannot explain, and the repair we are deliberately not performing because it would destroy the evidence."
canonical: https://sf3d.fi/blog/the-bottleneck-wasnt-the-hardware-again
html: https://sf3d.fi/blog/the-bottleneck-wasnt-the-hardware-again
author: Petri Korhonen, SF3D AI Lab (Sunrise Software Oy, Tampere, Finland)
site: https://sf3d.fi (index of all articles in Markdown: https://sf3d.fi/llms.txt)
numbers: measured on the lab's own hardware unless the text marks them as held for the AMD R9700 whitepaper; the Evidence block at the end of an experimental article names the run, the date, n and the pre-registration each claim rests on
---
<p class="standfirst">When a machine is slow, or stops, or just will not produce, the hardware is the first thing anyone reaches for. It is the most satisfying suspect in all of computing — you can see it, you can blame it, and best of all you can buy a new one. Three times this summer the hardware was the obvious answer. Three times we measured first. It was innocent every time, and the third case is still open: we do not know what stopped the machine, we know precisely what did not, and we are deliberately <em>not</em> doing the repair that would make the symptom go away.</p>

<aside class="primer">

**PRIMER**

Somnus is a research loop on one GPU box: a local model proposes candidate research questions from gaps between existing findings, a second local model screens them, and the best are escalated to cloud models that try to refute them. Claims that survive go into a store, and a local judge decides whether one claim follows from another; those edges are the structure most parts of this series measure.

</aside>

## First: the night that produced nothing

The machine ran for eleven hours without interruption. Ninety-nine generation cycles, a hundred and seventy-eight screening passes, twelve rounds of synthesis. In the morning the count of new ideas it had produced was **zero**.

Not a crash. Not an error in the log. It had done all the work and there was nothing at the end of it.

The reflex here is memory. It is always memory. More RAM, a bigger model, another card — the machine is clearly straining, so give it more machine.

More RAM would have produced exactly as many ideas: none. The generator's job is to find *missing connections* between things the archive already believes. It was drawing on `236` verified items, which admit a finite number of meaningful pairs — `4,569` of them — and it had already proposed every single one. It kept a record of which pairs it had seen, so it dutifully skipped them all and wrote nothing. The machine was not straining. It had finished, and nobody had told it.

The fix cost nothing and involved no hardware. The archive holds more than its verified conclusions: it also holds untested hypotheses, and records of past sessions, and those live in the same mathematical space as the verified ones — the same coordinates, directly comparable. We let the sampler draw on those too. New candidates went from zero to `4,406`.

<aside class="plain">

**IN PLAIN TERMS**

A researcher who has read every book in a small library and cross-referenced every pair of them will produce nothing new tomorrow, and a faster desk will not help. They do not need more equipment. They need a wider shelf — and it turned out we already owned one and had not been letting them near it.

</aside>

There is a footnote to that day worth keeping, because it is the same lesson pointed at ourselves. A dry run reported that only `18` of `343` stored hypotheses passed the filter, which looked like a catastrophe — most of the shelf apparently dead. Checked directly against the database, all of them were alive. The `18` was an artefact of the filter one of us had just written. The build diary falsified its own conclusion inside a day, which is roughly the rate we would like to keep.

## Second: two cards, one wrong assumption

The box has two GPUs. So when the next question was how to make the model generate faster, the architecture practically wrote the hypothesis for us.

The technique is speculative decoding: a small, fast model guesses the next several tokens, and the large model checks the whole batch in one pass instead of producing them one at a time. Cheap guesses, expensive verification, done in bulk. It works — that part is not in doubt.

The tempting version was to split it across the two cards. Small model on one, large model on the other. Two cards, two jobs, parallel work. It is such a clean picture that not trying it would have felt like negligence.

The measurement took one afternoon and came back split cleanly in two:

- **Both models on the same card:** a large gain on code prompts (the ratio is held for the hardware paper). The technique delivers.
- **Split across the two cards: no gain.** The transfer across the bus between them, once per round, ate everything the trick had bought.

"Two cards means the work should be divided between two cards" is an intuition, not a result, and for this workload it was simply false. What made it cheap to find out was that we asked the machine instead of arguing about it.

And then the good part, which arrived out of the wreckage rather than the plan. While measuring the failed version we found a variant that needs no second model at all: the large model speculates against **its own recent output** — repeated phrasings, repeated identifiers, the ordinary redundancy of structured text — and verifies its guesses the same way. No draft model, no second card, no transfer.

We did not take that on faith either. It went into production and was measured in the field on real screening calls: `1.82×` faster, no quality change, no new component. The flag now sits on two live service units. (The throughput series itself belongs to the hardware paper; the blog gets the method.)

One trap from that deployment is worth publishing, because it is the kind that gets "fixed" by someone helpful. The first call after a server restart always reports no speculation and baseline speed, every time, because the pattern cache is empty and has nothing to guess from yet. That is correct behaviour that is indistinguishable from a broken feature. It is now written down, which is the only defence.

<aside class="math">

**MATH BOX**

#### Why the two-card version was capped before it ran

Speculative decoding's gain has a known shape (Leviathan et al., ICML 2023; Chen et al., 2023). The draft proposes γ tokens; the target verifies them in a single forward pass; with a per-token acceptance rate α, the expected number of tokens accepted per round is (1 − α<sup>γ+1</sup>) / (1 − α). Divide that by the cost of one round and you have the speed-up. The whole method is the bet that a round costs barely more than the one target step it replaces.

Splitting the two models across devices adds a device-to-device transfer to every round. That term is **serial** — it happens between the draft and the verify, so nothing overlaps it — and, crucially, it does not shrink as γ grows: proposing more tokens per round amortises it but never removes it. When it becomes comparable to the target's own forward pass, the denominator swallows the numerator and the ratio falls below one. The technique then costs more than it saves, which is exactly what the measurement showed.

This is Amdahl's law wearing a 2023 costume (Amdahl, 1967): a serial term added inside the loop you are trying to accelerate bounds the result no matter how good the accelerated part gets. The single-card configuration has no such term, which is why the same technique paid off there.

</aside>

<figure>
<img src="/images/blog8-serial-term.svg" alt="One round of speculative decoding drawn twice. Top, same card: a draft segment of gamma tokens followed directly by one verify pass, giving plus 74 percent. Bottom, split across cards: the same draft and verify with a transfer segment wedged between them, marked serial and annotated that it does not shrink as gamma grows, giving no gain. Below, the accepted-tokens-per-round expression." width="880" height="420" />
<figcaption>The transfer is not slow because the bus is slow. It is fatal because it is <em>serial</em> and sits inside the round — so proposing more tokens amortises it and never removes it.</figcaption>
</figure>

## Third: the machine that stopped without saying anything

Twice in one morning the box froze. Not crashed — froze, in complete silence, after `87 minutes` the first time and `40` the second.

There was no kernel panic. No out-of-memory kill. No GPU fault, no machine-check exception, no disk error. Nothing reached the disk at all: the kernel's log simply stops, seventy minutes before the first freeze and forty before the second, and resumes at the next boot.

The ordinary version of this story ends "we replaced the motherboard and the problem went away." That ending is available to us and we are not taking it, for a reason worth stating plainly: **a swap that replaces the board, the processor, the memory and the disk in one operation cannot tell you which of them it was.** If the symptom disappears you have bought silence, not an answer, and you will believe a cause you never tested. We wrote that into the procedure *before* the parts were ordered, precisely so nobody could reason backwards from a happy outcome later.

So instead, each suspect was refuted individually.

**The bus.** Error counters on all four devices: zero, zero, zero. Both links negotiated at full width and full speed. Refuted.

**A memory leak.** Sixteen consecutive model loads and unloads. Memory returned each round — between `592` and `625 MB` back every cycle — and the kernel's own allocator grew by `3 MB` across all sixteen. Refuted.

**The disks.** Both drives pass, every failure counter at zero, one of them reporting `100%` of its life remaining. Refuted.

And then the good suspect, the one we expected to be right. A specific driver path was hogging the processor for over ten milliseconds at a time, its counter climbed with every model swap, and it had appeared seventy minutes before the first freeze. It runs on a shared work queue, so if it ever wedged there, everything behind it would stop — and stop *quietly*, with no panic, which matched the symptom exactly.

We drove that counter deliberately to `131` and the machine did not so much as stutter. Refuted, and it was the one we wanted.

<aside class="plain">

**IN PLAIN TERMS**

Four suspects, four alibis, and the case is still open. That is an uncomfortable place to stop, and it is a much better place than the alternative on offer — swapping four parts at once, watching the symptom vanish, and telling ourselves we fixed it. We would have learned nothing and believed something.

The engineering work went into making the *next* occurrence readable instead of making this one disappear.

</aside>

What replaced the guessing was instrumentation. The kernel's last words now go out over the network to a remote server as they are written, so a freeze that never reaches the local disk still leaves a record somewhere else. We validated that channel with a real shutdown and watched the full sequence arrive intact — which is the step that turns it into an instrument rather than a hope. Because the channel is proven, **silence is now data**: if the machine stops and nothing at all comes out of a link we know works, the processor stopped in one piece, and that points at power or hardware rather than software.

One blind spot got smaller at a known price. The professional card supports error-correcting memory on the GPU, and it was switched on. The cost is exactly `2.000 GiB` of video memory (`31.860 → 29.859 GiB`, `6.28%`), with no measurable slowdown, and in exchange memory errors on that card are now visible instead of silent. The consumer card in the same box has no such capability, so that half of the machine stays dark.

<aside class="nerd">

**NERD BOX**

#### The instrument that had to be built twice

The obvious way to ship kernel messages off a box that is about to die is to send them over the network you already trust — in our case a private overlay network that everything else uses. It does not work, and the reason is structural rather than a misconfiguration: the overlay presents a virtual tunnel interface, and the kernel's emergency logging path deliberately bypasses the normal networking stack to survive a dying system. It needs a real driver on a real interface. So the emergency channel runs on the physical port, addressed directly, outside the overlay that carries everything else.

Alongside it a sampler wakes on every boot from a scheduled job and writes twice — a small UDP datagram to the remote collector every five seconds, and a local line to disk, flushed. Each sample carries link state and error counters, video and system memory, allocator internals, pressure-stall figures, disk queue depth, the top three processes and which model currently holds the professional card. It is `797 bytes`, sized to fit inside the overlay's smaller packet limit in one piece.

Two channels, two failure modes: the sampler tells you the state of the world up to the last five seconds, the emergency log tells you what the kernel thought at the end. A freeze that defeats both is itself a finding, and a narrower one than we had before.

</aside>

<figure>
<img src="/images/blog8-three-suspects.svg" alt="Three cases in rows. The eleven-hour night: suspect more RAM, struck through; refuted by 11 hours of cycles yielding 0 new candidates with all 4,569 pairs already proposed; actually a saturated archive, 0 to 4,406. The two-card speed-up: suspect splitting across GPUs, struck through; refuted by plus 74 percent on one card against no gain when split; actually a serial transfer, with draft-free speculation shipping at 1.82 times. The silent freeze: suspect replace the motherboard, struck through; refuted by AER counters at zero, memory returned every cycle, both disks passing, and the driver counter driven to 131 with no crash; still open, with an instrument built instead." width="880" height="570" />
<figcaption>Three unrelated failures, one available mistake. The third row is the one worth staring at: every suspect cleared, and the case still open — which is the honest end of the story rather than a failure of it.</figcaption>
</figure>

## What survives

The three cases have nothing technical in common. A saturated archive, a bus transfer, and a silent stop are unrelated failures in unrelated layers. What they share is the shape of the mistake that was available in each one:

1. **Hardware is the suspect that requires the least thought and offers the most action.** That combination is exactly why it should be the one you check with the most discipline, not the least.
2. **A hypothesis the architecture suggests is still just a hypothesis.** Two cards made the split version feel like the obvious design. One afternoon of measurement was cheaper than one week of building it properly first.
3. **The best result of the summer came out of a falsified hypothesis, not a confirmed one.** We would never have found the single-model variant if the two-card version had quietly half-worked. A clean negative sends you looking; an ambiguous positive keeps you tuning.
4. **A repair that changes four things at once is not an experiment, and its success is not evidence.** If you cannot resist making the symptom go away, at least write down beforehand that you will not be entitled to a conclusion.
5. **Validate the instrument, and silence becomes a measurement.** An untested logging channel that stays quiet tells you nothing at all. A tested one that stays quiet has told you where the fault is not.

The freeze has not recurred. We still do not know what it was, and the trap is set for the next one.

<aside class="reachout">

**FOR RESEARCHERS**

#### Working on something like this?

Two open questions we would trade notes on. First: has anyone measured the cross-device penalty for speculative decoding as a function of draft length — we have one workload's answer and would like to know where the crossover sits for others. Second, and more useful to us: if you have chased a silent, log-less freeze on a multi-GPU workstation to an actual root cause, we would like to hear what it turned out to be — ours is still open, with the bus, memory, disks and the obvious driver path all refuted by measurement: [pete@sunrisesoftware.app](mailto:pete@sunrisesoftware.app).

</aside>

## Sources & further reading

- Y. Leviathan, M. Kalman & Y. Matias, *Fast Inference from Transformers via Speculative Decoding* (ICML, 2023) — the acceptance-rate/cost model the two-card analysis rests on.
- C. Chen et al., *Accelerating Large Language Model Decoding with Speculative Sampling* (2023) — the concurrent formulation, with the same round-cost structure.
- G.M. Amdahl, *Validity of the single processor approach to achieving large scale computing capabilities* (AFIPS, 1967) — why a serial term inside the accelerated loop bounds the whole result.
- J.R. Platt, *Strong Inference* (Science, 1964) — the discipline of refuting suspects one at a time rather than confirming a favourite.
- J. Gray, *Why Do Computers Stop and What Can Be Done About It?* (Tandem TR 85.7, 1985) — the classic taxonomy of silent and transient faults, and why the ones that leave no trace are their own category.
- J.L. Hennessy & D.A. Patterson, *Computer Architecture: A Quantitative Approach* — the standing argument that measurement precedes optimisation, in the form most engineers first met it.

<aside class="evidence">

**EVIDENCE**

#### Three cases, three records · 21.7.–6.8.2026

- speculative decoding (case 2): `SOMNUS-KOE1-SPEC-DECODE-TULOKSET` v1, 2026-07-21T10:13Z — `$0`, about `35 min`, llama.cpp `b9775` (HIP), ROCm `7.2.0`; target model on the R9700, draft model on the 7900 XT; two prompts (code, prose) × two repeats per configuration; raw data `docs/gpu-optimointi/koe1-results.jsonl`. Claim: the draft on the same card beat the draft on the second card throughout, and the two-card hypothesis was falsified at this scale; the prose prompt got slower with speculation — status: measured. The throughput series and the ratios are held with the AMD R9700 whitepaper.
- headline number: the model-free `ngram-mod` speculation went into production on 2.8.2026 and measured `1.82×` on real screening calls, lossless by construction at temperature 0 — source Somnus `LN-20260721-spec-decode-ja-ngram-mod` — status: measured
- the freezes (case 3): Somnus `LN-20260802-vakausjakso-ja-instrumentointi`; `SOMNUS-RUNBOOK-X670E-VAIHTO` v2 (2026-08-06T14:27Z, the board swap, held). Claim: two unexplained freezes on 2.8.2026 at uptimes `87` and `40 min`; no OOM, no panic, no GPU error, no MCE; PCIe hypothesis refuted (AER `0/0/0`, links at full speed); memory-leak hypothesis refuted (`16` model swaps, `+592–625 MB` returned per cycle, slab `+3 MB`); the userptr counter went `4 → 131` without a crash, so that correlation fell; GPU ECC switched on: VRAM `31.860 → 29.859 GiB` (`2.000 GiB`, `6.28%`), no measurable slowdown — status: measured; the cause is open.
- the saturated archive (case 1): the counts in the text (`11 h`, `99` cycles, `178` screens, `12` syntheses, `236` verified items, `4,569` pairs, `0 → 4,406`) come from the run log of the night in question as recorded in the article; no Atlas record of that night was found when this block was written — status: recorded, not re-verified here.
- not shown: a draft model larger than `0.6B` across cards (the record names it as the next test); the freeze's cause.

</aside>

---

<p class="sig">— SF3D</p>
