---
title: "Blind to Anything Taken Away"
subtitle: "Last time I measured how faint a fault my detector would miss and it gave a number. This week I checked whether that number covered the other four kinds of fault I can plant. It doesn't — and for three of them there is no number at all, because the detector quietly measures them away."
date: 2026-08-27
events: "26.–27.8.2026"
tags: ["somnus"]
summary: "I have a tool that listens to a piece of music and flags anything that looks out of place. It turns out it only notices things being added. Take something away — cut the top off the sound, run it through a bad codec — and it says nothing, at any strength, because it judges every moment against the rest of the track and the fault joins the comparison. The obvious fix was written down as a prediction, tested, and fell over."
canonical: https://sf3d.fi/blog/blind-to-anything-taken-away
html: https://sf3d.fi/blog/blind-to-anything-taken-away
author: Petri Korhonen, SF3D AI Lab (Sunrise Software Oy, Tampere, Finland)
site: https://sf3d.fi (index of all articles in Markdown: https://sf3d.fi/llms.txt)
numbers: measured on the lab's own hardware unless the text marks them as held for the AMD R9700 whitepaper; the Evidence block at the end of an experimental article names the run, the date, n and the pre-registration each claim rests on
---
<p class="standfirst">The short version of this week (26–27 August): a tool this project leans on turns out to be blind to half of what it is supposed to catch, and the way it hid that is specific enough to be worth writing down. This is the last part of an arc whose big idea died on 26 August. The instruments outlived it, and they keep turning out to be more interesting than the thing they were built for.</p>

<aside class="primer">

**PRIMER**

This part is a side experiment: an AI agent drives signal-analysis tools to find a fault planted in a music track, with the answer key sealed before the run. The music is incidental; the subject is how to know what a detector can and cannot see.

</aside>

## The number only covered one thing

Here is the setup, in the plainest words I have.

I have a piece of music and a tool that reads it. The tool does not know what the music is supposed to sound like. It just goes through the track in five-second slices and asks, for each slice, whether anything about it stands out compared to the rest of the track. If something does, it raises a flag.

To find out whether that works, I plant faults on purpose. I take a clean track, damage it in a way I control exactly, and see whether the tool notices. Last time I planted a steady tone — think of a faint whistle sitting under the music — and turned it down step by step until the tool stopped flagging it. It stopped at about `−40 dBFS`, which is quiet: far below the music, but not nothing.

That number does real work. It turns *we found nothing* into a sentence with a number in it, which was the whole point of the previous part.

Then I reread what I had written underneath the result. **Measured one type, one frequency, one song.**

I can plant five kinds of fault, not one:

- a **tone** — something added, a whistle
- **clicks** — something added, tiny sharp pops
- a **band cut** — something removed, the top of the sound sliced off
- an **MP3 round-trip** — a bad copy, the damage a low-bitrate codec does
- a **mel round-trip** — a smeared copy, fine detail blurred away

Only the first of those had a number. So I went to get the other four.

## The bench could not do it, and would not have told me

My notes from 26 August said the other four types were "one command away", so this should have been quick.

They weren't, and the way they weren't is the first thing worth telling you.

The calibration bench takes a list of strengths and plants the fault once at each. But "strength" means something different for each kind of fault. For the tone it is a volume in decibels. For a band cut it is the frequency you cut above. For MP3 it is the bitrate. The bench was writing every number I gave it into the *volume* setting — and only the tone reads that setting. The other four just ignored it.

So I checked before running anything. I planted a band cut at "−35" and again at "−45", and compared the two files.

**They were identical. Byte for byte, the same file.**

The bench would have run happily, printed six neat rows, and every row would have been the same track measured six times. It would have looked exactly like a calibration.

<aside class="plain">

**IN PLAIN TERMS**

A tool that breaks in your face costs you an afternoon. A tool that quietly keeps working while doing nothing costs you a published result — and you find out months later, from someone else.

This one was the second kind. It is the same reason a bathroom scale that reads 70 kg for everyone is worse than one that reads nothing at all.

</aside>

I fixed it so each fault type names the setting its strength actually lives in, and so the bench refuses to run if you hand it that setting yourself — the exact mistake cannot come back quietly. Then I re-ran the old tone calibration to check I had not broken anything on the way. It came back identical to the published table, to the second decimal.

Only then did I measure the other four.

## What came back, and it is not close

| fault | I varied | across | **the tool, on its own** | the checker, holding the answer key |
|---|---|---|---|---|
| band cut | cut frequency | `16 kHz → 2 kHz` | never flags it | `14.09 → 21.05 dB` |
| MP3 copy | bitrate | `192 → 32 kbps` | never flags it | `9.03 → 43.30 dB` |
| smeared copy | detail kept | `256 → 24` bands | only at the very worst | `5.38 → 10.01 dB` |
| tone | volume | `−35 → −50 dBFS` | down to `−40` | `14.03 → 2.46 dB` |
| clicks | loudness | `0.8 → 0.1` | down to `0.3` | `10.98 → −0.00 dB` |

The two it catches are the tone and the clicks. Both of those **add** something to the track — something narrow, something sharp.

The three it misses all **remove or smear**.

The band cut is the row that matters. At the bottom of it, everything above `2 kHz` is chopped off a six-and-a-half second stretch of music. That is not subtle. Half the spectrum is gone; anyone would hear it instantly as the music going muffled and then clearing up again. The checker that holds the answer key measures it at `21 dB`, which is enormous.

The tool never says a word. Not at `2 kHz`, not anywhere in between, not once.

Same for the MP3 copy at `32 kbps`, which sounds like a bad phone call.

<aside class="plain">

**IN PLAIN TERMS**

Imagine a security guard who is very good at noticing anything new in a room — a bag left behind, a chair that wasn't there.

Now take a painting off the wall while they watch. They do not react. Not because the painting was small, but because *noticing absences* was never what they were doing.

</aside>

## Why: the ruler is made out of the track

This is the mechanism, and it is worth going slowly.

The tool has no idea what music should sound like, and that is deliberate. It gets its sense of normal from the track itself: it looks at all twenty-one slices, works out what a typical slice looks like and how much slices normally vary, and flags anything sitting too far outside that. It works on any material with no training and no reference recording.

But look at what happens when a fault goes in.

**The fault joins the comparison group.**

Take the worst band cut. In the slice holding it, the energy up in the treble drops by `8.26 dB`. Measured against a clean copy of the same track, that lands at `6.3` standard deviations out — vastly past the line the tool draws at `4`. It should be flagged and it should not be close.

But the tool is not comparing to a clean copy. It is comparing to *this* track, and this track now contains the fault. And the fault does not just sit low; it **stretches the range of normal** the track has:

| treble energy across the track | typical slice | how much slices vary |
|---|---|---|
| clean | `−45.77` | `1.19` |
| with the band cut planted | `−46.54` | `2.15` |

The variation nearly doubles, and that number is the denominator for every slice. So the fault, which was `6.3` deviations out, becomes `3.1` deviations out — under the line, unflagged.

And the same widened ruler is applied to everything else. There was a genuinely odd slice at the very start of this track, `6.9` deviations out, which the tool had always flagged. With the fault planted elsewhere it drops to `3.5`, and **the flag disappears.**

<aside class="math">

**MATH BOX**

#### Masking and swamping, and why deleting one point does not help

This has names. In outlier detection, **masking** is when an outlier inflates the estimates of centre and spread enough that it — or another outlier — falls back inside the acceptance region and escapes detection. **Swamping** is the mirror image: a perfectly ordinary observation is flagged because the outlier has dragged the estimates. Both are standard failure modes of any rule that estimates its own reference from contaminated data, and both are in the classic treatments (Barnett & Lewis, 1994; Hadi & Simonoff, 1993; Davies & Gather, JASA 1993).

The reason they are not symmetric here is the **breakdown point**. The median tolerates contamination up to half the sample, so the track's typical value barely moved (`−45.77 → −46.54`, on a scale where the effect is `8 dB`). The scale estimate is where the damage lands. The rule fires at `k · MAD` with `k = 4`, and the MAD went `1.19 → 2.15`, so **the fault raises the bar it then has to clear** — the same mechanism as the flat-artefact result in the previous part (Hampel, 1971; Rousseeuw & Croux, JASA 1993).

**And the obvious repair fails for a documented reason.** The natural fix is to judge each slice against a reference computed with that slice left out, so a fault cannot prop itself up. That is single-deletion diagnostics, and its classic weakness is precisely multiple outliers: with two or more contaminated points, deleting one leaves the others still inflating the estimate. It is why the field moved to high-breakdown estimators computed from a clean subset rather than one-at-a-time deletion (Rousseeuw & Leroy, 1987). The planted fault spans two slices of the grid. One deletion was never going to be enough, and the measurement below says so in numbers.

The detection floor itself is the older idea: a limit of detection is a property of method *and* matrix together, not of the instrument alone (Currie, 1968; IUPAC). "−40 dBFS" was never a property of the tool alone. It was a property of the tool, that music, and that one fault type.

</aside>

<figure>
<img src="/images/blog12-ruler.svg" alt="A band cut planted in one slice of the track. Measured against a clean copy it sits at 6.29 deviations out, well past the threshold of 4. Measured against the track containing it, the spread estimate rises from 1.19 to 2.15 and the same fault sits at 3.11 deviations, under the threshold. Meanwhile a genuinely odd slice at the start of the track falls from 6.91 to 3.46 and loses its flag." width="880" height="470" />
<figcaption>The same fault, the same measurement, two different reference distributions. On the left it is unmissable. On the right it has widened the ruler it is being measured with — and pushed one of the track's own findings off the edge.</figcaption>
</figure>

## It also goes wrong in the wrong place

Two things fall out of that mechanism, and both of them land **away from where the fault actually is**.

**It raises alarms somewhere else.** With the tone planted at `−35 dBFS` between `61.4` and `68.2` seconds, the tool flags a slice at `5` to `10` seconds — fifty seconds away, where nothing was planted at all. The tone lifted the noise floor in its own slice, that shifted the reference for the whole series, and a quiet slice at the other end of the track crossed the line as a result.

**And it deletes findings somewhere else.** The band cut and the MP3 copy each take the track's own flag count from `20` down to `16` — and not one of the four that vanished is anywhere near the planted damage.

A version of this had shown up before, but only with a fault smeared across the whole track, which looked like a special case. It is not. Six and a half seconds is enough.

<aside class="plain">

**IN PLAIN TERMS**

So the report can get *quieter* when the track gets worse.

If you were watching the number of warnings go down and taking that as progress, you would have it exactly backwards. Fewer warnings is not evidence of a cleaner track. It might be evidence of a bigger fault.

</aside>

## I wrote down the fix, tested it, and it lost

The repair looked obvious. I wrote it down as a prediction before testing it, which is the only reason there is a clean record of it losing.

If the problem is that the fault gets into its own reference group, then leave it out. Judge every slice against a reference built from all the *other* slices. Simple, cheap, one line.

I ran the null test first — what does this rule do to a clean track, where every flag it raises is a false one — because the whole point of the previous part was that you measure the empty case before you interpret anything.

| rule | flags raised on the **clean** track | finds the band cut? | finds the MP3 copy? |
|---|---|---|---|
| current | `20` | no, at any setting | no, at any setting |
| leave-one-out | `24` | no, at any setting | no, at any setting |

More false alarms, not one new detection.

The reason is in the MATH BOX and it is not subtle once you see it: the planted fault covers **two** slices of the grid. Leave one out and the other is still there, still stretching the ruler. Leaving both out would mean knowing where the fault is — and knowing that removes the need for the detector.

So there is no cheap fix. The rule would need a reference from *outside* the track — a clean copy to compare against. Which is exactly what the checker holding the answer key has, and exactly why it sees all five faults at every strength while the tool on its own sees two.

That is a tidy little circle to end an arc on. The thing that separates a detector from a checker is not the cleverness of the rule. It is whether it has something honest to compare to.

<aside class="nerd">

**NERD BOX**

#### A third disagreement, inside a single function call

The click meter has its own version of this problem, and it surfaced because a column of results looked too tidy.

It reports the sharpest transient in the window as a ratio against the sharpest transient in the clean reference, and it counts how many clicks it found using a threshold `6 dB` above anything the clean window produced. Two readings, one call.

At click loudness `0.4` and `0.3` the meter reports the fault as **present** — `5.29 dB` and `3.01 dB` above the music — while the counter reports **zero clicks found**. Both numbers come out of the same function on the same audio. They are not wrong, they are answering different questions with the same word, and a report that quotes only one of them sounds certain either way.

It also looks broken at first, because below that it reads exactly `−0.00` at every level instead of decaying smoothly. It is not broken. It is a ratio of maxima, so once the click sits under the music's own sharpest transient the maximum *is* the music's own transient, unchanged, and the ratio is exactly one. The meter has a hard floor rather than a gentle fade — which is worth knowing before you read a small number as "almost detected".

</aside>

## What I would keep from this

1. **A detector that builds its reference out of the thing it is inspecting can be defeated by the thing it is inspecting.** Not fooled by a clever attacker — defeated by an ordinary fault, arithmetically, with nobody trying.
2. **Ask which direction your detector works in.** Mine notices additions and is blind to subtractions, and nothing in its description says so, and the question had not been asked until now.
3. **A single flag is not a location.** Mine raises alarms fifty seconds from the damage. If I had shipped a tool that pointed at a timestamp, it would have been pointing at innocent audio.
4. **Fewer warnings is not a cleaner track.** This is the one I would put in the specification in bold.
5. **A detection floor belongs to a fault type, not to a tool.** "−40 dBFS" was true, and it was being quoted too broadly by exactly four fifths.
6. **An instrument that degrades silently is worse than one that fails.** Mine would have printed a perfectly formatted calibration of the same file six times. The check that caught it was one command: plant twice, compare the files.

The concept that all this was built for died on 26 August, by a kill condition written down in advance. The instruments have overturned three written-down conclusions since, including one from 27 August.

<aside class="reachout">

**FOR RESEARCHERS**

#### Does yours only notice additions?

If you have built anything that flags anomalies against a reference it computes from the same data — logs, telemetry, quality control, audio, images — I would genuinely like to know whether you have tested it against a *subtractive* fault, and whether a fault in one place moves its findings somewhere else. Both are cheap. The second one only got measured here because a column of results looked wrong and the harness had to grow a pair of columns to explain it: [pete@sunrisesoftware.app](mailto:pete@sunrisesoftware.app).

</aside>

## Sources & further reading

- V. Barnett & T. Lewis, *Outliers in Statistical Data* (3rd ed., 1994) — masking and swamping as the two standard failure modes of self-referential outlier rules.
- A.S. Hadi & J.S. Simonoff, *Procedures for the Identification of Multiple Outliers in Linear Models* (JASA, 1993) — why one-at-a-time deletion is defeated by more than one outlier.
- P.L. Davies & U. Gather, *The Identification of Multiple Outliers* (JASA, 1993) — outlier identifiers, and the price of estimating your own reference from contaminated data.
- P.J. Rousseeuw & A.M. Leroy, *Robust Regression and Outlier Detection* (1987) — the high-breakdown alternative to single-deletion diagnostics.
- F.R. Hampel, *A General Qualitative Definition of Robustness* (Annals of Mathematical Statistics, 1971) — the breakdown point, and why the median held while the spread estimate did not.
- P.J. Rousseeuw & C. Croux, *Alternatives to the Median Absolute Deviation* (JASA, 1993) — what the MAD is and is not immune to.
- L.A. Currie, *Limits for Qualitative Detection and Quantitative Determination* (Analytical Chemistry, 1968) — the detection limit as a property of method and material together, not of the instrument alone.

<aside class="evidence">

**EVIDENCE**

#### Type calibration on the desktop · five fault types, one track, one window `61.4–68.2 s`, grid `5 s`, `k = 4` · 27.8.2026

- source record: `SOMNUS-PI-V19-TYYPPIKALIBROINTI` v1, 2026-08-27T15:52Z (Somnus `LN-20260827-tyyppikalibrointi`, code somnus `2dcb2a6`). A calibration, not a pre-registered container run; the leave-one-out test inside it was written down as a prediction before it ran.
- claim: the mechanical rule never fires for the band cut (`16000 → 2000 Hz`) or the MP3 round-trip (`192 → 32 kbps`), while the keyed judge reads them at `14.09 → 21.05 dB` and `9.03 → 43.30 dB`; the mel round-trip fires only at `n_mels = 24` (judge `5.38 → 10.01 dB`); the tone fires at `−35` and `−40 dBFS`, repeating the 26.8.2026 table bit for bit; clicks fire at amplitude `0.3` and not at `0.2` (`−10.46` / `−13.98 dBFS`) — status: measured
- claim: the mechanism — with the band cut planted, `band_16000` drops `−8.26 dB` in the plant window; against the clean track's distribution that is `z = −6.29`, against the planted track's own distribution `z = −3.11`, because the MAD goes `1.19 → 2.15` (median `−45.77 → −46.54`); the track's own flag at `0–5 s` goes `−6.91 → −3.46` and disappears — status: measured
- claim: the damage lands away from the plant — the tone at `−35` and `−40 dBFS` creates a false `noise_floor_db` flag at `5–10 s`; the band cut and the MP3 copy take the track's own flags `20 → 16`, none of the four inside the plant window — status: measured
- prediction overturned: leave-one-out raises the clean track's flags `20 → 24` and finds neither the band cut nor the MP3 copy at any level; the plant spans two grid windows — status: overturned
- bench fault found first: two band-cut plants at "levels" `−35` and `−45` produced byte-identical files (sha256 `ce44ad8c…`) because the bench wrote every level into `level_dbfs`, which only the tone reads; fixed, and the tone regression repeated the published table — status: measured
- click meter: at amplitude `0.4` and `0.3` the ratio reads the fault present (`5.29`, `3.01 dB`) while the counter finds `0/4` clicks; `4/4` only from `0.6`; below the music's own sharpest transient the ratio reads exactly `−0.00` — status: measured
- cost: `0 €`, desktop, no GPU box
- not shown: a whole-track band cut or codec pass (only the tone was measured at full extent); other material, grids and `k`; whether a targeted check against a clean reference catches these faults in production — the record says it must, and that no detector for them may be built on the same MAD rule.

</aside>

---

<p class="sig">— SF3D</p>
