Tampere, Finland · 61.5°N
From world records
to world models.
Independent AI research lab running multi-agent cognition systems on AMD RDNA4 hardware. Two decades ago we froze silicon with liquid nitrogen to break records. Now we run it 24/7 to break assumptions.
01Lab Notes
ALL POSTS →Blind to Anything Taken Away
I have a tool that listens to a piece of music and flags anything that looks out of place. It turns out it only notices things being added. Take something away — cut the top off the sound, run it through a bad codec — and it says nothing, at any strength, because it judges every moment against the rest of the track and the fault joins the comparison. The obvious fix was written down as a prediction, tested, and fell over.
What a Negative Result Is Worth
Three tests in sequence, each closing the one before it. A blind run on clean material returned no false find. A level sweep turned 'we found nothing' into a sentence with a number in it — and produced three different thresholds, not one. And a limit written down on the strength of an argument fell to the measurement it should have had first, revealing something worse than the limit it replaced: an artefact that makes the track score cleaner.
The Concept That Killed Itself
R = 0.73, below the line we had signed. The concept is falsified and the queue behind it does not open. And the losing side hit every field of a sealed answer key — the frequency to a tenth of a hertz, below the resolution of the instrument that found it — while the winner named neither the type nor the level. Both numbers are true, and the reason is that the score counts something the work does not produce in proportion.
Threads That Don't Lose Each Other
A proposal written at nine in the evening was a line in a shared status file the next morning and in production that evening — three sessions, under a day, none of them ever talking to another. The same day, two of them wrote to the same table and collided. This is the machinery that made both of those outcomes fine, including the parts of it that admit they guarantee nothing.
The Bottleneck Wasn't the Hardware, Again
Hardware is the most satisfying suspect in computing: visible, blameable, replaceable. This part is three cases where it was the obvious answer and the wrong one — including a freeze we still cannot explain, and the repair we are deliberately not performing because it would destroy the evidence.
Three Budget Raises and the Wrong Lever
A short one about a reflex. When a machine runs out of room, you give it more room. We did that three times to a model whose reasoning was being truncated, and the number never moved. The fix was a knob that pointed the other way — and the only reason we can call this an experiment rather than a lucky guess is that the answer key was sealed before the run.
The Question Is Part of the Instrument
Part 4 ended with 136 forced pairs and no edges: aiming at the right target was not enough. This part puts the question itself under the microscope — four forms, four preregistered runs, locked answer keys, and a chain that named a theorem our own key did not contain. The question turns out to be an instrument component, and preregistration is the only reason any of it was visible.
Abduction Produces Generalities
In production, the abduction stage has never produced an explanation that survived validation — fifteen clusters, zero passes. Killing every alternative explanation, one paid run at a time, led to a planted control, a fabrication leak that grows with model capability, and a map: the corpus is an archipelago, and abduction needs a continent.
The Judge Is Part of the Instrument
A 17× speed anomaly exposed an unrecorded judge swap. A pre-registered test then killed our favourite hypothesis about the judge's errors — and field data revealed the real one: the fast judge misses genuine inferences the careful judge confirms unanimously. Structure is not a property of a corpus. It is a property of a corpus–judge pair.
The Value Function That Reads the Shape
To focus on one topic, the machine had to price every candidate question against seven sub-questions — and the obvious way to do that silently collapses all seven into one average direction. The theorem behind that collapse, the fix that reads a shape instead of a sum, and the discipline that catches it before a euro is spent.
02Signal Feed
FULL FEED →A curated stream of AI research, models, agents and hardware — assembled by an autonomous agent, every item linking to its source. Open the feed →