---
title: "Threads That Don't Lose Each Other"
subtitle: "Several AI sessions work on this project at once. None of them remembers the others, none can be in a meeting, and each starts from nothing. The thing that keeps them from colliding is not communication — it is a small set of rules about when you are allowed to read and when you are allowed to write."
date: 2026-08-27
events: "16.–23.8.2026"
tags: ["somnus"]
summary: "A proposal written at nine in the evening was a line in a shared status file the next morning and in production that evening — three sessions, under a day, none of them ever talking to another. The same day, two of them wrote to the same table and collided. This is the machinery that made both of those outcomes fine, including the parts of it that admit they guarantee nothing."
canonical: https://sf3d.fi/blog/threads-that-dont-lose-each-other
html: https://sf3d.fi/blog/threads-that-dont-lose-each-other
author: Petri Korhonen, SF3D AI Lab (Sunrise Software Oy, Tampere, Finland)
site: https://sf3d.fi (index of all articles in Markdown: https://sf3d.fi/llms.txt)
numbers: measured on the lab's own hardware unless the text marks them as held for the AMD R9700 whitepaper; the Evidence block at the end of an experimental article names the run, the date, n and the pre-registration each claim rests on
---
<p class="standfirst">Here is a coordination problem with the usual answer removed. Several workers are on the same project at the same time. None of them remembers the others. None of them can be in a meeting, send a message, or ask a question and wait. Each one starts from nothing, works for an hour or a day, and stops — and the next one is not the same worker. You cannot solve this with communication, because communication requires two parties awake at once and there is no such moment. What is left is writing things down. That sounds like the weak substitute for talking. It is the stronger one, and this part is about why.</p>

<aside class="primer">

**PRIMER**

Somnus is a research loop on one GPU box: a local model proposes candidate research questions from gaps between existing findings, a second local model screens them, and the best are escalated to cloud models that try to refute them. Claims that survive go into a store, and a local judge decides whether one claim follows from another; those edges are the structure most parts of this series measure.

</aside>

## The situation, honestly stated

The workers are AI sessions — some running in an editor on my machine, some in a browser, some scheduled. Several projects share them, and one physical machine with two GPUs is shared too. A session that starts on Tuesday has no memory of Monday's session. It has the repository, and whatever anyone wrote down.

Left alone, that arrangement fails in a specific and boring way. Two sessions do the same work twice. One builds on a conclusion the other has already retracted. Someone reads a document that was true last month and confidently acts on it. None of these are dramatic failures — they are quiet, and you find out days later.

The instinct is to add a coordinator: one thread that knows everything and hands out work. We did not do that, for a reason that is easy to state and was learned the hard way. **A coordinator is a thread like any other.** It also forgets, it also reads stale documents, and when it is wrong everything downstream is wrong in the same direction. What we built instead has no centre.

## What replaced the meeting

Five mechanisms. Each is small, and each earns its place by one property rather than by being clever.

**A status stream.** One short piece of writing per project that says where things stand. Every session reads it before it does anything and updates it before it stops. That is the whole protocol, and its strength is the ordering: *read first, write last*. A session that skips the read acts on a month-old picture. A session that skips the write has, from the next session's point of view, not happened.

**An artifact channel.** Handoffs are documents with stable names, not things one session remembers and the next hopes to be told. Immutable and versioned — a new write with the same name supersedes the old one rather than editing it, so the previous version is still there when someone needs to know what changed. One name per topic, which sounds like filing pedantry and is not: stack two unrelated messages on one name and the newer one hides the older in every default listing.

**Advisory leases.** The shared machine has a sign-up sheet. Before a long run you claim it — the resource, why, and for how long — and before that you look at what others have claimed.

**A decision log.** Append-only, and it records something most logs do not: what was *recommended*, what was *chosen*, and — separately, and given explicitly rather than worked out from the other two — **why the two stand where they do.** Agreement, disagreement, or the clock running out are three different things that produce the same pair of entries.

**Fetch-first, read-before-write.** Get the current state before you touch anything, and merge into what is there rather than over it. With one sharpener that does most of the work: **if your write would produce version N+2 where you expected N+1, someone wrote while you were thinking.** Stop, read, merge. That is a stop sign a machine can see.

<aside class="plain">

**IN PLAIN TERMS**

The difference between a meeting and a shared notebook is that a meeting only works if everyone is in the room, and the notebook works precisely when they are not.

A conversation you missed is gone. A line someone wrote in the notebook is still there next week, and it is still there for the person after that, and it does not care that nobody introduced you.

</aside>

<figure>
<img src="/images/blog9-mechanisms.svg" alt="Five coordination mechanisms as rows, each with the one property that earns its place. Status stream: read first, write last — a session that skips the write did not happen. Artifact channel: documents, not memory — immutable and versioned, one name per topic. Advisory lease: a sign-up sheet, not a lock — expires by itself, and an empty list is not proof the machine is idle. Decision log: why, as its own field — recommendation, choice and reason are three separate facts. Fetch-first: expected N+1 and got N+2 means someone wrote while you were thinking." width="880" height="530" />
<figcaption>Each row's right-hand column is the property that makes a quiet failure loud. The one in red is the only place where the mechanism warns you about itself.</figcaption>
</figure>

## The day it proved itself, twice

**17 August, in one direction.** At nine the previous evening, a session working in chat wrote down a proposal: it should be able to read the project's verified-claim store directly, because working from someone else's summary costs three specific things — verification breaks at the summary, questions get written blind, and connections between claims never get found because nobody can see both.

The next morning a different session, working in the editor, read that proposal and put it into the status stream as a task for whichever thread got there next. It was not that session's job. It filed it and moved on.

That evening a third session had built it: three read-only tools against the claim store, placed on the data server rather than the control plane, with the reasoning for the placement written down. And it carried the proposal's methodological caveat all the way into the tool descriptions themselves — duplicate checks before a run, free browsing only after — so the discipline would not depend on anyone remembering it.

**Proposal to production in under a day, across three sessions, none of which ever addressed another.**

**The same day, in the other direction.** Two sessions added rows to the same decision-log table and hit a merge conflict. That is the failure the whole arrangement is supposed to make survivable, and it was: both sets of rows were kept, all four of that day's decisions survived, nothing was silently dropped.

The second story is the more useful one. A coordination scheme that has never been tested by a collision has not been tested.

<aside class="math">

**MATH BOX**

#### Why the lease expires, and why that is not laziness

The sign-up sheet is deliberately **not a lock**, and the reason is a theorem rather than a shortcut.

In an asynchronous system where participants can fail and you cannot tell a crashed worker from a slow one, no protocol reliably reaches agreement — the FLP impossibility result (Fischer, Lynch & Paterson, 1985). A hard lock inherits that problem directly: if the holder dies while holding it, the resource is locked forever, and no amount of waiting distinguishes "still working" from "gone".

A **lease** is the standard escape: a lock that carries an expiry, so a holder that vanishes releases it by doing nothing (Gray & Cheriton, 1989). It converts an unsolvable agreement problem into a time-bounded one, and the price is stated in the original paper — during the window you get exclusion, and outside it you get availability, but you never get both without a clock you trust. Ours expires by itself, caps at four hours, and — deliberately — computes expiry *at read time* rather than in a cleanup job, so there is no background task whose silent death would turn a forgotten row into a permanent block.

Two more standard pieces, both chosen rather than stumbled into. **Acquiring is a single conditional statement, not read-then-write** — a compare-and-set, which is the primitive that can resolve this class of race at all (Herlihy, 1991); read-then-write would contain the exact race the lease exists to narrow. And the version-jump stop sign is **optimistic concurrency control** (Kung & Robinson, 1981): do the work assuming no conflict, then validate at commit time that the version you read is still the current one, and redo rather than overwrite when it is not.

None of this is novel — the problem is old, the answers are standard, and the only real design work was noticing which old problem we had.

</aside>

## The mechanism that admits it guarantees nothing

The most important line in the whole arrangement is a disclaimer written into the lease tool's own description:

> A mechanism that hints at a guarantee it does not have is worse than no mechanism at all.

The reader of that description is usually a machine, and a machine will act on what the interface implies. So the interface says, in its own text, that the lease is advisory, that Atlas enforces nothing, and that two readers can see "free" in the same second.

There is a second admission next to it, and it is the one I would most want a future session to read:

> An empty lease list does not prove the machine is idle. A thread that did not take a lease is invisible here.

That sentence exists because the bookkeeping can only ever describe the threads that participate in it. Anything else is a blind spot, and the honest move is to name it in the place where someone would otherwise draw the wrong conclusion.

<aside class="plain">

**IN PLAIN TERMS**

It is a sign-up sheet on a door, not a lock. It cannot stop anyone. What it removes is the specific situation where someone walks in assuming the room is empty *because nobody told them otherwise* — and it says on the sheet itself that a blank sheet is not proof the room is empty.

Claiming more than that would be worse than having no sheet at all, because people trust locks.

</aside>

## The rules that came from mistakes

Three of these conventions exist because something went wrong, and all three are recorded with the mistake attached.

**Fix the links in the same commit that moves the files.** Documents were reorganised into new folders and the references to them were not updated. A later session followed a broken path, found an outdated table, inferred the project's state from it, and wrote that wrong state forward into a shared record. The damage was not the broken link. It was that a *confident* conclusion was drawn from it and passed on.

**Machine state does not belong in a document channel.** Covered in the box below — a proposal was accepted in principle and rejected in form, which is a better outcome than either a yes or a no.

**And the one worth the most.** In a written exchange between two threads, one of them recorded this about a status entry that had been wrong:

> Example 1 is my error: I wrote into the status stream that two stages were open, having read a broken README. You corrected it by hand.

That is a thread documenting its own bad write, by name, in the shared channel, in the middle of proposing improvements. It matters more than the mechanisms. A coordination scheme where being wrong is embarrassing produces threads that quietly paper over their mistakes, and then the next session inherits a clean-looking record that is false.

<aside class="nerd">

**NERD BOX**

#### State is not a document, and the arithmetic says so

A proposal came in to publish the machine's live condition — GPU memory, which model holds which card, whether the box is reachable — into the shared artifact channel. The reasoning was sound: other threads need it, and the channel is where shared things go.

The answer accepted the need and rejected the location, with a number. The channel is immutable and versioned, built for documents a human reads. A machine writing every five minutes produces about `288` versions a day. It would drown the "new to you" feed, fill the read-receipt table — receipts are per row id, and every version is a new id — and make version history useless for the documents that actually need it.

The distinction underneath is the one worth keeping: **machine state wants one current value that gets replaced; a document channel offers an unbroken history that never gets replaced.** Those are opposite access patterns, and the fact that both are "shared information" is not enough to put them in the same store. So the telemetry layer gets its own table on the same pattern as the lease — one live row per machine, overwritten — and, at time of writing, **it has not been built.** The lease was built first, because of an observation in the proposal itself: telemetry tells you what a machine is doing, and it cannot tell you what someone intends to do next.

One more detail from that exchange, adopted without argument: the *server* writes the machine's status, not the machine. A box that has crashed writes nothing, so a reader cannot distinguish "the machine is down" from "the reporter is broken" — whereas an outside prober can write `reachable: false`, which is a fact rather than a silence.

</aside>

<figure>
<img src="/images/blog9-state-vs-document.svg" alt="Two stores side by side. Left, the document channel: a stack of versions v1 to v4, never replaced, written by people a few times a day, where history is the product. Right, the state table: one live row per machine, overwritten, written by a machine constantly, where only the current value matters. Below, the arithmetic — a write every five minutes is about 288 versions a day, which drowns the feed, fills the read-receipt table and makes version history useless." width="880" height="450" />
<figcaption>The same information can be shared and still belong in two different places. The tell is whether you need the last value or all of them.</figcaption>
</figure>

## 23 August, in this session

The arrangement is not a finished story, so here is its most recent test, from 23 August.

A session — this one — was asked to work out where a set of measurement instruments should go in the product, and recommended a specific piece of work in a specific file. Before proposing it, it read that repository's recent history and found commits from **eleven minutes earlier**: another session was in that exact file, right then, doing adjacent work. The recommendation was correct and the timing was not, so the work was recorded and frozen instead of started.

Two things about that are worth keeping. The collision was avoided by reading current state before writing — the plainest of the five rules, applied to a repository rather than to the status stream. And the decision log entry says `disagreed`, because what was chosen differed from what was recommended, with the prose immediately explaining that this was a **timing override rather than a disagreement about the content**.

That distinction only survives because the log demands the reason as its own field instead of inferring it from the two choices. Comparing "recommended" with "chosen" tells you whether they match. It never tells you why — and a year from now, why is the only part anyone will need.

## What survives

1. **When communication is impossible, ordering replaces it.** Read first, write last. Nearly everything else is a refinement of that one rule.
2. **A coordinator is a thread like any other.** It forgets what the others forget and it reads the same stale documents. Distributing the state beats centralising the authority.
3. **State the limits inside the mechanism, not in the documentation nobody opens.** The lease says it is advisory in its own description, because its reader is usually a machine that will believe the interface.
4. **Record why, separately from what.** Recommendation and choice are two facts; the reason they differ is a third one, and it is the one that ages best.
5. **A collision is a test result, not a failure.** The 17 August merge conflict is better evidence that the arrangement works than the 17 August success is.
6. **Write down your own bad writes.** A record that is embarrassing to correct becomes a record that is quietly wrong.

Atlas — the system holding the status stream, the channel, the leases and the log — has more to it than this: an architecture, a set of connectors, a gate that separates what has been proposed from what has been accepted. Those are their own parts. This one is only the practical story: what happens when several workers who cannot talk are pointed at the same project, and what it takes for that to be fine.

<aside class="reachout">

**FOR RESEARCHERS**

#### Working on something like this?

If you run several agent sessions against one codebase or one machine, the interesting question is not the protocol — the distributed-systems answers are decades old and they port over unchanged. It is what you had to write into the *tool descriptions* to keep an agent from over-trusting them, since an agent reads an interface far more literally than a person does. If you have found a phrasing that reliably stops a model from treating an advisory mechanism as a guarantee, I would like to see it: [pete@sunrisesoftware.app](mailto:pete@sunrisesoftware.app).

</aside>

## Sources & further reading

- M.J. Fischer, N.A. Lynch & M.S. Paterson, *Impossibility of Distributed Consensus with One Faulty Process* (JACM, 1985) — why a hard lock cannot be made safe against a holder that vanishes.
- C. Gray & D. Cheriton, *Leases: An Efficient Fault-Tolerant Mechanism for Distributed File Cache Consistency* (SOSP, 1989) — the expiring lock, and the consistency-for-availability trade it names explicitly.
- H.T. Kung & J.T. Robinson, *On Optimistic Methods for Concurrency Control* (ACM TODS, 1981) — validate at commit that the version you read is still current; the version-jump stop sign in one sentence.
- M. Herlihy, *Wait-Free Synchronization* (ACM TOPLAS, 1991) — why the acquire is a compare-and-set rather than a read followed by a write.
- L. Lamport, *Time, Clocks, and the Ordering of Events in a Distributed System* (CACM, 1978) — the founding argument that ordering, not simultaneity, is what shared work actually requires.
- D. Parnas, *On the Criteria To Be Used in Decomposing Systems into Modules* (CACM, 1972) — the older form of the same idea: coordinate through stable written interfaces rather than through knowing what the other party is doing.

<aside class="evidence">

**EVIDENCE**

#### Two threads, one day: the crystal read tools and the decision-log collision · 16.–17.8.2026

- source records: the proposal `SOMNUS-KIDELUKUOIKEUS` v1 (written 16.8.2026 in the evening), its result `SOMNUS-KIDELUKUOIKEUS-TULOS` v1 (2026-08-17T17:54Z), the cross-work note `SOMNUS-ATLAS-RISTIINTYO` v1 (2026-07-31T12:11Z), the lease message `SOMNUS-MSG-ATLAS-VARAUS` v1; Somnus `LN-20260817-kidetyokalut-ja-ristiintyo`. No pre-registration: this part reports a process, not an experiment.
- claim: proposal to production in under a day across three sessions that never talked — written 16.8. at nine in the evening, a TODO in the status stream on the morning of 17.8., in production on the evening of 17.8. (quaesitor PR #72, `v2.13.27`): three read-only tools `list_crystals`, `get_crystal`, `search_crystals` (pgvector search server-side) on the existing read-only view, same token, same rate limit, no new infrastructure — source `SOMNUS-KIDELUKUOIKEUS-TULOS` v1 — status: measured
- claim: live verification from the box — `tools/list` returns `5` tools; `search_crystals("IIR cascade cumulative phase at low frequencies", 3)` returned similarities `0.539 / 0.486 / 0.451` for crystals `128 / 125 / 123`, the old IIR-phase material, i.e. the duplicate check the tool was built for; the corpus stood at crystal `300` — source same — status: measured
- claim: the same day two threads appended rows to the same decision-log table and collided; the merge lost no content, `4` rows kept — source `LN-20260817-kidetyokalut-ja-ristiintyo` — status: measured
- claim: `288` versions a day is arithmetic (one status write every five minutes); the `4 h` lease cap is the mechanism's declared setting — status: stated, not measured
- not shown: the commit timestamps behind "eleven minutes earlier" are in the repositories, not in the channel; whether the rules scale beyond a handful of threads has not been measured.

</aside>

---

<p class="sig">— SF3D</p>
