Software factory

Nodes that write software

An orchestrator owns a requirements.md, splits it into tickets, hands each one to a worker backed by a node's own model, and closes a ticket only when CI goes green. Ten nodes without that last part are ten conflicting branches.

The loop

requirements.md
      │  split into tickets with a dependency DAG
      ▼
  overlap check   — abort if two tickets declare the same file
      │
      ▼
  one batch = tickets whose files are disjoint, run in parallel
      │
      ├── worker: asks the model, parses ```file blocks,
      │           writes ONLY inside its allowed paths
      ▼
  CI gate — the tests run
      │
      ├── green → the ticket is done, never attempted again
      └── red   → picked up by the next run
      │
      ▼
  append-only run log, so a cron resumes instead of redoing

The parts that are load-bearing

The gate must not be written by the model

The tests that judge the work are seeded before any worker starts and belong to no ticket. If the same model wrote both the code and its tests, a green run would prove nothing — the cheapest way to pass is a test that always passes. The overlap check does not cover this: it stops two tickets from clashing, and the gate is nobody's ticket.

Conflicts are made impossible, not resolved

Two tickets are never allowed to declare the same file — the run aborts before anything is assigned. That is what lets a batch run in parallel with no merge step: the union of what the workers produce cannot collide.

It also has a cost worth naming: a requirement that cannot be split into disjoint file sets — a refactor, a schema change, a shared lockfile — has no path through this design today.

Every limit lives in the harness, none in the prompt

Steps and tokens are checked before spending, not after. The per-tool timeout is smaller than the task's, so a hung tool fails on its own instead of taking the task down. Retries are for transient failures only: a 400 or a failing test is never retried, and a timeout gets one retry rather than three — a ten-minute timeout is not a 503.

What the model has to be able to do

Measured against a Qwen3 4B served locally on CPU:

 Result
Emits parseable file blocksYes — including a variant syntax of its own, which the parser had to learn
Writes two files in one ticketYes
Respects the allowed paths on its ownNo. It tried a README.md outside its ticket; the check rejected it
Solves the task rather than copying the prompt's exampleYes — once the example stopped being a solution to the task
Two concurrent requests on one nodeYes, 109 s and 163 s side by side

A 1B model could not follow the format reliably. Treat roughly 4B as the floor.

The jail was exercised, not merely present. The model attempted the file outside its ticket. The result is not a model politely respecting a list — it is a check catching an attempt, which is the only version of that result worth anything.

Four things the prompt cannot do without

Each came from a run that failed, and removing any one of them breaks it in a different way:

  1. The example must carry real code. Replaced with a placeholder comment, a small model returned nothing at all: a filler comment is not a mould, and a mould is what guides it.
  2. One path in the whole prompt. With an example path different from the ticket's, the model wrote to the example's — reasonably, it was the one sitting in the "this is how a path looks" slot.
  3. The example must not be the answer. When the example solved the task, a model that merely copied it looked exactly like one that had understood — destroying the only thing being measured.
  4. No meta-instructions. Two lines explaining the example — "do not copy it", "copy the first line exactly" — made a model that had been writing correct code emit broken multilingual text about someone unable to follow instructions. The prompt says what to deliver; it never discusses itself.

Crossing machines is built

Was the missing piece; now has a coordinator, a worker, and tests. A worker no longer needs the coordinator's local disk: orchestrator/coordinator.mjs sends task:assign over the Protomux channel to worker/task-accept.mjs (run standalone via worker/serve-tasks.mjs), the workspace crosses as a Hyperdrive read sparsely over the connection the two nodes already hold — no DHT announce per ticket — and results come back inline in the same message for the common case. See the task protocol for the full shape, what changed from the original sketch, and what is still open (a coordinator-election story and prompt-injection measurement — the retry ceiling and the global budget shipped in this same merge).

What is not built

Closed since the first draft of this page. scripts/nightly-build.mjs is the cron wrapper: it runs the coordinator once, logs the night, and exits non-zero when a human needs to look. --max-attempts (default 4) escalates a ticket as ticket:blocked instead of reassigning it forever. --budget caps cumulative tokens for the whole project, checked before every wave, not per task.
MissingConsequence
Node specialisationThe pieces exist — signed manifest, an unused tool allowlist, a RAG index that replicates as a hypercore — but nothing uses them to make a node good at one kind of work.
Tools beyond writing filesNo shell, no git. The worker writes files and the coordinator runs the tests; that is the whole toolset.
Prompt-injection measurementsThe prompt states that ticket text and file contents are data, not instructions. That has never been tested — even now that the context drive widens what a worker reads.
A signed security.acceptsTasksThe manifest schema is frozen (additionalProperties: false, generated outside this repo); authorization for the task protocol is node config (--allow / --worker) instead, for now — see the protocol's authorization section.
Next: the task protocol — message types, authorization, idempotency and failure handling for the worker that is no longer a child process but another machine.