Software factory / Protocol

The task protocol

qvac/task/v0 — how one node hands a unit of software work to another, gets the result back, and decides whether to trust it. The protocol ships over OTA to every node, and two nodes on different versions is the normal case, not the rare one.

Status: built, phase 1. orchestrator/coordinator.mjs, worker/task-accept.mjs, worker/serve-tasks.mjs, orchestrator/task-protocol.mjs, orchestrator/context-drive.mjs and orchestrator/mirror.mjs exist and are tested: message-layer unit tests, a mirror + CI test, an idempotency/resume test, a real corestore-replication test over a non-DHT connection, an end-to-end run over a fake swarm, and a one-off smoke test (scripts/smoke-real-swarm.mjs) over a real Hyperswarm/DHT connection between two independent node identities. Two things in this page changed from the original design because they ran into a real constraint — the manifest is frozen — and are called out where they happen: no task:ack, and authorization is node config rather than a signed manifest field, for now. Everything else here is what shipped. See the factory overview for the single-machine loop this grows out of.

The problem this solves

The orchestrator today spawns its workers as child processes:

spawn(process.execPath, ['worker/run.mjs', '--ticket', ...])

That assumes one machine. The worker writes to its local workspace, the orchestrator runs CI on its local workspace, and those are the same directory — which is the only reason any of it works. Ten nodes under that design are ten machines writing to ten disks that nobody joins.

Three things have to cross the boundary: the assignment (here is a ticket), the context (here is the tree you are working against), and the result (here is what I produced). Each takes a different route, and the reasons are below.

The shape

┌─ COORDINATOR (orchestrator/coordinator.mjs) ┐
│  requirements.md → tickets                   │
│  publishes workspace/ as a read-only         │
│    Hyperdrive — in its OWN corestore ────────┼──┐
│                                               │  │
│  task:assign  ────────────────────────────────┼──┼──▶ Protomux, qvac/task/v0,
│  ◀── task:accept / task:progress              │  │      over the connection the
│  ◀── task:result (files INLINE ≤ 1 MiB,       │  │      two nodes already hold —
│       or a driveKey for the overflow case)    │  │      no new DHT announce
│                                               │  │
│  re-validates → mirrors → runs CI            │  │
└────────────────────────────────────────────────┘  │
                                                    │ same connection replicates
   ┌────────────────────────────────────────────────┘ this drive too — mounted
   ▼                                                   read-only, fetched sparsely
┌─ WORKER (worker/task-accept.mjs, via serve-tasks.mjs) ┐
│  accepts only a key on --allow                         │
│  reads context from the coordinator's                  │
│    drive — only the paths it opens                     │
│  infers over HTTP to the LOCAL gateway                  │
│    (never the engine in-process — see below)            │
│  validates paths against allowedFiles                    │
│  returns files inline; only over the ceiling               │
│    does it write to ITS OWN drive and return a driveKey     │
└────────────────────────────────────────────────────────────┘

Why the assignment goes over Protomux

The connection already exists and is already authenticated. Hyperswarm introduces the peers through the DHT and holepunches, so there is no inbound port to open and NAT is not in the way — a point worth stating because it is the usual reason a design like this ends up needing a queue. It does not.

The control channel is also built for exactly this kind of extension. It carries one JSON message type on purpose, so that a node running an older build ignores a type it does not know and keeps going. Adding a message is backwards-compatible by construction. qvac/swarm.mjs now dispatches any task:* message to a listener either side registers (addTaskListener/sendTask); it never invents a second channel.

Why the context goes over a drive — and never a second connection to get it

A worker that can only see the ticket text can write new files and nothing else. Real work means modifying code that already exists, which means reading the tree. Hyperdrive downloads sparsely, by path: the coordinator publishes a 40 MB workspace and a worker that opens four files transfers four files. This is the case where a drive earns its place.

What changed from the first sketch of this page: the drive is created in the coordinator's own corestore and replicates over the connection the coordinator and worker already hold as marketplace peers — the same thing that already carries directory and file replication (corestore.replicate(socket) in qvac/swarm.mjs). There is no per-ticket swarm.join on a fresh discovery key. Skipping that saves the 2–17s (38s tail, measured) a fresh DHT announce costs, on every ticket — the single largest latency item this protocol controls, next to the inference itself. See orchestrator/context-drive.mjs.

Why the result usually travels inline, not over a second drive

The control channel shares one multiplexed stream with hypercore replication and with the tokens of live inference — sending bulk bytes over it would compete with what the node is serving to somebody else. A separate Hyperdrive per ticket avoids that, but for a handful of KB-sized source files it is more latency (create → replicate → discover → sparse-fetch) and more code than the problem needs.

So task:result carries the files' bytes inline, up to a 1 MiB total ceiling — comfortably under the 16 MiB frame limit NoiseSecretStream enforces below Protomux, and enough for a batch of source files with room to spare. Only past that ceiling does a file go on the worker's own Files drive (still replicated over the same existing connection, same as the context drive) and travel as { path, hash, drive: true } plus a driveKey instead of inline content. See orchestrator/mirror.mjs's INLINE_CEILING.

The store-and-forward objection still does not apply, for a narrower reason than originally stated. Hypercore is not store-and-forward: the sender must be online while the receiver downloads. That is fine here because the overflow drive is only used while the worker process — worker/serve-tasks.mjs, a long-lived companion process, not a one-shot child — is up and actively seeding the result it just produced; it is not claiming the node is 24/7 in general, only that this process is alive for exactly as long as the coordinator needs it to be.
Why authorization and this drive choice are node config, not manifest fields. The signed manifest schema is frozen (manifest-v0.json, additionalProperties: false, generated from a zod schema in a package outside this repo) — adding security.acceptsTasks or security.maxConcurrentTasks there needs a coordinated schema-version bump, not a quick edit. Until then: a coordinator is told which workers to use with --worker <hex-key,...>, and a worker is told which coordinators to trust with --allow <hex-key,...>. Both still enforce "being on the topic is not enough" — just as node config instead of as an advertisement. Coordinator.workers() also honours security.acceptsTasks if a future manifest version carries it, so this does not need to change again the day the schema catches up.

The trust boundary, and where it moved

A remote worker's jail protects the worker's owner, not the coordinator. Today the path check runs inside the worker, and the worker is trusted because the orchestrator spawned it. Across machines, that check runs on the other party's hardware. A modified — or merely buggy — worker can put anything into its own drive.

Two consequences, and neither is optional:

  1. The coordinator fetches by declared path, never "everything you have". It knows the ticket's allowedFiles because it assigned them. Anything else in the worker's drive is not fetched and not looked at. A file the ticket produced but did not declare is lost — which is correct: a ticket that needs to produce a file declares it.
  2. The jail runs twice. In the worker (worker/task-runner.mjs, via orchestrator/security.mjs's validateWrite), where it catches the honest mistake; and in the coordinator on arrival (orchestrator/mirror.mjs, the same validateWrite), where it catches the dishonest one. Same function, two sides of the boundary — exercised by test/coordinator-e2e.mjs, which sends a worker a spec that tries an out-of-scope file and checks it never reaches the coordinator's workspace, with the rejection logged as a violation event either side catches it on.

Authorization

Being on the topic is enough to be a peer. It must not be enough to make another machine work and write to its disk.

ControlWhereWhat it does
worker/serve-tasks.mjs --allow <hex-key,...>Node config, worker sideOff by nothing to advertise: the process only exists if you run it, and it accepts task:assign from nobody until a key is on this list.
orchestrator/coordinator.mjs --worker <hex-key,...>Node config, coordinator sideWhich connected peers this coordinator will actually place work on — see the manifest note above for why this is config today, not a signed field.
security.acceptsTasks / security.maxConcurrentTasksSigned manifest — not yetWhat the design originally proposed. Coordinator.workers() already honours these if a peer's manifest carries them, so the day the frozen schema gets a version bump this activates with no further change — it is just not the primary path today.
This is where x402 attaches later. The worker's check — "do I accept work from this key?" — is a single point where a payment verification drops in without reopening the protocol. Building it now costs three small things. Retrofitting it onto a protocol already deployed over OTA costs a migration across every node.

Messages

All messages ride the existing single JSON message of qvac/node/v0, with a type in the task: namespace. Every one carries protocol: "qvac/task/v0"; a receiver that does not recognise the version replies task:reject with reason: "unsupported-protocol" rather than guessing.

task:assign coordinator → worker

FieldTypeDescription
protocolstring"qvac/task/v0"
attemptIdstringUnique per assignment, not per ticket. The idempotency key — see below.
ticketIdstringStable across attempts. What the run log keys on.
specstringWhat to build. Treated as data by the worker, never as instructions.
allowedFilesstring[]Exactly the paths this ticket may write. Enforced on both sides.
contextDrivestringHex key of the coordinator's read-only workspace drive.
contextPathsstring[]Optional. Files worth reading first. A hint, not a limit — the worker may open anything in the drive.
limitsobject{ maxSteps, maxTokens, toolTimeoutMs, taskTimeoutMs }. Advisory: the worker enforces its own and may be stricter.
deadlinenumberUnix ms after which the coordinator stops waiting. The worker should give up on its own before this.
spec and allowedFiles must agree. Measured against a 4B model: given a spec asking for three deliverables while the ticket permitted one file, the model did not cheerfully write extra files — it looped on the contradiction for 253 s and produced nothing. That is worse than a violation, because a violation at least gets caught and logged. Whatever splits requirements into tickets has to hold this invariant.

task:accept · task:reject worker → coordinator

Answered immediately, before any work starts, so the coordinator learns quickly that a node will not take the job and can place it elsewhere.

FieldTypeDescription
attemptIdstringEchoed back.
acceptedboolean
reasonstringOn refusal: not-authorized · at-capacity · unsupported-protocol · no-model · tasks-disabled · busy-elsewhere (a duplicate attemptId, refused idempotently rather than started twice)
etaMsnumberOptional, on acceptance. A rough estimate, so the coordinator's timeout is informed rather than arbitrary.

task:progress worker → coordinator

Optional heartbeat, and the only thing that distinguishes a node thinking hard from a node that fell over. Bytes generated, chunks, time to first token — the same numbers --log-inference already prints locally. A worker that stops sending these for longer than the coordinator's patience is treated as gone.

task:result worker → coordinator

FieldTypeDescription
attemptIdstringOnly the live attempt's result is accepted.
okbooleanWhether the worker believes it produced something usable. Not a verdict — CI decides.
filesobject[]{ path, hash, bytes, content } per file, content inline as UTF-8 text — up to a 1 MiB total across the whole result. hash carries its algorithm inline (sha256:<hex>), the same shape the attestation format uses. A file over the ceiling instead carries { path, hash, bytes, drive: true }, no content.
driveKeystringPresent only if at least one file is drive: true. Hex key of the worker's own Files drive, replicated over the same connection the task messages ride — no separate discovery.
rejectedobject[]Paths the worker's own jail refused. Reported, not hidden: it is the count that says how often something tried to step outside its scope. The coordinator logs these as violation events with side: "worker", distinct from ones its own re-check on arrival finds (side: "coordinator").
usageobject{ steps, tokens, tokenSource }. tokenSource is provider when the model counted and gateway when it was estimated from bytes — the two are never shown as the same thing.
reasonstringOn failure, what happened: limit-reached · no-blocks · reasoning-unclosed · engine-error · context-unavailable · timed-out
There is no task:ack. The earliest sketch of this page had one, sent once the coordinator finished fetching and verifying, doubling as the worker's cue to release its task slot. With results inline under the ceiling there is nothing left to seed once task:result is sent, so nothing to acknowledge — the worker releases its slot the moment it replies. The overflow drive case does not change this either: the worker keeps its Files drive open regardless (it is the same drive qvac-node send/fetch already uses), so there is no separate lifecycle to close out per ticket.

Idempotency

The case that forces it: a ticket goes to node B, B goes quiet, the coordinator times out and reassigns to C — and then B comes back and delivers. Two results, one ticket.

LayerMechanismWhat it prevents
1 · AttemptThe coordinator mints an attemptId per assignment (Coordinator.live, a ticketId → live-attemptId map) and accepts a result only for the attempt currently live for that ticket.A superseded worker's late delivery overwriting fresher work. Discarded with a log line — never silently. Coordinator._onTaskMessage; tested directly in test/coordinator-idempotency-test.mjs.
2 · ContentFiles carry hashes; a mismatch rejects the whole result rather than writing anything (orchestrator/mirror.mjs).A partially-trusted mirror. Makes retrying safe: identical bytes verify and land the same way every time.
3 · StateThe run log is append-only and done() reads the last event per ticket.ticket:done becoming a counter instead of a state. Replaying it is harmless. This already worked on one machine — the requirement was not to break it.
4 · Resumeresult:received is appended to the run log the moment a valid task:result arrives, before the mirror or CI run. Coordinator.resume() replays any ticket left at that state on startup.A coordinator that dies between accepting a result and closing the ticket reassigning the work and paying for the inference a second time — instead it re-applies the already-logged result, with no worker contacted. Tested in test/state-resume-test.mjs and test/coordinator-idempotency-test.mjs.
The key is per attempt, not per ticket. Keyed per ticket, a reassignment reuses it and the second delivery is rejected as a duplicate — which would mean a legitimate retry could never deliver at all.

Capacity, and a gauge that does not lie

A node working on a ticket is using its GPU. If that does not show up in what it advertises, the manifest is lying to whoever is choosing where to route.

Count inferences in flight, not tasks open. A task's inference calls go through the same path as any other request and occupy a slot for their duration. Between inferences — parsing, writing files — the node genuinely is not busy. Counting per inference is both simpler and more truthful than reserving a slot for the whole task.

This is exactly why worker/task-accept.mjs calls the LOCAL gateway over HTTP for every completion instead of the engine in-process: the gateway's store.beginRequest/endRequest is the one place this count is kept, and a task's inference goes through it like any other request. An in-process engine call would bypass that counter and the manifest would start lying the moment a node picked up task work — see the discussion of this in the repo's own audit of this design.

Willingness to commit is a separate number — maxConcurrentTasks, still node config today rather than a manifest field, per the note above. A node with three request slots should not accept fifty tasks and interleave them.

When things break

FailureWhat happens
Worker never acceptsShort timeout on task:assign; the coordinator offers the ticket to the next node in its pool.
Worker goes silent mid-taskNo task:progress past the deadline ⇒ the attempt is abandoned and reassigned with a fresh attemptId. Layer 1 handles the ghost that comes back.
Connection drops, then reconnects before the deadlineHyperswarm reconnects on its own, and messages resume flowing under the same peer key — nothing special has to happen for a task:progress/task:result that arrives after a reconnect. Not built: there is no task:status message to actively probe an attempt's state after a drop; today the coordinator only finds out via the next message that arrives, or via the progress deadline.
Files do not match their hashesRejected wholesale. A partially trusted result is worse than none.
A path outside allowedFilesNot written on the coordinator's side even if it somehow arrived. Logged as a violation event — this is the number that says how often a worker tried to step outside its ticket.
The coordinator process dies after accepting a resultresult:received is already in the run log (idempotency layer 4) — the next coord.run() mirrors and re-runs CI from the log, no worker contacted.
CI redThe ticket stays open for the next run. Unresolved: nothing caps how many times a ticket may be reattempted, so today it would be retried forever.

Versioning

The channel was built so that adding a message type is backwards-compatible: an older node ignores what it does not recognise. That covers new messages. It does not cover a changed field in an existing one.

With OTA running, two nodes on different versions is the normal state, not the exception. Any design that assumes both sides upgraded together is a design that breaks on the day it ships.

Open questions

  1. Closed since this page was first drafted: a retry ceiling (--max-attempts, default 4, escalates as ticket:blocked) and a global token budget (--budget, checked before every wave) both shipped in orchestrator/coordinator.mjs, and scripts/nightly-build.mjs is now the cron wrapper these seven-day runs use.
  2. Coordinator election. The role is requested by one node today. What happens if two claim the same requirements.md is undefined — and it is the kind of undefined that shows up as two workers doing the same ticket. Unchanged by this build: orchestrator/coordinator.mjs assumes it is the only coordinator for its requirements.md.
  3. Prompt injection. The worker's prompt states that spec text and file contents are data, not instructions. That the jail catches a path that tries to escape its ticket is tested (test/coordinator-e2e.mjs) — but whether hostile text sitting in a context file can talk a model into something the jail does not check for (a misleading comment, a subtly wrong implementation, tool use if any is ever added) has never been measured. The context drive widens what a worker reads, which is exactly where such text would arrive.
  4. An active task:status probe. Not built. Today a dropped connection is discovered passively — by the next message that does or does not arrive, or by the progress deadline — rather than by the coordinator asking. Before reassigning (and paying for a fresh inference), a short probe to a still-connected-but-silent worker could recover a result that in fact already finished; see "When things break" above.
  5. The manifest schema. security.acceptsTasks and security.maxConcurrentTasks are still not real fields — see the note by the authorization table. The code already reads them if they show up, so this is a schema change to coordinate with manifest-v0.json's source package, not a code change here.