Research / Synthesis

What to build first

The synthesis across PolyLink, DeServe, AERIA and DCBM: one build item above everything else, three decisions that block the rest, and a short list of things not to build.

The synthesis, in one paragraph

Four independent groups, four disciplines, none citing each other — and the three that have to match a request to a machine all leave that role centralized, two admitting it in writing. AERIA is the most revealing: it doesn't leave the central coordinator in by accident — the auctioneer is the design, and the paper details exactly what it needs to know to function: budget, deadline, accuracy requirement, and which model each user wants. Not an accidental leak — it's the algorithm's required input. It can't be encrypted, because the algorithm has to read it. It can only be moved somewhere nobody watching can see it.

Build first: the signed inference receipt

No hesitation on this one. Every provider signs every response: pubkey, model_id from the manifest, hash(prompt), hash(output), tokens in/out, quantization, timestamp, nonce.

Three reasons, in order:

  1. It's the atom the other five things depend on. Billing (the tokens), verification (the hashes), reputation (the per-provider series), arbitration (non-repudiation), auction clearing (knowing what was actually delivered). DeServe calls it the only per-request cost of its entire optimistic scheme: one signature.
  2. It costs almost nothing. One Ed25519 signature per request. No enclave, no blockchain, no token needed to start.
  3. It's the one piece that can't be added later. Requests that already happened can't be retroactively signed. Every day without this is unauditable history that doesn't come back.

After that, in order: measurement shaped the way it will live inside the enclave → model-identity probes (canaries) → a per-provider reserve price → windowed clearing inside the enclave → cascade routing. The first four are inputs to the fifth — doing them now means the day confidential matching ships, the enclave already has everything it needs.

One callout worth stating explicitly: canaries only work well in an architecture like this one. PolyLink can't inject them undetectably because its API server is visible and the worker knows exactly where every request comes from. When the enclave does the matching and decouples identity, a provider cannot tell a probe from a real user — the privacy layer is what makes verification possible. Not a happy coincidence — a design argument worth writing down explicitly.

What has to be decided

Three block everything else:

DecisionRecommendation
D1 What's the operating economy denominated in? Separate the payment rail (stablecoin) from the incentive asset (token). The entire DCBM apparatus exists to solve a problem created by denominating the real economy in a volatile asset — and the paper never considers the simpler alternative. With this split, a PID controller is never needed just to keep the lights on. See DCBM → the deeper objection.
D2 Who is the target provider? Already-idle hardware, not investors buying hardware to join. DeServe's math: buying a GPU for this never pays back (~4.7-year break-even at mining-opportunity-cost pricing). See DeServe → the economics.
D3 Optimistic verification, or a committee? Optimistic, with the enclave as arbiter. PolyLink permanently reserves 30% of all revenue for validators, directly contradicting "the margin stays with whoever contributes the compute." DeServe leaves arbitration pluggable because it has nothing to fill it with. An enclave-based design does. See DeServe → verification.

Four more are design-level, not blocking:

What not to build

❌
Naive cross-device sharding

7 tok/s on a 14B in PolyLink's own numbers. Route to whoever holds the whole model instead.

❌
A permanent validator committee

PolyLink's 30% permanent tax and inverted-quality incentive. Use optimistic verification with the enclave as arbiter.

❌
PID stabilization of the token

DCBM's own Theorem IV.1 forbids it at pre-seed liquidity depth. A later-phase problem, not a now problem — and denominating the operating economy in stablecoin avoids needing it at all.

❌
ZK proofs of inference

zkLLM: 803 seconds per proof in the numbers PolyLink itself cites. Fails DeServe's own rejection criteria — see below.

❌
A serving engine from scratch

DeServe's pipeline + KV-offload engine already exists (github.com/CoLearn-Dev/deserve) as a starting point for the batch lane. Verify it's still usable before relying on it — it's January 2025 research code.

Adopt DeServe's three rejection criteria as a permanent filter for any verification mechanism under consideration: does it load heavy compute onto the provider? Does it put data on-chain? Does it let anyone challenge at will? A "yes" to any one disqualifies it. That filter alone eliminates zkLLM, opML, and spML.