DeServe
"DeServe: Towards Affordable Offline LLM Inference via Decentralization." UC Berkeley — Dawn Song's group. arXiv 2501.14784v1, cs.DC, 4 Jan 2025. PDF. A systems paper, not a crypto paper — read in full, and the single most useful paper of the four for PyrusLLM's architecture decisions.
The thesis, and it's a strategic pivot
DeServe starts from a concession PolyLink never makes: a decentralized network cannot win on latency. Full stop. So don't compete there — compete on offline serving (batch inference), where the metric is throughput and the user accepts waiting. The precedent they cite is OpenAI's batch API: a 50% discount in exchange for up to 24 hours of wait.
The entire paper follows from that one decision, and it's arguably the single most valuable idea across all four papers.
The economics (Table 2) — the hardest number in the set
8 equivalent GPUs, reference price $0.90/1M tokens (Llama-70B via Together.ai):
| Compute source | Cost ($/h) | Throughput to break even |
|---|---|---|
| GCP 8× L4 | 13.88 | 4,283 tok/s |
| RunPod 8×4090 | 5.52 | 1,704 tok/s |
| io.net 8×4090 | 3.69 | 1,139 tok/s |
| Mining-rate 8×4090 | 0.35 | 108 tok/s |
A 40x gap between cloud pricing and hardware valued at mining opportunity cost.
Now cross that against what their own system achieves: 445 tok/s. Run the numbers:
- Against mining rate (108) → 4x margin. Closes.
- Against io.net (1,139) → 445 < 1,139. Does not close.
- Against cloud → nowhere close.
A conclusion the paper doesn't spell out, and matters a lot here: a decentralized 70B pipeline is only profitable if the hardware owner values their time at mining opportunity cost (~$0.044/GPU-hour), not at rental price. Follow the math: a 4090 costs ~$1,800; at $0.044/h that's a ~41,000-hour payback ≈ 4.7 years. In other words: buying hardware to serve inference on this network never pays back.
That corrects the standard supply-side segmentation ("crypto node operators, server owners, and investors with the capacity to acquire hardware"). The math says the "investor who buys hardware" segment doesn't close — the one that closes is "already has the hardware, sitting idle," where the marginal cost is electricity. Those are two different pitches, and only one of them is true.
The technical core
1. Inter-layer (pipeline) parallelism, not tensor parallelism. vLLM with TP failed outright between us-east and us-west — didn't even produce a number. TP exchanges tensors on every layer; over a WAN, that's dead. Pipeline only passes the activation between stages.
2. KV cache offloading. KV cache is offloaded to CPU RAM using two global page pools (G0/G1) in double buffering, overlapping the swap with compute over full-duplex PCIe. Memory per microbatch:
M'_B = (M_KV − 2·M_G)/N_B + M_G, with M_G = W × T_S
The important part is the second term: it guarantees a memory floor per microbatch, no matter how many microbatches you pack in.
3. Microbatch scheduling. Network latency injects bubbles into the pipeline. They're filled by adding more microbatches: if latency is ½·T_S across 4 stages, add 2 more microbatches.
The two techniques need each other. Without offloading, doubling
microbatches halves each one's memory and the optimization cancels out. With the
M_G floor in place, it scales.
Results (Table 4, output tok/s)
| centralized <1ms | east–west 58.4ms | sim. 64ms | sim. 256ms | |
|---|---|---|---|---|
| vLLM (tensor parallel) | 253.0 | failed | / | / |
| vLLM (pipeline) | 89.1 | 37.3 | 36.1 | / |
| DeServe (pipeline) | 194.6 | 138.4 | 133.7 | / |
| DeServe (optimized) | 445.2 | 434.1 | 456.8 | 442.9 |
The remarkable part isn't the 6.7x–12.6x headline from the abstract. It's that the last row is flat: 445 centralized → 434 at 58 ms → 443 at 256 ms. Throughput independent of latency. With enough microbatches in flight, WAN latency disappears entirely — and it beats centralized vLLM with TP outright (445 vs 253).
The price of that is in Table 3: at batch 1, 66.6 ms per instance; at batch 256, 0.537 ms. Throughput comes from massive batches → hundreds of requests need to be in flight simultaneously. No single user gets a fast answer. Which is exactly why this only works offline.
A correction to the prior sharding conclusion
The PolyLink notes say "don't build sharding, the evidence is damning." Needs refining: naive sharding is dead; well-pipelined sharding works, but only in batch mode. PolyLink does sequential sharding with no batching and gets 7 tok/s on a 14B. DeServe does deep pipelining with KV offload and gets 445 tok/s on a 70B, cross-coast. The difference isn't the network — it's queue engineering. What no technique recovers is per-request latency.
Section 6 — verification, and it beats TIQE
Three design principles, stated explicitly, that work as hard constraints:
- Verification must not load heavy compute onto the provider. Rules out opML and spML (need deterministic floating-point arithmetic → slow down inference) and zkLLM (too expensive).
- Input and output stay off-chain by default. The gas cost of storing results on-chain can exceed the cost of the inference itself. opML fails here.
- The provider must not be forced to answer challenges from anyone → a DoS vector. opML fails here too.
Their solution: optimistic. The miner signs the result; the honest path is 100% off-chain and costs the provider only a signature; the user can send the output to third parties for verification if they want; if fraud is detected, they invoke arbitration and take the stake. The arbitration module is explicitly pluggable.
Compare the cost to PolyLink: PolyLink pays a validator committee on every batch — β = 0.3, i.e. 30% of all revenue leaves the provider's pocket permanently. DeServe pays zero on the happy path. A pitch built on "the margin stays with whoever contributes the actual compute" is directly contradicted by a permanent 30% validator tax.
What they admit (§6.3, Open Problems)
Verbatim: contract deployment and resource-to-task matching remain centralized. Fair task distribution is unexplored. And correctness-verification criteria are not well-defined.
Berkeley and PolyU — two independent groups — leave the same hole: the matching layer keeps an owner. Two independent confirmations of the same architectural gap.
What to take to PyrusLLM
1. A two-lane product. Today's design describes a single, interactive lane. Add the batch lane:
| Online lane | Batch lane | |
|---|---|---|
| SLA | seconds | hours |
| Routing | 1 node with the whole model | multi-node pipeline |
| Metric | TTFT / latency | aggregate tok/s |
| Price | premium | −50% |
| Tolerates intermittent hardware | no | yes |
That last row is the one that matters most: idle hardware is intermittent, and batch is the only mode where that's not a liability.
2. Launch confidential matching in the batch lane first. The open question about a per-request enclave round trip disappears at batch scale: with an hours-long SLA, "300 ms of enclave overhead against a 24-hour budget" is statistical noise. That converts the main technical risk into a roadmap with its first step already solved.
3. Optimistic verification, with the enclave as arbiter. Adopt DeServe's scheme (provider signature, 100% off-chain honest path, stake, arbitration only on dispute) and resolve its expensive part with what's already there: the enclave is the arbiter. No opML-style bisection, no ZK needed — the TEE re-executes or adjudicates. DeServe leaves the arbitration module explicitly pluggable because it has nothing to fill it with. PyrusLLM does.
4. Sign outputs, not just the manifest. Today the node signs its manifest. Extend that to signing every result (hash of prompt + hash of output + model id + nonce). It's the only per-request cost of the optimistic scheme, and it's what makes any later arbitration possible. Cheap now, impossible to retrofit.
5. The three principles as a rejection filter. For any verification mechanism under consideration: does it load compute onto the provider? Does it put data on-chain? Does it let anyone challenge at will? A "yes" to any one disqualifies it. This filter alone eliminates zkLLM, opML, and spML.
6. KV offloading helps even on a single node.
M_G = W × T_S. Consumer GPUs are memory-bound, not compute-bound.
Offloading KV cache to RAM with overlap buys more batch size on a 4090 or an M3 —
applies to the online lane too.
7. The engine already exists: github.com/CoLearn-Dev/deserve,
PyTorch + FlashInfer. A starting point for the batch lane, not a from-scratch build.
Caveat: January 2025 research code, likely unmaintained — verify before depending on
it.
Where the two stand
| PolyLink | DeServe | |
|---|---|---|
| Throughput / systems | naive, 7 tok/s | strong, 445 tok/s |
| Provider economics | inverted-incentive formulas | real cost table |
| Quality / reputation | TIQE, measured | absent |
| Verification | permanent committee, 30% tax | optimistic, ~0% tax |
| Privacy | none | none (linkability is a feature) |
| Decentralized matching | no (central API server) | no (admitted) |
Nearly complementary — and the two blank cells in both columns, privacy and ownerless matching, are, quite literally, PyrusLLM's product.