Research / Paper notes

PolyLink

"PolyLink: A Blockchain Based Decentralized Edge AI Platform for LLM Inference." PolyU Hong Kong + China Mobile (HK). arXiv 2510.02395v1, cs.CR, 1 Oct 2025. PDF. 8 pages, read in full.

What it is

PolyLink is a decentralized marketplace for LLM inference on heterogeneous edge devices, settled on-chain. It is, in substance, the same product as PyrusLLM, published as an academic system. Three claimed contributions:

  1. Inference on heterogeneous workers, both single-device and cross-device (sharding via EdgeShard).
  2. TIQE (Trustless Inference Quality Evaluation) — verifying inference integrity without a central authority.
  3. An incentive model with three paid roles: worker, validator, and model provider.

Real deployment: 20 devices, 10 workers, across Hong Kong, Guangzhou, Shenzhen and Kanazawa. ERC-20 contracts on Sepolia.

How it works

Architecture. An API server receives requests, batches them by model and routes each batch to the worker serving that model. Validators score afterward. The blockchain stores invoke/inference/score records.

TIQE — the core of the paper. Per epoch:

Incentives. R_worker = α(1−β)·R_batch, R_validators = [β + (1−α)(1−β)]·R_batch. The model provider is paid a Model Usage Fee, charged to the worker.

The numbers that matter

MeasurementResult
Cross-encoder: degradation detectionTP 66%, FP 2%
LLM-as-Judge: detectionTP 98%, FP 12%
Cross-encoder latency10 ms (batch 1) → 250 ms (batch 128)
LLM-judge latency10–30 s per batch
LLM-judge cost~$0.04 per batch of 128 → ~$0.0003/request
14B on RTX 4090 (1 device)17.3 s latency, 55.7 tok/s
14B cross-device (3080Ti + 4060Ti)160 s, 7.1 tok/s
7B on MacBook M3 Pro52.8 s, TTFT 17.6 s, 5% failure rate

Conclusions

1. Cross-device sharding does not work as a product. 160 s and 7 tok/s versus 17 s on a single 4090 — a 9x penalty for splitting the model across devices. The paper admits it in its own limitations section ("network communication latency, so the number of devices is limited"). Takeaway: don't build sharding. Route to whoever holds the whole model. That's an entire engineering front saved, with published empirical evidence to justify skipping it.

2. The "decentralized" paper has a central operator — the same one PyrusLLM's own design identifies as the thing to avoid. The API server that batches and routes runs on a single 4-core cloud instance. It sees who asks and who answers. That's exactly the coordinator whose absence is PyrusLLM's differentiator. Published academic state of the art did not solve this — real, citable competitive positioning.

3. TIQE conflates two different problems. It verifies response quality, not which model actually ran. The judge prompt (Fig. 3) literally says "rate the quality of the following LLM inference result" — never mentions what model was supposed to run. A worker who announces a 14B and quietly serves a 1.5B with decent-sounding answers still scores well. That's why the cross-encoder only catches 66% — it's measuring relevance, not provenance. The methods that actually attack this (TOPLOC, SPEX — LSH over activations) are cited in the paper and then discarded.

4. The validator incentive is inverted. Look at the formula: the validator payout β + (1−α)(1−β) is strictly decreasing in α. At α=1 they earn β=0.3; at α=0 they earn 100% of the batch. And slashing only punishes deviating from the median, not scoring low. So a committee that scores uniformly low — through collusion, or simply shared bias — earns more, and nobody deviates. Figure 7 draws this without seeming to notice. It's the most serious flaw in the design.

5. The user always pays in full. C_inference does not depend on α. If the answer is garbage, the user still pays, and that money goes to the validators. No refund, no discount. Economically unsustainable on the demand side.

6. The most trusted signal is a centralized API. LLM-as-a-Judge, which carries the larger weight λ, is invoked against DeepSeek's commercial API. The "trustless" protocol depends on a single external provider that can go down, reprice, or silently swap models underneath it. A direct contradiction of the decentralization claim.

7. A 12% false-positive rate is expensive when α multiplies pay. 1 in 8 honest batches gets penalized. With R_worker = α(1−β)R_batch that is pure variance injected into an honest provider's income — exactly what breaks the "your hardware is a yield-bearing asset" pitch.

8. A 5% failure rate. The paper presents this as a success ("below 5%"). As a product SLA it's 1 in 20 requests that never completes. No retry, no failover in the design.

Ideas for PyrusLLM

Verifiable quality (the most useful challenge this paper informs)

Separate the two questions instead of conflating them the way PolyLink does:

Canaries are where PyrusLLM's architecture beats this paper. PolyLink can't inject undetectable probes because its API server is visible and the worker knows exactly where every request comes from. In PyrusLLM's design, the enclave does the matching and decouples identity — the provider cannot tell a probe from a real user. The privacy layer is what makes verification work. That's a strong argument, and exactly the kind of synergy worth stating explicitly.

Run the cross-encoder inside the enclave. MiniLM-L6-v2 is ~22M parameters, 10 ms on CPU. If a confidential job is already metering requests, scoring rides along for free. That eliminates the entire validator committee — no stake, no slashing, no median consensus, no VRF, none of the three attack surfaces analyzed above. The enclave is the validator, and its code is public.

Random spot-checks, but genuinely unpredictable. PolyLink derives "a random point in the epoch" from the chain — predictable in principle. An enclave holds secrets and can decide, in a way nobody can game, which requests to audit.

The economics check out. $0.0003 per judged request. Auditing 2–5% of traffic with an LLM-judge costs next to nothing next to the value of the inference itself.

Economic layer

Don't use α as a pay multiplier. That's the paper's costly mistake (12% FP × variable income). Instead:

A bad provider loses traffic, not money per request. Fairer, more stable, and the incentive points the same direction without punishing statistical noise.

Never let a role earn more when quality drops. A design rule that follows directly from flaw #4.

Define what happens to a bad response. PolyLink charges anyway. A discount or credit tied to the score is a real product differentiator and disciplines the supply side.

Model provider fee, and honest framing

The model provider as a third paid role (the Model Usage Fee) is the genuinely novel idea PolyLink claims, and PyrusLLM lacks it today — the economy currently has hardware providers and users, nothing rewarding whoever brings a good fine-tune to the network. Cheap to add now, expensive to retrofit.

Verification is post-hoc, not preventive — TIQE runs 10–30 s after the response, and the user has already received it. That's fine — it's the only viable design — but say so explicitly: quality guarantees are statistical and reputational, not per-request. More credible in a pitch than promising verification it can't actually deliver.

Retry and failover. The 5% failure rate is the metric PyrusLLM's gateway should attack that PolyLink doesn't. Cheap differentiation.

For the pitch

PolyLink is a citable academic baseline with real hardware and published numbers: latency, throughput, cost, detection accuracy. "Published state of the art does X, we do Y" beats comparing against an abstraction. Its related-work table (zkLLM: 803 s per proof; SVIP: needs a TTP; TOPLOC/SPEX: LSH) is a ready-made map of verification options.

PyrusLLM's application to Vela flags an open risk: "a round trip to the enclave per request may not be viable." Two other papers in this set — DeServe and AERIA — turn that from an open question into a concrete plan.