Research / Paper notes

AERIA

"Dynamic Pricing for On-Demand DNN Inference in the Edge-AI Market." University of Exeter (Songyuan Li, Jia Hu, Geyong Min) + HUST + China University of Petroleum. arXiv 2503.04521v2, in review at IEEE Trans. Mobile Computing. PDF. Code: github.com/songyuanli/AERIA.

A different genre from the first two: auction theory and mechanism design. It's the gap flagged in PolyLink's pricing — an undefined "market adjustment" coefficient δ — filled in with an entire formal mechanism.

What it proposes

AERIA (Auction-based Edge infeRence prIcing mechAnism). Three actors:

The technical piece underneath is a multi-exit DNN: intermediate classifier heads: if confidence at an early exit clears threshold σ_i, inference stops there. Shallow layers run on the user's device; deep layers offload to the edge.

The mechanism

Phase I — each user computes their minimum demand locally, in parallel, O(g):

ā_i,t = F^edge_i(s̄) / (t_i − T^dev_i − T^net_i)

The minimum edge FLOPS needed so the part that doesn't run locally finishes in the time budget remaining. Users are incentivized to ask for the minimum, because competition is by bidding density d_i = β_i / ā_i,t — budget per unit of resource.

Phase II — the COCO auction, three steps:

  1. ROSO — omniscient single-price auction. Orders by density, fills capacity, computes Π̃ = max_i d_[i]·Σ_{k≤i} ā_[k]. The theoretical ceiling on revenue. Not truthful (needs perfect information); used only as a benchmark.
  2. CENTRE — randomized consensus estimate: f(ϖ) = y^(⌊log_y ϖ − ε⌋ + ε), ε uniform on [0,1], δ = Φ̃/(Φ̃−ζ). The key property: the resulting revenue target is independent of the bids — lying about β_i can't move it. That's where incentive compatibility comes from. Guaranteed competitive ratio: (1/ln y)(1/δ − 1/y).
  3. REAL — Moulin-Shenker cost sharing. Iteratively drops bidders who fall below the provisional price p' = Π̂/Φ̂, recomputing, converging on a uniform price p_t = Π̂/Φ̂ for everyone who wins. That's where envy-freeness comes from.

Three proven properties: revenue-competitiveness, truthfulness, and fairness.

The numbers

Full mechanism runtime~3 ms, flat at any scale (i9-10920X)
ComplexityO(N log N + C₂ + Ñ log Ñ)
Revenue advantage vs. baselines~60%
Simulated scale1,500 users (Shanghai Telecom trace), 39–112 bids per slot
Billing slot τ1 hour (Azure's minimum)
Edge cluster3,930 GFLOPS ≈ three i7-13700K, 450 W
Real tracesOntario hourly electricity prices (Sep 2022); Ghent 4G/LTE, 7–60 Mbps
Multi-exit accuracy loss vs. full modelAlexNet −1.97%, ResNet34 −2.47%, VGG16 −3.77%
ME-BART vs. full BART+1.10% (an improvement)
Samples exiting early (BART)18%–48%, depending on σ

Conclusions

1. The 3 ms figure is the number that changes the design. An auction over ~100 bids clears in 3 milliseconds of CPU. That means market clearing isn't the bottleneck — it fits inside an enclave without a second thought.

That's the pattern that answers PyrusLLM's own open question about a per-request enclave round trip — the hypothesis of "batching matches, precomputing assignments." This paper is that hypothesis, formalized: don't match per request, do clearing per window inside the enclave. One enclave invocation per window (they use 1 hour; PyrusLLM could use 1–10 seconds), N requests matched at once. Enclave overhead amortizes across every bid in the window instead of being paid N times over. That turns a stated risk into a problem with a published mechanism, proven properties, and open-source code.

2. This paper describes exactly why the market needs the enclave. Look at what the auctioneer has to see to function: every user's budget, latency requirement, accuracy requirement, and which model they want. That's a full participant profile. Not a side effect — it's the algorithm's required input.

This is the most concrete validation of the underlying thesis, and the third independent one: PolyLink leaves matching on a cloud server, DeServe admits matching "remains centralized," and this paper details exactly what sensitive data that role necessarily accumulates. All three pieces of published state of the art leave open the hole PyrusLLM closes.

3. The auction prices patience, on its own. Since d_i = β_i/ā_i,t and ā_i,t = F^edge/(t_i − T^dev − T^net), a more relaxed t_i shrinks the denominator → fewer resources requested → higher bidding density → wins more often, more cheaply.

Patient users win — exactly the online/batch split DeServe argues from a systems angle, except here it emerges from the pricing mechanism instead of being hardcoded. No need for two products with two fixed tariffs: the batch lane forms itself, and the market sets the discount instead of a hardcoded "−50% for 24h."

4. The same mechanism has a hole on the other axis. The same trick works on σ_i, and there it's a bug: declaring a looser accuracy requirement than you actually want raises your bidding density and makes you win more — and then you get a worse answer. Incentive compatibility is only proven over β_i (budget). t_i and σ_i are left uncovered, and Phase I literally tells the user to minimize ā_i,t to be more competitive. If PyrusLLM adopts this, that's the first manipulation to expect.

5. The uniform price is the most importable design decision. A single p_t for everyone, non-discriminatory. Formally it buys envy-freeness; practically it buys something more valuable: it's explainable. "Everyone paid the same this window, and this was the price" is a sentence a user understands and can audit. Against that, PolyLink has an undefined δ and DeServe uses a fixed Together.ai reference price.

6. The floating reserve price fits what's already known from DeServe. p^res_t = (1+γ)·p^r_t / f_e, tied to the hourly electricity price. From the DeServe analysis: the real marginal cost for an idle-hardware provider is electricity. So: a floor per provider, floating with their local electricity tariff, guaranteeing a minimum margin γ. Neither of the other two papers has a reserve price; without one, a provider can end up serving at a loss during a tariff spike. A small, necessary feature.

7. The "overthinking" finding is the one worth the most money. ME-BART came out 1.10% more accurate than full BART — not less, more. The explanation: running every layer on easy samples overfits the inference. Between 18% and 48% of samples finished early with acceptable quality.

Translated to a business model: for a third to half of requests, running the big model is spending more and sometimes answering worse.

Critiques

It's on the wrong side of the market. The objective function (9) is max Π(U_t) — the service provider's revenue. The auctioneer is a seller extracting maximum surplus from users. A pitch built on "price is set by an open market instead of three corporations, the margin stays with whoever contributes the actual compute" is contradicted by importing AERIA as-is — that makes the platform the extractor.

But there's a useful escape: the mechanism is objective-agnostic. Consensus estimate + cost sharing work identically maximizing social welfare, or the compute provider's revenue instead of the intermediary's. And here's the interesting part: if the auction runs in the enclave with published code, anyone can verify what it's actually maximizing. An auction whose objective is auditable is a different kind of economic object than one a company merely promises to run fairly.

It's one-sided, with homogeneous supply. There is one edge cluster with capacity f_e. Only users bid. PyrusLLM has many providers with different GPUs, different costs, different models and different reputations — that's a double auction, and by Myerson-Satterthwaite you cannot simultaneously have truthfulness, budget balance, and efficiency. AERIA sidesteps the whole problem. It's half an answer, not a drop-in.

The cost model doesn't transfer to LLMs. This is the most important critique. All latency is modeled as T = F/f — FLOPs divided by FLOPS. That assumes a single forward pass. An autoregressive LLM is r sequential passes, with a KV cache, bound by memory bandwidth, not FLOPS. Its natural resource unit is "FLOPS allocated" and its price is "$/unit of FLOPS"; for LLMs the natural unit is tokens (and GPU memory for the KV cache).

And the "LLM" in the paper is actually BART doing German→English translation on Multi30K — a 2019 encoder-decoder with a few hundred million parameters. Calling it "a representative large language model" is a stretch. Everything is simulated, no GPUs — the cluster is three CPUs.

Take the mechanism, not the units. No verification, no trust, no privacy, no reputation, no blockchain, zero real deployment.

What to take to PyrusLLM

1. Windowed clearing inside the enclave. The concrete pattern: accumulate bids over a short window, run a uniform-price auction inside the enclave, return assignments. One round trip per window, not per request. This is the answer to the open enclave-latency question, with a formal mechanism, proven properties, and a 3 ms cost — it stops being an open question and becomes a plan.

2. Uniform price per window + reserve price per provider. A single explainable, auditable, envy-free price per window; each provider declares its floor, tied to its electricity tariff. Cheap to implement, and neither PolyLink nor DeServe has either piece.

3. Cascade routing — the most profitable idea in the paper. The LLM equivalent of multi-exit isn't early-exit layers, it's model cascading: a small model on a cheap node first, a confidence check, escalate to the big model only if needed. The data supports it (18–48% resolved early, and sometimes better).

Cross this with DeServe's number: 445 tok/s against a 1,139 tok/s break-even at rental prices. The cheapest way to serve a 70B model is to not run it for a third of requests. A well-tuned cascade moves the break-even point more than any kernel optimization — and it's product work, not research.

4. If adopting the auction, close the σ axis. Charge for what the user asked for, but measure what they received — reconnecting with PolyLink's TIQE: the in-enclave cross-encoder can verify whether the cascade's output actually met the declared σ. A user who loosens their requirement to win cheap and then complains needs to show up in the data.

5. Decide explicitly what the auction maximizes, and publish it. A product decision, not a technical one — it defines what kind of company PyrusLLM is. With the enclave, this is the first time that declaration can be verifiable instead of a promise.

Where the map stands

PolyLinkDeServeAERIA
Throughput / systemsweakstrongn/a (simulated, CPU)
Provider economicsinverted incentivereal cost tablefloating reserve + margin γ
Price discoveryundefined δfixed reference priceproven auction
Quality / reputationTIQE, measuredabsentσ as a requirement, unverified
Verificationcommittee, 30% taxoptimistic, ~0% taxabsent
Privacynonenonenone — and defines the profile the auctioneer accumulates
Decentralized matchingnono (admitted)no (it's the auctioneer, by design)
Applies to autoregressive LLMsyesyesno — single-pass cost model