AERIA
"Dynamic Pricing for On-Demand DNN Inference in the Edge-AI
Market." University of Exeter (Songyuan Li, Jia Hu, Geyong Min) + HUST + China
University of Petroleum. arXiv 2503.04521v2, in review at IEEE Trans. Mobile
Computing. PDF. Code:
github.com/songyuanli/AERIA.
A different genre from the first two: auction theory and mechanism design. It's the gap flagged in PolyLink's pricing — an undefined "market adjustment" coefficient δ — filled in with an entire formal mechanism.
What it proposes
AERIA (Auction-based Edge infeRence prIcing mechAnism). Three actors:
- AI users (bidders): submit
R_i = {β_i, t_i, σ_i, w_i}— budget, latency requirement, accuracy requirement, which model. - AI service provider (auctioneer): decides the winners, how to partition the model, and the price.
- Edge infrastructure provider (hardware owner): rents at
p^r_t, which floats with the hourly electricity price.
The technical piece underneath is a multi-exit DNN: intermediate classifier heads: if confidence at an early exit clears threshold σ_i, inference stops there. Shallow layers run on the user's device; deep layers offload to the edge.
The mechanism
Phase I — each user computes their minimum demand locally, in parallel, O(g):
ā_i,t = F^edge_i(s̄) / (t_i − T^dev_i − T^net_i)
The minimum edge FLOPS needed so the part that doesn't run locally finishes in the
time budget remaining. Users are incentivized to ask for the minimum, because
competition is by bidding density d_i = β_i / ā_i,t —
budget per unit of resource.
Phase II — the COCO auction, three steps:
- ROSO — omniscient single-price auction. Orders by density, fills capacity, computes
Π̃ = max_i d_[i]·Σ_{k≤i} ā_[k]. The theoretical ceiling on revenue. Not truthful (needs perfect information); used only as a benchmark. - CENTRE — randomized consensus estimate:
f(ϖ) = y^(⌊log_y ϖ − ε⌋ + ε), ε uniform on [0,1],δ = Φ̃/(Φ̃−ζ). The key property: the resulting revenue target is independent of the bids — lying about β_i can't move it. That's where incentive compatibility comes from. Guaranteed competitive ratio:(1/ln y)(1/δ − 1/y). - REAL — Moulin-Shenker cost sharing. Iteratively drops bidders who fall below the provisional price
p' = Π̂/Φ̂, recomputing, converging on a uniform pricep_t = Π̂/Φ̂for everyone who wins. That's where envy-freeness comes from.
Three proven properties: revenue-competitiveness, truthfulness, and fairness.
The numbers
| Full mechanism runtime | ~3 ms, flat at any scale (i9-10920X) |
| Complexity | O(N log N + C₂ + Ñ log Ñ) |
| Revenue advantage vs. baselines | ~60% |
| Simulated scale | 1,500 users (Shanghai Telecom trace), 39–112 bids per slot |
| Billing slot τ | 1 hour (Azure's minimum) |
| Edge cluster | 3,930 GFLOPS ≈ three i7-13700K, 450 W |
| Real traces | Ontario hourly electricity prices (Sep 2022); Ghent 4G/LTE, 7–60 Mbps |
| Multi-exit accuracy loss vs. full model | AlexNet −1.97%, ResNet34 −2.47%, VGG16 −3.77% |
| ME-BART vs. full BART | +1.10% (an improvement) |
| Samples exiting early (BART) | 18%–48%, depending on σ |
Conclusions
1. The 3 ms figure is the number that changes the design. An auction over ~100 bids clears in 3 milliseconds of CPU. That means market clearing isn't the bottleneck — it fits inside an enclave without a second thought.
That's the pattern that answers PyrusLLM's own open question about a per-request enclave round trip — the hypothesis of "batching matches, precomputing assignments." This paper is that hypothesis, formalized: don't match per request, do clearing per window inside the enclave. One enclave invocation per window (they use 1 hour; PyrusLLM could use 1–10 seconds), N requests matched at once. Enclave overhead amortizes across every bid in the window instead of being paid N times over. That turns a stated risk into a problem with a published mechanism, proven properties, and open-source code.
2. This paper describes exactly why the market needs the enclave. Look at what the auctioneer has to see to function: every user's budget, latency requirement, accuracy requirement, and which model they want. That's a full participant profile. Not a side effect — it's the algorithm's required input.
This is the most concrete validation of the underlying thesis, and the third independent one: PolyLink leaves matching on a cloud server, DeServe admits matching "remains centralized," and this paper details exactly what sensitive data that role necessarily accumulates. All three pieces of published state of the art leave open the hole PyrusLLM closes.
3. The auction prices patience, on its own. Since
d_i = β_i/ā_i,t and ā_i,t = F^edge/(t_i − T^dev − T^net), a
more relaxed t_i shrinks the denominator → fewer resources requested →
higher bidding density → wins more often, more cheaply.
Patient users win — exactly the online/batch split DeServe argues from a systems angle, except here it emerges from the pricing mechanism instead of being hardcoded. No need for two products with two fixed tariffs: the batch lane forms itself, and the market sets the discount instead of a hardcoded "−50% for 24h."
4. The same mechanism has a hole on the other axis. The same
trick works on σ_i, and there it's a bug: declaring a looser
accuracy requirement than you actually want raises your bidding density and makes you
win more — and then you get a worse answer. Incentive compatibility is only
proven over β_i (budget). t_i and σ_i are left uncovered,
and Phase I literally tells the user to minimize ā_i,t to be more
competitive. If PyrusLLM adopts this, that's the first manipulation to expect.
5. The uniform price is the most importable design decision. A
single p_t for everyone, non-discriminatory. Formally it buys
envy-freeness; practically it buys something more valuable: it's
explainable. "Everyone paid the same this window, and this was the
price" is a sentence a user understands and can audit. Against that, PolyLink has an
undefined δ and DeServe uses a fixed Together.ai reference price.
6. The floating reserve price fits what's already known from DeServe.
p^res_t = (1+γ)·p^r_t / f_e, tied to the hourly electricity price. From
the DeServe analysis: the real marginal cost for an idle-hardware provider
is electricity. So: a floor per provider, floating with their local
electricity tariff, guaranteeing a minimum margin γ. Neither of the other two papers
has a reserve price; without one, a provider can end up serving at a loss during a
tariff spike. A small, necessary feature.
7. The "overthinking" finding is the one worth the most money. ME-BART came out 1.10% more accurate than full BART — not less, more. The explanation: running every layer on easy samples overfits the inference. Between 18% and 48% of samples finished early with acceptable quality.
Translated to a business model: for a third to half of requests, running the big model is spending more and sometimes answering worse.
Critiques
It's on the wrong side of the market. The objective function (9)
is max Π(U_t) — the service provider's revenue. The auctioneer is a
seller extracting maximum surplus from users. A pitch built on "price is set by an
open market instead of three corporations, the margin stays with whoever contributes
the actual compute" is contradicted by importing AERIA as-is — that makes the
platform the extractor.
But there's a useful escape: the mechanism is objective-agnostic. Consensus estimate + cost sharing work identically maximizing social welfare, or the compute provider's revenue instead of the intermediary's. And here's the interesting part: if the auction runs in the enclave with published code, anyone can verify what it's actually maximizing. An auction whose objective is auditable is a different kind of economic object than one a company merely promises to run fairly.
It's one-sided, with homogeneous supply. There is one
edge cluster with capacity f_e. Only users bid. PyrusLLM has many
providers with different GPUs, different costs, different models and different
reputations — that's a double auction, and by
Myerson-Satterthwaite you cannot simultaneously have truthfulness, budget balance,
and efficiency. AERIA sidesteps the whole problem. It's half an answer, not a
drop-in.
The cost model doesn't transfer to LLMs. This is the most
important critique. All latency is modeled as T = F/f — FLOPs divided by
FLOPS. That assumes a single forward pass. An autoregressive LLM is
r sequential passes, with a KV cache, bound by memory bandwidth, not
FLOPS. Its natural resource unit is "FLOPS allocated" and its price is
"$/unit of FLOPS"; for LLMs the natural unit is tokens (and GPU
memory for the KV cache).
And the "LLM" in the paper is actually BART doing German→English translation on Multi30K — a 2019 encoder-decoder with a few hundred million parameters. Calling it "a representative large language model" is a stretch. Everything is simulated, no GPUs — the cluster is three CPUs.
Take the mechanism, not the units. No verification, no trust, no privacy, no reputation, no blockchain, zero real deployment.
What to take to PyrusLLM
1. Windowed clearing inside the enclave. The concrete pattern: accumulate bids over a short window, run a uniform-price auction inside the enclave, return assignments. One round trip per window, not per request. This is the answer to the open enclave-latency question, with a formal mechanism, proven properties, and a 3 ms cost — it stops being an open question and becomes a plan.
2. Uniform price per window + reserve price per provider. A single explainable, auditable, envy-free price per window; each provider declares its floor, tied to its electricity tariff. Cheap to implement, and neither PolyLink nor DeServe has either piece.
3. Cascade routing — the most profitable idea in the paper. The LLM equivalent of multi-exit isn't early-exit layers, it's model cascading: a small model on a cheap node first, a confidence check, escalate to the big model only if needed. The data supports it (18–48% resolved early, and sometimes better).
Cross this with DeServe's number: 445 tok/s against a 1,139 tok/s break-even at rental prices. The cheapest way to serve a 70B model is to not run it for a third of requests. A well-tuned cascade moves the break-even point more than any kernel optimization — and it's product work, not research.
4. If adopting the auction, close the σ axis. Charge for what the user asked for, but measure what they received — reconnecting with PolyLink's TIQE: the in-enclave cross-encoder can verify whether the cascade's output actually met the declared σ. A user who loosens their requirement to win cheap and then complains needs to show up in the data.
5. Decide explicitly what the auction maximizes, and publish it. A product decision, not a technical one — it defines what kind of company PyrusLLM is. With the enclave, this is the first time that declaration can be verifiable instead of a promise.
Where the map stands
| PolyLink | DeServe | AERIA | |
|---|---|---|---|
| Throughput / systems | weak | strong | n/a (simulated, CPU) |
| Provider economics | inverted incentive | real cost table | floating reserve + margin γ |
| Price discovery | undefined δ | fixed reference price | proven auction |
| Quality / reputation | TIQE, measured | absent | σ as a requirement, unverified |
| Verification | committee, 30% tax | optimistic, ~0% tax | absent |
| Privacy | none | none | none — and defines the profile the auctioneer accumulates |
| Decentralized matching | no | no (admitted) | no (it's the auctioneer, by design) |
| Applies to autoregressive LLMs | yes | yes | no — single-pass cost model |