DCBM
"A Control Theoretic Approach to Decentralized AI Economy
Stabilization via Dynamic Buyback-and-Burn Mechanisms." FLock.io + University of
Oxford (+ Manchester, Newcastle). arXiv 2601.09961v1, cs.GT, 15 Jan 2026.
PDF. Contact: hello@flock.io.
This one closes the map from a different angle: it is not about inference at all. The other three discuss how to serve, how to verify, and how to price. This one discusses the monetary policy of the token everything above gets paid in — directly relevant to a provider earning $QVAC.
What it proposes
DCBM (Dynamic-Control Buyback Mechanism): treat the token economy as a dynamical system and regulate it with a PID controller, instead of the static buyback-and-burn rules the rest of the industry uses.
The starting argument is sound: static buyback is procyclical. It buys back a fixed % of revenue, so it buys a lot during a bull market (high revenue) and little during a bear market — precisely when support is needed most. And threshold rules ("buy if price drops X%") are bang-bang control, which induces oscillation.
The mechanism
State: S_k = [P_k, P̄_k, T_k] = token TWAP, target price (EMA),
treasury (in stablecoin). Error: e_k = ln P̄_k − ln P_k. Positive =
undervalued = needs support.
Control law with clamped integration (the integral term only accumulates while the treasury isn't saturated, to prevent windup):
u_k = K_p·e_k + K_i·Σ ê_j·Δt + K_d·(e_k − e_{k−1})/Δt
Sigmoid actuator with a "circuit breaker" γ:
J_k = T_k · γ · tanh(max(0, u_k))
Treasury: T_{k+1} = max(0, T_k + R_acc(k) − J_k − C_ops)
Plant (constant-product AMM, x·y = K). Linearized in
logs: p_{k+1} = p_k + α_k·J_k + ξ_k, with α_k = 2/y_k. The
market is a discrete integrator whose gain is inversely proportional to
liquidity depth.
The two good ideas
1. Asymptotic solvency. Since J_k is a
multiplicative fraction of the treasury (γ·tanh(·) < 1),
in the worst case T_{k+1} = T_k(1−λ_k) with λ_k < 1 →
geometric decay → the treasury trends toward zero but never reaches
it. Bankruptcy is structurally impossible, not a calibration matter.
The trick is one line and worth isolating on its own: spend a fraction, never an amount.
2. "Whale in a Puddle" (Theorem IV.1). Since
α = 2/y_k, as liquidity falls the plant gain tends to infinity. The Jury
stability criteria give: K_p + K_i > 0,
K_d < (2−αK_p)/α, K_i < (4−2αK_p)/α — the stability
region shrinks as the pool dries up. In plain terms: any
price-support mechanism becomes unstable in a thin market.
This theorem is the one that applies most directly right now — and argues against adopting any of this today. More on that below.
The numbers
Table I (6 scenarios × 6 models), high-volatility regime:
| Model | σ_P ↓ | Churn ↓ | Gas ↓ |
|---|---|---|---|
| No Buyback | 0.89 | 28.4% | – |
| Fixed-Rate | 0.75 | 24.2% | ≈21k |
| Threshold | 0.55 | 19.5% | ≈15k |
| RL (PPO) | 0.45 | 11.5% | ≈720k |
| MPC (Oracle) | 0.26 | 7.2% | >40M |
| DCBM | 0.30 | 8.1% | ≈28k |
The column that decides is the last one. MPC beats DCBM on every stability metric and costs over 40 million gas — unexecutable on-chain. RL comes close at 720k, and is a black box. DCBM captures ~90% of MPC's benefit at ~0.1% of the cost. The paper's real contribution isn't "PID is optimal" — it's "PID is the best thing that fits the execution budget."
Countercyclical treasury: in a bear scenario, Fixed-Rate contracts 12.4% while DCBM contracts only 1.8%; in a bull scenario, DCBM accumulates 11.2% instead of spending.
Ablation (Table II) — textbook, and confirms each term does its job: P-only leaves 4.2% permanent error; PI-only removes it (0.1%) but introduces 25.4% overshoot; PD-only is fastest (85 blocks) but doesn't correct drift; full PID gives the best MSE (0.008).
Adversarial (Table III) — maps ML attacks onto economic manipulation (FGSM = flash crash, PGD = sustained manipulation, C&W = minimum-cost saturation):
| Defense | ASR @ε=1% | Treasury drain |
|---|---|---|
| Threshold | 88.4% | 22.1% |
| Threshold @ε=5% (PGD) | 100% | 72.3% — insolvent |
| PID-Standard | 42.1% | 8.5% |
| PID-Standard @ε=5% (PGD) | 94.1% | 48.2% |
| PID-AdvTrain | 12.5% | 3.2% |
| PID-Cert | 4.2% | 1.2% |
Two readings that matter. One: an un-hardened PID is still drainable — 94.1% success rate at a 5% budget. Two, and this is the transferable one: PID-Cert (hard Lipschitz bounds + saturation) beats PID-AdvTrain (learned robustness). Structural constraints beat well-tuned parameters, under adversarial conditions.
Critiques — and here they're serious
1. Section V is written in future tense. Verbatim: "We will use a dual-data approach," "We will implement and simulate five baseline models," "The main experiments will involve running the simulation," "We will use Monte Carlo methods, running each scenario 1,000 times." Section VI then presents results as fait accompli. The methodology never converted from a proposal into a report.
2. It contradicts itself on how many baselines were run. §V.B: "five baseline models." §V.D: "each model (2 baselines, 1 proposed)." §VI.A: "against three baselines." Table I: five. Four different figures, in one paper.
3. No code. PolyLink published a repo. DeServe published a repo. AERIA published a repo. This one publishes nothing — every number comes from a simulator that doesn't exist publicly.
4. The promised "real data integration" never appears. Hugging Face data and Bittensor/Render prices are announced; no figure, table, or result in the results section ever cites a real data point. Everything is synthetic jump-diffusion.
5. Table III is too clean. A perfect monotonic hierarchy across 4 defenses × 3 attacks × 3 budgets, no inversions, no noise, and — unlike Table I — no confidence intervals. Real adversarial sweeps don't come out this tidy.
6. The abstract's headline 66% is measured against doing nothing. "reduced price volatility by approximately 66% compared to the No Buyback baseline": 0.89 → 0.30. Against the realistic baseline (Threshold, 0.55), the improvement is 46%. The abstract leads with the straw-man comparison. (The churn figure, 19.5% → 8.1%, is correctly measured against Threshold.)
7. Undisclosed conflict of interest. The first affiliation is FLock.io, a decentralized-AI company with its own token, and the paper's contact is its corporate address. A paper concluding that decentralized-AI tokens need an active buyback-and-burn stabilizer, signed by an issuer of one, has a stake in that conclusion. No competing-interests statement.
8. Nobody mentions the regulatory elephant. A treasury systematically buying back its own token to support the price, funded by protocol revenue, is — under nearly any securities framework — exactly the kind of activity that draws scrutiny. The paper brushes the topic once ("disclosure and market integrity") and moves on.
The deeper objection: the problem is self-inflicted
The entire motivation is: providers need stable income and users need stable cost, therefore stabilize the token.
There's a simpler answer the paper never considers: denominate in stablecoin.
If the provider is paid in USDC and the user pays in USDC, native-token volatility is irrelevant to the operating economy. The entire DCBM apparatus exists to solve a problem created by the earlier decision to denominate the real economy in a volatile asset.
PyrusLLM's own design already earns $QVAC per served request via an x402 payment — and x402 is a rail built for stablecoin payments. This paper forces an explicit design question: what's denominated in what.
| Layer | Denomination | Why |
|---|---|---|
| Payment rail (what users pay, what providers are paid) | stablecoin | the provider plans capex/opex against their electricity bill, not the market |
| Incentive asset (staking, reputation bonds, slashing, governance, bootstrap subsidy) | native token | volatility is tolerable here, and the upside is the incentive |
Split this way, volatility hits the speculative layer, not the operational one, and a PID controller is never needed just to keep the lights on. This is the opposite conclusion from the paper's, and it looks like the correct one here.
And even a token with a stabilization layer couldn't work today anyway:
Theorem IV.1 rules it out. Pre-seed, without liquidity,
α = 2/y_k is huge → any controller lands in the unstable region. This
mechanism is a later-phase problem, not a now problem. Knowing that in advance avoids
building it too early — which is the mistake the paper's framing invites.
What to take anyway
1. "Spend a fraction, never an amount" — apply this to any treasury-funded program, not just tokens. A two-sided bootstrap subsidy (early providers need income before there's demand, demand needs providers before there's a reason to add hardware) should be a fraction of the treasury, never a fixed amount per provider. Consequences: the program cannot go broke (a structural guarantee, not a calibration), it self-throttles as reserves fall, and self-accelerates when they're full. Applicable today — no token, no blockchain, no PID. The single most useful idea in the paper.
2. A C_ops term in the conservation equation. Their
correction to the "Treasury Paradox": models that assume the treasury grows on every
inflow are wrong — fixed operating costs have to be subtracted. Any treasury model
needs a C_ops term and a time-to-ruin figure — obvious once stated,
systematically omitted elsewhere.
3. TWAP, never spot, for anything that moves money. A low-pass filter attenuates single-block manipulation by 1/N. This generalizes beyond price to any metric that conditions a payment, including reputation and quality scores — if a reputation score reads an instantaneous value it's manipulable; time-weight it. Connects back to the PolyLink flaw where validators earn more when quality falls.
4. Structural bounds, not calibrated parameters. PID-Cert > PID-AdvTrain > PID-Standard, and even PID-Standard drained at 94%. Apply the lesson to every parameter in the system — slashing rates, reputation decay, price limits, reward multipliers. Bound them by construction.
5. Derivative on the measurement, not the error. Avoids "derivative kick," where the loop panics because the target changed rather than the state. Applies to any feedback loop — e.g. a reputation score shouldn't spike because the quality threshold moved.
6. Execution budget beats the theoretically best solution. MPC dominates on every metric and is unusable above 40M gas — the same shape as the enclave decision: the mechanism that fits the confidential-compute budget wins, not the theoretically optimal one. AERIA's 3 ms is this same argument seen from the other side.