Abstract
UHI turns the devices people already own into one AI inference cloud, routed by SLA and paid per job on its own chain.
Every device is a worker. Every job is a wage.
Inference, not training, is now most of the AI compute bill: about two-thirds of it in 2026 [P1], against a capital build-out of $400–450 billion this year [P2] that is forecast to reach $1.7 trillion a year of data-centre spending by 2030 [P3]. Yet the workloads that most products run every second, speech to text, text to speech, summarisation, translation, embeddings and small-model chat, fit on the neural processor of a phone. There are 6.9 billion smartphones in use [P4], roughly 1.5 billion laptops and tablets [P5], about 1.9 billion people who play games on a PC and 650 million on a console [P6]. Most of that hardware is idle most of the day, and its owners earn nothing from it.
UHI connects the two. Five parties take part. Requesters submit jobs with a service level: latency, redundancy, region, privacy. Device operators run a small app that executes jobs inside a budget they set for battery, heat, hours and data, and are paid per job. Model providers publish small, efficient models with a device spec and an SLA envelope, and earn a royalty on every job that runs their model. Validators check the work with sampled redundant execution, output hashing, canary jobs and hardware attestation, and produce blocks. Routers match jobs to devices and compete on SLA attainment. Every settled job splits its price 70/15/10/3/2 between device, model provider, validators, router and treasury, in one block on a sovereign Substrate chain with two-second blocks and near-zero fees paid by the protocol.
The result is an inference provider with an SLA an enterprise can sign, whose data centre is everyone's pocket. The first customer is us: QuickDial's AgentBox and Vartalaap already run millions of minutes a month of speech and small-model inference on commodity compute, and UHI is the network that workload moves onto first.
1 The problem
1.1 Inference is the bill
The economics of AI have inverted. Training a frontier model is a one-off cost paid by a handful of labs; serving models to users is a cost paid by everyone, every second, forever. Inference now accounts for about two-thirds of AI compute demand [P1], and the physical build-out behind it is the largest capital programme in the history of the technology industry: $400–450 billion of AI-related capex in 2026 [P2], and data-centre capital spending forecast to reach $1.7 trillion a year by 2030 [P3].
That spending buys GPUs in data centres, and GPU capacity is rationed. Every product team that ships a voice agent, a transcription feature, a translation layer or a summariser pays a cloud price that reflects the scarcity of that capacity, not the cost of the arithmetic.
1.2 The work that matters fits on an NPU
The arithmetic is small. The workloads that most products run, by volume of jobs rather than by headline, are:
| Workload | Typical model size | Runs on a 2024 flagship phone |
|---|---|---|
| Text to speech (TTS) | 80 M – 1 B parameters | Yes [P7] |
| Speech to text (STT) | 40 M – 1.5 B parameters | Yes [P7] |
| Summarisation, rewriting, small chat | 1 – 4 B parameters, 4-bit | Yes [P7] |
| Translation | 0.6 – 3 B parameters | Yes [P7] |
| Embeddings | 30 M – 350 M parameters | Yes [P7] |
| Vision: classify, OCR, detect | 5 M – 500 M parameters | Yes [P7] |
Every flagship phone and every new laptop now ships with a neural processing unit of 35–45 TOPS: Apple's M4 at 38 TOPS [P8], Qualcomm's Snapdragon X Elite at 45 TOPS [P9], and Microsoft's Copilot+ PC programme requires 40 TOPS as the floor [P10]. A gaming graphics card is an order of magnitude above that. On our own products, a CPU-only server runs the full speech-in, model, speech-out loop and returns the first syllable of a reply 253–392 ms after the caller stops talking [M1]. A phone NPU is a better inference processor than the CPU we measured on.
Our assessment, which we label an estimate because quality is workload-specific: for TTS, STT, summarisation, translation and embeddings, models of 1–3 billion parameters now match the cloud-served quality that products actually ship with, when run on a 2024 or newer device [E1].
1.3 The supply is already paid for
| Device class | Installed base | Source |
|---|---|---|
| Smartphones | 6.9 billion | [P4] |
| Laptops and tablets | ~1.5 billion | [P5] |
| PC gamers (people, not rigs) | ~1.9 billion | [P6] |
| Console gamers | ~650 million | [P6] |
| AI-capable share of the installed base, end-2026 | under 20%, rising every replacement cycle | [P11] |
A device is in use for a small part of the day. A phone is on a desk or a bedside table for most of its life; a gaming rig is idle outside gaming hours; a console sits in standby. We estimate the idle fraction at about 90% [E2], from typical daily screen-time surveys (3–5 hours a day for phones) and from the simple observation that people sleep. The owners of this hardware have already paid for it, already pay to power it, and already pay to connect it. It earns nothing.
1.4 What is missing
The compute exists. The demand exists. What is missing is the thing that lets an enterprise buy one from the other: a guarantee. An enterprise does not buy "some phones somewhere". It buys a p95 latency, a redundancy level, a region, a privacy mode and a price, and it wants a counterparty who is accountable when those are missed. Distributed compute networks to date sell raw GPU hours to developers. Nobody sells SLA-backed inference from consumer devices to enterprises. That is the product.
Nobody sells SLA-backed inference from consumer devices to enterprises. That is the product.
1.5 The market, sized
All three figures are estimates with the assumption chain stated [E10]; the primary anchors are published figures.
| Measure | Figure | How it is built |
|---|---|---|
| TAM: AI inference compute spend, 2026 | $283.5 billion | two-thirds [P1] of the $400–450 billion AI capex mid-point [P2], as a proxy for what is spent to serve inference |
| SAM: small-model inference bought through an API, US and EU, 2026 | $9.6 billion | the TTS, STT, translation, small-LLM and embeddings API markets summed, × 62% US and EU share, × 70% purchasable through a third-party API rather than bundled inside a platform |
| SAM in 2029 | $17.5 billion | 22% a year growth, in line with edge-AI forecasts (verify) |
| SOM: UHI's year-3 exit ARR | $51 million | the model's month-36 job value × 12, which is 0.29% of the 2029 SAM |
The SAM is deliberately narrow: it is the inference that products already buy by API today, from vendors whose list prices UHI undercuts, not the share of all AI that could in principle run on a phone.
2 The network
UHI is five parties and one loop. Each party is defined by who they are, what they do, and why they show up; a network only works if the last column is true for everyone in it.
2.1 Requesters
Who. Enterprises, developers and consumer apps; retail users indirectly, through the apps they use. At launch: voice-AI, contact-centre and media companies that buy TTS, STT and summarisation by the million minutes. First among them, our own products.
What. Submit jobs, each with a model, a payload and a service level: an SLA tier (section 4), a region, a privacy mode and a price ceiling. Pay in dollars by card or invoice, or in UHI.
Why. Price: 15% of cloud list at the Batch tier, which is 6.7× cheaper; 4.2× at Standard and 2.7× at Realtime [E3, section 12]. Elasticity: the network has no queue because capacity is a function of how many devices are online, not of how many GPUs were bought last quarter. Privacy: a job can be pinned to a continent, to a metro, to the requester's own fleet, or to the end user's own device.
2.2 Device operators
Who. Anyone with a phone, tablet, laptop, desktop, gaming rig, console, watch or glasses. Later: fleets, telcos and OEMs who ship the node app pre-installed.
What. Install the UHI node app. Set a budget: a battery floor below which the device will not work, a thermal cap, hours of the day, a data allowance, and whether to work on cellular. Keep the device online. The app downloads the models the router expects to use on that device class and keeps them warm when the budget allows.
Why. Income from hardware they already own, in a currency they can hold or cash out, paid per job within seconds of the job finishing. A phone that is kept busy at its reference utilisation earns about $2 a day [E4, section 12], which pays its own plan. A gaming rig earns about $15 a day, which pays its electricity and then some. The network admits devices only as fast as demand can keep them busy (section 6.4), so an active device earns at that rate rather than sharing a thin stream with an idle crowd. The app never does anything the operator did not budget for; that rule is the whole relationship.
2.3 Model providers
Who. ML teams who make models small and efficient: quantised, distilled, device-tuned. The quantisation and small-model community that already publishes to open model hubs; speech labs; the teams behind compact TTS, STT and translation models.
What. Publish a model to the models pallet with a device spec (minimum RAM, NPU TOPS, operating systems, runtime) and an SLA envelope (p95 latency and throughput per device class, measured by the network's own benchmark jobs). Keep it maintained.
Why. Distribution to millions of devices without running any infrastructure, and recurring income: a royalty on every job that runs their model, 15% of the job price at the full rate. Loyalty: a provider's royalty rises with sustained SLA attainment and volume (section 7), so the teams whose models actually work on real devices earn more over time. Treasury grants fund the porting work.
2.4 Validators
Who. Staked operators running the chain: at launch a permissioned set of 21 drawn from infrastructure partners and the team, opening to nominated proof-of-stake over the first year.
What. Verify work: sampled redundant execution, output-hash agreement, tolerance checks for non-deterministic outputs, canary jobs with known answers, and attestation where the device hardware offers it. Approve payouts. Produce blocks.
Why. Block rewards and a 10% validation fee on every settled job. Slashed for approving work that is later proven bad.
2.5 Routers
Who. The matching layer. UHI Labs operates the only router at launch; the role opens to permissioned partners at month 18 and to open, staked routers at month 24 (section 6).
What. Score devices by capability, warm model, measured latency, reputation, price and redundancy; dispatch; retry; meet the SLA.
Why. A 3% routing fee. Routers compete on SLA attainment, and the protocol rotates traffic toward the best; a router that misses its tiers loses traffic before it loses anything else.
2.6 The treasury
The treasury is a chain account that receives 2% of every settled job and a share of the initial token allocation. It is spent only by on-chain vote on the Treasury track (section 10): grants to model providers to port and tune models, incentives to onboard devices in a region or a device class the network is short of, and audits of the runtime and the node app.
3 The job loop
3.1 In prose
- Submit. A requester's SDK submits a job: model, SLA tier, region, privacy mode, price ceiling, and the payload. The payload is encrypted before it leaves the requester (section 3.4). The router records a commitment to the job on-chain: a hash of the request, the tier, and an escrow of the price ceiling from the requester's balance.
- Match. The router picks devices that already hold the model warm, inside the tier's latency budget, with the reputation the tier requires, in the right region, and dispatches to
nof them, wherenis the tier's redundancy (1, 2 or 3). - Execute. Each device decrypts the payload in-process, runs the model on its NPU, GPU or CPU, and produces the result plus a proof: the output hash, signed start and end timestamps, and a hardware attestation if the platform provides one.
- Deliver. The first device to finish sends the encrypted result straight to the requester. The requester has its answer at this point, before anything is settled.
- Verify. Validators compare the proofs from the
ndevices: hashes agree (or fall inside the tolerance band for non-deterministic workloads), canaries pass, timings sit inside the tier. Disagreements are resolved by a tie-break execution. - Settle. The
settlementpallet closes the job in one block: the escrow is released and the price splits 70/15/10/3/2 to device operators, model provider, validators, router and treasury. The device operator sees the payout in the app two to four seconds after the job finished.
- Device operators 70%
- Model provider 15%
- Validators 10%
- Router 3%
- Treasury 2%
3.2 As a sequence
- 1Requester → Routerjob: model · tier · region · privacy · ceiling · encrypted payload
- 2Router → Ledgercommit: hash(job) · tier · escrow(ceiling)
- 3Router → Devicesdispatch to n devices with the model warm
- 4Devices → Requesterencrypted result, direct from the first device to finish
- the requester has its result here, before anything is settled
- 5Devices → Routerproof: output hash · signed timings · attestation
- 6Router → Validatorsbatch of proofs
- 7Validators → Ledgerapprove: hashes agree · canaries pass · timing in tier
- 8Ledger ↻ Ledgersettle in one block: 70 / 15 / 10 / 3 / 2
- 70 / 15 / 10 / 3 / 2 · gold is the wage
- 9Ledger → DevicesUHI credited to the devices
- 10Ledger → Requesterescrow difference refunded
- settlement closes 2 to 4 seconds after the job finished · gold is money, green is the job in flight
3.3 What is on-chain and what is not
| On-chain (the ledger) | Off-chain |
|---|---|
| Job commitment: hash of the request, model id, tier, region, price ceiling, escrow | The payload and the result, encrypted end to end |
| Dispatch record: which devices, which router, when | Model weights (distributed through the node app and a content-addressed cache) |
| Proofs: output hashes, signed timings, attestation digests | The computation itself |
| Approval and settlement: the split, as net balance changes per account per batch | The router's scoring state (its root hash is published each epoch) |
| Reputation updates, SLA attainment per router, device and model | Job records older than the 7-day dispute window (retained by archive nodes, pruned from state) |
Individual jobs are settled in batches. A batch carries a Merkle root over its job records plus net balance deltas per account, so the state the chain keeps per job is a few bytes, not a record. Anyone can audit a job by asking an archive node for the record and checking it against the root.
3.4 Encryption of payloads
A job's payload is never visible to the router or to the validators.
- The requester's SDK generates a per-job symmetric key and encrypts the payload with an authenticated cipher (XChaCha20-Poly1305 at launch).
- The job key is wrapped to the public key of each device the router selects (X25519). The wrapping happens in the SDK, after the router returns its selection and before dispatch, so the router only ever carries ciphertext plus metadata: model id, payload size, tier.
- Each device decrypts in-process, runs the model, and encrypts the result to the requester's public key. The output hash in the proof is over the plaintext result, salted with the job id, so validators can compare hashes across devices without seeing the result.
- For the sampled redundant execution and canary checks in section 5, validators compare hashes; they do not decrypt. Where a tie-break execution is needed, it is dispatched as a fresh job to a validator-selected device with the same wrapping, and the validator still sees only the hash.
Device keys are generated on the device and, where the platform supports it, kept in its secure element (Android StrongBox or TEE-backed Keystore, Apple Secure Enclave, Windows TPM). Rotation is monthly and on any reputation event.
3.5 Result before settlement
The requester is never waiting for the chain. The result arrives first; settlement follows one to two blocks later.
The requester is never waiting for the chain. The result arrives directly from the first device to finish, over a QUIC connection brokered by the router (with a relay for devices behind restrictive NATs). Settlement follows one to two blocks later. If verification later fails, the requester has already received the result, the job is refunded from escrow, and the device that produced the bad result is the one that loses: reputation first, stake if it is a Realtime device. The requester is charged only for verified work.
4 Service levels
An SLA tier is a contract the requester chooses per job and the network commits to as a whole. Three tiers exist at launch.
4.1 The tiers
Prices are indicative and set on the Market track: the Batch price of a job is 15% of the cloud list price for the same unit of work, Standard is 1.6× Batch and Realtime 2.5× Batch, so UHI is 6.7×, 4.2× and 2.7× below cloud list at the three tiers [E3]. Latency is measured from the router receiving the job to the requester receiving the result, so it includes dispatch, compute and delivery. Batch is for backfill: embedding a corpus, transcribing an archive, generating a night's worth of audio. Standard is the default for most product features. Realtime is for conversation: a voice agent's reply, a live caption.
4.2 How attainment is measured
Attainment is the fraction of jobs in a rolling window that met both the tier's latency bound and verification. It is kept on-chain by the sla pallet at three grains: per router, per device, per model, each over 24-hour and 7-day windows.
Timing is evidenced by three signatures: the router signs the moment it received the job, the device signs its start and end of compute, and the requester's SDK signs receipt of the result. Validators compute latency from the router's and the SDK's timestamps; the device's timestamps are used to separate compute time from transit time, so a slow network path is attributed to the router's choice of device, not to the device's work. Clock skew is bounded by the router's measured round-trip to each party and jobs with inconsistent timestamps are flagged, not paid.
4.3 Penalties and retries
- A job that misses its tier is not paid at that tier. A Realtime job that returns within 1 s is paid at the Standard rate; one that returns later is not paid at all and the requester is refunded. A Standard job that misses pays at Batch; a Batch job that misses is retried until it clears or the requester's deadline passes.
- The router retries automatically: a replacement device is dispatched the moment a device's expected finish time passes the budget, without waiting for a timeout.
- If the miss is attributable to routing (the router chose a device outside its own measured latency budget, or with a cold model), the router's 3% on that job is forfeited to the requester's refund. If it is attributable to the device (compute slower than its envelope), the device's reputation takes the hit.
- Repeated misses move a device down a tier automatically: a Realtime device whose 24-hour attainment falls below 0.95 is routed Standard work until it recovers.
4.4 Regional pinning
A job carries a region: any, a continent, a country, or a metro. The router filters by it before scoring. Region is established three ways and the strictest wins: the operator's declaration at install, the device's IP geolocation, and round-trip-time triangulation from the router's probe points, re-measured daily. A device whose measured position disagrees with its declaration is routed as any and its operator is told why.
4.5 Privacy modes
| Mode | What it guarantees |
|---|---|
| Standard | Payload and result encrypted end to end; only the executing devices ever hold plaintext, in memory, for the duration of the job. Nothing is retained. |
| Region-locked | Standard, plus execution only on devices in the named region. |
| Fleet | Standard, plus execution only on devices the requester has enrolled: its own offices, its own employees' laptops, its own kiosks. |
| On-device-only | The job runs on the end user's own device, inside the requester's app, with the network's model and the network's SDK. Nothing leaves the device. The chain settles the model provider's royalty and records the job; there is no payout to a device operator because the user is the operator. |
On-device-only is the mode that lets a healthcare app, a bank or a government use UHI's models with no data movement at all, and it is how model providers earn from on-device apps they could never have billed before.
5 Verification
Devices are untrusted. The verification design assumes a device may be lazy (returning a cheap wrong answer), faulty, or hostile, and makes each of those unprofitable.
5.1 Sampled redundant execution
Standard and Realtime jobs run on 2 and 3 devices respectively; every job in those tiers is its own cross-check. Batch jobs run once, and 5% of them, chosen by a verifiable random function seeded from the block hash so neither the router nor the device can predict which, are re-executed on a second device. A Batch device therefore faces a 1-in-20 chance per job of being checked, and a single failed check costs more than twenty jobs' earnings (section 5.6).
5.2 Output-hash agreement for deterministic workloads
Where the workload is deterministic, the proofs' output hashes must match exactly. Determinism is engineered, not assumed: the model package pins the runtime, the quantisation, greedy decoding with a fixed seed, and a canonical serialisation of the output. Embeddings, classification, OCR and greedy-decoded STT and small-LLM jobs fall in this class on a single device class. Across device classes (an Apple NPU against an NVIDIA GPU) floating-point results can differ in the last bits, so hash agreement is required within a device class and tolerance bands apply across classes.
5.3 Tolerance bands for non-deterministic workloads
TTS, sampled text generation and some vision workloads do not produce identical outputs on two devices. For these the proof carries a compact fingerprint rather than a hash, and validators compare fingerprints against a per-model tolerance band set by the provider and confirmed by the network's benchmark jobs:
- TTS: a perceptual audio hash of the waveform plus a speaker-and-content embedding; agreement means embedding cosine similarity above the model's band (0.90 at launch) and duration within 5%.
- Embeddings across device classes: cosine similarity above 0.995.
- Sampled text: a sentence-embedding distance plus a length check; and for any job where the requester needs exactness, it can request greedy decoding and get hash agreement instead.
- Vision: intersection-over-union for boxes, label agreement for classes.
The fingerprint is computed by the device from the plaintext result and included in the proof, so validators never need the result itself.
5.4 Canary jobs
The validation pallet keeps a pool of jobs with known answers, contributed by model providers at publication and refreshed by validators from verified past jobs. Between 1% and 3% of each device's dispatches are canaries, indistinguishable from real jobs because they are real jobs: same model, same encryption, same price. A failed canary is the strongest evidence the network has, and it is weighted accordingly in reputation.
5.5 Attestation where available
Where the platform offers it, the node app includes a hardware-backed attestation in each proof: Android Play Integrity and Key Attestation, Apple App Attest, Windows TPM-backed attestation, and on desktop GPUs and servers the vendor's confidential-computing attestations. Attestation proves that the proof was produced by an unmodified node app on a genuine device; it does not prove the computation was correct, which is what the checks above are for. Attested devices earn a reputation bonus and are the only devices eligible for Realtime.
5.6 Reputation and decay
Every device, router and model has a reputation in [0, 1] held in the reputation pallet. It rises slowly with verified jobs and falls sharply with failures: a failed canary or a hash disagreement costs roughly the gain of 50 successful jobs [E5, a launch parameter]. Reputation decays toward a prior of 0.5 with a half-life of 30 days of inactivity, so a device cannot bank a score and then go rogue months later, and a new device cannot buy its way to Realtime with a burst of easy Batch jobs: tiers require a minimum count of verified jobs in that tier's conditions as well as a minimum score.
5.7 Slashing
Realtime devices and all validators and routers post stake. Stake is slashed by the validation pallet on proof: 10% for a proven bad output, 100% for a proven Sybil or collusion pattern (section 11). Slashed stake goes half to the treasury and half to the party that proved the fault, which is how independent auditors are paid.
6 Routing
6.1 The scoring function
For a job j and a candidate device d, after hard filters (capability: the model's device spec is met; region; tier eligibility by reputation, stake and attestation), the router computes
- W warm model
- 1 if the model is loaded in memory · 0.3 if cached on disk · 0 if it must be fetched
- L latency fit
- clamp(1 − rttest(d, j) / budget(tier), 0, 1), from the router's rolling measurements
- R reputation
- the device's score in the
reputationpallet, in [0, 1] - P price
- 1 − ask(d) / ceiling(j); a device may set an ask above the floor
- D diversity
- 0 if a device already chosen for this job shares an owner, network or metro with d, else 1
Hard filters run first: capability (the model's device spec is met), region, and tier eligibility by reputation, stake and attestation. The router dispatches to the top n by score, where n is the tier's redundancy.
The router dispatches to the top n by score. The weights are a launch proposal and are set on the Market track. A router may deviate from the published function, but its attainment is what it is judged on, and the function is published so that device operators can see why they do or do not receive work.
6.2 Router competition and traffic rotation
More than one router can register, each with stake. The sla pallet records attainment per router. Requesters who do not pin a router (the default) are assigned routers in proportion to attainment raised to the fourth power over the trailing 7 days, so a router at 0.99 attainment receives about 1.5× the traffic of one at 0.90 and about 16× the traffic of one at 0.50. Below 0.80 a router receives no default traffic until it recovers on pinned traffic. Competition is on attainment; the fee is capped at 3% on the Market track, so routers cannot win on price at the devices' expense.
6.3 Progressive decentralisation
| Phase | Routers | When |
|---|---|---|
| 1 | UHI Labs operates the sole router; the scoring function is published | Testnet to month 12 |
| 2 | Three to five permissioned routers run by partners under the same published function | Month 18 |
| 3 | Open registration with stake and slashing; rotation by attainment | Month 24 |
6.4 Admission is paced to demand
A device only earns if it is busy, so the router does not admit supply faster than demand can keep it busy. Devices register and sit on a waitlist. The admission rule is: active devices are admitted so that realised utilisation stays at or above 50% of reference utilisation, where reference utilisation is the busy share of online hours the earnings figures assume (a phone busy 20% of its 16 online hours, a rig 30% of 18, section 12.1). When the trailing-7-day realised utilisation of active devices in a region and device class is above the target, the router admits the next registrations from the waitlist for that class and region, in order of capability (NPU TOPS, RAM, mains power) and then registration time; when it falls below, admission pauses until demand catches up. Registered devices keep their place, run benchmark and canary jobs so their capability and region are measured before they earn, and are told their position.
The rule is why the network states earnings for a busy device at reference utilisation rather than as a network average, and why the roadmap (section 13) carries two counts: registered, which is the waitlist and the growth engine, and active, which is the number that earns.
The router role is the one part of the network that is centralised at launch, and we say so. It is centralised because the first year is about proving SLA attainment, and attainment is easier to prove with one accountable operator. It is decentralised on a date, not on a sentiment.
The router is centralised at launch, and we say so. It is decentralised on a date, not on a sentiment.
7 Models and providers
7.1 Publication
A model is published to the models pallet as a package: weights by content hash, runtime and quantisation, a device-spec envelope (minimum RAM, NPU TOPS, OS and runtime versions, storage), and an SLA envelope per device class (p95 latency and throughput for a reference input). The provider's envelope is a claim; the network's benchmark jobs, run on real devices of each class at publication and weekly thereafter, are the measurement, and it is the measurement the router uses.
7.2 Royalty basis points
The model provider's share of every job is 15% of the job price at the full rate, which is 1,500 basis points. A newly published model starts at 900 bps and earns its way to 1,500 through the loyalty curve. The difference between a provider's current rate and 1,500 bps accrues to the treasury's provider-grants fund, so the money stays in the provider community.
7.3 The loyalty curve
- V volume
- jobs served by the model in the trailing 90 days
- V* saturation volume
- 5,000,000 jobs at launch, after which volume adds nothing more
- A attainment factor
- 0 at 0.90 attainment, 1 at 0.98 and above: a model at 0.90 earns nothing above the base, a model at 0.98 earns the full curve
Worked example. A speech lab publishes a 300 M-parameter TTS model. In its first 90 days it serves 2.5 million jobs at an attainment of 0.97. Then V/V* = 0.5, A = (0.97 − 0.90)/0.08 = 0.875, L = 0.4375, and the royalty is 900 + 600 × 0.4375 = 1,162 bps, or 11.6% of each job's price. At 10 million jobs a month priced at 0.2¢ each (section 9.6), the model's jobs are worth $20,000 a month, and the provider's royalty is $2,324 a month against $1,800 at the base rate and $3,000 at saturation [E6: both the volume and the price are illustrative]. If attainment drops to 0.92 the next quarter, A falls to 0.25 and the royalty falls to 975 bps, which is the point: the curve pays for models that keep working on real devices.
- attainment 0.98: the full curve
- attainment 0.95
- attainment 0.92
7.4 Grants
The treasury funds porting and tuning. A grant is proposed on the Treasury track with a target: a model, a device class, an SLA envelope to hit, and a benchmark that proves it. Grants are paid in tranches against the benchmark, and the provider council (section 10) has a veto on grants to providers, so that the people who know the work judge the work.
8 The chain
8.1 Why a sovereign L1
Settlement is the product. A job that pays 0.2¢ cannot carry a 2¢ gas fee, and a network settling thousands of jobs a second cannot wait for someone else's block space, or re-price itself every time an unrelated application congests a shared chain. UHI runs on its own Substrate-based chain so that fees, block time, throughput and upgrade policy are set by the people the network pays. Bridges give that chain liquidity later; settlement never leaves it.
8.2 Block production and finality
Two-second blocks, with deterministic finality from GRANDPA on top of slot-based block production. At launch 40 validators run a permissioned set, each with 500,000 UHI at stake; the set grows to 100 under nominated proof-of-stake over the following two years. A job is settled one to two blocks after approval, which is why "settles in seconds" is a description and not a slogan.
8.3 Fees
Substrate charges weight-based fees: a transaction pays for the compute and storage it actually uses. UHI sets the weight fee for job settlement near zero and has the protocol pay it: settlement batches are submitted by validators, whose 10% share of every job is what funds the work. Requesters, device operators and model providers never see a transaction fee on a job. Ordinary user transactions (transfers, staking, votes) pay a small weight-based fee as on any Substrate chain, and that fee is burned.
8.4 The six pallets
| Pallet | What it holds and does |
|---|---|
jobs | The market: job commitments, escrow, dispatch records, the price ceiling and the floor per tier |
sla | The tier registry, timing evidence, attainment per router, device and model, penalties and re-tiering |
reputation | Scores for devices, routers and models; the update rules; decay |
models | The provider registry: model packages, device-spec and SLA envelopes, benchmark results, royalty rate and the loyalty curve |
settlement | The split; batch settlement by Merkle root and net deltas; USD-to-UHI conversion at settlement; the burn |
validation | Sampling, canary pool, tolerance bands, attestation verification, dispute window, slashing |
Governance, staking, balances, treasury and the bridge are the standard Substrate pallets, which is one reason to use Substrate: five years of audited pallets for everything that is not specific to UHI.
8.5 Batching and throughput
Validators submit one settlement batch per router per block. A batch carries the Merkle root of its job records and the net balance deltas per account. At 2 s blocks and a batch ceiling of 5,000 jobs per router per block, one router's settlement capacity is 2,500 jobs a second, and capacity scales with routers [E7: a design target, to be measured on testnet]. The model's month-36 volume of 86 million jobs a day [E10] averages about 1,000 jobs a second, so one router covers the average and the peaks are what the second and third routers are for.
8.6 Forkless upgrades
The runtime is a Wasm blob stored on-chain and upgraded by a governance vote, so the payout split, the tiers, the loyalty curve, the fee schedule and the pallets themselves can change without a hard fork and without a client release. The economics in this document are a launch proposal precisely because changing them is cheap.
8.7 Bridges, later
The chain is sovereign; the token is not meant to be stranded. At month 24 the roadmap adds bridges to Polkadot and Ethereum for liquidity, with settlement remaining on the UHI chain. The bridge is a way for value to move, not a way for jobs to move.
9 The token
9.1 What UHI is for
UHI is four things, and nothing else:
- Settlement. Every job settles in UHI. Requesters may pay in dollars; the
settlementpallet converts at the moment of settlement, so the supply side always earns the native asset and can hold or cash out. - Stake. Validators, routers and Realtime-tier devices post UHI, which can be slashed.
- Loyalty. A model provider's royalty rate and a device's tier are earned through performance; both are recorded and paid in UHI.
- Vote. UHI is the vote on the Protocol and Treasury tracks and one of the weights on the Market track.
9.2 Demand is priced in dollars
A requester is quoted in dollars per unit of work: per thousand characters of TTS, per minute of STT, per thousand tokens, per thousand embeddings. The quote is fixed when the job is submitted. At settlement, the dollars are converted to UHI at a reference rate produced by the chain's oracle (a median over several venues, section 11), and the UHI is split. A customer never needs to hold, buy or think about the token. This is a deliberate choice: enterprise demand must not depend on a price chart.
9.3 Supply
Genesis supply of 1,000,000,000 UHI at mainnet, which is the token generation event. Allocation, as a proposal [E8]:
| Allocation | Share | Amount | Schedule |
|---|---|---|---|
| Supply-side pool: device onboarding, provider grants, validator rewards | 40% | 400 M | Four years, declining (section 9.4): 65% to device onboarding, 20% to provider grants, 15% to validators |
| Treasury | 20% | 200 M | 10% at genesis, the rest linear over 48 months; spent by Treasury-track vote only |
| Team and founders | 18% | 180 M | 12-month cliff, then linear over 36 months |
| Investors (pre-seed and seed) | 12% | 120 M | 6-month cliff, then linear over 24 months |
| Community and liquidity | 10% | 100 M | 30% at genesis, the rest linear over 24 months: testnet participants, early design partners, liquidity |
The supply-side pool is part of the genesis supply, so the maximum supply is the genesis supply (section 15 lists this as an open decision).
9.4 Emission schedule
Supply-side emissions bootstrap the network before job revenue can: a device earns an onboarding bonus for its first verified jobs in each tier, a model provider earns a publication bonus when a model passes benchmark on a new device class, and both earn a per-job top-up that declines as the schedule runs. The schedule is front-loaded and ends.
| Year from mainnet | Released from the pool | Share of the 400 M pool | Cumulative |
|---|---|---|---|
| 1 | 160 M UHI | 40% | 160 M |
| 2 | 120 M UHI | 30% | 280 M |
| 3 | 80 M UHI | 20% | 360 M |
| 4 | 40 M UHI | 10% | 400 M |
Of each year's release, 65% goes to device onboarding, 20% to model-provider grants and 15% to validator block rewards. In the model, emissions are 0.08× job fees by month 36 [E10]: usage, not emissions, is the main operator income by then.
After year four there are no emissions to the supply side; device operators and model providers earn only from jobs. The schedule is governable, but only downward: a Protocol-track vote can slow emissions, and cannot increase the pool.
9.5 Burn, treasury and stake
- Burn. A quarter of the treasury's share of every job, which is 0.5% of the job's price, is burned at settlement, and all ordinary transaction fees are burned. So usage, not speculation, sets the floor: the more jobs the network settles, the more UHI leaves circulation.
- Treasury. The remaining 1.5% of every job's price, plus the 20% genesis allocation, spent only by vote.
- Validator stake. 500,000 UHI per validator at launch, set by the Protocol track; the active set is the top 100 by stake plus nominations once the set opens.
- Realtime device stake. 500 UHI, set by the Market track and sized to about a month of Realtime earnings for a rig at the model's price assumption. The stake is small enough that a gamer can post it from earnings and large enough that cheating on one job is never worth it.
9.6 One TTS job, in cents
A requester submits a 400-character TTS job, about 25 seconds of speech, at the Standard tier. Cloud list prices for neural TTS run $15–30 per million characters [P12], so the same job costs 0.6–1.2¢ in the cloud, 0.8¢ at the mid-point. UHI's Batch price is 15% of that, 0.12¢, and Standard is 1.6× Batch: 0.192¢, which the arithmetic below rounds to 0.2¢ ($0.002) [E9].
| Line | Share | Cents |
|---|---|---|
| Requester pays | 0.200¢ | |
| Device operators (2 devices at Standard, split equally) | 70% | 0.140¢ (0.070¢ each) |
| Model provider (at the full 1,500 bps) | 15% | 0.030¢ |
| Validators | 10% | 0.020¢ |
| Router | 3% | 0.006¢ |
| Treasury, of which a quarter is burned | 2% | 0.004¢ (0.003¢ kept, 0.001¢ burned) |
At 0.07¢ per Standard job, a device needs about 1,430 verified jobs to earn a dollar. A phone that runs one of these jobs every 8 seconds while busy [E4] earns about $0.32 for each busy hour. The economics of the network are the economics of volume: small amounts, settled correctly, millions of times.
10 Governance
Three tracks, all on-chain, OpenGov-style, with time-locked conviction voting.
| Track | Decides | Who votes |
|---|---|---|
| Protocol | Runtime upgrades, validator set size, slashing rules, emission slowdowns | Staked validators and token holders |
| Market | SLA tier definitions, payout split, routing-fee cap, model-provider loyalty curve, stake sizes, scoring weights | Every party, weighted: operators by earned jobs, providers by served jobs, requesters by spend, validators by stake |
| Treasury | Grants to model providers, onboarding incentives, audits | Token holders, with a provider council veto on provider grants |
10.1 Conviction voting
A vote can be cast with no lock at 0.1× weight, or locked for one to six lock periods (a period is 28 days) for 1× to 6× weight. Locking is the cost of conviction: a holder who wants six times the say on a payout-split change holds the token through roughly six months of its consequences.
10.2 Market-track weighting
The Market track is the one that sets the price of the work, and it is weighted by work done, not by tokens held. Each party's bloc holds a fixed share of the track's weight (operators 40%, providers 20%, requesters 20%, validators 20%) and within a bloc, weight is proportional to that party's 90-day record: jobs earned, jobs served, dollars spent, stake. A fund that buys tokens buys a voice on the Protocol and Treasury tracks; it cannot buy the price of a job. The people who do the work govern the price of the work.
A fund that buys tokens buys a voice on the Protocol and Treasury tracks. It cannot buy the price of a job.
10.3 Provider council veto
A council of seven model providers, elected by providers weighted by served jobs, can veto any Treasury-track grant to a model provider. The veto is a safeguard against the treasury funding models that do not work on devices, judged by the people who make models work on devices.
10.4 Emergency track
A technical committee of nine (five from the validator set, two from UHI Labs, two from the provider council) can, with six signatures, pause settlement or whitelist an emergency runtime fix for 72 hours. Every emergency action is put to a Protocol-track vote retroactively; a committee whose emergency action is voted down is dissolved and re-elected.
11 Security and privacy
11.1 Threat model
| Threat | What an attacker gains | Defence |
|---|---|---|
| Lazy device returns a cheap wrong answer | Payment without compute | Redundancy in Standard and Realtime; 5% random re-execution in Batch; canaries; a failed check costs ~50 jobs of reputation; Realtime stake slashed 10% |
| Sybil operator runs many fake devices | Collect onboarding emissions; win redundancy slots and agree with itself | Attestation required for emissions above a floor; diversity term in routing keeps a job's n devices on different owners, networks and metros; emissions per device capped and declining; proven Sybil slashed 100% |
| Collusion between devices on the same job | Agree on a wrong hash | Diversity term; canaries indistinguishable from real jobs; VRF-chosen spot checks; tie-break execution on validator-chosen devices |
| Router bias favours its own devices or starves others | Capture more of the 70% | Published scoring function; attainment rotation; routers staked and slashable; devices see their own score inputs; open routers at month 24 |
| Validator collusion approves bad batches | Steal escrow; protect colluding devices | GRANDPA finality with 2/3 honest assumption; permissioned set at launch chosen for independence; NPoS with slashing; archive nodes let anyone re-verify a batch against its root and raise a dispute within 7 days, paid from the slashed stake |
| Oracle manipulation moves the USD-to-UHI rate | Over- or under-pay the supply side | Median of several venues with outlier rejection; rate bounded to a daily move; settlement in dollars is fixed at submission so only the conversion is exposed; oracle feeders are staked validators |
| Payload privacy: router, validator or device operator reads data | Data theft | End-to-end encryption to the executing device; validators see hashes and fingerprints only; keys in secure hardware where available; on-device-only and Fleet modes for data that must not move; nothing retained after the job |
| Device owner safety: the app harms the device or the bill | Not an attack, but the failure that kills supply | Hard budgets enforced in the app: battery floor (default 30%, never below 20%), thermal cap read from the OS, hours, data allowance with a cellular off switch, and a kill switch in the notification tray. The app exceeding a budget is a bug we treat as a security incident |
11.2 Device owner safety in practice
The node app reads the platform's thermal and battery APIs and stops accepting work before the OS would throttle; it does not run on battery below the floor; it prefers Wi-Fi and never exceeds the data allowance; it never runs while the device is in active use unless the operator opts in. A phone on the bedside table at night, charging, on Wi-Fi, is the canonical worker. Everything about the defaults is set for that device.
11.3 Audits
Runtime and node app audits are a standing Treasury-track line item. The runtime is audited before mainnet and after every upgrade that touches settlement or validation; the node app is audited before each platform's store release.
12 Economics for each party
12.1 Device operators
Earnings are stated for a busy device at reference utilisation: a device that admission (section 6.4) has let in because demand keeps it busy for its reference share of its online hours. They are estimates from the financial model [E4, E10], net of electricity, and they are not a network average: a registered device on the waitlist earns nothing until it is admitted, which is the point of the rule. The chain is: online hours × busy share = busy hours; busy hours × job value per busy hour × the 70% operator share, less power at 16¢/kWh [P13]. Job value per busy hour is the model's blended workload (TTS, STT, summarisation, embeddings, translation, small-LLM chat) at 15% of cloud list, with each class's tier mix (phones half Batch and half Standard; rigs half Realtime) and its throughput against a phone (laptop 1.5×, rig 3×, console 1.5×, watch 0.08×).
| Device class | Online hours | Busy share | Busy hours | Job value per busy hour | Operator share, 70% | Power | Earns per day, net |
|---|---|---|---|---|---|---|---|
| Phone (2024+ flagship, charging, Wi-Fi) | 16 | 20% | 3.2 | $0.92 | $2.06 | 5 W, $0.00 | ~$2 |
| Laptop (NPU, on mains) | 10 | 25% | 2.5 | $1.74 | $3.05 | 35 W, $0.01 | ~$3 |
| Gaming rig or desktop (RTX-class GPU) | 18 | 30% | 5.4 | $4.16 | $15.72 | 350 W, $0.30 | ~$15 |
| Console (current generation) | 6 | 30% | 1.8 | $1.63 | $2.05 | 150 W, $0.04 | ~$2 |
| Watch or glasses (charging, Batch only) | 4 | 50% | 2.0 | about $0.06 | about $0.08 | negligible | cents |
Two things the table does not say. First, the admission target is 50% of reference, so in the model's base case an active device runs at half of its reference utilisation and its realised earnings are about half of the table: a phone $1.03 a day, a laptop $1.52, a rig $7.71 at month 36 [E10]. Raising the target to 100% would pay active devices the full table and admit half as many; that is a Market-track decision (section 15). Second, the model caps utilisation at launch: the node app allows 50% of reference utilisation at mainnet and relaxes to 100% by month 24 as battery and thermal handling matures. The table will be replaced by measured testnet figures at month 6 [M, pending].
Worked example: one laptop, one day. A laptop with an NPU, on mains and Wi-Fi, admitted to the Standard tier, holding one warm TTS model and fed by the router at its reference utilisation [E4].
| Line | Figure |
|---|---|
| Online | 10 hours |
| Busy share at reference utilisation | 25%, so 2.5 busy hours |
| Job value per busy hour (blended workload, 15% of cloud list, laptop tier mix and 1.5× phone throughput) | $1.74 |
| Job value for the day | $4.36 |
| Operator share, 70% | $3.05 |
| Power: 35 W for 2.5 busy hours is 0.09 kWh at 16¢/kWh [P13] | $0.01 |
| Net to the operator | $3.04 a day, about $91 a month at 30 days |
At about 0.2¢ a Standard job, the laptop's $4.36 of job value is roughly 2,200 jobs, each settled inside a batch a few seconds after it finished. The operator sees one number rise through the day.
12.2 Requesters
UHI's price is set against cloud list: the Batch price of a job is 15% of the mid-point of the published list prices for the same unit of work [E3], Standard is 1.6× that and Realtime 2.5×. Cloud list prices are published figures at the mid-point of the vendor ranges in the model's Sources sheet [P12, P14; verify current list prices].
| Workload | Unit of work (one job) | Cloud list, mid-point | UHI Batch (15%) | UHI Standard (24%) | UHI Realtime (37.5%) |
|---|---|---|---|---|---|
| TTS, neural | 400 characters | $0.0080 | $0.0012 | $0.0019 | $0.0030 |
| STT | 1 minute of audio | $0.0100 | $0.0015 | $0.0024 | $0.0038 |
| Summarisation | 2,000 input tokens | $0.0006 | $0.00009 | $0.00014 | $0.00023 |
| Embeddings | 1,000 tokens | $0.0001 | $0.000015 | $0.000024 | $0.0000375 |
| Translation | 1,000 characters | $0.0200 | $0.0030 | $0.0048 | $0.0075 |
| Small-LLM chat | 1,500 tokens in and out | $0.0006 | $0.00009 | $0.00014 | $0.00023 |
| One voice minute (1 STT minute + 1 TTS turn + 0.6 LLM turn + 0.2 summary) | 2.8 jobs | $0.0185 | $0.0028 | $0.0044 | $0.0069 |
So a requester pays 6.7× less than cloud list at Batch, 4.2× less at Standard and 2.7× less at Realtime. The spread across vendors is wider than the spread across tiers: list prices for the same TTS minute differ by 2× between vendors, which is where the "5–20×" of the pitch comes from and why this document quotes the mid-point. The Standard voice minute at $0.0044 is above our measured CPU cost of $0.0025 a minute on QuickDial's own infrastructure [M2], so the price covers the compute with room for the split. The second thing requesters buy is elasticity: capacity that scales with devices online, so a product launch does not wait on a GPU quota.
12.3 Model providers
From section 7: 900 to 1,500 bps of every job, earned through attainment and volume. At the full rate on a model serving 100 million jobs a month at 0.2¢, the royalty is $30,000 a month [E6]; the same model on a model hub today earns nothing per inference.
12.4 Validators
10% of every job, plus 15% of the supply-side pool as block rewards over four years. At the model's month-36 volume of about 2.6 billion jobs a month and $4.3 million of job value [E10], the validation pool is about $430,000 a month across the set; at month 12 it is about $4,800. Validator economics follow network volume, which is why block rewards exist for the first years.
12.5 Routers
3% of every job, capped by the Market track. Routing is a low-margin, high-volume service by design; its reward is traffic, and traffic follows attainment.
12.6 UHI Labs and the protocol
The protocol's take is thin, and it declines. UHI Labs earns the 3% routing fee while it runs the routers, a share that falls from month 24 as routers decentralise, plus royalties on the models it publishes itself (60% of the provider pool at mainnet, 20% by month 36 as outside providers arrive). The treasury keeps 1.5% of every job. Together that is a take of 5–6% of job value at month 36 [E10], about $1 million of protocol revenue in year three on $31 million of job value, and UHI Labs is not profitable on fees alone inside 36 months. That is by design: value in the network accrues to the token, through the burn and the stake, and to the treasury, through its 20% of genesis and its share of every job, which is the structure every working distributed-infrastructure network has converged on. The company's job is to make the number of settled jobs large.
13 Roadmap
Two device counts, on purpose. Registered devices are the waitlist: everyone who installed the node app. Active devices are the ones admission has let in because demand keeps them busy at reference utilisation (section 6.4), earning the figures in section 12. Active is the number that matters; registered is the growth engine waiting behind it. The counts and ARR are the financial model's base case [E10]; ARR is gross job value in the month × 12.
| Month | Milestone | Registered (waitlist) | Active (earning) | ARR |
|---|---|---|---|---|
| M6 | Testnet with the six pallets; node app on rigs and phones (macOS, Windows, Android; iOS in TestFlight); Batch tier; our own voice AI as the first workload; measured earnings replace the estimates in section 12 | 1,000 | testnet | — |
| M12 | Mainnet with 40 validators; Standard tier; three design partners as requesters; USD billing with conversion at settlement; the first ten models in the registry | 10,000 | ~500 | ~$0.6 M |
| M18 | Realtime tier with staked devices; the model-provider program with treasury grants; permissioned routers | 50,000 | ~3,500 | ~$5 M |
| M24 | Open, staked routers; bridges to Polkadot and Ethereum; validator set opening toward 100 | 200,000 | ~8,000 | ~$13 M |
| M36 | USA to global: Europe, India and Latin America; OEM and telco pre-installs | 1,000,000 | ~50,000 | ~$50 M |
Every active device is busy, which is the promise the brand makes: in the base case at the 50% admission target it earns about half of the reference figures in section 12.1, and the Market track can trade the active count for the rate. Demand is the lever: ARR and the active count are both linear in requester volume and in the discount to cloud list.
Supply comes first from gamers and developers with rigs (the highest TOPS per device, mains power, always on), then phones through the node app. Demand comes first from our own products, then from three design partners in voice AI, contact centres and media. Model providers come from the quantisation and small-model community, with treasury grants for the porting work.
14 Team
Purushottam "Puru" Chaudhary, Founder & CEO. AI inference engineer. Built the STT, LLM and TTS models and the production platform behind QuickDial AI and Vartalaap, where speech and small-model inference run on commodity compute at telephony latency. Fifteen years in engineering at GE HealthCare, S&P Global, Bristol Myers Squibb and State Street. Author of The Inference Mechanic.
Ankur Kapoor, Co-founder, Product & Operations. Fifteen years in financial-services technology: at TCS contracted to GE Capital, at SVB, and at USAA, where he is a Product Manager. Cornell Johnson MBA, 2020. Co-founder of and investor in QuickDial.
UHI Labs is the founding operator: it writes the runtime and the node app, runs the launch router and the first validators, and hands each of those roles to the network on the schedule in section 6.3 and section 13.
15 What is not yet decided
A litepaper that reads as finished is hiding something. These are the open decisions, with the direction we lean and what decides them.
| Decision | Where we lean | What decides it |
|---|---|---|
| Phones on iOS | Apple limits background compute; an iOS device may only work while charging with the app open, or inside Background Processing windows. The earnings table assumes Android-class availability for phones | The M6 node app on iOS, measured over real nights |
| Indicative prices | Batch at 15% of cloud list, Standard 1.6× and Realtime 2.5× Batch; a 400-character TTS job at 0.12¢, 0.19¢ and 0.30¢ | Testnet cost and demand data; the Market track after mainnet |
| Maximum supply | The 400 M supply-side pool sits inside the 1,000 M genesis supply, so the cap is 1,000 M; the model's Token sheet currently mints the pool on top of genesis and must be reconciled | Counsel and the Protocol track before mainnet |
| Stake sizes | 500 UHI for a Realtime device, 500,000 UHI for a validator | Observed misbehaviour rates on testnet; the token's price at mainnet |
| Admission target | Realised utilisation at or above 50% of reference, which pays an active device about half of the reference figures; 100% would pay the full figures to half as many devices | Testnet: how far below reference a device can sit before operators churn |
| Oracle design | A median over several venues, bounded daily move, fed by staked validators | Which venues list UHI at mainnet; audit |
| USD billing | UHI Labs as merchant of record for card and invoice payment, converting at settlement | Legal and regulatory review by jurisdiction |
| Token legal structure | Not decided; counsel engaged before any sale or distribution | Counsel; the Alliance programme |
| Consensus details | Slot-based block production with GRANDPA finality, 2 s blocks; Aura at launch | Testnet performance; validator set size |
| Tolerance bands per model | Set by the provider, confirmed by benchmark jobs; 0.90 cosine for TTS at launch | Per-model benchmarks on real devices |
| Batch spot-check rate | 5% | The measured rate of lazy devices on testnet; must stay above the point where cheating pays |
| Emission top-up formula | Per-job top-up that declines with the schedule; onboarding bonus per tier | Device growth against the M12 and M18 targets |
| Bridges | Polkadot first, Ethereum second, at M24 | Liquidity and security review at the time |
| Compliance path | SOC 2 for the router and billing; a HIPAA-eligible configuration through Fleet and On-device-only modes | Design-partner requirements |
| Requesters as routers | Allowed at phase 3, with the same stake and attainment rules | Phase 2 results |
16 References
Published figures [P], measurements [M], estimates [E]. verify marks a reference where the primary source is still being confirmed before the 1.0 release of this document.
- P1 Inference as roughly two-thirds of AI compute demand in 2026. Deloitte, TMT Predictions 2026. verify: confirm the exact figure and wording in the published report.↩ 4 1 2 3
- P2 AI-related capital expenditure of $400–450 billion in 2026. Deloitte, TMT Predictions 2026. verify: confirm the range as published.↩ 4 1 2 3
- P3 Data-centre capital expenditure forecast of $1.7 trillion a year by 2030. Dell'Oro Group, data center capex five-year forecast, 2025. verify: confirm the year and the figure.↩ 1 2
- P4 6.9 billion smartphones in use. GSMA Intelligence / Ericsson Mobility Report, 2025–2026. verify: distinguish smartphone connections from unique devices in use.↩ 3 1 2
- P5 About 1.5 billion laptops and tablets in use. Derived from IDC and Gartner shipment data over a five-year replacement cycle. verify: an installed-base figure from a single primary source.↩ 1 2
- P6 About 1.9 billion PC gamers and about 650 million console gamers. Newzoo, Global Games Market Report, 2025. verify: figures are people, not devices; confirm the edition.↩ 1 2 3
- P7 Model sizes and on-device feasibility: published model cards for compact TTS (Kokoro-82M, Piper), STT (Whisper base/small, Parakeet), small LLMs (Qwen3-1.7B, Gemma 3n, Phi-4-mini, Llama 3.2 1B/3B), translation (NLLB-600M, MADLAD) and embeddings (bge-small, nomic-embed), together with their published on-device demonstrations.↩ 1 2 3 4 5 6
- P8 Apple M4 neural engine, 38 TOPS. Apple, May 2024 announcement.↩ 1
- P9 Qualcomm Snapdragon X Elite NPU, 45 TOPS. Qualcomm, 2023–2024 product documentation.↩ 1
- P10 Copilot+ PC requirement of a 40+ TOPS NPU. Microsoft, May 2024.↩ 1
- P11 AI-capable devices under 20% of the installed base at end-2026. Derived from IDC and Gartner forecasts of GenAI smartphone and AI PC shipment shares against the installed base. verify: a single primary figure for installed-base share.↩ 1
- P12 Cloud neural TTS list prices of $15–30 per million characters: Google Cloud Text-to-Speech, Amazon Polly, OpenAI, ElevenLabs, published price lists; the model uses the $20 mid-point. verify current list prices at publication.↩ 1 2 3 4
- P13 US average residential electricity price of about 16¢/kWh. US Energy Information Administration, Electric Power Monthly, 2025–2026. verify the current month's figure.↩ 1 2 3
- P14 Cloud list prices for STT ($0.36–1.44 per hour of audio), summarisation and small-LLM chat ($0.10–0.60 per million tokens), embeddings ($0.10 per million tokens) and translation ($20 per million characters): Deepgram, Google Cloud, Amazon, OpenAI and DeepL published price lists, mid-points as used in the model's Sources sheet. verify current list prices at publication.↩ 1
- P15 Conviction voting multipliers (0.1× unlocked to 6× at the longest lock). Polkadot OpenGov documentation.
- P16 Substrate: weight-based fees, forkless Wasm runtime upgrades, FRAME pallets. Parity Technologies, Substrate and Polkadot SDK documentation.
- P17 Distributed compute networks (Render, Akash, io.net) and their revenue and market capitalisation. verify: current figures from each network's published statistics and public market data; used in section 1.4 only as context.
- M1 Speech-in to first-audio-out of 253–392 ms on a CPU-only server running STT, a small LLM and TTS. Measured on QuickDial AgentBox production, 2026, endpoint-to-first-syllable, over 8 vCPU instances.↩ 1
- M2 Compute cost of $0.0025 per voice minute for the STT, small-LLM and TTS loop on CPU. Measured on QuickDial production infrastructure, 2026, from instance cost against minutes served.↩ 1
- E1 Small models at 1–3 B parameters match shipped cloud quality for TTS, STT, summarisation, translation and embeddings on 2024+ devices. Our assessment from building and serving these workloads; quality is workload-specific and is to be confirmed per model by the network's benchmark jobs.↩ 1
- E2 Devices idle about 90% of the day. From screen-time surveys of 3–5 hours a day of phone use, and typical gaming and console use of 1–3 hours a day.↩ 2 1
- E3 Requester savings of 6.7× at Batch, 4.2× at Standard and 2.7× at Realtime against cloud list mid-points. From the model's pricing rule, Batch = 15% of list, Standard 1.6× Batch, Realtime 2.5× Batch, against [P12] and [P14].↩ 1 2 3
- E4 Device-operator earnings for a busy device at reference utilisation. From the financial model's Devices sheet: online hours and reference utilisation by class (phone 16 h at 20%, laptop 10 h at 25%, rig 18 h at 30%, console 6 h at 30%, watch 4 h at 50%), throughput against a phone, tier mix, the blended workload at 15% of cloud list, the 70% operator share, and power at [P13]. To be replaced by testnet measurements at M6.↩ 1 2 3 4 5 6
- E5 A failed verification costs about 50 successful jobs of reputation. A launch parameter of the
reputationpallet, set so that a 5% spot-check rate makes a lazy Batch device unprofitable.↩ 1 - E6 Model-provider earnings in the worked examples. Illustrative volumes at the indicative Standard price.↩ 1 2
- E7 Settlement capacity of 2,500 jobs a second per router. A design target from the batch ceiling and block time, to be measured on testnet.↩ 1
- E8 Token allocation and emission schedule. The launch proposal; governable downward on the Protocol track.↩ 1
- E9 Standard-tier price of 0.192¢ for a 400-character TTS job, rounded to 0.2¢ in the worked example. From the model's pricing rule against the $20 per million characters mid-point in [P12].↩ 1
- E10 The financial model:
uhi-net/uhi-model,UHI-Model.xlsx, dashboard exported 2026-10-09. Device counts, ARR, job volumes, take rate, TAM, SAM and SOM are its base-case outputs; every input is on its Assumptions sheet and every anchor on its Sources sheet with a verify flag. TAM is two-thirds of the AI capex mid-point; SAM is the sum of the TTS, STT, translation, small-LLM and embeddings API markets × 62% US and EU × 70% purchasable by API, growing 22% a year; SOM is the model's own month-36 ARR.↩ 1 2 3 4 5 6 7 8
Glossary
- Attainment. The fraction of jobs in a window that met their tier's latency bound and passed verification.
- Attestation. A hardware-backed statement that a proof came from an unmodified node app on a genuine device.
- Batch, Standard, Realtime. The three SLA tiers: minutes at 1× redundancy; under 1 s at 2×; under 300 ms at 3×.
- Canary. A real job with a known answer, dispatched to test a device.
- Conviction voting. Voting with weight that rises with the length of time the voter locks their tokens.
- Device operator. A person or organisation that runs the node app on hardware they own and is paid per job.
- Envelope. A model's published limits: the device spec it needs and the SLA it meets on each device class.
- Escrow. The requester's price ceiling held by the
jobspallet until the job settles or is refunded. - Fingerprint. A compact representation of a non-deterministic result (a perceptual hash or an embedding) used for tolerance-band comparison.
- Job. One unit of inference work with its model, payload, tier, region, privacy mode and price.
- Loyalty curve. The function that raises a model provider's royalty from 900 to 1,500 basis points with attainment and volume.
- Model provider. A team that publishes a model to the network and earns a royalty on every job that runs it.
- Node app. The application a device operator installs; "node" is the engineers' word for a device in the network.
- Pallet. A module of a Substrate runtime; UHI adds six.
- Proof. What a device returns with a result: the output hash or fingerprint, signed timings, and an attestation where available.
- Reputation. A score in [0, 1] per device, router and model, raised by verified work, lowered by failures, and decayed by inactivity.
- Requester. The party that submits a job and pays for it.
- Router. The party that matches jobs to devices and is paid 3% for meeting the SLA.
- Settlement. The on-chain step that releases escrow and splits a job's price 70/15/10/3/2.
- Burn. The quarter of the treasury's share of every job, 0.5% of the price, that is destroyed at settlement.
- Admission. The router's rule for letting registered devices start earning: only as fast as demand keeps active devices at or above half of reference utilisation.
- Registered and active. A registered device has installed the node app and sits on the waitlist; an active device has been admitted and earns.
- Reference utilisation. The busy share of online hours that the earnings figures assume for a device class: 20% for a phone, 30% for a rig.
- SLA tier. A named service level: latency bound, redundancy, device eligibility, region.
- Treasury. The chain account that receives 1.5% of every job (2% less the burn) and funds grants, onboarding and audits by vote.
- UHI. The chain's native asset: settlement, stake, loyalty and vote. Also the network, and universal high income in lower case when expanded.
- Validator. A staked operator that verifies work, approves payouts and produces blocks.