AI Usage β€” Open-Token Tracker

Weekly OpenRouter token flows: open vs closed weights, China vs US origin, price collapse, and who actually monetizes the tokens. Source: openrouter.ai rankings API, snapshotted weekly (Mondays) by python/fetch_open_tokens.py. Baseline 2026-07-22.

Open vs closed β€” share of weekly tokens

Open-weight = model publishes weights (Hugging Face id present in OpenRouter catalog). Series builds weekly from the snapshot history.

Total tokens vs crossover hurdle scenarios

Actual OpenRouter tokens/wk (solid) vs the three regime trajectories from the 7/22 baseline: β‰₯10x/yr = 2027 crossover (spot explosion) Β· 4x/yr = the hurdle (central 2028 crossover) Β· 2.5x/yr = glut regime (efficiency cancels growth). Where the Monday dots land IS the neocloud thesis scoreboard.

Model origin β€” CN vs US vs EU share of weekly tokens

Curated author-origin map in the fetcher; unknown authors (small fine-tune shops) counted as other.

Google disclosed token volume β€” the longest public series

Monthly tokens processed across Google surfaces, as disclosed by management (each point = an actual disclosure event). The slope of this line vs the ~4x/yr hurdle is the single best long-baseline demand read. Q2'26 update (7/29 call): Google switched metrics β€” disclosed "model APIs ~22B tokens/min, up from 16B a quarter ago" = ~0.95Q/mo API-ONLY (+37.5% q/q β‰ˆ 3.6x annualized β€” just below the 4x hurdle). Not comparable to this all-surfaces series, so plotted separately in the ledger; all-surfaces series awaiting next comparable disclosure.

Cloud backlog wars β€” contracted RPO stacked ($B)

Remaining performance obligations from filings (AWS from earnings-call Q&A). The demand-certainty picture: $2.43T of contracted cloud revenue after the Q2'26 prints (MSFT $678B Β· ORCL $638B Β· GCP $514B Β· AWS $496B β€” +$132B q/q Β· CRWV $104B, 8/11 print, +$25B early-Q3 commits not yet included). Jassy 7/30: "will still not have enough capacity to meet all the demand we have in 2026… also true in 2027." Update quarterly as prints land.

Effective compute stock β€” cumulative AI chip sales in H100-equivalents

Supply-side denominator for the crossover hurdle: Epoch's cumulative chip-sales stock (Nvidia + Google TPU + AMD + Amazon + Huawei + Cambricon, FLOP-weighted to H100e), with an editable algorithmic-efficiency multiplier on top β€” effective compute = hardware stock Γ— algo gains. If tokens must keep pace with EFFECTIVE supply, the demand bar is higher than the GW view suggests.

Frontier datacenter buildout β€” tracked IT power (GW)

Epoch's satellite + permit tracking of ~78 frontier AI datacenters (Colossus, Stargate sites, Fairwater, etc.), aggregated as a time series of IT power β€” solid = observed online, dashed = announced/under-construction phases on Epoch's timelines. This is the physical-supply cross-check against our capex-derived GW builds.

Blended market price β€” $/1M output tokens (token-weighted)

Weighted by each model's weekly volume. The open/closed gap is the arbitrage funding the migration.

Top models by weekly tokens

Monetization per token β€” open labs vs frontier benchmark

Question the table answers: how many dollars does each lab extract per trillion tokens served per year? Open labs computed from router flow + disclosed ARR; frontier rows from their own disclosed throughput + reported revenue.
Read with two biases in mind: (1) open-lab rows use OpenRouter flow only β€” their first-party API traffic is invisible here, so true token volume is higher and true $/1T is LOWER than shown (these are upper bounds); (2) frontier revenue is blended API + consumer subs. Both biases WIDEN the real gap. Sources: DeepSeek ~$220M & Mistral ~$400M ARR (State of Open Source AI V1.0, Jul-26); Zhipu ~$1B (IPO disclosures Jan-26); MiniMax ~$300M (mgmt May-26); OpenAI ~6B tokens/min API (DevDay Oct-25) + ~$29B 2026e revenue (press); Anthropic ~$9B exit-2025 run-rate (press), throughput undisclosed. Update quarterly.

Implied inference GW β€” from global tokens to datacenter power (editable)

Bottoms-up: global token demand β†’ decode-equivalent GPU-seconds β†’ GPUs β†’ GW, vs the announced buildout pipeline. Every input is editable; provenance in the note below. OpenRouter is a routing sliver β€” the global number is built from lab/platform disclosures.
Global tokens/month build (mid-26): Google disclosed 1.3Q/mo Oct-25 all-surfaces (assume ~2Q by now) + OpenAI ~6B tok/min API Oct-25 plus consumer (~0.5-1Q) + Fireworks 40T/day + Baseten 30T/day (UBS 7/21 β‰ˆ 2Q/yr combined) + China first-party (Doubao 200M DAU, DeepSeek, Qwen β€” ~1-2Q) + Meta/Anthropic/rest. 5Q/mo is a CENTRAL estimate; stress 3-8. Decode share from our OpenRouter data (completion β‰ˆ heavily prompt-dominated mixes; agentic re-reads inflate prefill). Throughput: batched decode on H100/B200-class serving frontier MoE β‰ˆ 500-1,500 tok/s/GPU aggregate. kW all-in: ~700-1,000W chip + server/network/cooling. Pipeline: OpenAI+Anthropic "tens of GW" (UBS checks) + Meta 14GW + ORCL ~10GW + hyperscalers/neoclouds. Buildouts also serve TRAINING β€” this module sizes inference only.

Token disclosure ledger β€” what the top AI companies have actually said

CompanyDisclosureWhen / whereβ‰ˆ Q tokens/mo
ByteDance (Doubao)180T/day (was 50T β†’ 120T May-26)Jun-26, Volcano Engine FORCE conf5.5
Google (Gemini all surfaces)1.3Q/mo (480T May-25 β†’ 980T Jul-25)Oct-25 disclosure; ~2Q est mid-26~2.0
Fireworks40T+/dayJul-26, UBS Private AI conf1.2
Baseten30T/day ("may exceed OpenAI API")Jul-26, UBS Private AI conf0.9
OpenAI~6B tok/min API (+ consumer undisclosed)Oct-25 DevDay0.26 API + ~0.5 est
OpenRouter (our weekly scrape)61T/wklive, this page0.27
Microsoft (Azure AI/Foundry)100T/qtr (Apr-25, dated). 7/29 call: 100K Foundry customers; customers at 1T-token annualized run-rate up 4x y/y; no absolute volumeFY26Q4 call 7/29~0.2-0.3 est
Amazon (AWS)no token disclosure; AI revenue run-rate >$25B + chips >$25B, both triple-digit growth (Jassy 7/30)Q2'26 callunmeasured
DeepSeek (V4 Flash, first-party)18.4T/moMozilla 2026 open-source report0.02
Anthropicno aggregate token disclosure foundrev-implied only (~$9B run-rate)est 0.1-0.3
Baidu (ERNIE) / Alibaba (Qwen) / Meta / xAIusers disclosed, tokens notβ€”unmeasured
VISIBLE TOTAL~10.5-11
Doubao alone β‰ˆ half the visible total β€” China first-party inference is the biggest single block and was invisible to Western routing data. Calculator default set to 11Q/mo (visible floor); unmeasured (Meta internal, Apple, Qwen/ERNIE, enterprise on-prem) argues true global is meaningfully higher.
SENSE-CHECK ANCHORS (external, for auditing): Google disclosed 0.24 Wh per median Gemini prompt (Aug-25 technical paper) ≈ 864 J/query ≈ ~1-2 J/token on a FLEET-MEASURED basis (includes idle capacity, host CPU, small-batch latency serving) vs this model’s 0.47 J/token spec-basis — reproduce the measured basis by setting MFU ≈ 4%. Altman: 0.34 Wh per ChatGPT query (Jun-25) corroborates. IEA: AI datacenters ~155 TWh in 2025 ≈ ~18 GW average draw for ALL AI; if inference is 50-80% of AI compute, implied inference fleet ≈ 9-14 GW — reproducible here with MFU 4% + global tokens ~15Q/mo (the disclosed-only 5Q build is a floor; Google alone was 1.3Q/mo back in Oct-25). AUDIT PRESETS: spec-optimistic (5Q, MFU 12%) → ~2 GW today, break-even ~22x; measured-realistic (15Q, MFU 4%) → ~8-10 GW today, break-even ~5-6x.