AI Usage β Open-Token Tracker
Weekly OpenRouter token flows: open vs closed weights, China vs US origin, price collapse, and who actually monetizes the tokens.
Source: openrouter.ai rankings API, snapshotted weekly (Mondays) by python/fetch_open_tokens.py. Baseline 2026-07-22.
Open vs closed β share of weekly tokens
Open-weight = model publishes weights (Hugging Face id present in OpenRouter catalog). Series builds weekly from the snapshot history.
Total tokens vs crossover hurdle scenarios
Actual OpenRouter tokens/wk (solid) vs the three regime trajectories from the 7/22 baseline: β₯10x/yr = 2027 crossover (spot explosion) Β· 4x/yr = the hurdle (central 2028 crossover) Β· 2.5x/yr = glut regime (efficiency cancels growth). Where the Monday dots land IS the neocloud thesis scoreboard.
Model origin β CN vs US vs EU share of weekly tokens
Curated author-origin map in the fetcher; unknown authors (small fine-tune shops) counted as other.
Google disclosed token volume β the longest public series
Monthly tokens processed across Google surfaces, as disclosed by management (each point = an actual disclosure event). The slope of this line vs the ~4x/yr hurdle is the single best long-baseline demand read. Q2'26 update (7/29 call): Google switched metrics β disclosed "model APIs ~22B tokens/min, up from 16B a quarter ago" = ~0.95Q/mo API-ONLY (+37.5% q/q β 3.6x annualized β just below the 4x hurdle). Not comparable to this all-surfaces series, so plotted separately in the ledger; all-surfaces series awaiting next comparable disclosure.
Cloud backlog wars β contracted RPO stacked ($B)
Remaining performance obligations from filings (AWS from earnings-call Q&A). The demand-certainty picture: $2.43T of contracted cloud revenue after the Q2'26 prints (MSFT $678B Β· ORCL $638B Β· GCP $514B Β· AWS $496B β +$132B q/q Β· CRWV $104B, 8/11 print, +$25B early-Q3 commits not yet included). Jassy 7/30: "will still not have enough capacity to meet all the demand we have in 2026β¦ also true in 2027." Update quarterly as prints land.
Effective compute stock β cumulative AI chip sales in H100-equivalents
Supply-side denominator for the crossover hurdle: Epoch's cumulative chip-sales stock (Nvidia + Google TPU + AMD + Amazon + Huawei + Cambricon, FLOP-weighted to H100e), with an editable algorithmic-efficiency multiplier on top β effective compute = hardware stock Γ algo gains. If tokens must keep pace with EFFECTIVE supply, the demand bar is higher than the GW view suggests.
Frontier datacenter buildout β tracked IT power (GW)
Epoch's satellite + permit tracking of ~78 frontier AI datacenters (Colossus, Stargate sites, Fairwater, etc.), aggregated as a time series of IT power β solid = observed online, dashed = announced/under-construction phases on Epoch's timelines. This is the physical-supply cross-check against our capex-derived GW builds.
Blended market price β $/1M output tokens (token-weighted)
Weighted by each model's weekly volume. The open/closed gap is the arbitrage funding the migration.
Top models by weekly tokens
Monetization per token β open labs vs frontier benchmark
Question the table answers: how many dollars does each lab extract per trillion tokens served per year? Open labs computed from router flow + disclosed ARR; frontier rows from their own disclosed throughput + reported revenue.
Read with two biases in mind: (1) open-lab rows use OpenRouter flow only β their first-party API traffic is invisible here, so true token volume is higher and true $/1T is LOWER than shown (these are upper bounds); (2) frontier revenue is blended API + consumer subs. Both biases WIDEN the real gap. Sources: DeepSeek ~$220M & Mistral ~$400M ARR (State of Open Source AI V1.0, Jul-26); Zhipu ~$1B (IPO disclosures Jan-26); MiniMax ~$300M (mgmt May-26); OpenAI ~6B tokens/min API (DevDay Oct-25) + ~$29B 2026e revenue (press); Anthropic ~$9B exit-2025 run-rate (press), throughput undisclosed. Update quarterly.
Implied inference GW β from global tokens to datacenter power (editable)
Bottoms-up: global token demand β decode-equivalent GPU-seconds β GPUs β GW, vs the announced buildout pipeline. Every input is editable; provenance in the note below. OpenRouter is a routing sliver β the global number is built from lab/platform disclosures.
Global tokens/month build (mid-26): Google disclosed 1.3Q/mo Oct-25 all-surfaces (assume ~2Q by now) + OpenAI ~6B tok/min API Oct-25 plus consumer (~0.5-1Q) + Fireworks 40T/day + Baseten 30T/day (UBS 7/21 β 2Q/yr combined) + China first-party (Doubao 200M DAU, DeepSeek, Qwen β ~1-2Q) + Meta/Anthropic/rest. 5Q/mo is a CENTRAL estimate; stress 3-8. Decode share from our OpenRouter data (completion β heavily prompt-dominated mixes; agentic re-reads inflate prefill). Throughput: batched decode on H100/B200-class serving frontier MoE β 500-1,500 tok/s/GPU aggregate. kW all-in: ~700-1,000W chip + server/network/cooling. Pipeline: OpenAI+Anthropic "tens of GW" (UBS checks) + Meta 14GW + ORCL ~10GW + hyperscalers/neoclouds. Buildouts also serve TRAINING β this module sizes inference only.
Token disclosure ledger β what the top AI companies have actually said
| Company | Disclosure | When / where | β Q tokens/mo |
|---|---|---|---|
| ByteDance (Doubao) | 180T/day (was 50T β 120T May-26) | Jun-26, Volcano Engine FORCE conf | 5.5 |
| Google (Gemini all surfaces) | 1.3Q/mo (480T May-25 β 980T Jul-25) | Oct-25 disclosure; ~2Q est mid-26 | ~2.0 |
| Fireworks | 40T+/day | Jul-26, UBS Private AI conf | 1.2 |
| Baseten | 30T/day ("may exceed OpenAI API") | Jul-26, UBS Private AI conf | 0.9 |
| OpenAI | ~6B tok/min API (+ consumer undisclosed) | Oct-25 DevDay | 0.26 API + ~0.5 est |
| OpenRouter (our weekly scrape) | 61T/wk | live, this page | 0.27 |
| Microsoft (Azure AI/Foundry) | 100T/qtr (Apr-25, dated). 7/29 call: 100K Foundry customers; customers at 1T-token annualized run-rate up 4x y/y; no absolute volume | FY26Q4 call 7/29 | ~0.2-0.3 est |
| Amazon (AWS) | no token disclosure; AI revenue run-rate >$25B + chips >$25B, both triple-digit growth (Jassy 7/30) | Q2'26 call | unmeasured |
| DeepSeek (V4 Flash, first-party) | 18.4T/mo | Mozilla 2026 open-source report | 0.02 |
| Anthropic | no aggregate token disclosure found | rev-implied only (~$9B run-rate) | est 0.1-0.3 |
| Baidu (ERNIE) / Alibaba (Qwen) / Meta / xAI | users disclosed, tokens not | β | unmeasured |
| VISIBLE TOTAL | ~10.5-11 |
Doubao alone β half the visible total β China first-party inference is the biggest single block and was invisible to Western routing data. Calculator default set to 11Q/mo (visible floor); unmeasured (Meta internal, Apple, Qwen/ERNIE, enterprise on-prem) argues true global is meaningfully higher.
SENSE-CHECK ANCHORS (external, for auditing): Google disclosed 0.24 Wh per median Gemini prompt (Aug-25 technical paper) ≈ 864 J/query ≈ ~1-2 J/token on a FLEET-MEASURED basis (includes idle capacity, host CPU, small-batch latency serving) vs this model’s 0.47 J/token spec-basis — reproduce the measured basis by setting MFU ≈ 4%. Altman: 0.34 Wh per ChatGPT query (Jun-25) corroborates. IEA: AI datacenters ~155 TWh in 2025 ≈ ~18 GW average draw for ALL AI; if inference is 50-80% of AI compute, implied inference fleet ≈ 9-14 GW — reproducible here with MFU 4% + global tokens ~15Q/mo (the disclosed-only 5Q build is a floor; Google alone was 1.3Q/mo back in Oct-25). AUDIT PRESETS: spec-optimistic (5Q, MFU 12%) → ~2 GW today, break-even ~22x; measured-realistic (15Q, MFU 4%) → ~8-10 GW today, break-even ~5-6x.