The napkin math is public now
Moonshot put a three-trillion-parameter model on Hugging Face, and the most interesting number in the release is not a benchmark. When dozens of providers race to serve the same public artifact, the price they converge on is a measurement: the first checkable floor under what a frontier token costs to serve. So I built the napkin, and it re-checks itself in CI.
Moonshot AI is putting Kimi-K3 on Hugging Face today: three trillion parameters, native MXFP4, about a terabyte and a half of weights. The release even had theater. A countdown page, a Hacker News thread refreshing itself into rumor, a 404 minutes before zero. Nobody is running this thing at home, and the benchmark arguments will grind on all week. The number I care about is not on any leaderboard. When dozens of hosting providers race to serve the same public artifact, the price they converge on becomes a measurement, and one of the most closely guarded numbers in this industry, what a frontier-class token actually costs to serve, just turned into arithmetic anyone can check.
A closed price is a sealed envelope
Buy tokens from a closed frontier API and you see exactly one number: the price per million tokens. Everything that composes it is sealed. You do not know how many parameters the model has, what precision it is served at, what the hardware behind it costs, what the gross margin is, or whether the price is being held down to buy market share. So the industry's longest-running argument runs in circles. One camp says the labs lose money on every token and the prices are venture-subsidized theater. The other points at SemiAnalysis estimating Anthropic's API gross margin above 80 percent, and at a leaked DeepSeek memo that reportedly priced tokens to recover hardware cost in ten months. Every one of these claims is plausible. None of them is checkable, because every operand in the division problem is private. A closed price is a policy, and a policy can encode anything: a margin, a subsidy, a moat, a hope.
Three trillion parameters of price discovery
An open-weights release at this scale changes the shape of the argument, because it fixes the artifact. Every provider that hosts K3 is serving the same three trillion parameters at a known precision from the same public file. The top comment in the release thread framed it before the weights even landed: wherever median third-party pricing settles will tell us what it costs to serve a 3T-class model, and from there you can finally start to sanity-check the subsidy claims. The hosts have no training bill to recover and no frontier narrative to protect. They are selling compute against identical competitors, which is the one market structure that grinds a price down toward marginal cost.
This experiment has already run once, one tier down. GLM 5.2 went up in mid-June; six weeks of provider competition later, the cheapest listed price on OpenRouter had fallen by roughly 45 percent. The mechanism is not new today. The artifact under it is, for the first time, the size of the frontier.
The napkin goes public
My favorite thing about the thread is that the napkin math happens in the open, with inputs attached. A terabyte and a half of weights just barely fits on one eight-GPU B200 node, so realistic serving starts at sixteen GPUs once you want context room and throughput. One commenter priced a used quad-socket Xeon with up to three terabytes of RAM at under thirty thousand dollars, drawing about forty-eight dollars a month of electricity at a 7.5-cent rate, for five or six tokens a second on overnight jobs. Another shot back that you would spend a hundred times more on electricity than the same tokens cost through an API. A third posted a $135-a-month build at a 12.4-cent rate. Every number in that exchange can be argued with. That is the change. The argument now contains numbers, the numbers have owners, and anyone can redo them.
What the napkin says
At middle assumptions, a rented sixteen-GPU node serves K3-sized tokens for about $1.90 per million, and the same node swings between $10.00 and $0.74 as you drag the assumptions from conservative to optimistic. That spread is the honest error bar on public information, and it will narrow fast as real providers post real prices. The sovereignty rig lands near $67 per million, roughly thirty-five times the rented floor, and the ratio is the real lesson. A mixture-of-experts model reads only a small slice of its weights per token, so a provider batching hundreds of concurrent requests amortizes that trillion-parameter file across all of them at once. A single-user box loads the whole model to feed one stream. Shared hardware wins on price per token by structure.
Which means running it yourself buys sovereignty, not economy. Thirty-five times the floor is what it costs for your prompts to never leave the building, and the thread's detour through EU firms, the CLOUD Act, and the missing confidential-computing GPU provider is a catalog of buyers for whom that premium is rational. The premium is now a number too.
Competition cannot tell you what a model cost to train. It is very good at telling you what one costs to serve.
What the floor still cannot tell you
It is worth being precise about what this instrument measures, because it is one number, not a lab's income statement. Training cost stays invisible, and reinforcement learning blurs even that line, since modern post-training burns most of its compute on inference-shaped rollouts. Closed model sizes stay unknown, so mapping the K3 floor onto any specific closed price still requires an assumption about scale that nobody outside the lab can check. Providers can quietly serve a degraded quantization, though OpenRouter listing precision per provider and running rolling benchmarks has made that harder to hide. And a price read in release week is not a settled price; the GLM curve needed six weeks to find its level. My own napkin carries the same humility: the hourly costs trace to public listings, but the throughput column is an assumption band, and the board treats it as one. Use the floor for its order of magnitude. Do not use it to compute any particular company's margin; it cannot do that, and neither can the thread.
One more sealed number, opened
This site keeps returning to one shape: a system becomes accountable at the moment its numbers become checkable by people outside it. Ninety-Seven Percent Off traced the gray market that grows where closed access meets people locked out of it. The Hugging Face field notes mapped the commons where K3 now sits, and the live observatory watches it hourly. Today the commons reached the frontier tier and took one sealed number with it. The price of a frontier token used to be a claim you could only believe or doubt. As of this morning it is an argument, and anyone can bring arithmetic.
Get the next one
An occasional note when something genuinely new ships here — essays, free tools, projects. No schedule, no filler, easy out.
Need something like this built?
I design and ship AI tools, full-stack apps, and data pipelines — end to end, to production. Tell me the problem in a sentence; I'll give you an honest read on fit within a day.
Work with me →