What it actually costs to serve a token — six cost layers from silicon to R&D — and how that compares to the price vendors charge. A first-principles estimate, every input a slider.
What it costs a vendor to serve one token — built up from silicon to R&D — compared against what they charge. This is the fully-loaded cost underneath the API price on the Impact Calculator: physical serving here is the same quantity as its "token cost". Every number is an editable estimate, not a disclosure. Data anchors as of —.
Fully-loaded and physical-only cost per 1M tokens across the year baselines, at the current slider settings. The unit cost falls several-fold as throughput, batching and silicon improve — directionally tracking the experience curve.
Top-down from a fleet's average IT power (MW): annualised cost of running the fleet that exists, plus its CO₂e and water. Pick a lab:
Root everything in accelerator-seconds; the rest is unit conversions + accounting. Per 1M tokens:
Physical serving = silicon + facility + energy + network — the same quantity as the "token cost" on the Impact Calculator. Soft = training + R&D, dominated by the two undisclosed inputs. Onsite-water ceiling is latent heat: rejecting 1 kWh by evaporation needs ≈1.48 L (hfg ≈ 2.44 MJ/kg); real WUE 0.2–1.9 reflects how much heat is rejected evaporatively.
Physical chain is robust; soft-cost and per-company inputs are explicitly estimates. Every value is an editable anchor in js/cost-model.js. Compiled from public sources + first-principles derivation — verify before load-bearing use. Runs entirely in your browser; no network calls.