Skip to the show

A scrolling comedy special about thermodynamics

DRINK · THINK · BURN

What your chatbot actually costs in water and watts — and the post-transformer, post-von-Neumann physics trying to make the whole bill disappear.

Inspired by Milan Patel's water-bottle bit on Don't Tell Comedy. Numbers as of Aug 2026 — receipts at the end.

▶ Watch the Institute host perform it

Transformer LLMs on GPUs (today) Energy-based models on post-von-Neumann silicon Comedy / baseline
scroll ↓

Cold open · the premise

Every question you ask a chatbot drinks a full 500 mL bottle of water. Picture a person doing that at your desk, one bottle per question, and you would have words for them.

the reenactment
hey, is oat milk actually better for the planet?
assistant is thinking (and hydrating)
1 bottle / question
0:00thinking…
(time-lapse — we don't have the budget to make you actually wait)
Yeah, they're pretty good for you.
ok now say it again but like a pirate
assistant is reaching for a second bottle…

That's the premise. It kills, because it feels true.
Now let's audit the premise. Bring your receipts and a poncho.

Act I

DRINK 💧

First, the fact-check the comedian never asked for: per question, today's frontier chatbots sip about a quarter of a milliliter — not a bottle. The bit is off by roughly 2,000×. Comedy has never owed anyone error bars.

Gemini · median text prompt
water per prompt, direct
0.26mL
Google's own disclosure (Aug 2025) ✓ measured*
ChatGPT · average query
water per query — "a fifteenth of a teaspoon"
0.32mL
Sam Altman's blog, methodology: vibes △ vendor
GPT-3 era · full lifecycle
answers per 500 mL bottle
10–50
UC Riverside — includes the water behind the electricity ✓ academic est.
Training GPT-3 · once
freshwater evaporated
700,000L
UC Riverside — the bar tab before the first question ✓ academic est.

*"Measured" with an asterisk you could park a truck under: Google counts direct cooling water only — not the water evaporated making the electricity, which the UC Riverside team found is most of it. The bottle is smaller than the bit says; it is also not the whole bottle.

How far does one 500 mL bottle actually go?

Prompts answered per bottle of water, by accounting method. Log scale — every gridline is a 10× jump.

prompts per bottle · log scale · hover a row for the fine print

0.0 bottles sipped by the world's chatbots since you started scrolling
napkin math: ~2.5B ChatGPT prompts/day × 0.32 mL ≈ 18.5 bottles/sec 🍼

So no — your one question is not a bottle. But two and a half billion questions a day is ~800,000 liters of direct water. Daily. The comedian was wrong about the sip and right about the drinking problem.

Act II

THINK 🧠

Why does thinking cost anything at all? Because of a 1945 floor plan. A von Neumann machine keeps memory over here and compute over there — so a transformer answers you by marching essentially every weight in the model across a bus, for every single token. Your chatbot is not thinking. It's commuting.

MEMORY all the weights live here COMPUTE the math happens here round trip × billions of weights × every token you generate "…are chia seeds good for me?" → the entire library walks across the chip. Twice.
Share of AI runtime energy spent moving data, not doing math
~90%
IBM Research: "compute energy is still around 10% of modern AI workloads" ✓ measured
Energy to fetch one number from DRAM vs multiplying it
~170× more
~640 pJ per 32-bit DRAM access vs ~3.7 pJ per FP32 multiply — Horowitz, ISSCC 2014 ✓ measured
The math is the cheap part. The commute is the bill.

And that 37-second pause in the bit? On stage it's a man chugging water. In silicon it's DRAM round-trips. Either way: the pause is the expensive part.

Now the flip. An energy-based model doesn't march weights around to compute your answer. It learns an energy landscape where good answers sit in the valleys — then asks physics to roll downhill. On a GPU you have to simulate that rolling (expensive). On a thermodynamic sampling unit, the rolling is literal: the circuit's own thermal noise does the sampling, natively, in analog.

a fine answer the answer energy ↑ = wrong · valleys = right · noise = the search algorithm

Thermal jiggle isn't a bug to cool away — it's the compute. GPUs spend real megawatts generating fake randomness; a TSU is made of real randomness and just… uses it.

Act III

BURN 🔥

Zoom out from your one little prompt and the numbers stop being cute. Data centres are on track to double their electricity use by 2030 — to roughly Japan's entire annual consumption — with AI-focused facilities tripling.

Data centre electricity, worldwide

TWh per year · IEA (2026). Single series; the 2030 point is the IEA base-case projection.

hover for values · projection striped, because the future hasn't been measured yet

US electricity already going to data centres (2023)
4.4%
Lawrence Berkeley National Lab ✓ measured
LBNL's range for 2028
6.7–12%
Same report — the wide bars are doing a lot of work ✓ projection
ChatGPT's day, at ~2.5B prompts × 0.34 Wh
~850MWh/day
≈ a day of electricity for ~29,000 US homes ✎ napkin

So here's the actual headline: a wave of post-von-Neumann hardware wants to delete the commute instead of widening the road. The full tour, honesty gauges included:

runs energy-based models
Extropic · thermodynamic sampling units
10,000× claimed
Probabilistic bits in bog-standard CMOS sample distributions using their own thermal noise. Roadmap X0 → Z1 → Z1.5; a $75M letter of intent with the US Dept. of Commerce (Jul 2026) to build them at an American foundry. Doesn't even need cutting-edge wafers.
△ claimed — simulation-backed, by people selling the chip
They would also like you to know the Second Law of Thermodynamics is on their side. It has never lost.
runs conventional nets
IBM · NorthPole
73× measured
Memory woven directly into compute — near-memory architecture. 73× more energy-efficient (and 47× faster) than the comparison GPUs on 3B-parameter LLM inference.
✓ measured — peer-reviewed lineage
The weights stopped commuting. Productivity soared.
runs spiking nets
Intel · Loihi 2 / Hala Point
~100× on sparse loads
1.15 billion silicon neurons that only fire when something happens — the brain's oldest efficiency hack. Up to 15 TOPS/W; shines on sparse, event-driven workloads.
△ vendor-measured
Your brain runs on ~20 watts and still remembers embarrassing things from 2009. Efficiency is possible.
runs matmuls
Photonics · Lightmatter et al.
10×-ish, lab→fab
Matrix multiplication performed by interference of light itself. No electrons harmed. Interconnects shipping first; full photonic compute still graduating.
○ lab stage
Finally, a computer where "speed of light" is the spec sheet, not the marketing.

Claimed energy-efficiency multipliers vs a GPU baseline

Log scale. One of these numbers is measured on real silicon; one is a prophecy with a foundry deal. The badges say which.

× more energy-efficient than GPU baseline · log scale · hover a row for provenance

Finale

The bit, re-run on better physics

Same heckler, same question, two machines. On the left, today. On the right, the claimed post-transformer future — an energy-based model relaxing on a thermodynamic chip.

transformer · GPU · today
are chia seeds good for me?
0:00thinking…
Yeah, they're pretty good for you.
energy: ~0.34 Wh ≈ the calories in ~35 chia seeds · water: ~0.32 mL direct
energy-based model · TSU · claimed
are chia seeds good for me?
Good for you, chia seeds are. Hmm. Yes.
already did the Yoda one. unprompted. it was cheaper than not doing it
energy: ~0.00003 Wh ≈ 1⁄300th of one chia seed · water: a bead of fog (claimed)
Cost of asking whether chia seeds are good for you, in chia seeds
~35seeds
transformer on GPUs · 0.34 Wh vs ~6.8 mWh per seed ✎ napkin
Same question on a thermodynamic chip — one seed now funds
~300answers
at the claimed 10,000× · the question would finally answer itself △ claimed
Disclaimers, honestly delivered: the 10,000× is a simulation-backed claim from people who profit if it's true (though the US Dept. of Commerce liked it enough to sign a $75M letter of intent). The 73× is measured, peer-reviewed-adjacent, and real silicon. The 0.26 mL is a tech giant grading its own homework and skipping the thirstiest subject (the water behind the electricity). And the comedian's 500 mL per question was off by three orders of magnitude — yet he still made the truest point in this entire document: nobody should need 40 seconds and a bottle of water to find out that chia seeds are fine. No plastic was single-used in the making of this infographic. Someone is going to pee in the error bars later, so they're not technically single-use either.

Encore

The receipts 🧾

Every number above, its epistemic status, and where it came from. This table is the accessible twin of every chart on this page.

NumberWhat it isStatusSource
0.24 Wh · 0.26 mL · 0.03 gCO₂eMedian Gemini text prompt: energy, direct cooling water, emissions ("less TV than 9 seconds")self-reported measurement; excludes training, networks & indirect waterGoogle (Aug 2025), DCD
0.34 Wh · 0.32 mLAverage ChatGPT query — "about one-fifteenth of a teaspoon"vendor blog post, no methodology publishedAltman (Jun 2025)
500 mL per 10–50 answersGPT-3-era water per medium-length responses, lifecycle incl. water behind electricity, varies by site & seasonacademic estimateUC Riverside, "Making AI Less Thirsty"
700,000 LFreshwater evaporated training GPT-3, on-site estimateacademic estimatesame paper
~1,900 / bottleGemini prompts per 500 mL bottle (500 ÷ 0.26), direct water onlynapkin, from Google's figurederivation shown
~2.5 B prompts/dayChatGPT daily volume, mid-2025 — powering the ticker (~18.5 bottles/sec, ~800 kL/day, ~850 MWh/day ≈ 29,000 US homes)vendor figure + napkin chainOpenAI via press
~90% / ~10%Share of AI runtime energy spent on data movement vs actual computemeasured, industry researchIBM Research
~640 pJ vs 3.7 pJ32-bit DRAM fetch vs FP32 multiply (45 nm) — the ~170× commute taxmeasured, canonicalHorowitz, ISSCC 2014
485 → ~950 TWhGlobal data centre electricity 2025 → 2030 (≈3% of global demand; ≈ Japan's annual use; AI-focused DCs tripling; +17% in 2025)IEA base caseIEA (2026)
4.4% → 6.7–12%US electricity share to data centres, 2023 → 2028 rangenational-lab analysisLBNL (Dec 2024)
73× · 47×IBM NorthPole vs comparison GPUs on 3B-param LLM inference: energy efficiency · speedmeasured on real siliconIBM Research
1.15 B neurons · 15 TOPS/W · ~100×Intel Hala Point (Loihi 2) neuromorphic system; advantage on sparse, event-driven workloadsvendor-measuredIntel (2024)
10,000×Extropic's claimed energy advantage for denoising thermodynamic models (EBMs) on TSUs vs GPU samplingclaim — simulation-backed, pre-production siliconExtropic, VKTR
$75 MNon-binding letter of intent, US Dept. of Commerce CHIPS R&D office → onshore TSU fabrication (X0 → Z1 → Z1.5)announced Jul 29, 2026Extropic
~35 seeds · ~1⁄300thChia-seed cost of one answer: one seed ≈ 1.2 mg ≈ 5.8 cal ≈ 6.8 mWh; 0.34 Wh ÷ 6.8 mWh ≈ 35; claimed EBM 0.034 mWh ≈ seed⁄300napkin, USDA calorie dataderivation shown