steve.net
>_ terminal

Cost Series 01 — The Bill Is Not the Token Count

Working notes for a planned episode — an outline, not an article. Details will shift when it gets scripted. Series index: The Bill.

The thesis episode: why "minimize tokens" optimizes the wrong variable, and what to optimize instead.

The hook

Everyone's first instinct for a big LLM bill is "use fewer tokens." That instinct is now wrong — or at least, it's aiming at the cheapest thing on the invoice. Caching and model tiering quietly broke the rule that tokens ≈ cost, and once you see the real cost function you'll sometimes send more tokens to pay less.

What the viewer walks away with

Beats

Demo ideas

Source material

Open questions