~/steve.net — topics/ai/fleet-series/09-llm-router · theme: Commodore 64

Fleet Series 09 — The LLM Router: Local Models as a Utility

Working notes for a planned episode — an outline, not an article. Details will shift when it gets scripted. Series index: Building the Fleet.

Deep dive on the LLM router: the homelab model fleet behind one OpenAI-compatible endpoint — the power supply the rest of the fleet plugs into.

The hook

Every agent in the fleet talks to "the model" through one URL on the local network. Behind that URL: a router, a rack of local models, aliases, and auto-routing. Swap a model, nothing upstream changes. That one seam is what makes an always-on agent fleet economically sane.

What the viewer walks away with

Beats

Demo ideas

Source material

Open questions

Borrowed line — cost acceleration

The loop-engineering crash course makes the Jevons-paradox point crisply: better loops + cheaper tokens = more runs = faster burn. Cheap inference doesn't shrink the bill — it changes what you dare to run in a loop.

theme: apple synthwave amber c64 arcade tek
steve.net · page: topics/ai/fleet-series/09-llm-router
? keys try: curl steve.net or: ssh steve.net or: shell mode or: reading view