Token Economy
Model routing and token economics — for builders who ship on a budget.
Every AI-assisted build is really a stream of very different jobs — planning, hard reasoning, routine construction, mechanical edits, and review — and the current model landscape prices each of them wildly differently. A flagship reasoning model can cost 50–100× a lightweight one per token; subscriptions meter you in compute-hours, messages, or raw tokens depending on the vendor; and the quieter levers — prompt caching (~90% off repeated input), batch discounts, reasoning-effort budgets, and per-model context windows — often decide your bill more than the model name does.
Token Economy is the practical playbook for navigating all of it. Get the routing right and the savings aren't marginal: teams routinely cut spend by half or more while shipping the same work, because most of a session is cheap work that never needed a premium model. Each part maps model-and-effort to task type across Anthropic (Fable / Opus / Sonnet), OpenAI (GPT-5.6 Sol / Terra / Luna) and the fast-moving challengers (GLM, Kimi, DeepSeek, Qwen), grounded in current, sourced pricing we refresh weekly. It's written for solo builders on a budget — and for the teams who'll turn it into policy — and it doubles as the thinking behind a forthcoming background routing harness that automates the whole discipline. Read the 8 parts in order, or jump to whatever you need.
Why One Model for Everything Is the Most Expensive Habit in AI Coding
Reading the Meter: How Anthropic, OpenAI, and the API Actually Bill You
Effort Levels Are a Budget Dial: When to Turn Reasoning Down
Prompt Caching Is Free Money Most Builders Leave on the Table
Planning Models vs Coding Models: Don't Write Architecture With Your Code Model
Build-a-Budget: A Personal Routing Policy for Solo Devs on Pro or Plus
Scaling Routing to a Team: Shared Policy, Guardrails, and Spend Governance