Tokenomics by Design: How the 4M Framework Turns AI Cost Chaos into Discipline
Here is a pattern that should worry any finance leader tracking AI spend: the price of a single token keeps falling, year after year, and yet the monthly invoice keeps climbing regardless. That is not a contradiction; it is the defining economic puzzle of enterprise AI right now, and understanding why it happens is the first step towards actually fixing it.

Unit price and total enterprise AI spend increasingly move in opposite directions.
Why Falling Prices Haven’t Meant Falling Bills
Independent benchmarking shows the cost of processing a million tokens has fallen by well over three hundred times since 2020, with further declines widely expected through the end of the decade. If unit price told the whole story, AI budgets should be shrinking in step. They are not, because demand for intelligence does not merely keep pace with falling prices; it accelerates. Multi-step agentic workflows, in particular, can consume five to thirty times more tokens per completed task than a single chatbot exchange, and that multiplier is exactly what has blown through carefully planned budgets at organisations that assumed unit-price trends would translate directly into total-spend trends. Analyst firms across the board now agree that this is an architecture and governance problem, not something that will simply resolve itself as prices continue to fall.
Introducing the 4M Framework
To help enterprises close that gap deliberately rather than by accident, a structured approach has emerged that organises the available cost levers into four pillars, each building on the one before it: Measure, Minimize, Match, and Memoize.

The 4M Framework is designed as a continuous, cyclical discipline rather than a one-time project.
Measure: You Cannot Optimise What You Cannot See
Every durable cost discipline starts with instrumentation, and token spend is no exception. A workable measurement programme tracks input and output token counts separately, the specific model and provider handling each request, latency and retry behaviour, and a business-context tag identifying which team or feature generated the request. Surveys of enterprise AI spend consistently find that a large share of organisations misjudge their own AI budget by a wide margin, largely because they are still applying infrastructure-era metrics, such as licence seats, to a workload that bills by consumption instead. The goal is not to minimise cost per token in isolation; it is to understand cost per successful outcome, since a cheap model that needs three retries can easily end up costing more than a pricier model that gets the answer right the first time.
Minimize: Cut What Never Needed to Be There
Once visibility exists, the fastest wins usually come from simply trimming what should never have been sent or generated in the first place. Because output tokens are typically priced several times higher than input tokens, controlling response length is often the single highest-leverage change available to a team that has not yet touched its architecture at all. Bloated system prompts, retrieved documents nobody actually needed, and unconstrained free-form output all add up quickly, and pruning them, alongside setting explicit output limits and shifting non-urgent work to asynchronous batch processing, can produce savings within an afternoon rather than a quarter.
Match: Stop Sending Every Request to the Most Expensive Model
Perhaps no single habit inflates enterprise AI bills more reliably than routing every request, regardless of difficulty, to the newest and most capable model on offer. The Match pillar treats model selection as a decision made at runtime rather than once at the start of a project. A routing layer that classifies incoming requests by complexity and sends routine work to smaller, cheaper models, reserving frontier-tier capability for genuinely hard cases, has been shown to cut costs substantially while preserving the bulk of the quality a flagship model would have delivered, because the requests that most need that capability are still the ones receiving it.
Memoize: Stop Paying for the Same Answer Twice
A great deal of real-world AI traffic is not actually new: the same system instructions, the same reference documents, and even differently worded versions of the same underlying question recur constantly. Prompt caching addresses the first category, storing the computed representation of a stable prompt prefix so it never needs reprocessing; several major providers report cost reductions of up to ninety percent on cached tokens. Semantic caching goes further, recognising when two differently phrased questions are asking the same thing and serving a cached answer accordingly. Used together, caching and complexity-aware routing compound rather than simply add, because each is chipping away at a different source of waste.
Making It Stick
None of these four pillars survives on technical merit alone. They need a governance layer around them: budget guardrails, cross-functional ownership that pairs engineering decisions with financial accountability, and show-back or chargeback practices that make each team genuinely responsible for the cost its own usage generates. Organisations that treat this as a phased rollout, tackling visibility and quick wins first before layering in caching, routing, and full automation, consistently outperform those that try to do everything at once and end up stalling on all four fronts simultaneously.
Conclusion
The 4M Framework does not promise a silver bullet, and it should not; anyone offering one has not looked closely enough at how AI workloads actually behave in production. What it offers instead is a repeatable discipline: measure honestly, cut what is unnecessary, match effort to difficulty, and never pay twice for the same piece of work. Applied consistently, that discipline is what separates organisations that scale AI sustainably from those left explaining, quarter after quarter, why the bill keeps outrunning the plan.
References
- FinOps Foundation – FinOps for AI Overview – https://www.finops.org/wg/finops-for-ai-overview/
- Anthropic – Prompt Caching Documentation – https://platform.claude.com/docs/en/build-with-claude/prompt-caching
- Amazon Bedrock – Prompt Caching – https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html
- Amazon Bedrock – Intelligent Prompt Routing – https://aws.amazon.com/bedrock/intelligent-prompt-routing/
- FinOps Foundation – Framework Maturity Model – https://www.finops.org/framework/maturity/