AI Provider Cost Updates 2026-07-31 | Token Economics News

Anthropic

Claude Code Now Lets You Cap Subagents and Actually Enforce Your Budget

Anthropic just shipped a session budget enforcement update in Claude Code v2.1.217 that finally puts real teeth behind spend limits. The release adds a cap on how many subagents can run concurrently, which stops runaway parallel workloads from quietly draining your budget.

On top of that, the max-budget behavior got a proper fix: once you hit your spend limit, new subagent spawns get denied outright, and any background agents still running get halted automatically. For anyone responsible for keeping AI spend in check, this is a big deal — it means you can set a ceiling and actually trust that Claude Code will respect it, instead of finding out after the fact that costs blew past your limit.

Bedrock Billing Finally Matches the Model You’re Actually Using

Claude Code v2.1.218 fixes a gateway spend metering bug that was causing incorrect billing for certain Bedrock setups. Specifically, application-inference-profile ARNs and other mapped upstream model IDs are now billed at the rates of the model you actually configured, not some mismatched default.

This might sound like a small technical fix, but for FinOps teams doing chargeback or cost reconciliation, it’s huge. Inaccurate metering means inaccurate invoices and messy internal cost allocation. With this fix, the numbers you see should now line up with what you’re really spending, making budget tracking and cross-team billing far more trustworthy.

Claude Opus 5 Lands as the New Default, With Clear Pricing to Match

Claude Code v2.1.219 introduces Claude Opus 5 as the default Opus model, bringing a 1M token context window along with fast mode pricing set at $10 per million input tokens and $50 per million output tokens.

Beyond just a model swap, this release bundles in workflow and budget control improvements too. Having clear, upfront pricing for the new default model makes capacity planning and cost forecasting much easier — you know exactly what you’re paying before you scale up usage, and the expanded context window gives you more headroom for complex tasks without needing to break work into smaller, more expensive calls.

Google Cloud

Gemini 3.6 Flash Goes Global, But Get Ready to Wave Goodbye to 3.5 Flash

Gemini 3.6 Flash is now available in the global region within Gemini Enterprise, while Gemini 3.5 Flash has been scheduled for removal from that same region.

This kind of model consolidation matters a lot for anyone planning budgets or usage across large user bases. Wider global availability of the newer model means more consistent performance and pricing wherever your teams or customers are located, but the looming removal of 3.5 Flash also means it’s time to start migration planning now rather than scrambling later. If you’re relying on 3.5 Flash in the global region, this is your heads-up to map out a transition path before it disappears.

See What Gemini Is “Thinking” in Real Time, Now Generally Available

Transparent thinking has reached general availability in Gemini Enterprise, giving you a live view of the model’s reasoning and tool activity as it works.

Alongside the visibility boost, this feature also improves time to first token, which translates to snappier perceived performance for end users. For teams running AI workloads at scale, better observability into what the model is doing — combined with faster response times — can help with both debugging costly reasoning chains and optimizing latency-sensitive applications without needing to guess what’s happening under the hood.

OpenAI

GPT-5.6 Luna is now 80% less expensive

 Luna now costs $0.20 per million input tokens and $1.20 per million output tokens.

GPT-5.6 Terra is now 20% less expensive

 now costs $2 per million input tokens and $12 per million output tokens.

Fast mode for GPT 5.6 Sol

Up to 2.5x faster responses.  introducing Fast mode in the API, which replaces our Priority Processing offering. For GPT‑5.6 Sol, Fast mode now delivers up to 2.5× faster speeds than Standard processing at twice the price, with no change in intelligence. Fast mode is backward compatible: requests tagged priority will automatically use Fast mode. For new requests, pass service_tier=”fast”.

FinOps Weekly
FinOps Weekly
Articles: 219