AI Provider Cost Updates 2026-08-21 | AI Economics News
Anthropic
Attribute gateway API spend to individual users
Claude Code now helps teams connect gateway usage to the people generating it. An opt-in gateway setting forwards signed-in user identity, allowing organizations to attribute upstream API spend to individual users.
Claude Code also surfaces gateway spend-limit details when usage reaches an enforced cap. Users can see the limit and its reset time, giving cost owners clearer information for spend tracking and limit management.
Restore prompt caching for gateway-routed Claude Code sessions
Claude Code v2.1.237 fixes prompt caching when sessions use an LLM gateway or custom base URL. Restored caching reduces repeated input-token processing for those sessions.
That can improve effective cost for workloads routed through gateways or custom endpoints, while also helping repeated sessions avoid processing the same input unnecessarily.
Cut Claude Code’s built-in API skill context overhead
Claude Code now loads API reference documentation only when it’s needed. This reduces the context cost of the built-in claude-api skill from more than 200,000 tokens to approximately 25,000.
The lower prompt overhead can improve cost efficiency for repeated sessions. It also reduces the amount of context loaded before the reference documentation is required.
Reuse prompt caches across forked and parallel agents
Claude Code v2.1.232 improves prompt-cache reuse for multi-agent workflows. Forked subagents now inherit the full conversation and prompt cache, and background execution is enabled by default for non-teammate agent spawns.
Reusing cached prefixes can reduce duplicate input processing and improve throughput. For teams managing multi-agent usage, that can make repeated workflows more efficient.
OpenAI
See estimated lifetime credit usage for each Codex chat
OpenAI now shows thread-level cost data for eligible Codex conversations. Enterprise workspaces with credit-based billing can view estimated lifetime credit usage for individual Codex chats, including usage in ChatGPT Desktop and the Usage & billing settings.
The figures are intended as a planning aid rather than an invoice. They give cost owners better workload-level visibility when reviewing how credits are being used.
Speed up GPT-5.6 Sol workloads with Ultrafast mode
OpenAI introduced Ultrafast mode for GPT-5.6 Sol in limited preview. The API service tier delivers up to 14 times the speed of Standard processing for selected customers.
The higher-throughput tier may improve efficiency for latency-sensitive production workloads. Teams evaluating it will need to consider its pricing and capacity terms alongside the speed benefit.
Plan Gemini Enterprise seats with new subscription caps
Gemini Enterprise now has clearer seat limits for subscription planning. Direct seat-based subscriptions support up to 25 seats for self-serve or resold accounts, while invoiced accounts can have up to 1,000 seats.
These limits affect procurement, subscription planning, and license governance. Teams can use the caps to assess which subscription approach fits their organization’s size and purchasing process.
Track Gemini developer usage with new monitoring dashboards
Gemini Enterprise’s AI developer tools now include usage metrics for better cost visibility. The generally available capabilities cover developer adoption, active users, token consumption, and API call volumes.
The metrics are available through Cloud Monitoring and logging. That gives teams responsible for provider costs more visibility into utilization, workload governance, and opportunities for cost optimization.
Add Gemini 3.7 Flash as a regional model option
Gemini 3.7 Flash is now generally available across global, US, and EU regions. The model can also be used by Agent Designer workflow agents.
Its availability across multiple regions gives organizations another model option when deciding where to place workloads. It also adds flexibility for teams evaluating workload efficiency across regions.
Use Gemini 3.6 Flash in US and EU multi-regions
Gemini 3.6 Flash is now generally available in Google’s US and EU multi-regions. Allowlist requirements have been removed for the US multi-region.
The new Flash model type gives organizations another potentially efficient option for high-volume workloads. That can help cost teams evaluate model choices alongside workload volume and regional placement.
Run large-scale code optimization with AlphaEvolve
Gemini Enterprise adds a distributed AlphaEvolve solution for large optimization experiments. The containerized, high-performance computing solution runs evolutionary code-optimization workloads on Google Cloud.
It can improve resource utilization and support experiments that exceed the limits of a single machine. This is useful for teams managing the resources and efficiency of large optimization workloads.