AI Provider Cost Updates 2026-09-25 | AI Economics News

Anthropic

Opus 5.5 brings clear token prices to Claude Code

Claude Code now defaults to Opus 5.5, with published rates to help you plan model costs.

The model has a 1-million-token context window. Its listed rates are $4 per million input tokens, $20 per million output tokens, and $0.20 per million cache-read tokens.

That gives teams concrete rates to use when estimating workload costs and comparing input, output, and cached-token usage.

Auto mode can avoid billed classifier overhead

Claude Code now uses a server-side classifier by default in supported configurations.

The change applies to supported API, Enterprise, Bedrock, Vertex, Foundry, and gateway configurations. The server-side classifier doesn’t charge for classifier overhead, unlike the local fallback.

For teams using auto mode, this can avoid those billed overhead costs where the server-side classifier is supported.

Slack cost and token totals are more reliable after worker restarts

Claude Code fixed a bug that could inflate cost and token totals in Slack reply footers.

The totals could appear many times too high after a cloud worker restarted. The fix improves usage and spend visibility for affected sessions, giving teams more dependable figures to review.

Add labels to Claude telemetry across users or environments

Claude apps gateway now supports configurable resource attributes for telemetry.

Teams can add fixed labels to telemetry from Claude Desktop and login sessions. Those labels can help organize observability data and attribute it across users or environments.

Google Cloud

Check Gemini Code Assist access before renewing subscriptions

Google Cloud changed which subscriptions include Gemini Code Assist.

New and renewing Gemini Enterprise Standard and Plus subscriptions no longer include Gemini Code Assist features.

This entitlement change is worth factoring into developer-tooling cost and subscription planning.

Gemini 3.8 Flash is now the default in Gemini Enterprise

Gemini 3.8 Flash is generally available and enabled by default in the Gemini Enterprise app.

The rollout covers the global, US, and EU regions. For teams using the app, the default model has changed, so it’s useful to account for that when reviewing model usage.

Bring D&B Risk Analytics into Gemini Enterprise workflows

Gemini Enterprise now offers the D&B Risk Analytics data store.

The data store is generally available and enables natural-language workflows for third-party and counterparty risk assessment. Teams handling those workflows can use Gemini Enterprise to work with this risk data through natural-language interactions.

OpenAI

Get more value from GPT-6 prompt caching

OpenAI added higher cache hit rates, explicit breakpoints, diagnostics, and controls for GPT-6 prompt caching.

These updates give teams more ways to manage and inspect caching behavior. Better cache hit rates can reduce token spend and latency, making this a practical update for teams looking to optimize GPT-6 workloads.

Choose GPT-6 models with lower token prices for Work and Codex

GPT-6 Sol and GPT-6 Luna are available for ChatGPT Work and Codex at lower token prices than their GPT-5.6 predecessors.

Sol is aimed at more complex coding and agentic work, while Luna is designed for high-volume tasks. Teams can match workloads to these model options and compare costs against the previous GPT-5.6 models.

Ringg reports 90% lower costs with GPT-5.6

Ringg says its multilingual agents cost 90% less with GPT-5.6 than with GPT-4.1.

Ringg uses GPT-5.6 for agents across voice, chat, WhatsApp, and web. The company also says its agents resolve up to 65% of customer calls.

For teams evaluating model economics, this offers a concrete example of reported cost differences across models and a range of customer-service channels.

GPT-6 Astra helped cut research time and cost in half

Parallel used GPT-6 Astra to cut labor-market research time and cost by half.

OpenAI reports that Parallel used the model to research and synthesize labor-market data, completing the work in half the time and at half the cost compared with its prior models.

It’s a useful workload-efficiency example for teams assessing the cost and time impact of a model change.

Plan ChatGPT for Word usage with model-based token billing

OpenAI launched ChatGPT for Word with token-based usage for Business, Enterprise, and Edu.

For those plans, usage is billed at the API rates for the selected model. That gives teams a clear billing basis to consider when planning use of this new productivity surface.

Compare GPT-6 Sol and Luna by capability and cost

OpenAI introduced GPT-6 Sol and GPT-6 Luna as frontier models with different balances of capability and cost.

The models offer another choice for teams weighing model capability against cost. That distinction can help inform model selection for different workloads.

FinOps Weekly
FinOps Weekly
Articles: 259