GCP Updates 2026 January – April | Cloud Provider News

April 23, 2026

GKE workload right-sizing recommendations are in the FinOps hub

Google Cloud added GKE workload right-sizing and optimization recommendations into the FinOps hub — view the release notes.

This brings GKE workload recommendations (right-sizing and optimization) into the central FinOps hub so teams can find cluster optimization insights alongside other cost signals.

Also, centralizing these recommendations helps reduce overprovisioning across clusters and environments, making it easier to act on container-level savings opportunities.

Gemini Cloud Assist and Optimization page now explain cost changes

The Cloud Hub Optimization page and Gemini Cloud Assist chat panel now provide explanations for recent cost changes — read the notes.

Gemini Cloud Assist can now surface contextual explanations for recent cost fluctuations and tie them to resource usage, helping teams quickly pinpoint drivers behind bill changes.

Furthermore, those explanations speed troubleshooting for unexpected charges and reduce time-to-answer for FinOps teams who need to justify or remediate cost spikes.

Application Monitoring surfaces AI resource metrics

Cloud Monitoring’s Application Monitoring now shows AI resource performance metrics, including token usage and error rates — check the release notes.

Application Monitoring added visibility into AI-specific metrics so teams can see token usage and AI error rates alongside traditional app metrics.

Consequently, you can monitor AI consumption and performance in the same dashboards you use for apps, making it easier to correlate AI spend with app behavior.

April 16, 2026

BigQuery’s optimized mode for managed AI functions cuts LLM token use and latency

BigQuery introduced an optimized mode for managed AI functions (for AI.IF, AI.CLASSIFY, AI.SCORE) that reduces LLM token consumption and query latency on large datasets.

This optimized mode lowers model inference costs by reducing token use while also speeding up queries used for generative or AI-enhanced analytics.

As a result, organizations using BigQuery for LLM-augmented analytics can reduce per-query cost and improve turnaround time for large-scale AI queries.

Hyperdisk ML disks are GA across multiple machine series for very-high throughput

Hyperdisk ML disks are generally available for A3 Ultra, C4D, N4, and N4A machine series, offering very-high throughput (up to 2 TiB/s) storage for ML training workloads.

Those disks are designed for large-scale training and can dramatically reduce job runtime by delivering much higher I/O throughput.

Consequently, faster training runs can lower total compute time and associated cost, especially for large foundation-model work.

Cloud Monitoring topology map gives a dynamic app-level view to prioritize fixes

Cloud Monitoring can now show a single dynamic topology map of App Hub applications and registered services with error rates and P95 latencies between services.

That unified view highlights open incidents and latency hotspots so SREs can triage and prioritize remediation more quickly.

Therefore, teams can avoid unnecessary overprovisioning by using the map to guide targeted fixes and capacity planning.

April 2, 2026

Cloud Billing launches scenario modeling for CUD recommendations to help you simulate savings

Google Cloud made scenario modeling for Committed Use Discount (CUD) recommendations generally available.

The feature lets customers simulate spend-based and resource-based CUDs to model purchase scenarios that maximize savings. So FinOps teams can test different commitment sizes and durations to find cost-optimal commitments before committing budget.

TPU7x (Ironwood) is GA

Google Cloud announced Cloud TPU TPU7x (Ironwood) generally available.

TPU7x is a next-gen TPU for large-scale AI training and inference that provides performance improvements suitable for demanding ML workloads.

March 27, 2026

Monitoring telemetry quotas are now regional and byte/request-based, plan observability costs per region

Google Cloud updated Cloud Monitoring telemetry quotas to regional metric-ingestion limits and byte/request-based quotas (e.g., up to 60,000 metric-ingestion requests per minute per region).

The quota model moved from a single global quota to per-region limits measured in bytes and requests.

Additionally, this change affects monitoring scale planning for multi-region deployments since quotas and ingestion costs now need regional consideration.

March 19, 2026

BigQuery can auto‑deploy open models to Vertex AI with cost controls and immediate undeploy

BigQuery ML can automatically deploy open models to Vertex AI endpoints with automatic resource management, reservation affinity (Compute Engine reservations), and automatic/immediate undeployment to save costs.

The flow provides managed serving with affinity to reservations and the ability to undeploy endpoints automatically, which helps control serving spend.

Additionally, automatic undeployment and resource management reduce the risk of forgotten endpoints driving ongoing costs.

BigQuery adds a query execution heatmap to spotlight slot‑time hotspots

BigQuery introduced a visual heatmap that maps SQL query text to the execution graph and highlights steps consuming the most slot‑time.

The heatmap helps you find which query stages are using the most slots so you can optimize those steps to reduce consumed slot‑time.

Moreover, faster identification of hot spots accelerates query tuning and can meaningfully lower interactive and analysis costs by reducing slot usage.

March 12, 2026

Conversational Analytics: agent query cost labels and partition optimizations, better cost attribution for semantic analytics

BigQuery conversational analytics now labels agent-generated queries in job history and improves partition support.

The update tags queries generated by agents so you can identify and attribute those jobs in job history, helping you separate agent-driven cost from user or batch queries.

Additionally, it improves handling for partitioned tables and BigQuery ML functions to reduce query costs and speed up production semantic analytics. This helps you tune partitioning strategies that directly lower query costs.

SecOps: Data Processing Pipelines (Preview) to filter and redact before ingestion, cut SecOps ingestion costs

Google SecOps introduced Data Processing Pipelines in Preview to filter, transform, and redact data before ingestion.

These pipelines let you drop unwanted events and redact sensitive fields before they’re ingested into SecOps (and SecOps SIEM), so you’re not paying for storage and processing on data you don’t need.

Moreover, by reducing the volume and sensitivity of ingested telemetry, teams lower both storage and downstream processing costs while meeting data governance needs.

VECTOR_SEARCH alternate syntax (Preview), faster single-vector queries, lower vector search cost

BigQuery added an alternate VECTOR_SEARCH syntax (Preview) to speed up single-vector queries.

The new syntax is optimized for single-vector searches, improving query performance compared with the older form, which can reduce query runtime and cost for common single-query patterns.

Additionally, faster queries mean less slot time and lower per-query charges for vector workloads at scale.

Cloud Storage Rapid Bucket (zonal buckets) Generally Available, zone-local buckets for better locality and lower egress

Cloud Storage announced Rapid Bucket GA (zonal buckets) to place storage in a zone for better I/O and locality.

Zonal (Rapid) Buckets give you storage placed in a single zone to optimize I/O and data locality for AI/ML and high-scale analytics workloads, minimizing cross-region egress.

Moreover, by keeping storage and compute in the same zone you can improve performance and avoid cross-region transfer costs when your compute runs in that zone.

March 5, 2026

BigQuery GA: monitor cross-region replication latency and egress bytes in Cloud Monitoring

BigQuery added GA metrics for cross-region replication latency and network egress bytes in Cloud Monitoring.

These metrics give visibility into replication performance and egress volumes so teams can find high-cost transfers and optimize data placement.

Compute Engine: managed compute.managed._ constraints for org-wide VM policy enforcement (GA)

Compute Engine introduced managed replacement constraints (compute.managed._) for Organization Policy Service, with Policy Simulator and dry run support.

That enables centralized enforcement to limit VM types and configurations across an organization to control cloud spend and compliance.

Compact placement policies GA for Flex-start VMs, colocate VMs for lower cross-node cost

Google Cloud made compact placement policies generally available for Flex-start VMs (AI Hypercomputer / Compute Engine).

These policies let you colocate VMs to minimize network hops and improve latency-sensitive AI/ML workload efficiency.

You can reduce cross-node communication overhead and improve utilization, which helps lower operational and networking costs for distributed workloads.

February 19, 2026

Carbon Footprint methodology update: AI inference emissions allocated at SKU level

See Google Cloud’s Carbon Footprint methodology update (effective Jan 2026) for AI inference allocation.

Google updated Carbon Footprint calculations to allocate AI inference emissions at the SKU level following the AI energy/emissions framework, starting Jan 2026.

And this change increases reported emissions for AI-powered services (for example, Vertex AI) while improving transparency for sustainability reporting.

Compute Engine instance flexibility for bulk VM creation (GA) to reduce failed launches

Read about instance flexibility becoming generally available for bulk VM creation.

Compute Engine made instance flexibility generally available for bulk VM creation, letting you provide a list of acceptable machine types and allowing the platform to provision based on capacity and quota.

And that reduces failed launches and manual rework when provisioning many VMs at once, improving large-scale provisioning resiliency.

Hyperdisk Exapools GA for massive pooled block storage and predictable performance

Learn about Hyperdisk Exapools GA for very large block storage needs.

Google announced Hyperdisk Exapools generally available, enabling purchase of bulk storage+performance from 500 TiB up to 5 EiB and sharing across up to 500,000 disks.

And this pooled model supports massive AI/ML workloads with predictable performance and the potential for unit-cost savings for very large storage users.

February 13, 2026

Compute flexible CUDs expanded to all Cloud Billing accounts to simplify committed discounts

Google Cloud automatically migrated Cloud Billing accounts to the new spend‑based compute flexible committed use discounts (CUDs), broadening CUD coverage across Compute Engine, GKE, and Cloud Run SKUs. Google moved accounts to a spend‑based flexible CUD model that applies across multiple compute SKUs, simplifying how commitments are consumed.

Additionally, broader coverage across Compute Engine, GKE, and Cloud Run makes it easier to apply committed discounts to actual usage patterns.

Capacity Planner adds Cloud Storage egress bandwidth and Spot GPU visibility

Google Cloud’s Capacity Planner preview now shows Cloud Storage egress bandwidth usage and GPUs attached to Spot VMs. The update gives teams visibility into egress bandwidth and which GPUs are attached to Spot VMs so you can forecast bandwidth and GPU capacity needs.

Plus, that helps avoid quota bottlenecks and better plan capacity for spot-based GPU workloads.

Cloud Monitoring MCP server and OTLP metric ingestion (Preview) to improve observability pipelines

Google Cloud added a Cloud Monitoring MCP server (preview) and OTLP metric ingestion paths via OpenTelemetry Collectors. The MCP server lets agents and AI apps interact with time series data, while OTLP ingestion via OpenTelemetry Collectors gives additional ingestion paths for metrics.

Also, these changes improve observability pipelines and integration flexibility for telemetry sources.

February 6, 2026

Carbon Footprint corrected Cloud Run emissions

Important: Google Cloud fixed incomplete Cloud Run emissions data for Nov–Dec 2025 in Carbon Footprint and advised customers to backfill transfers to see corrected data.

Google updated Carbon Footprint to correct previously missing Cloud Run emissions for that period so reported emissions now reflect actual usage.

Consequently, customers who rely on Carbon Footprint for sustainability accounting should backfill transfers to ensure their historical reporting is accurate.

January 30, 2026

Vertex AI Search lets you change pricing models to match forecasting needs

Vertex AI Search now supports two pricing models and switching between them.

You can choose general pay‑as‑you‑go or a configurable monthly subscription for apps and data stores, and change between models as needed.

Additionally, that flexibility affects predictability versus consumption-based spend, which matters for forecasting and chargeback.

Therefore, product owners can pick the model that best balances predictable costs and variable usage for their search workloads.

Cloud SQL fast clone (same zone) goes GA for MySQL and PostgreSQL

Cloud SQL added GA support for fast clone within the same zone for MySQL and PostgreSQL.

Fast clone gives near‑instant cloning for dev/test and analytics workloads instead of full volume copies, speeding environment provisioning.

Also, clones are cheaper and faster than full copies, which reduces storage costs and time‑to‑test for CI workflows.

N4A (Axion/Arm) machine family is generally available

Google Compute announced GA for the N4A machine family powered by Axion Arm processors.

N4A supports 1–64 vCPUs, up to 512 GB memory, and standard/highmem/highcpu and custom types, adding another price‑performance option based on Arm architecture.

Moreover, that means workloads that can run on Arm may realize different price‑performance tradeoffs useful for FinOps evaluations.

January 23, 2026

Committed‑use recommendations support more machine types

Google Cloud expanded resource‑based Committed Use Discount (CUD) recommendations to support additional machine series.

Recommendations are available in the FinOps hub, Recommender API, and via BigQuery exports for programmatic analysis.

Backup & DR Service cost reports generally available

Google Cloud Backup & DR Service launched cost reports in GA to provide resource‑specific billing insights.

The reports expose backup‑related spend so teams can analyze retention, replication, and protection costs.

BigQuery: Gemini Cloud Assist surfaces job‑history analysis (preview)

BigQuery previewed Gemini Cloud Assist to analyze job history and surface slow or resource‑intensive queries.

The feature helps identify which queries are high cost or slow, speeding FinOps investigations into query-driven spend. That lets teams prioritize query optimization and tune cost‑heavy SQL patterns more quickly.

Cloud Monitoring: Application Monitoring dashboards show associated trace spans,link traces to costly ops

Cloud Monitoring dashboards for Application Monitoring now display associated trace spans for registered App Hub applications.

This improves observability by helping teams correlate trace spans to slow or costly operations shown in dashboards.

 

January 9, 2026

View future availability for GPU VMs, H4D VMs, or TPUs

Generally available: You can view future resource availability before you create a future reservation request in calendar mode. This action helps increase the likelihood that Google Cloud approves your request.

 

January 2, 2026

GKE gains In-place Pod Resize and Writable cgroups to improve resource efficiency

GKE (Kubernetes 1.35) went GA with In-place Pod Resize (change CPU/memory without restart) and Writable cgroups.

In-place resizing helps tighten right‑sizing by letting you adjust pod resources without downtime, reducing waste and disruption.

FinOps Weekly
FinOps Weekly
Articles: 211