GCP Updates 2026-09-17 | Cloud Provider News
Google Cloud makes it easier to right-size reservations as demand changes
Compute Engine reservations can now switch between shared and single-project configurations. Organizations can change how reserved capacity is scoped when their usage patterns shift.
A reservation can be converted between a shared configuration and one limited to a single project. This gives teams more flexibility to redistribute reserved capacity across projects instead of leaving it tied to an area with lower demand.
Vertex AI Search lets teams turn off unnecessary query add-ons
Vertex AI Search query add-on specifications are now generally available, allowing applications to disable optional add-ons for individual programmatic requests.
Teams can choose which search features each request needs rather than applying every available add-on by default. When enhanced features aren’t necessary, disabling them can reduce the cost of individual requests.
Cloud Run jobs can wait for lower pricing on non-urgent work
Cloud Run jobs can now defer non-urgent task execution for up to 12 hours to take advantage of reduced pricing.
The option gives batch workload owners an additional scheduling lever when a job doesn’t need to start immediately. Teams can defer eligible work instead of requiring immediate execution.
Storage Intelligence Advisor is now generally available for Cloud Storage
Google Cloud has made Storage Intelligence advisor generally available for monitoring and managing Cloud Storage environments across organizations, folders, and projects.
The centralized view supports storage governance and helps teams identify opportunities to optimize Cloud Storage at scale. Bringing visibility across multiple levels can make it easier to review storage environments consistently.
GKE Agent Substrate reduces idle resource use for agents
Agent Substrate on Google Kubernetes Engine can suspend idle agents and restore them on demand. When an agent is suspended, its memory and local files are snapshotted, then restored with sub-second latency.
By reducing the resources consumed by idle agents, the capability can increase the number of concurrent agents supported on each machine. This is aimed at large-scale agent workloads where many agents may not be active continuously.