Google Routes Inference on KV-Cache Pressure: The Multi-Cluster GKE Gateway, Measured

"Google benchmarked the multi-cluster GKE Inference Gateway at 17,000 nodes: near-linear throughput, under 1% routing overhead, spillover driven by KV-cache pressure instead of round-robin l

Multi-cluster GKE Inference Gateway article card with the official Google identity and Netics branding.
Multi-cluster GKE Inference Gateway article card, Netics editorial.

TL;DR

  • On September 21, 2026, Google benchmarked the multi-cluster GKE Inference Gateway at 17,000 compute nodes across the US and Europe, behind one global virtual IP.
  • The routing signal reads memory, not packets: the EPP exposes KV-cache utilization, and past a 40% threshold traffic spills to the next healthy region automatically.
  • Measured cost: under 1% overhead at 99.5% of direct local throughput, with near-linear scaling 0.72 → 2.10 req/s across one to three clusters at 99.9% success.
  • Round-robin treats every request as equal; LLM requests are not. Agentic 100k to 800k+ token contexts exhaust memory before compute — routing must read KV-cache pressure.
  • Caveats: Google's own benchmark, one MoE model (SGLang) on Google clusters, feature in preview since March 2026 — evidence, not a launch, no SLA.
  • The reusable model: config cluster outside the request path, KV-cache utilization first-class, LeaderWorkerSet-aware routing to leader pods; claims tagged in the manifest below.

On September 21, 2026, the Google Cloud blog published Global AI routing with <1% overhead on multi-cluster GKE Inference Gateway, signed by GKE product manager Fisayo Feyisetan and software engineer Sina Chavoshi. The headline numbers — 17,000 compute nodes, near-linear scaling across three regions, under 1% routing overhead — deserve attention, but what the gateway routes on matters more. It reads how full each region's KV cache is and sends each request where memory remains, and that choice separates a network load balancer from an inference routing plane.

Official Google Cloud blog figure: multi-cluster GKE Inference Gateway topology — config cluster outside the request path, target clusters reporting KV-cache utilization (Sept 21, 2026 post)
Official Google Cloud blog figure (2026-09-21): multi-cluster GKE Inference Gateway topology — the config cluster holds routing configuration but sits outside the request path; each target cluster runs its own EPP and reports KV-cache utilization back to the load balancer.

The routing plane is where inference economics are won

Round-robin at the network layer treats every request as equal: "Round-robin load balancing is inadequate for serving LLMs because it treats every request as equal," as the post states. Heavy prompts saturate GPU compute cores, long generations stress memory bandwidth, and long-context conversations eat VRAM until the engine can no longer schedule anything new. Agentic workloads with 100k to 800k+ token contexts consume memory faster than any previous AI traffic, and a routing layer blind to that pressure strands expensive compute.

Our coverage has been building toward this layer: NVIDIA Topograph topology scheduling places work on devices inside a node, while the gateway sits above it, deciding which region takes the request — chasing the same objective as any scheduler: maximize "intelligence per dollar" instead of letting capital idle while requests queue elsewhere.

Three regions, one virtual IP — and a telemetry signal underneath

The deployment spans three clusters: us-east5 as the config cluster, with us-west8 and europe-west4 serving. Clients see none of that geography — requests hit a single global virtual IP, and the gateway decides in real time which cluster serves each one.

Official Google Cloud blog figure: routing overhead comparison — <1% overhead, 99.5% of direct local throughput (Sept 21, 2026 post)
Official Google Cloud blog figure (2026-09-21): the measured trade-off — routing through the gateway added less than 1% overhead, delivering 99.5% of a direct local cluster call.

The decision engine is the Endpoint Picker Proxy. The EPP reads the KV-cache token utilization natively exposed by inference engines and emits it as a load-balancer metric; when a region runs hot on that signal, Google writes, "it spills traffic to the next healthy region." The threshold is a parameter you choose — this deployment crossed at 40% KV-cache utilization. Routing can also key off queue depth or running concurrency, while the config cluster holds the routing configuration outside the request path.

The measured trade-off: under 1% overhead for a global tier

Every team evaluating a global routing tier asks the question the post poses: "How much throughput am I giving up for cross-region capability?" Measured answer: routing through the gateway added less than 1% overhead, delivering 99.5% of the throughput of a direct, local cluster call. The number deserves respect and a caveat in the same sentence: it was produced on Google's own deployment, one leading MoE foundation model served with SGLang. Under 1% is real evidence that a global entry point can be nearly free, not a promise for every fleet, model mix, or accelerator. The post names the signal, the threshold, and the measurement, so the result is reproducible on your own hardware.

Near-linear scaling at 17,000 nodes

The harder test is scale. Every client request originated from us-east5, so added clusters sat at maximum distance from the traffic source; if the gateway could not distribute load across that distance, throughput would flatten as hardware grew. Instead, scaling held near-linear:

Official Google Cloud blog figure: near-linear scaling across three regions — 0.72, 1.40, then 2.10 requests per second while keeping a 99.9% success rate (Sept 21, 2026 post)
Official Google Cloud blog figure (2026-09-21): near-linear throughput growth from one to three clusters (0.72 → 2.10 req/s) at a 99.9% success rate under heavy multi-client concurrency.
Fleet topologyRequest throughputToken throughputSuccess rate
1 cluster (us-east5-a)0.72 req/s2,898 tok/s99.87%
2 clusters (+ us-west8-a)1.40 req/s6,380 tok/s99.95%
3 clusters (+ europe-west4-b)2.10 req/s8,457 tok/s99.90%

Google reports "a near-linear throughput boost while maintaining a 99.9% success rate under heavy multi-client concurrency." Read it precisely: 99.9% is the success rate observed in this test, not a general service-level agreement.

Why round-robin fails when memory is the bottleneck

Round-robin did not become wrong; it became insufficient for the workloads dominating accelerator fleets. Below memory saturation, network-layer round-robin remains perfectly adequate — an honest counterpoint worth stating before teams over-engineer. Agentic traffic exhausts memory long before compute: LLM requests run for minutes and hold VRAM for their duration, until the engine can no longer schedule anything new.

The mechanics are concrete: as the primary region climbed toward its high-bandwidth memory limits, the gateway detected saturation when the cluster crossed its 40% KV-cache utilization threshold and routed the overflow onward. "No operator intervention was needed." The fleet-wide lesson stands apart from any gateway: plan connection limits and timeouts for AI-scale latency early, or expect aborted connections.

Netics claim-vs-check visual (round-robin assumption vs KV-cache-aware routing)
Netics editorial visual: the round-robin assumption versus KV-cache-aware routing — reading memory pressure (40% threshold spillover) instead of treating every request as equal.

Copy the operational model: config cluster, leader pods, telemetry

Distributed inference engines under tensor parallelism serve only from the master; the post puts it plainly: "only the master (rank-0) pod serves the API." GKE already directs local traffic to leader pods with standard Service selectors and LeaderWorkerSet (LWS), and the multi-cluster gateway integrates with that foundation — routing global traffic to the correct regional services so cross-region balancing respects multi-node topologies out of the box. Building custom proxies to replicate that is the work Google argues against.

The March 17, 2026 preview announcement names the resources: InferencePool groups pods sharing hardware and model configuration; InferenceObjective names models and sets serving priorities; GCPBackendPolicy on a GCPInferencePoolImport configures CUSTOM_METRICS or in-flight request limits. For teams standardizing on native Kubernetes constructs, this is the pattern worth copying — the same principle behind our read of Kubernetes 1.37 container storage hardening: keep the control plane small, declarative, and portable while data paths grow.

What Google's benchmark proves — and what it leaves open

Vendor benchmarks are necessary and insufficient, and this one is honest about its shape. The numbers are Google's, at 17,000 nodes, serving one leading MoE model with SGLang. They prove the architecture can behave like one fleet at scale; they do not prove those numbers transfer to heterogeneous fleets running multiple model families on mixed GPU and TPU generations — the post claims runtime-, model-, and accelerator-agnostic architecture while demonstrating a single workload.

Netics comparison-table visual
Netics editorial visual: network-layer round-robin versus the multi-cluster Inference Gateway — routing signal, scope, measured overhead, and saturation behavior.

Two dates help. The gateway entered preview on March 17, 2026 with model-aware routing already promised; September 21 is a benchmark of that preview, not a launch. The 99.9% success rate is a test outcome, not a contractual SLA, and the 40% threshold is a parameter you choose. The useful takeaway is the instrument list — KV-cache utilization first-class, queue depth, in-flight requests — because memory-aware routing is a property of telemetry quality, and telemetry is a property of your stack. Measure on your own fleet before promising anyone else's numbers. New infrastructure analysis lands on neticslabs.com.

Sources

Source: Global AI routing with <1% overhead on multi-cluster GKE Inference Gateway — cloud.google.com/blog/products/containers-kubernetes/gpu-and-tpu-utilization-with-multi-cluster-gke-inference-gateway, 2026-09-21. Companion: "Introducing multi-cluster GKE Inference Gateway: Scale AI workloads around the world" — cloud.google.com, 2026-03-17.