Google Routes Inference on KV-Cache Pressure: The Multi-Cluster GKE Gateway, Measured
"Google benchmarked the multi-cluster GKE Inference Gateway at 17,000 nodes: near-linear throughput, under 1% routing overhead, spillover driven by KV-cache pressure instead of round-robin l
TL;DR
- On September 21, 2026, Google benchmarked the multi-cluster GKE Inference Gateway at 17,000 compute nodes across the US and Europe, behind one global virtual IP.
- The routing signal reads memory, not packets: the EPP exposes KV-cache utilization, and past a 40% threshold traffic spills to the next healthy region automatically.
- Measured cost: under 1% overhead at 99.5% of direct local throughput, with near-linear scaling 0.72 → 2.10 req/s across one to three clusters at 99.9% success.
- Round-robin treats every request as equal; LLM requests are not. Agentic 100k to 800k+ token contexts exhaust memory before compute — routing must read KV-cache pressure.
- Caveats: Google's own benchmark, one MoE model (SGLang) on Google clusters, feature in preview since March 2026 — evidence, not a launch, no SLA.
- The reusable model: config cluster outside the request path, KV-cache utilization first-class, LeaderWorkerSet-aware routing to leader pods; claims tagged in the manifest below.
On September 21, 2026, the Google Cloud blog published Global AI routing with <1% overhead on multi-cluster GKE Inference Gateway, signed by GKE product manager Fisayo Feyisetan and software engineer Sina Chavoshi. The headline numbers — 17,000 compute nodes, near-linear scaling across three regions, under 1% routing overhead — deserve attention, but what the gateway routes on matters more. It reads how full each region's KV cache is and sends each request where memory remains, and that choice separates a network load balancer from an inference routing plane.

The routing plane is where inference economics are won
Round-robin at the network layer treats every request as equal: "Round-robin load balancing is inadequate for serving LLMs because it treats every request as equal," as the post states. Heavy prompts saturate GPU compute cores, long generations stress memory bandwidth, and long-context conversations eat VRAM until the engine can no longer schedule anything new. Agentic workloads with 100k to 800k+ token contexts consume memory faster than any previous AI traffic, and a routing layer blind to that pressure strands expensive compute.
Our coverage has been building toward this layer: NVIDIA Topograph topology scheduling places work on devices inside a node, while the gateway sits above it, deciding which region takes the request — chasing the same objective as any scheduler: maximize "intelligence per dollar" instead of letting capital idle while requests queue elsewhere.
Three regions, one virtual IP — and a telemetry signal underneath
The deployment spans three clusters: us-east5 as the config cluster, with us-west8 and europe-west4 serving. Clients see none of that geography — requests hit a single global virtual IP, and the gateway decides in real time which cluster serves each one.

The decision engine is the Endpoint Picker Proxy. The EPP reads the KV-cache token utilization natively exposed by inference engines and emits it as a load-balancer metric; when a region runs hot on that signal, Google writes, "it spills traffic to the next healthy region." The threshold is a parameter you choose — this deployment crossed at 40% KV-cache utilization. Routing can also key off queue depth or running concurrency, while the config cluster holds the routing configuration outside the request path.
The measured trade-off: under 1% overhead for a global tier
Every team evaluating a global routing tier asks the question the post poses: "How much throughput am I giving up for cross-region capability?" Measured answer: routing through the gateway added less than 1% overhead, delivering 99.5% of the throughput of a direct, local cluster call. The number deserves respect and a caveat in the same sentence: it was produced on Google's own deployment, one leading MoE foundation model served with SGLang. Under 1% is real evidence that a global entry point can be nearly free, not a promise for every fleet, model mix, or accelerator. The post names the signal, the threshold, and the measurement, so the result is reproducible on your own hardware.
Near-linear scaling at 17,000 nodes
The harder test is scale. Every client request originated from us-east5, so added clusters sat at maximum distance from the traffic source; if the gateway could not distribute load across that distance, throughput would flatten as hardware grew. Instead, scaling held near-linear:

| Fleet topology | Request throughput | Token throughput | Success rate |
|---|---|---|---|
| 1 cluster (us-east5-a) | 0.72 req/s | 2,898 tok/s | 99.87% |
| 2 clusters (+ us-west8-a) | 1.40 req/s | 6,380 tok/s | 99.95% |
| 3 clusters (+ europe-west4-b) | 2.10 req/s | 8,457 tok/s | 99.90% |
Google reports "a near-linear throughput boost while maintaining a 99.9% success rate under heavy multi-client concurrency." Read it precisely: 99.9% is the success rate observed in this test, not a general service-level agreement.
Why round-robin fails when memory is the bottleneck
Round-robin did not become wrong; it became insufficient for the workloads dominating accelerator fleets. Below memory saturation, network-layer round-robin remains perfectly adequate — an honest counterpoint worth stating before teams over-engineer. Agentic traffic exhausts memory long before compute: LLM requests run for minutes and hold VRAM for their duration, until the engine can no longer schedule anything new.
The mechanics are concrete: as the primary region climbed toward its high-bandwidth memory limits, the gateway detected saturation when the cluster crossed its 40% KV-cache utilization threshold and routed the overflow onward. "No operator intervention was needed." The fleet-wide lesson stands apart from any gateway: plan connection limits and timeouts for AI-scale latency early, or expect aborted connections.

Copy the operational model: config cluster, leader pods, telemetry
Distributed inference engines under tensor parallelism serve only from the master; the post puts it plainly: "only the master (rank-0) pod serves the API." GKE already directs local traffic to leader pods with standard Service selectors and LeaderWorkerSet (LWS), and the multi-cluster gateway integrates with that foundation — routing global traffic to the correct regional services so cross-region balancing respects multi-node topologies out of the box. Building custom proxies to replicate that is the work Google argues against.
The March 17, 2026 preview announcement names the resources: InferencePool groups pods sharing hardware and model configuration; InferenceObjective names models and sets serving priorities; GCPBackendPolicy on a GCPInferencePoolImport configures CUSTOM_METRICS or in-flight request limits. For teams standardizing on native Kubernetes constructs, this is the pattern worth copying — the same principle behind our read of Kubernetes 1.37 container storage hardening: keep the control plane small, declarative, and portable while data paths grow.
What Google's benchmark proves — and what it leaves open
Vendor benchmarks are necessary and insufficient, and this one is honest about its shape. The numbers are Google's, at 17,000 nodes, serving one leading MoE model with SGLang. They prove the architecture can behave like one fleet at scale; they do not prove those numbers transfer to heterogeneous fleets running multiple model families on mixed GPU and TPU generations — the post claims runtime-, model-, and accelerator-agnostic architecture while demonstrating a single workload.

Two dates help. The gateway entered preview on March 17, 2026 with model-aware routing already promised; September 21 is a benchmark of that preview, not a launch. The 99.9% success rate is a test outcome, not a contractual SLA, and the 40% threshold is a parameter you choose. The useful takeaway is the instrument list — KV-cache utilization first-class, queue depth, in-flight requests — because memory-aware routing is a property of telemetry quality, and telemetry is a property of your stack. Measure on your own fleet before promising anyone else's numbers. New infrastructure analysis lands on neticslabs.com.
Sources
Source: Global AI routing with <1% overhead on multi-cluster GKE Inference Gateway — cloud.google.com/blog/products/containers-kubernetes/gpu-and-tpu-utilization-with-multi-cluster-gke-inference-gateway, 2026-09-21. Companion: "Introducing multi-cluster GKE Inference Gateway: Scale AI workloads around the world" — cloud.google.com, 2026-03-17.