Valkey 9.1 Turns the Redis Licensing Shock Into a Real Upgrade Path
Memorystore for Valkey 9.1 went GA on September 25 with up to 3x QPS at microsecond latency. The upgrade path runs through the 2024 Redis license shock.
TL;DR
- Google Cloud's page marks Valkey 9.1 GA; its structured datePublished is 2026-09-25 while the live page displays September 26, 2026.
- The mechanism: lock-free SPMC, MPSC, and SPSC queues replace round-robin socket assignment and list polling, plus dynamic scaling (30% CPU ignition, queue-depth auto-scaling).
- The governance story: Redis Inc.'s 2024 license shift pushed Google Cloud and other leaders to back Valkey under the Linux Foundation; two years later the fork is a managed upgrade path.
- New capabilities: database-level ACLs for environment isolation, CLUSTERSCAN, and new atomicity commands (HGETDEL, MSETEX, HSETEX).
- Read "up to 3x" as a vendor-chosen, benchmark-based comparison.
- Before migrating: audit command coverage, ACL configuration, and latency targets against your own workloads.
Google Cloud's live page displays September 26, 2026; its structured datePublished field says 2026-09-25. The post announces general availability of Memorystore for Valkey 9.1, with a headline claim of "up to 3x queries per second (QPS) at microsecond latency compared to Memorystore for Redis Cluster." Two stories sit under that sentence — how a cache gets three times faster, and the 2024 licensing shift that created the project. Teams choosing a cache layer today need both.
The 2024 license shock is the reason Valkey 9.1 exists
Google's own timeline starts in 2024, when Redis Inc. shifted its licensing away from the permissive open-source BSD license to a dual-license model. Google Cloud and other technology leaders responded by backing Valkey, an open-source alternative governed by the Linux Foundation. Procurement teams should sit with that sequence: the project exists because a license changed, not because a benchmark demanded it. Two years on, the fork is a managed service with a faster engine than the original — a viable exit after a license shift, though a young one: Valkey is about two years old, a maturing path rather than a fresh one. Google's positioning is nonetheless confident: "Valkey is particularly compelling for organizations scaling AI and microservices to handle millions of concurrent users."
The claim is backed by shipped controls, most usefully database-level access control. ACL rules previously applied globally across an instance; Valkey 9.1 restricts user access at the specific numeric database level. Google's example is a three-environment cluster — a production user scoped to db=0, a staging user to db=1, a dev user to db=2 — isolating workloads without key prefixes and at no additional cost, environment separation without running three clusters.

Operationally, that separation depends on both sides of the connection being right: the client must select the intended numeric database and the ACL must restrict its user to that database. Verify db=0/1/2 mappings in each connection pool, then test a denied cross-environment access before treating the arrangement as isolation. This is a configuration boundary, not a substitute for application authorization or separate failure domains.
The mechanism behind the 3x: queues instead of polling
Google describes the old design plainly: previously, Valkey assigned client sockets to I/O threads statically in a round-robin fashion, requiring the main thread to continuously poll lists of pending clients. Busy-wait list polling is the cross-thread CPU waste that shows up at high QPS. Valkey 9.1 replaces list polling with a lock-free, multi-queue messaging architecture: the main thread dispatches read and write jobs into a single-producer multi-consumer (SPMC) queue, where free workers pull tasks on demand (work-stealing, no starvation); completed work returns through a multi-producer single-consumer (MPSC) queue popped instantly, eliminating busy-wait list iteration; single-producer single-consumer (SPSC) queues handle thread-affine memory cleanup and high-volume epoll offloading. Scaling is dynamic too, through what Google calls an ignition phase: when main-thread CPU usage crosses 30%, the engine activates the first background I/O thread before queue bottlenecks form, and queue-depth auto-scaling moves the worker count on real-time SPMC backlog, "ensuring extra cores are used only when needed and parked when idle." That mechanism is the argument for 3x.

That is the mechanism Google connects to the performance claim, not a universal result for every service workload. Validate with your own request mix, connection count, and latency percentiles under production-like concurrency.
Scaling, the honest reading, and a scan that survives failovers
Scanning keys across a large cluster used to mean querying nodes individually — an approach that was neither cluster- nor failover-aware, so scans could miss keys, return duplicates, or fail if slot migrations or node failovers occurred mid-scan. CLUSTERSCAN fixes this with a topology-aware cursor that encodes the current slot, a fingerprint of the local hashtable, and the local cursor. Two strategies: a sequential full-cluster scan for simple scripts, and a parallelized mode where the SLOT argument partitions the cluster's 16,384 slots across multiple workers.
That creates a useful acceptance test beyond the vendor benchmark: run sequential and SLOT-parallel scans against the same keyspace, compare returned key sets, and repeat during a slot migration or node failover. Record duplicate, missing, and cursor-resume outcomes separately. A demo proves the command runs; reconciliation under topology change is what tells a team whether its own scan consumer is safe.

Now the inconvenient part. The 3x figure is "up to 3x," measured against Memorystore for Redis Cluster — the comparison the vendor chooses against its own previous service. The source carries the caveat itself: the sizing figures are based on open-source benchmarks, with real gains varying by workload. Customer traction is vendor-stated: Google says over 95% of the top 100 Google Cloud customers already rely on Memorystore for demanding, high-throughput workloads.
What else ships in 9.1: commands, sizes, and the 9.0 base
HGETDEL reads a hash field and deletes it atomically in a single network round-trip — the pattern behind single-use authentication tokens and short-lived session state. MSETEX sets multiple keys with a single, shared expiration time, removing multi-command pipeline overhead. HSETEX gains NX and XX conditional flags for writes that must not clobber an existing value. Valkey 9.1 also builds on 9.0 from Google Cloud Next '26: native JSON support and Bloom filters for AI and vector workloads, plus six new node sizes — Custom-Pico (1.25 GB), Custom-Micro (2.5 GB), and Custom-Mini (3.5 GB), available only for cluster mode disabled environments, up to Highmem-XXLarge with 110 GB RAM and 16 vCPUs per node.
What to check before moving a Redis Cluster workload to Valkey 9.1
Migration is now a supported, generally available workflow: provision the target instance with your shard count and node sizing, establish continuous dual-sync online replication, validate data synchronization against replication metrics, then execute the cutover by switching application connection endpoints. "This workflow is generally available." The checks are the usual managed-Redis list, with a licensing twist. Audit command coverage: every Redis command your application uses must exist in the Valkey command set, kept in full in the docs. Decide the ACL configuration before cutover so environment separation is in place on day one. Set latency targets from your own workload, not the 3x headline — the same latency-economics argument we made when GPT-6 prompt caching launched.
Take a concrete case — hypothetical, but the shape is common. A European e-commerce SME runs checkout session state on self-managed Redis OSS and pays an engineer to watch failovers; the dual-license model made its procurement team uneasy, and a managed Valkey path removes both the license question and the operational load, with DB-level ACLs separating staging from production inside one cluster. That journey is as much a procurement decision as a performance one.
None of this argues against migrating — it argues for checking first. The next time a vendor quotes 3x, look for the mechanism before the chart. For Valkey 9.1 it is checkable: lock-free queues, a 30% ignition threshold, ACL scoping per database — a release a team can evaluate against its own workloads, which is exactly the position a license shift should put you in. Teams making that call can read how Netics works on cache-layer performance and open-source governance decisions.

Sources
The facts in this piece are drawn from Unlock 3x QPS and microsecond latency with Memorystore for Valkey 9.1 (cloud.google.com/blog, by Ankit Sud and Jacob Murphy, published 2026-09-25). Command coverage and ACL configuration details are summarized from the same post; the Memorystore for Valkey documentation maintains the full supported command and ACL lists.
Source: Unlock 3x QPS and microsecond latency with Memorystore for Valkey 9.1 — cloud.google.com/blog, 2026-09-25.