In March 2024, Redis Ltd. changed the Redis license from the permissive BSD model to a dual RSALv2/SSPLv1 model. I had customers ask me about it within forty-eight hours. By month’s end, the Linux Foundation had announced Valkey, a hard fork backed by AWS, Google Cloud, Oracle, Ericsson, and Snap. Over the following year, Dragonfly picked up serious adoption from production engineering teams, and Microsoft Research dropped Garnet on the world with genuinely surprising performance claims.
After twenty years building distributed systems, I have watched this pattern before: a dominant open-source project changes its licensing, the community forks or migrates, and then a wave of architectural challengers uses the disruption to make their case. What makes this one different is that Dragonfly and Garnet are not just license-safe alternatives. They are rethinking the core architectural decisions Redis made in 2009 when single-threaded event loops and gigabit ethernet were reasonable constraints. Multi-core CPUs, NVMe SSDs, and 100 Gbps networks change the calculus.
This article is not a rehash of the license change drama. It is a practical guide to the three paths available to teams running Redis today, based on what I have seen work in production and where the traps are.
What the License Change Actually Did
Redis Ltd.’s shift to SSPL and RSALv2 was aimed squarely at AWS, Azure, and GCP, who were running Redis as a managed service and contributing almost nothing back upstream. That is a real grievance. The practical effect on self-hosted users was largely symbolic, but the SSPL’s requirement that you open-source all infrastructure management code if you offer the software as a service was too poisonous for major cloud providers to accept.
Within a week of the announcement, the Linux Foundation accepted the Valkey fork, starting from Redis 7.2.4, the last version under BSD. Within months, AWS ElastiCache, Google Cloud Memorystore, and several others had pledged Valkey support. In May 2025, Redis Ltd. partially reversed course by releasing version 8.0 under a tri-license that adds AGPLv3 as an option alongside SSPL and RSAL. The AGPLv3 path means you can use Redis 8.0 under a genuine open-source license, but AGPLv3 still carries network-copyleft requirements that some legal teams treat as a hard block.
The net result in mid-2026 is a market with three credible paths: stay on Redis (possibly upgrading to 8.0 with AGPLv3), migrate to Valkey, or switch to a ground-up rewrite like Dragonfly or Garnet. Each path has a different risk profile.

Understanding the Architectural Problem Dragonfly Is Solving
Redis’s single-threaded command processing is a feature, not a bug, for certain workload shapes. A single thread means no lock contention, simple reasoning about atomicity, and predictable latency when your dataset fits in RAM and your command rate is moderate. The problem is that most production systems hitting Redis hard are running on machines with 32 to 128 CPU cores, and Redis uses one of them.
Dragonfly takes a shared-nothing approach. Each CPU core runs its own event loop with its own memory shard. A key is assigned to a shard via consistent hashing. Multi-key operations that cross shards use a “multi-executor” fiber coordination layer. The result is a system that scales vertically with core count in a way Redis fundamentally cannot.
The memory story is equally interesting. Dragonfly uses a custom hash table called Dashtable that delivers better cache locality than Redis’s dictionary implementation. According to Dragonfly’s published benchmarks, a dataset of 20 million string keys consumes roughly 880 MB in Dragonfly versus about 4.2 GB in Redis 7.4. I would attribute those numbers as Dragonfly claims rather than independently verified fact, but the structural reasons for the gap are real: Redis carries significant per-key metadata overhead from its object encoding system, and Dashtable’s design amortizes that overhead differently.
In practice, teams migrating from large sharded Redis Cluster deployments to Dragonfly report replacing clusters of a dozen or more nodes with two or three Dragonfly instances. The operational simplification is the real draw for most of them. Redis Cluster is genuinely painful to operate: slot migration, resharding, cross-slot multi-key operations, and the coordination overhead of running a distributed consensus system just for cache invalidation all add up.
Dragonfly’s license is BSL 1.1 (Business Source License), which converts to Apache 2.0 after four years. You can self-host it freely in production as long as you are not selling a competing managed in-memory datastore service. For application teams running their own infrastructure, this is a non-issue.
Garnet: Microsoft Research’s Different Bet
Garnet comes from Microsoft Research and takes a different architectural approach. Where Dragonfly is a C++ rewrite built around fibers and Dashtable, Garnet is built on .NET and focuses on tiered storage as its primary differentiation.
The core insight in Garnet is that NVMe SSDs are now fast enough to blur the line between memory-tier and storage-tier data. Garnet supports a tiered storage model where hot data lives in DRAM and cold data spills to NVMe, all behind the same RESP (Redis Serialization Protocol) interface. Redis does have persistence via RDB snapshots and AOF, but it is not designed for transparent tiering where cold keys migrate to disk automatically during runtime.
Garnet is MIT licensed and already used internally by Microsoft for services including Azure Resource Manager. It supports a large fraction of the Redis API surface, including strings, sorted sets, bitmaps, and HyperLogLog, which covers the bulk of real-world Redis usage. The .NET runtime also gives it a mature story around memory management, garbage collection tuning, and tooling ecosystem. However, as of late 2025, Garnet releases were still marked preview or beta, and it does not yet carry the same production battle-testing as Dragonfly or Valkey.
Garnet is the right bet to watch if your use case involves datasets that are larger than available DRAM but still benefit from in-memory speed for hot data. Think session stores for applications with hundreds of millions of users where the active session set is a fraction of the full dataset, or feature stores where you need sub-millisecond access to the hot features but the tail is rarely touched.
Valkey: The Conservative Migration Path
For teams that want out from under the Redis license uncertainty but want minimal architectural risk, Valkey is the conservative choice. It is a direct fork of Redis 7.2.4, developed under the Linux Foundation with major cloud provider backing. The API is identical to Redis 7.2, your existing Lettuce, StackExchange.Redis, or redis-py client code will connect without changes, and your operational runbooks transfer directly.
Valkey has made progress beyond its Redis origins. By mid-2025, the project added multi-threading improvements and backported performance work that was happening in Redis 8.x. The main limitation compared to Dragonfly is that it is still fundamentally the Redis architecture with incremental improvements rather than a ground-up redesign. If you are currently running a sharded Redis Cluster and the operational complexity is painful, Valkey does not solve that problem; it just solves the licensing problem.
For the general tradeoffs between Redis, Memcached, and Valkey as caching layers, the distributed caching overview covers the foundational decisions well. This article focuses on the newer architectural challengers and the production migration questions that come with them.
Deciding Which Path Fits Your Team
I use a simple decision tree when teams ask me what to do:
If you need minimal migration risk and a BSD-licensed drop-in: Go to Valkey. Switch the connection string, validate your Redis-specific commands still work, and move on. AWS ElastiCache, Google Memorystore, and Azure Cache for Redis have all announced Valkey support paths.
If you are running Redis Cluster today and the operational complexity is killing you: Look hard at Dragonfly. The vertical scaling story means a single beefy instance can replace a cluster that required constant resharding and node management. Teams report significant reductions in infrastructure-as-code complexity when they make this move.
If your dataset is larger than your available RAM and you want in-memory speed for hot data: Garnet’s tiered storage is worth evaluating. Accept that you are an early adopter and plan for some rough edges.
If you are already on Redis 8.0 and the AGPLv3 path is acceptable to your legal team: You can stay. Redis 8.0 delivers genuine performance improvements including async I/O and better multi-threading in the cluster path. The AGPLv3 copyleft applies to modifications you distribute, which for most application teams means nothing in practice.
The workload shape matters more than people admit. Dragonfly’s multi-threaded model delivers the biggest gains on write-heavy workloads with large key counts where the single-threaded Redis bottleneck is visible. Pure read-heavy caches with small datasets on fast networks see a smaller delta. Profile your actual workload before committing to a migration.

Running Dragonfly in Production on Kubernetes
I have helped three teams run Dragonfly on Kubernetes over the past year, and the operational story is genuinely simpler than Redis Cluster. Dragonfly exposes a single stateful deployment rather than a cluster of coordinating nodes. You run it as a StatefulSet with persistent storage for the RDB/AOF path if you need durability guarantees.
A typical production setup looks like a StatefulSet with one replica for standalone mode or two replicas in Dragonfly’s replication mode, which uses primary-secondary replication similar to Redis Sentinel. You front it with a ClusterIP service and configure your application’s connection pool to point there. If you need Kubernetes-native failover and leader election, the Dragonfly Operator handles promoting a replica when the primary fails. For teams already running databases on Kubernetes, this fits the same operator pattern used for PostgreSQL, MongoDB, and other stateful workloads.
Resource allocation for Dragonfly requires thought. Because it is multi-threaded, you should actually provision realistic CPU requests rather than the thin CPU slices that often work fine for single-threaded Redis. Dragonfly recommends at least 2 dedicated cores for production workloads of any significance. For Kubernetes resource management, this means setting CPU requests equal to limits for the Dragonfly container rather than relying on the burstable QoS class.

One operational nuance: Dragonfly’s SAVE and BGSAVE commands create snapshots, but the snapshot format is not compatible with Redis RDB files. This matters for migrations. If you need to migrate existing Redis data into Dragonfly, you must do it via RESTORE commands or by replaying the writes, not by copying the RDB file. I have seen teams bitten by this assumption.
The persistence semantics deserve careful review if you are using Redis for anything beyond pure caching. Dragonfly supports both RDB snapshots and AOF-style journaling, but the durability guarantees and fsync behavior need to be validated against your specific requirements. For pure caching workloads where the backing store is the source of truth, this is not a concern.
The Cost Story
Cost reduction is the most common driver I see for these migrations, and it deserves honest framing. The potential savings are real but context-dependent.
The most straightforward savings come from cluster simplification. If you are paying for a 12-node Redis Cluster on AWS or GCP because that is what you needed to handle your write throughput, and you can replace that with two Dragonfly instances, the infrastructure cost reduction can be substantial. The operational savings in engineering time are often larger than the compute savings, but they are harder to quantify.
Managed Redis pricing on major clouds is significantly higher than the equivalent compute for self-hosted Dragonfly. AWS ElastiCache for Redis pricing varies significantly by region, instance type, and whether you use reserved capacity, so I would not quote specific numbers here; the AWS pricing page is the source of truth. The point is that running Dragonfly or Valkey self-hosted in EC2 or EKS eliminates the managed service markup in exchange for operational responsibility.
For teams thinking about the FinOps angle, the conversation connects directly to cloud commitment pricing and whether your caching tier is appropriately sized. Over-provisioned Redis clusters are common because teams provision for peak throughput that only happens during traffic spikes. Dragonfly’s more efficient memory usage can reduce the instance size required to hold a given dataset, which compounds with reserved instance or savings plan discounts.
What I Would Not Do
I have seen teams make two consistent mistakes in this space.
The first is treating the migration as purely a drop-and-replace operation. Dragonfly is wire-compatible with Redis via the RESP protocol, which means your clients connect and your basic commands work. But production systems often rely on Redis-specific behaviors around cluster topology, client-side sharding, or specific edge cases in EVAL/Lua scripting. Validate your actual workload, not just a synthetic benchmark, before you declare success.
The second mistake is defaulting to the newest thing without having a clear problem to solve. If you have a healthy Redis 7.2 deployment, your cluster is manageable, and the Valkey fork is available on your managed service provider, migrating to Dragonfly for marginal performance gains on a workload that is not CPU-bound is a way to spend engineering time without much return. The right time to look at architectural alternatives is when you have a concrete problem: cluster operational complexity, memory costs, or throughput limits you are actually hitting.
Garnet’s Production Readiness in 2026
I want to be honest about Garnet’s current status. As of late 2025, it is a preview-quality project from Microsoft Research. It is impressive research software with real production deployment inside Microsoft, but the community tooling, operator support, and third-party battle-testing that Redis and Dragonfly have accumulated over years is not there yet.
The tiered storage differentiation is genuinely interesting and I expect Garnet to mature significantly over the next two years. If your use case is a good fit for persistent tiered storage at in-memory speeds, it is worth running a proof of concept. If you need a battle-tested Redis replacement today, Valkey or Dragonfly are the safer bets.
Microsoft has a history of open-sourcing research projects and then investing inconsistently in their production hardening. Garnet’s production use at Azure is a meaningful signal that it is not vaporware, but it is not the same as a broad ecosystem of operators, managed services, and community debugging experience.
The Connection to Broader Cache Architecture
The in-memory database choice does not exist in isolation. It is one layer of a caching architecture that typically includes application-level caching, distributed cache, and persistent storage. The database connection pooling patterns that matter for PostgreSQL have equivalents in how you configure connection multiplexing to your Redis or Dragonfly cluster. The same observability pipeline that handles your metrics and logs needs to ingest Dragonfly’s Prometheus endpoint.
I have also seen teams use Dragonfly specifically to solve the cloud egress cost problem. If you are running a read-heavy application that is hammering a database in another region or availability zone, a well-placed Dragonfly instance caching hot reads eliminates that cross-region traffic. The memory efficiency gains matter here: a smaller Dragonfly instance can cache the same hot dataset that required a larger Redis instance, which reduces the instance cost for the caching layer itself.
The serverless database ecosystem is also starting to include managed Dragonfly options through providers like Dragonfly Cloud, which offers a fully managed Dragonfly service that handles the operational complexity while preserving the vertical scaling benefits. For teams that want the architecture without the operational burden, this is worth evaluating alongside ElastiCache Valkey.
Where This Is Heading
The in-memory database market in 2026 is genuinely more competitive than it has been in a decade. Redis’s original BSD license and dominance meant there was little commercial incentive to build a fundamentally better alternative. The license change broke that equilibrium.
My expectation is that the market settles into two tiers over the next few years: Valkey becomes the default “just works” Redis-compatible layer in managed services across all major clouds, and Dragonfly captures the subset of teams with demanding throughput or memory efficiency requirements who are willing to self-manage their caching infrastructure. Garnet becomes an interesting option for the tiered storage use case as it matures.
The practical advice: audit your Redis usage now if you have not already. Identify whether you are actually hitting throughput or memory limits that would benefit from Dragonfly’s architecture. Determine whether your legal team has concerns about the SSPL or AGPLv3 paths for Redis. If you are running on a managed Redis service that now offers Valkey, there is a low-risk migration path available. If you are managing your own Redis Cluster and the operational complexity has become a tax on your team, Dragonfly is worth a serious evaluation.
Do not let the licensing drama force an architectural decision your workload does not need. But do not ignore the options that have emerged from it either.
Get Cloud Architecture Insights
Practical deep dives on infrastructure, security, and scaling. No spam, no fluff.
By subscribing, you agree to receive emails. Unsubscribe anytime.
