Data & Analytics

Kafka Tiered Storage: How KIP-405 Finally Decouples Compute From Storage in Production Kafka Clusters

KIP-405 tiered storage hit production-ready status in Kafka 3.9 and is now the default cost-optimization lever for long-retention Kafka clusters. Here is the architecture, the configuration reality, and how it compares to the diskless-first alternatives.

Diagram showing Apache Kafka tiered storage architecture with hot local disk tier and cold remote object storage tier

I have run Kafka clusters at scale for twenty years, and the single most consistent complaint I hear from platform teams is the same one it has always been: disk. Kafka brokers are hungry for local storage. Every byte of retention you want translates directly into broker disk you have to provision, protect, and pay for. If you want to replay six months of events for a new downstream consumer, you need six months of data on broker disks, full stop. That has always been the constraint.

KIP-405 (Kafka Tiered Storage) changes that constraint at the architectural level. It ships a hot local tier on broker disks for recent data alongside a cold remote tier in object storage for everything older, and the split is transparent to producers and consumers. The feature entered early access in Kafka 3.6 (late 2023) and reached production-ready status in Kafka 3.9. By the Kafka 4.x series it is a supported, stable capability that every serious Kafka deployment should be evaluating.

This article covers the architecture in depth, the configuration you actually need to get it running, the operational realities of maintaining a tiered cluster, and an honest comparison to the diskless-first alternatives that have been marketing themselves as the better answer to the same problem.

Why Broker Disk Was Always the Real Kafka Bottleneck

The standard advice for Kafka capacity planning goes something like this: calculate your peak ingestion rate, multiply by your retention period, multiply by your replication factor, and that is the raw disk you need across the cluster. Then add overhead for compaction, indexes, and broker OS. The number gets large quickly.

I have seen this play out at organizations running petabyte-scale event streams where 80 to 90 percent of the data sitting on broker disks had not been read in weeks. Consumers were reading tail data, producers were writing new events, and the cluster was paying for the full disk footprint of historical data that was cold in practice even if Kafka treated it as hot by design.

The deeper problem is that broker compute and broker storage are coupled. If you need more retention, you add broker disk, which often means adding broker nodes because you have hit the practical disk limit per node. Adding broker nodes changes partition distribution, triggers rebalances, and increases your ongoing operational surface. You are scaling compute to solve a storage problem.

Object storage, on the other hand, is effectively infinite and costs a fraction of the price of NVMe or even spinning disk on broker nodes. S3, GCS, and Azure Blob are designed for exactly this kind of cold-but-accessible data. The question KIP-405 answers is: how do you put older Kafka log segments in object storage without breaking the Kafka protocol, the consumer offset model, or the operational tools teams already rely on?

How KIP-405 Tiered Storage Actually Works

The architecture introduces three new abstractions into the Kafka broker: the RemoteLogManager, the RemoteStorageManager, and the RemoteLogMetadataManager.

Kafka tiered storage architecture showing local disk hot tier, RemoteLogManager, and object storage cold tier

The RemoteLogManager (RLM) is a background process running on every broker. It monitors local log segments for tiering candidates, coordinates uploads to remote storage, and manages the lifecycle of remote segments, including deletion when retention policies expire. Only the partition leader uploads segments; followers do not independently upload, which avoids duplicate data in remote storage. If a leader reassignment happens, the new leader picks up the RLM responsibility for that partition.

The RemoteStorageManager (RSM) is a pluggable interface that abstracts the actual object storage backend. You implement or choose a plugin that knows how to talk to S3, GCS, Azure Blob, or whatever object store you are using. The Apache Kafka project does not ship a built-in RSM implementation; you bring your own. Aiven maintains an open-source RSM implementation for S3-compatible stores. Confluent, MSK, and other managed platforms bundle their own implementations.

The RemoteLogMetadataManager (RLMM) tracks which log segments have been uploaded, what their offset ranges are, and where they live in remote storage. The default implementation stores this metadata in a dedicated internal Kafka topic (__remote_log_metadata), which means the metadata itself benefits from Kafka’s replication and durability guarantees. You can plug in an alternative RLMM backed by an external metadata store, but for most teams the built-in topic-backed implementation is the right choice.

The Segment Lifecycle

When a log segment on a broker fills and rolls over, the RLM evaluates whether it is a tiering candidate. The key controls are log.local.retention.bytes and log.local.retention.ms. These work like the familiar log.retention.bytes and log.retention.ms, but they define when data is eligible to move from local storage to remote, not when it is deleted entirely. You set the local retention shorter than the total retention, and the gap is filled by remote storage.

Once a segment is copied to remote storage and the RLMM records its metadata, the local copy becomes eligible for deletion. The broker shrinks its local disk footprint while keeping the data accessible. Consumers reading from an offset that falls in the remote tier transparently fetch from object storage via the RLM rather than reading from the local log. From the consumer’s perspective, nothing has changed: they call the standard Fetch API and get the data back.

Sequence diagram of RemoteLogManager copying segments from local disk to S3 and consumer fetch flow

What Tiered Storage Does Not Change

It is worth being precise about what the architecture does not change. The write path is unaffected. Producers still write to broker memory and local disk. The local tier is still the durable, low-latency write target. Tiered storage is entirely about the read path for historical offsets and the deletion of local segments after they are safely persisted in object storage. If your use case is dominated by tail reads, tiered storage will not change your producer-side latency at all.

Consumer reads of recent data are also unaffected. If the consumer is reading within the local retention window, the fetch happens entirely from local disk. Only reads that go back into the remote tier will see the round trip to object storage, which adds latency compared to local reads. For most operational consumers this is never triggered; it matters primarily for backfill, replay, and ad-hoc historical queries.

Configuration: What You Actually Need

Enabling tiered storage requires changes at the broker level and at the topic level. Broker-level configuration activates the subsystem; topic-level configuration controls which topics participate and what their local/remote retention split is.

The essential broker configurations are:

# Enable the tiered storage subsystem globally
remote.log.storage.system.enable=true

# Your RemoteStorageManager implementation
remote.log.storage.manager.class.name=io.aiven.kafka.tieredstorage.RemoteStorageManager
remote.log.storage.manager.class.path=/opt/kafka/plugins/aiven-tiered-storage/*

# RSM-specific config for S3
remote.log.storage.manager.impl.prefix=tiered.storage.
tiered.storage.storage.backend.class=io.aiven.kafka.tieredstorage.storage.s3.S3StorageBackend
tiered.storage.s3.bucket.name=my-kafka-tiered-storage

# Optional: tune the RLM thread pool (default: 4)
remote.log.manager.thread.pool.size=4
# How often the RLM scans for segments to copy (default: 30 seconds)
remote.log.manager.task.interval.ms=30000

# RLMM listener name (required for default implementation)
remote.log.metadata.manager.listener.name=PLAINTEXT

At the topic level, you override retention to establish the local/remote split:

kafka-topics.sh --alter \
  --topic my-topic \
  --config remote.storage.enable=true \
  --config local.retention.bytes=10737418240 \
  --config retention.bytes=107374182400 \
  --bootstrap-server localhost:9092

This example keeps 10 GB of data on local disk and retains 100 GB total, with the difference living in object storage. Setting local.retention.bytes=-1 disables local size-based eviction to object storage; setting local.retention.ms=-1 disables time-based eviction. You can use both to control exactly when segments move.

Apply configuration changes in a rolling restart to avoid a full cluster downtime window. Enable remote.log.storage.system.enable on one broker at a time, verify the RLM starts cleanly in each broker’s logs, and proceed.

Managed Platform Support

If you are running Kafka on a managed platform rather than operating your own brokers, tiered storage availability varies by vendor.

Amazon MSK added tiered storage support alongside its Kafka 3.6 support in late 2023. The MSK implementation integrates with S3. One constraint to know: MSK still requires EBS volumes per broker for the local tier; you cannot eliminate local disk entirely, but you can significantly reduce how much you need.

Confluent Cloud has offered its own decoupled storage architecture under the “Infinite Storage” and tiered storage brand for some time. Confluent’s managed service handles the RSM implementation internally. In 2024, Confluent announced Freight Clusters, a WarpStream-inspired S3-native compute model for cost-sensitive workloads. Both coexist in the Confluent Cloud product lineup.

Aiven for Apache Kafka supports tiered storage and has open-sourced its RSM implementation. Aiven is also developing Inkless, which goes further than tiered storage toward a fully diskless architecture based on KIP-1150 (Diskless Topics), though Inkless was still maturing as of 2026 and is distinct from the KIP-405 tiered storage that is generally available today.

For self-managed deployments, Aiven’s open-source RSM for S3-compatible stores is the most widely referenced community option. You can also write your own RSM if you have specific requirements around encryption, metadata, or non-S3-compatible backends.

Operational Realities: What Gets Harder

Tiered storage introduces operational complexity that teams often underestimate before deploying it. I want to be honest about this.

Object storage costs and request patterns. Kafka’s segment upload pattern is write-once, and reads from the remote tier generate GET requests against object storage. If you have a consumer doing a large replay across months of historical data, you will see a spike in object storage GET requests and the associated costs. Tools like S3 Intelligent-Tiering or lifecycle policies on the Kafka storage bucket can reduce the object storage bill for data that has not been touched in a while. The cloud object storage lifecycle management article covers those controls in depth. Budget for object storage costs before you set 2-year retention on every topic.

Monitoring the RLM. The RemoteLogManager exposes JMX metrics for upload lag, segment copy counts, and errors. You want to add these to your Kafka monitoring dashboards. An RLM that is falling behind on segment uploads means your local tier is accumulating data that should have moved to remote storage, which defeats the purpose of the feature. Watch RemoteLogManager.CopyThrottleTime, RemoteLogManager.ReadThrottleTime, and any error counters. The Prometheus and Loki observability stack article is a good starting point if you are building this monitoring layer.

Consumer offset behavior during rebalances. KIP-848 changed how consumer group rebalancing works in Kafka 4.x, as covered in the Kafka consumer groups and KIP-848 article. Rebalances that move partition ownership also shift RLM leadership. Pay attention to the transition period: a newly assigned leader needs to re-initialize its RLM state for the partition before it can continue uploads. In practice this is fast, but it is worth understanding so you do not mistake normal rebalance-related RLM pauses for upload failures.

Schema compatibility across long retention windows. If you are keeping data for months or years, schema evolution becomes critical. A consumer replaying a twelve-month-old segment needs to be able to deserialize messages written under a schema that may have changed significantly. This is exactly where Kafka Schema Registry compatibility modes earn their keep. With tiered storage enabling genuinely long retention, a disciplined schema compatibility strategy transitions from nice-to-have to essential.

Operational monitoring architecture for Kafka tiered storage showing RLM metrics flowing to Prometheus and Grafana

Kafka Connect and tiered storage. Sink connectors that are reading from committed offsets behave transparently with tiered storage. If a connector is configured to start from an old offset and that offset’s data is in the remote tier, the connector will fetch it from object storage via the broker’s RLM proxy. This is usually fine but adds object storage read latency to the connector’s initial catch-up. If you have a connector that regularly reads from very old offsets, factor that latency into your Kafka Connect production configuration.

Tiered Storage vs. Diskless-First Alternatives

This is where the architecture conversation gets interesting. KIP-405 tiered storage keeps local broker disks as the authoritative write path. Object storage is a secondary tier for cold data. That is a fundamentally different model from AutoMQ, WarpStream, and similar diskless-first platforms that treat object storage as the primary data store.

The Kafka alternatives comparison article covers these platforms in depth. Here I want to focus specifically on the storage architecture tradeoffs.

With Kafka tiered storage, your producers are writing to local NVMe or SSD on broker nodes. The write path is fast, the tail-read path is fast, and the historical-read path goes through object storage with the associated latency. With a diskless platform like WarpStream, every produce call goes to S3 as the primary store. That means higher write latency by design, but the cost profile for long-retention workloads can be substantially lower since there are no broker disks at all.

AutoMQ takes a middle position: it uses a local write-ahead log for durability on the hot path, then flushes asynchronously to S3. The local disk is a WAL buffer rather than the long-term store, which gives lower write latency than pure S3-first while still allowing scale-to-zero broker disk.

The right choice depends on your latency requirements and your access patterns. If you have latency-sensitive consumers that need millisecond tail reads, you want Kafka with tiered storage (or Redpanda, which has a similar tiered architecture). If you have a high-throughput, latency-tolerant pipeline where cost is the primary driver and you rarely read historical data, a diskless-first platform may deliver better economics. Note that the latency figures you see in vendor benchmarks for these platforms are vendor-published claims and should be independently verified for your specific workload before making an architecture decision.

The other consideration is operational continuity. If you have an existing Kafka deployment and are evaluating tiered storage, KIP-405 is an in-place upgrade. You continue running Apache Kafka, the same tooling, the same monitoring, the same Kafka Connect connectors, and the same KRaft-based cluster architecture. Migrating to a diskless-first platform is a different workload entirely and typically requires validating API compatibility, re-testing connector behavior, and accepting some production risk during migration.

Migration Path: Adding Tiered Storage to an Existing Cluster

I have seen teams approach this migration in two ways. The first is to enable tiered storage on new topics only, letting old topics continue under the existing retention model and migrating high-value topics selectively. The second is a cluster-wide rollout where you enable the RSM on all brokers in a rolling restart and then update topic configurations to enable tiering.

For most teams, I recommend the topic-by-topic approach first. Start with topics that have high retention requirements, high data volume, and few latency-critical consumers. Event log topics for audit trails, CDC event streams from Debezium-based change data capture pipelines, and replay buffers for downstream batch consumers are good candidates. These are the topics where the economics of tiered storage are most compelling and where the historical-read latency penalty is least likely to break an SLA.

Once you have validated the RSM plugin, the RLMM behavior, and your monitoring setup on a handful of topics, rolling out cluster-wide is straightforward. The main coordination task is ensuring the rolling restart for broker-level config changes is done carefully, one broker at a time, with health checks between each restart to confirm the RLM is functioning normally on the newly restarted broker.

KIP-405 in the Context of Kafka 4.x

Kafka 4.0 completed the removal of ZooKeeper, making KRaft the only supported cluster metadata path. This matters for tiered storage because KRaft simplifies controller failover and the RLM’s leader-tracking behavior. In earlier Kafka versions with ZooKeeper, leadership transitions involved ZooKeeper coordination that added latency to RLM handoffs. With KRaft, the controller and the RLM both operate within the Kafka cluster’s own consensus model, which makes failover cleaner.

If you are still running ZooKeeper-mode Kafka, migrating to KRaft should be on your roadmap independently of tiered storage, and the two initiatives pair naturally. You do the KRaft migration first (covered in the KRaft migration guide), stabilize, and then layer in tiered storage.

Kafka 4.x also introduced refinements to the tiered storage metrics API and improvements to the RLM’s handling of compacted topics. KIP-405 had explicit gaps in compacted topic support in early releases. By the 4.x series, the behavior for compacted topics with tiered storage is better specified, though the interaction between log compaction and remote storage is still worth testing carefully before enabling tiered storage on compaction-heavy topics in your cluster.

When Tiered Storage Is Not the Answer

Let me be direct about the cases where KIP-405 is not what you need.

If your primary problem is broker count and compute cost rather than storage cost, tiered storage will not help much. You are still running the same number of brokers; you are just shifting some disk cost to object storage cost.

If you have producers and consumers with extremely tight latency requirements for the tail of the log and very few historical reads, the local-disk-first architecture you already have is likely the right one. Do not add operational complexity of tiered storage for workloads where historical data access is rare.

If you are starting a new deployment and cost is the primary driver, evaluate the diskless-first alternatives seriously. Their economics for steady-state, high-volume, latency-tolerant pipelines can be compelling, and starting fresh means you do not have a migration problem.

If you are integrating Kafka with Apache Flink for stateful stream processing and your Flink jobs regularly perform long historical reads for state recovery or table-valued function lookups, understand that those reads will hit object storage with tiered storage enabled. Test this behavior under load before enabling tiered storage on topics your Flink jobs depend on.

The Practical Result

Twenty years of watching storage bottlenecks constrain Kafka deployments has made me appreciate what KIP-405 actually delivers. It is not a revolutionary architecture; it is a practical one. The write path is unchanged. The consumer protocol is unchanged. The tooling is unchanged. What changes is the economics of long retention at scale.

For teams running long-retention Kafka clusters, particularly those in compliance-heavy industries where you need months of event history for audit purposes, or data platform teams building replay-capable event streams that downstream Flink and Spark jobs consume on demand, tiered storage removes a real constraint that previously required uncomfortable trade-offs between retention and cost.

The production-ready status in Kafka 3.9 and the continued hardening in 4.x mean this is no longer a feature you need to approach with research-project caution. It is a supported, operational capability with managed-platform backing from MSK, Confluent, and Aiven. Evaluate it now if you have not already.