---
title: "Cloudflare K2: Serverless Event Streams on Object Storage"
description: "How Cloudflare K2 builds a durable, partitioned event log on top of R2 object storage, and how its batching model trades produce latency for scale and retention compared with Kafka."
slug: "cloudflare-k2-serverless-event-streams-on-object-storage"
published: true
read_time: 8
created_at: "2026-10-02 14:04:19.281 +0000 UTC"
updated_at: "2026-10-02 14:04:19.283 +0000 UTC"
author: "Typen"
author_url: "https://typen.blog/@typen"
tags:
  - "Featured"
  - "Backend"
  - "Cloud"
  - "Databases"
---

# Cloudflare K2: Serverless Event Streams on Object Storage

How Cloudflare K2 builds a durable, partitioned event log on top of R2 object storage, and how its batching model trades produce latency for scale and retention compared with Kafka.

![https://media.typen.blog/img/id/01a0fcee-1ebc-797c-98b4-601ffbd7e8d8](trendbot-1821398898.webp)

## Why decouple producers and consumers

In a traditional Remote Procedure Call (RPC) architecture, producers and consumers have to agree on two things: scale and time. If producers emit more data than consumers can handle, or if a downstream service goes offline, events are dropped. The problem gets worse when several independent consumers need to process the same data — an ecommerce backend might emit transaction events that both an analytics system and a fraud detection service need to read.

The standard fix is to insert a durable buffer between producers and consumers. Producers write into it, and consumers read at their own pace. Cloudflare's new K2 service is a serverless implementation of exactly that primitive, built on the Developer Platform.

K2 is a durable event streaming service. You send events to a stream, which stores them as an ordered log. Consumers read them in different ways: splitting reads across a set of consumers, or delivering every message to every consumer. It is fully serverless, scales to large data volumes, and supports long-term retention, so extended consumer downtime does not lose data. Under the hood, K2 implements a partitioned, durable log on top of R2 object storage.

## Why not just run Kafka?

K2 was originally built because Cloudflare needed a durable buffer at the edge, initially as the ingestion layer for Basin Pipelines. Pipelines uses a pull-based stream processing engine, so some other system has to store events before they are read, transformed, and written to R2. Because Cloudflare commits to never dropping events once they are accepted into the Pipelines Stream, that storage has to be durable over potentially long periods.

For most companies, the answer would be Apache Kafka. But Pipelines runs on Cloudflare's edge, which spans a large number of servers across more than 335 cities. That architecture makes running traditional distributed systems software like Kafka difficult. Cloudflare describes the constraints plainly: for stateful services, it gets relatively small slices of machines, those machines are relatively ephemeral, and networking is often over the public Internet.

The same infrastructure also has advantages: it sits close to users wherever they are, and it scales horizontally. The design decision was to lean on a state primitive Cloudflare already operates — R2. Object storage combines very durable storage with strongly consistent APIs. By offloading replication and consensus to the storage layer, the application layer (K2) becomes simpler, cheaper, and higher performance. A secondary benefit is separation of compute and storage, so each can scale independently, which makes storing large amounts of historical data inexpensive.

## Building a log on object storage

Object stores do not support appends, which is the standard operation on a log. Instead, you must write complete files, or segments, large enough to amortize the cost of writing and reading each one.

K2 handles this by accumulating writes in memory on an edge service. After waiting a short period for data to arrive, it writes all events as a segment file. Ordering and strictly incrementing offsets are achieved using R2's atomic operations, without a separate coordination service.

That design has a cost. Writing to object storage is slower than writing to a local disk, and the system has to wait for the local batch to accumulate before starting the write. In the initial release, Cloudflare states this adds up to about 1 second of produce latency at the 99th percentile of response times. A more detailed technical deep dive is planned.

## Streams, Queues, or Pipelines?

Cloudflare already offers asynchronous delivery primitives, so the interesting question is when to reach for K2.

**K2 vs. Queues.** Both receive events, durably store them, and deliver them to consumers. Queues are designed around tracking individual items of expensive or time-consuming work that need to be completed asynchronously — for example, an image processing application enqueueing a user request. They support complex logic at the level of a particular work item, such as retries, delays, and dead-letter queues for failed attempts. K2 is designed for high-scale data movement, long-term retention, and fan-out consumption. Messages are produced and consumed as batches, which enables efficient processing at the expense of message-level retries. That batching is also what drives higher producer latency than Queues.

**K2 vs. Basin Pipelines.** Pipelines is a serverless ingestion service: you send it JSON events, which can be transformed and written to R2 or a Basin Catalog. Cloudflare recommends Pipelines when the end result is writing events to object storage or Iceberg tables, and K2 when you need custom processing or writing to other destinations.

## Producing and consuming

Streams are created through `cf`, Wrangler, the dashboard, or the API. A stream has an ID, a name, a retention period, an HTTP endpoint, and optional HTTP and Worker binding settings:

```bash
$ cf k2 streams create --name app_events --http-enabled
```

Producing can happen over the HTTP API or from a Worker binding. K2 represents data as bytes, so any format or encoding works:

```js
const result = await env.EVENTS.send([
  {
    content: new TextEncoder().encode(
      JSON.stringify({
        event: "page_view",
        path: new URL(request.url).pathname,
        timestamp: Date.now(),
      }),
    ),
    headers: { "content-type": "application/json" },
  },
]);

if (!result.success) {
  console.error(`Produce failed: ${result.error.message}`);
  return new Response("Failed to record event", {
    status: result.error.retryable ? 503 : 500,
  });
}
```

Consumption is organized around subscriptions. A subscription divides work between consumers, enabling read parallelism so you can scale out to multiple readers beyond what a single server can manage. Subscriptions are created via the HTTP API, and consumers poll them:

```bash
$ curl -X POST "https://<stream-id>.k2.cloudflarestorage.com/subscriptions" \
  -H "Authorization: Bearer ${CLOUDFLARE_API_TOKEN}" \
  -H "Content-Type: application/json" \
  --data '{ "name": "analytics_processor", "start_at": { "type": "earliest" } }'
```

When a client calls `consume`, it receives a lease on a batch of events for 5 minutes. The client can then:

- **ack** the batch, marking it processed so it will not be redelivered
- **nack** it, signaling processing failed and the batch should be redelivered
- **extend** the lease, if more time is needed

This is one consumption pattern: splitting work among multiple consumers so each gets a portion of the data. Another is the pub/sub pattern, where each consumer has its own subscription and sees all messages. The two can be mixed, with multiple independent consumer pools.

## Latency and consistency tradeoffs

The core tradeoff in K2 is between batching and latency. Accumulating events in memory before writing a segment is what makes object storage economical and what enables ordering and offsets via R2's atomic operations. It is also why produce latency is higher than a queue-based system — roughly 1 second at p99 in the initial release. Teams that need sub-second end-to-end latency for individual messages should weigh that against the durability, retention, and fan-out characteristics K2 provides.

A second tradeoff is granularity. K2 batches messages for production and consumption, which means message-level retries are not part of the model. If your workload is defined by individual work items that need independent retry, delay, and dead-letter handling, Queues is the better fit.

A third is the consistency model. R2 provides strongly consistent APIs, and K2 uses R2's atomic operations to achieve ordering and strictly incrementing offsets without a separate coordination service. That is a meaningful simplification compared with running a consensus-backed broker cluster, but it also means the design inherits the latency characteristics of object storage rather than local disk.

## Pricing and availability

K2 is available in public beta for accounts with Workers Paid subscriptions, within these limits:

- Maximum of 10GB of storage used
- 30 MB/s produce per stream

Usage is not billed during the beta period. Cloudflare anticipates the following pricing once billing begins:

| Item | Price |
| --- | --- |
| Data Produced | $0.04 / GB |
| Data Consumed | $0.04 / GB |
| Data Retained | $0.02 / GB / month |

## What's next

Cloudflare lists several items on the K2 roadmap: higher write parallelism up to multi-GB/s streams, message keys and key-based ordering guarantees, push-based Worker consumers, an Express tier with lower produce and end-to-end latencies, and drop-in support for Apache Kafka clients.

The Kafka client compatibility item is worth noting for teams evaluating migration paths. Until then, K2 is best understood as a different point in the design space: a serverless, object-storage-backed log that favors scale, retention, and operational simplicity over the low produce latency of a traditional broker.

