#059: Data Architecture 101: Kappa (Real-Time Data)

architecture Feb 07, 2024

The Kappa architecture is the all streaming end of the data architecture spectrum. Every source is captured in real time, usually through a messaging queue or change data capture, and pushed into a central landing zone. From there the data moves through the warehouse layer, the data models and out to analytics, with business logic applied as the records travel rather than during a nightly run. There is no batch path at all, which means there is only one set of processing logic to maintain. The catch is complexity: not every source is easy to stream, the hosting is heavier than a plug and play batch loader, and most teams do not have a use case that justifies it.

Key takeaways

  • Kappa streams everything. Sources, the landing zone, the warehouse layers, the models and analytics all operate in near real time.
  • The series reads as an evolution: the modern data warehouse is all batch, Lambda mixes batch and stream, Kappa is all stream.
  • Removing the batch path removes the duplicated transformation logic that makes Lambda expensive to maintain.
  • The hard part is the sources. Not all of them can be configured for streaming or change data capture without real engineering effort.
  • Streaming infrastructure asks for general purpose programming rather than the out of the box connectors a batch tool gives you.
  • If stakeholders act on data once a day, real time changes no decisions. That is the test I apply before recommending this.

Why real time is so appealing

Given the choice between data updated once a day and data updated in real time, most of us pick real time. It is the holy grail of data engineering and the thing most teams say they would love to do.

All things are not equal, though. Different architectures demand different skill levels and infrastructure, which drives design and maintenance time, which drives total cost of ownership.

Just because something is technically possible does not make it the right choice for your team. In my experience it is usually not necessary, but you should still be able to recognize it and evaluate it honestly.

What a Kappa architecture looks like

At a high level the shape is simple. Your sources sit on the left, and every one of them is streamed or captured in real time into a central landing zone.

Call that landing zone a data lake if you like. The idea is one central place where everything arrives, the same role it plays in the other architectures.

Everything after the landing zone is also streamed

From there the data moves into the warehouse layer, then into data models, then into analytics. All of that movement happens through the stream.

Business logic gets applied as records come through. They are stored in one place, moved to the next, and surfaced, all in near real time.

What stays the same

You still have the same layers you would have anywhere else. A landing zone, a warehouse, data models, a reporting tool on the end.

What changes is the engine. Nothing waits for a scheduled run, so every layer is being updated continuously rather than refreshed in a window.

Where it sits in the series

  • Modern data warehouse. Everything batched on a schedule, once an hour or once a day.
  • Lambda. A mixture. Some sources pulled on a schedule, a stream running alongside that feeds the lake and analytics directly.
  • Kappa. All stream, no batch, one processing path from source to report.

Lambda splits processing into two places, real time and batch. Kappa collapses that back into one, which is its genuine advantage.

The upside of dropping batch

One path means one place for your logic. The duplicated transformation rules that make Lambda expensive to maintain simply do not exist here.

If you were already going to stream a large share of your sources, that consolidation is a real argument for going all the way.

Considerations before you commit

This sounds great in theory. Here is what to weigh before building it.

Complexity, starting with the sources

We would all love every source available in real time through a stream or change data capture. In practice, not all of them are easily configured that way.

The technical setup and ongoing maintenance to make that happen is substantial, and it is per source. Each system you connect is its own project.

Hosting and skill requirements

The architecture for hosting these components is more complicated than a traditional ingestion tool. Compare it to a batch tool that is plug and play, ships with built in connectors, and runs once or twice a day.

Kappa needs more than that. It generally means general purpose programming languages instead of out of the box configuration, which changes who you need on the team.

Cost of ownership

Higher skill requirements and heavier infrastructure show up as design time and maintenance time. That is what pushes up the total cost of owning the architecture, well past the tool invoices.

The use case test

None of this means you should not do it. It depends entirely on whether the use case warrants the maintenance and complexity required. Sometimes it does.

Here is the test I use. If your stakeholders only need data updated once a day or every couple of hours, and no new decision gets made in between, real time is not buying anything.

In that case stick with the modern data warehouse, or use Lambda to get a little of both.

Example architecture

Do not get attached to the specific tools. There are other ways to build this, and the point is to see the shape with recognizable names on it.

The ingestion side

Say you have a handful of sources feeding a Kafka stream, the messaging queue carrying records as they are created. Some of them might instead be change data capture coming off a database.

All of that gets pushed into Databricks, where a storage location serves as the landing zone.

Notice that the sources did not change. What changed is that each one now has to emit records continuously, which is the part that takes engineering work per system.

Processing and reporting

From there the data moves through the different layers of the Databricks instance using Apache Spark as the real time processing tool. Spark is fairly integrated with Databricks, which is why it fits here, though it is not the only option.

The final layer of data models gets surfaced through a reporting tool, in this example Looker. There is no batch loading anywhere in the diagram, and the reporting tools read live data.

What that means day to day

Any time a new record is generated, it comes through the stream, gets processed, and moves all the way along. There is no window to wait for and no run to kick off.

Compare that to the other two designs, where some or all of the data arrives on a schedule, and the difference in what your team has to operate becomes clear.

Key terms

Kappa architecture

A data architecture with a single streaming path from source to analytics and no batch loading, where every record is processed in near real time as it arrives.

Central landing zone

The one place every source lands before anything downstream reads it, often called the data lake, regardless of whether the data got there by batch or by stream.

Near real time

Data that reaches reporting seconds or less after it is created, as opposed to waiting for the next scheduled load.

Change data capture

Emitting a message for every insert, update and delete on a database so changes flow out continuously instead of being collected by the next extract.

Total cost of ownership

The full cost of an architecture over time, including skill requirements, infrastructure, design effort and ongoing maintenance, not just the tool invoices.

Common questions

What is the difference between Kappa and Lambda architecture?

Lambda keeps a batch path and a streaming path running side by side, which means two sets of transformation logic. Kappa removes the batch path entirely, so there is one code path, but every source has to be streamable. Lambda is the compromise, Kappa is the commitment.

Does Kappa mean I have no data warehouse?

No. You still have a landing zone, a warehouse layer and data models. The difference is that data moves between those layers through the stream rather than through scheduled jobs.

Is Kappa realistic for a small data team?

Rarely, and I usually find it is not necessary. The setup per source, the hosting and the programming skills needed add up to a cost of ownership that small teams feel quickly. Start batch, and add streaming where a specific decision depends on it.

What tools are used in a Kappa architecture?

A messaging queue such as Apache Kafka or change data capture for ingestion, a stream processor such as Apache Spark for the transformations, a platform such as Databricks to host the layers, and a reporting tool on top. Those are examples, not requirements.

How do I know if my team needs real time data?

Ask what decision changes because the data arrived sooner. If someone acts on it within minutes, there is a case. If the answer is that a dashboard looks fresher, the daily batch is doing its job.

Related reading

Final takeaway

Across the small data teams I work with, I have yet to see one where going fully real time was the bottleneck holding back decisions. Know what Kappa is so you can evaluate it, then be honest about whether anything downstream would actually change.

 

Additional Free Resources

Starter Guides & Checklists

Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.

Browse Resources