#058: Data Architecture 101: Lambda Strategy (Stream + Batch)

architecture Jan 03, 2024

The Lambda architecture is a data strategy that runs a batch path and a streaming path side by side. Most sources still get extracted and loaded on a schedule into the data lake, exactly as they would in a modern data warehouse. Alongside that, a stream, usually a messaging queue like Apache Kafka or change data capture on a database, carries records as they happen. That stream has two destinations: it writes into the data lake like everything else, and it also connects straight to analytics so reporting can show data in near real time. You get both speeds, and the price is two sets of transformation logic and a meaningful step up in complexity.

Key takeaways

  • Lambda is the middle option between all batch and all streaming. Some data arrives on a schedule, some arrives as it happens.
  • The streaming path is a messaging queue such as Apache Kafka, or change data capture that emits every insert, update and delete on a database.
  • The stream feeds two consumers: the data lake, so history is kept, and the analytics layer directly, so reporting is near real time.
  • The first real cost is duplicated logic. You maintain real time processing rules and batch transformation rules separately and keep them in agreement.
  • The second cost is complexity. Streaming needs different tools, different skill sets, more infrastructure and different data types to handle.
  • My recommendation is to start with the batch modern data warehouse and add streaming only when the business case is clear and the team accepts the maintenance.

The decision behind the architecture

A huge decision for any data team is how fast you want to refresh data. The common answer is a batch schedule: hourly, daily, whatever fits the reporting.

For some teams that is not fast enough. There is data they would rather see without waiting for the next batch window.

Lambda is the option that does not force you to choose. It combines batch scheduling and real time streaming in one architecture, and it is a reasonable answer when only part of your data needs to be fast.

What the Lambda approach looks like

Most of the components are the same ones you already know. Sources, a data lake, transformations, a warehouse, data models and analytics.

The change is that some components are split in two. One version of the path runs as a batch process, the other runs as a stream.

What counts as the stream

A stream can mean a messaging queue, which is the Apache Kafka style setup where producers publish records and consumers read them. It can also mean change data capture on a database, which sends out a message for every insert, update and delete on a record.

Either way, that information travels through a different process than your scheduled loads do.

The two paths

  • Batch path. Sources are extracted on a cadence and loaded into the data lake, then transformed into the warehouse and into data models.
  • Streaming path. Records flow into the data lake as they happen, so history is still captured in the same central place.
  • The direct connection. The stream also runs straight to analytics, so a report can show records seconds after they are created.

That direct line is the whole point of the design. Without it, a stream is just a faster way to fill the lake.

What stays the same

The batch half is the modern data warehouse you already know. Sources on a schedule, a lake that keeps history, transformations into a warehouse, models for reporting.

Lambda does not replace that design. It adds a second lane next to it, which is why it is a reasonable next step rather than a rebuild.

Processing the stream

Streamed records are usually JSON, though the exact shape depends on the tool. The transformation section needs something that can apply business logic to records in motion rather than to tables at rest.

Things to consider before you build it

This gives you the best of both worlds, and that phrasing hides two real costs.

1. You double up on transformation logic

You end up maintaining the same business rules in two places: one set in the real time processing layer, one set in the batch transformation layer.

Both have to stay in agreement. If a definition changes in one and not the other, the real time number and the daily number stop matching, and that lands back on the data team.

2. Complexity goes up across the board

Streaming is not a plug and play batch tool with built in connectors. It asks for more from the team and from the environment:

  • Different tools. Queues, consumers and stream processors rather than a scheduled loader.
  • Different skill sets. Often general purpose programming rather than SQL alone.
  • More infrastructure. There is more to host, monitor and keep running.
  • Different data types. Nested JSON payloads behave differently from the structured rows a batch load hands you.

All of that costs time and skill, which means it costs money. The question for your team is whether getting data faster is worth that investment, or whether you would be just as effective doing everything in a batch.

My default recommendation

Start with the batch modern data warehouse approach. Move to streaming only when you feel it is genuinely necessary and the team understands what it takes to keep it running.

Decide it per source

The useful version of this question is not whether your company needs real time data. It is which specific source needs it.

Usually one or two systems carry the events people want to see immediately, and the rest are perfectly fine on a daily load. Streaming only those keeps the duplicated logic to a small surface.

Example stacks

These are examples rather than recommendations. The point is to see the shape with familiar names attached.

Fivetran, Kafka, S3, Redshift and Spark

On the batch side, a tool like Fivetran connects to your sources and lands them in the data lake, which here is an Amazon S3 bucket. On the streaming side, Apache Kafka carries events from the sources.

That Kafka stream has two consumers. One is the data lake, the other is the analytics platform, so the same records land at two different cadences for two different purposes.

From the lake, data moves into Amazon Redshift as the warehouse. Apache Spark does the processing, both the real time transformation and the batch transformation, which at least keeps the two code paths in one tool.

Why Spark shows up here

Spark handles streaming and batch processing, and you write in languages like Python, so your options go beyond SQL. You can even use it for the extract and load step.

Think of it as another tool in the toolbox. In this example it is doing double duty so there is consistency across both paths.

Flink for streaming, dbt for the warehouse

Same architecture, different split. Apache Flink handles the real time processing, since that is what it is focused on, and dbt handles the warehouse transformations.

This is the mix and match version, and it is a fair trade. You use the better tool on each side, and you accept that you are managing two separate code bases.

Key terms

Lambda architecture

A data architecture that runs a scheduled batch path and a real time streaming path in parallel, with the stream feeding both the data lake and analytics.

Messaging queue

A streaming service such as Apache Kafka where source systems publish records and multiple downstream systems read them independently.

Change data capture

Reading a database's own log of inserts, updates and deletes so each change is emitted as a message instead of being picked up by the next scheduled extract.

Stream consumer

Any system that reads from the stream. In a Lambda setup there are usually two: the data lake for history and the analytics layer for live reporting.

Real time processing

Applying business logic to records as they move through the stream, rather than to tables already sitting in the warehouse.

Common questions

What is the difference between Lambda and Kappa architecture?

Lambda keeps both a batch path and a streaming path. Kappa drops batch entirely and streams everything, which removes the duplicated logic but raises the bar on every source you connect. Lambda is the practical middle ground when only some of your data needs to be fast.

Do I need Kafka to build a Lambda architecture?

No. Kafka is the common example of a messaging queue, but change data capture on an operational database gets you a stream too. The architecture cares that records arrive continuously, not which product delivers them.

Why does the stream write to the data lake as well as analytics?

So you keep history in the same central place as everything else. The direct line to analytics is for the live view, and the copy in the lake is what your warehouse models and backfills read from later.

Is Lambda worth it for a small data team?

Usually not as a starting point. Build the batch architecture first, then add a stream for the specific use case that justifies it. Adopting streaming for the whole stack before you have that case tends to cost more than it returns.

Can one tool handle both the batch and the streaming side?

Apache Spark can, which is why it appears on both paths in the first example. That keeps your logic in one language and one code base, though you are still writing and testing two processing flows.

Related reading

Final takeaway

In the teams I work with, the Lambda design earns its keep when one or two sources genuinely need to be live and the rest are fine on a daily load. Decide that per source, not for the whole stack, and you keep the complexity contained.

 

Additional Free Resources

Starter Guides & Checklists

Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.

Browse Resources