#062: What is Data Mesh? Explained
Apr 24, 2024Data mesh is not a technology and not a tool you can buy. It is a framework for organizing who owns data inside a company. Instead of one centralized data team being responsible for every integration, every pipeline and every model, each business domain owns its own: accounting owns accounting data, finance owns finance data, operations owns operations data. The central data team stops being the bottleneck in the middle and becomes the platform team underneath, providing the lake, the warehouse, the tooling and the permissions that let everyone else build. The word mesh describes what emerges, a network of domains building and referencing each other's data rather than funneling every request through one group.
Key takeaways
- Data mesh is a framework organizations attempt to follow, not a product. No vendor can sell you one.
- Domain means department or business unit. Accounting, HR, operations, finance, each one is a domain.
- Each domain configures its own sources, builds its own pipelines and lands its own data into the central lake.
- The data team's role shifts from building everything to providing infrastructure: the lake, the warehouses, permissions, tooling.
- The strongest argument for it is that the people who know the data best are the ones building the logic on top of it.
- The strongest argument against it is that you are asking non-specialists to build, monitor and govern pipelines well.
- dbt has embraced the pattern with multiple projects that reference each other, formalizing what used to be a package workaround.
What data mesh actually is
A quick search on the term returns an enormous number of articles about principles and implementations, and tools are now building the concept directly into their products. It is worth cutting through that first.
Data mesh is not a technology. It is a framework that organizations try to follow, which means two companies can both claim to be doing it and look completely different.
The structural change
You still have your usual sources. What changes is who is responsible for them.
In the centralized model, one data team handles every integration from every system. In a mesh, each domain configures its own sources inside the shared tooling and lands that data into a central destination, usually a data lake.
What "domain" means here
Domain is the word that trips people up, because it sounds more technical than it is.
When you hear domain, think department or business unit. A company might have a domain for each of these:
- Accounting
- Human resources
- Operations
- Finance
The idea is to silo responsibility along those existing lines instead of collapsing everything into one data team's backlog.
If you run a batch process today, in a mesh each domain would configure its own sources inside that tooling and own landing their data into the lake.
How the data team's role changes
This is the part that matters most if you are the data team, and it is easy to read it as a demotion. It is not.
You stop being responsible for knowing every business process in the company. You become responsible for the platform everyone builds on.
What the central team owns in a mesh
- Standing up and running the data lake and the warehouses
- Choosing and maintaining the shared tooling
- Handling permissions and access across domains
- Setting the standards and governance that keep domains compatible
Each domain is then responsible for actually setting up and creating the logic for its own sources. The split is infrastructure versus business logic.
Why it is called a mesh
You end up with a lot of domains working in separate spots, each one building its own models and its own analytics.
They do not stay separate. Domains pull from other domains to combine data and answer questions that cross departments, so the connections multiply in every direction. That network of cross-references is the mesh.
Ownership follows the data
The practical consequence shows up when something breaks. Instead of every issue routing to a central team, a problem with a given source belongs to the domain that owns it.
That domain investigates it, resolves it and changes whatever piping sits behind it. Everyone is owning their own piece and interacting with everyone else's.
The pros and cons
I think the idea is good in theory. Whether it catches on broadly is still being decided.
What works in its favor
- The people who know the data best build the pipelines. Nobody understands accounting data like the accounting team does.
- It removes an impossible expectation. A central team no longer has to know everything about every part of the business.
- Requests stop queuing in one place. Domains move at their own speed instead of waiting on a shared backlog.
What works against it
The hard part is technical. You are asking domain owners to set up pipelines, monitor them well and stay inside your standards.
That is a lot to ask of people whose job is accounting or operations. There are a lot of variables in play, and skills and appetite vary by team.
I have not yet seen this fully implemented at a company exactly as described. Teams are trying, but getting everybody on board is a large project and it takes time.
What an implementation looks like
There are plenty of good articles on tactical approaches, and teams are taking different routes. The clearest example I have seen is with dbt, which is a tool I use quite a bit, and a client recently asked me about it directly.
What I find interesting is that dbt has embraced the pattern completely and built it into the product. Their framing is the right one: it is not a single feature, it is a pattern.
Multiple projects that reference each other
The mechanic is that you create a number of separate dbt projects that can reference each other's models, with that cross-project reference formalized in the product.
This used to be a workaround. You would wire projects together with packages and various tricks, which worked but was never really designed for it.
How teams are structuring it
The trend is different teams owning different projects, often with a single enterprise or core project underneath them holding shared models.
Picture each team with its own dbt project, plus governance rules around it. Teams reference each other's projects, and model contracts define what a project promises to downstream consumers, which is its own conversation.
I think this makes the most sense for a transformation focused tool, since that is where the ownership question is sharpest. Whether other components of the stack adopt the same pattern is the interesting thing to watch.
Key terms
Data mesh
A framework where each business domain owns its own data sources, pipelines and models, and the central data team owns the platform underneath.
Domain
A department or business unit within a company, such as finance or operations, treated as the owner of its own data.
Domain ownership
The principle that the team closest to a data source is responsible for its pipeline, its logic and fixing it when it breaks.
Platform team
What the central data team becomes in a mesh: the group providing the lake, warehouses, tooling, permissions and standards for everyone else.
Cross-project reference
A formalized way for one dbt project to use models from another, which is how the mesh pattern is expressed in a transformation tool.
Common questions
Is data mesh a tool I can buy?
No. It is an organizational framework, so the real work is deciding who owns what and holding to it. Tools can support the pattern, dbt being the clearest example, but no purchase gets you a mesh.
Does a small data team need a data mesh?
Usually not. The problem data mesh solves is a central team drowning in requests from many departments, which is a problem of scale. With a handful of sources and one team, centralizing is simpler and faster.
What is the difference between data mesh and a data lake?
A data lake is a place where raw data lands. Data mesh is a decision about who is responsible for getting it there and modeling it afterward. Most mesh designs still use a central lake, they just distribute who feeds it.
Who handles governance in a data mesh?
The central data team, through standards rather than through doing the work. They define naming, quality expectations, access rules and contracts, and domains build inside those rules.
Has anyone fully implemented data mesh?
In my experience, not exactly as described. Plenty of teams have adopted pieces of it, especially domain-owned transformation projects, but the complete picture requires organizational buy-in that takes a long time to build.
Related reading
- Data Architecture 101: The Modern Data Warehouse
- Common Data Team Structures
- Data Architecture with dbt
- The 10 Key MDS Components: Part 1 (Essentials)
Final takeaway
Most of the teams I work with are nowhere near needing a mesh, and that is fine. What is worth stealing from it now is the ownership question: for every source you maintain, ask whether the people who understand that data should be the ones defining its logic.
Additional Free Resources
Starter Guides & Checklists
Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.