Why a Good Data Model Makes Your Team More AI-Ready
Sep 30, 2026One of the best ways to get your data ready for AI is something that's been around for 30, maybe even 40 years by now. It has nothing to do with new tech or adding features. It's establishing a well-structured data model around your data. That might not sound exciting, but it's going to pay dividends against whatever your AI goals are, in three areas. Accuracy, cost, and scalability.
Key takeaways
- A data model defines the entities and relationships up front, so the AI tool doesn't have to work out what a customer is or how two tables relate.
- If you hand the logic to a coding assistant and something is wrong, you won't be able to explain what happened. The AI tool isn't the one answering to your C-suite.
- Reasoning is one of the biggest costs in AI. Every relationship you've already defined is reasoning the model doesn't have to pay for.
- Less reasoning means you may be able to point at a cheaper model, potentially even an open source one.
- A clear model makes troubleshooting a pinpoint exercise instead of digging through views on views on views. That's true with or without AI.
- Structure your data so that you can use the technology, not the other way around.
Why accuracy comes down to your data model
At the end of the day, we're only as effective as the data we provide and the insights we can give a business. If that data isn't accurate, or we consistently get called out by stakeholders over inconsistencies, we have a problem.
It doesn't matter whether you're using AI or something else to get the job done. If it isn't correct, or it's causing inconsistencies, that's a problem.
Why self-service analytics still needs a data model
When I talk to teams about what they want to do with AI, a lot of it comes down to self-service analytics. Giving your data product to stakeholders in some capacity so they can ask questions in natural language and get results.
This is picking up steam now with AI, chatbots and agents, but the concept has been around for years. I remember being on some of my first teams 10 years ago and self-service analytics was already at the forefront. Different products, different tools, all saying you can type in a question and it will give you the result.
The technology is much further advanced now. But those questions still have to look at data to get results.
And if we're being honest, a lot of teams want to introduce something like this because they think it gets rid of the need to do the complicated thing. The data modeling. Thinking through all the relationships. Dealing with the messiness of their data. They can throw an AI agent at it and assume it will figure things out, so they don't have to do as much.
That's the holy grail, you could say, of what teams are hoping for with AI.
How defining relationships makes your data accurate
The way to overcome this, in any scenario, whether it's self-service or reports or just data feeds, is to create a data model. Establish and define the relationships of the entities within your data. Say that this is what a customer represents. This is the actual event the business cares about. Rather than relying on AI to figure it out.
You can be intentional and clear about the business logic. You can still use AI coding assistants to build things faster, but you should be directing it. Establishing what the rules are and what the relationships are, so that by the time it gets passed to the next AI in the chain, whether that's a self-service tool or another integration, you can be confident that what's going out is accurate because you've already filtered and defined the logic.
This is something a lot of teams are skipping over or moving very quickly through for the sake of speed, and it's costing them the accuracy of their data. It feels good because you're moving quickly. But you're really just hurrying up to deliver poor results.
AI, self-service analytics and integrations are all great initiatives. This is a newer version of tech that's been around a while. But it all comes down to trust and reliability of the data. If you're not in control of that, you're going to have a bad experience, and you're going to lose the trust of the business.
Why you have to explain the logic, not the AI
I've worked with multiple clients over the past year who leaned very heavily into AI coding assistants. That's fine. We're all using them, and honestly you should be at this point to expedite yourself.
But they were abdicating their thinking of what the logic is, throwing it at the agent and letting it build. Then when they were asked to explain the logic, because something inevitably was incorrect, they didn't know what happened.
That's a slippery slope. Because it isn't going to be the AI tool that has to answer to your manager, to your C-suite, to your customers about the data integrity of what's going on. It's going to be you. You have to be able to explain it, or your engineers do.
Data modeling is one of the best ways to take control of your data while still getting the benefits of AI, because what you're serving up is something you've vetted, you understand, and you've established. Then it can take it, run with it, and be more effective with what you're giving it.
How a good data model lowers your AI costs
AI tools put a lot of technology at your fingertips. You can code quickly and connect things. But it isn't free. There are real costs for the reasoning, for the execution, for actually using it.
I've spoken with several close friends and former colleagues who work in finance departments, and they pretty much universally talk about how expensive tokens are becoming. For some, that line item is costing more than hiring somebody else. These are things that will get worked out over time.
Why less reasoning means you can use a cheaper model
If you have a solid data model, you've already defined the entities, established the facts and dimensions and the relationships. You don't have to think so much about how to use it. It's connecting keys together.
Relate that back to cost. One of the biggest costs is reasoning. Throwing a pile of information at something, saying figure this out for me, and having it internalize all of that. That amount of reasoning requires a more expensive model, which means you pay more.
On the other side, if you've already established a lot of this through clear modeling, you might be able to get away with a much cheaper model, because it doesn't need to do as much reasoning. You've already reasoned it out through the relationships. So you can point it at something less expensive, perhaps even an open source model.
I haven't personally explored that option yet and it's still very new. But in theory, if you've established those relationships with the goal of keeping cost down while still getting the output of the technology, why wouldn't you be able to use something like that? Just like with open source tools, there's a trade-off with maintenance and whether it's worth it. It's the idea of structuring your data so that you can use technology, and not the other way around.
Why a clear model cuts your maintenance time
The other note on cost is maintenance, which is nothing new. We've talked about this with open source versus cloud-hosted tools and with data concepts in general. The biggest cost is often maintenance. The human resource time for building and maintaining things.
If your data model is clear, it takes less time to debug and figure things out, because you know where to look. That gets expedited further with coding assistants. The agent will know where to look too, because everything is well organized. If it's clear for you, it's going to be clear for AI, which means everybody moves faster.
Without that, there's much more reasoning that needs to happen. You might be sorting through views on views on views, or fact tables joining other fact tables, or no modeling at all. Maybe a pile of stored procedures. The time it takes to find where something is happening curves in the wrong direction, and the same is true for an AI tool. It might try to solve it for you, and we all know that sometimes it fixes a problem without fixing what you wanted it to, and it can feel correct in the moment.
With a clear model you understand, troubleshooting becomes pinpoint precision instead of digging through the entire code base across multiple layers every time.
How a data model helps your team scale
When you have established conventions, things can be repeatable. When things are repeatable, your AI coding assistant is much more effective, just like another human would be. And using an AI tool alongside you is going to be faster.
So you spend less time reinventing the wheel. But that's only possible if you've already established conventions, what things mean, where things live, and ideally written it into a style guide. That's something you can point a coding assistant at and it will understand what to do and how to repeat it.
Think about scalability going in two directions. On one hand you've established the conventions and the clarity of your model, and it's easy for something to plug in and know what to do without reasoning too much. On the other, you rely completely on an agent to look at your raw data and figure it all out every single time. That requires more reasoning and more cost, and it keeps scaling in the wrong direction as more data and more complicated requests come in. You become more and more reliant on it, until you don't know what's going on. That's hard to recover from.
So think about how structured your data model is today, and how reliable it is for everyday purposes, before throwing AI tools at it and hoping it solves your problems.
I hear about this a lot right now. I've had calls with agencies and consulting companies trying to promote AI services to businesses, not even data related. In each case they run into the same thing. The company wants these things, promises get made, and then they look under the hood and the data isn't clear. They realize they need to structure the data before they can offer the service. A lot of the time they won't say that part directly to the customer, because customers don't want to hear it. But we know as data people that it matters.
There's a lot of talk about AI. Behind the scenes, the real work is still on modeling and structuring the data, the same as it's been for years. This is nothing new. It's a new way of seeing it, and hopefully it gives more importance to it.
Why tribal knowledge doesn't scale to AI
This one is especially important for smaller teams where one or two people control everything.
There's potentially so much information in your head that you're building on assumptions you already know. The AI agent doesn't know them, and neither does the next person on the team.
If you can define that knowledge into a clear model and clear documentation, everything scales the right way. If you don't, and then you leave the company or somebody else joins, it's much harder to explain. Or maybe it isn't another person. Maybe it's another AI agent you want to connect to the data. You need to get it out of your head and into some structure.
We've been talking about the tactical side of building as a developer, but what we're building here is for a business. I always come back to that. We're building data products and analytics for a business to use to make better decisions. AI just happens to be how it might be happening nowadays. The goal is to serve the business and to scale in a way where you can do more, serve more use cases, and deliver on requests faster.
Why a good data model isn't tied to one tool
When you create a good data model, it isn't limited to one tool. It isn't that you build a good model and can only use it on Snowflake, or only on BigQuery, or that it only relates to dbt or dbt Cloud.
Your data is agnostic to all of that. It can plug into different tools, but your structure and your definitions of what are facts and what are dimensions live outside of any one tool.
That gives you more options and leaves you prepared to use whichever tool makes the most sense, and to scale easily with it. Semantic views, semantic layers, whatever it is, all of it is built on top of a clear, well-defined data model that can plug into different use cases.
Key terms
Data model
A defined structure for your data that says what the entities are, what they mean to the business, and how they relate to each other.
Entity
A thing the business cares about, like a customer, an order or a product. Defining it means deciding what counts and what doesn't.
Fact
A record of something that happened, like a sale or a shipment. The measurable events you want to count.
Dimension
The descriptive context around a fact, like who the customer was or which product was involved.
Reasoning
The work an AI model does to figure something out before it answers. It's the part you pay the most for.
Semantic layer
A definition layer on top of your warehouse that tells tools what your metrics and entities mean. It's only as good as the model underneath it.
Tribal knowledge
Logic and assumptions that only live in someone's head. An AI agent can't read it and neither can the next person on your team.
Common questions
Can't an AI agent just figure out my data model for me?
It will try. But it has to reason its way through the relationships every single time, which costs more and gives you results you can't verify. And when a number comes back wrong, you're the one who has to explain why.
Does a data model actually reduce AI costs?
Yes, in two ways. Relationships you've defined up front are reasoning the model doesn't have to do, which means fewer tokens. And with less reasoning required, you may be able to use a cheaper model rather than the top-of-the-line one.
Do I need a semantic layer or a data model first?
The model first. A semantic layer is a set of definitions pointing at your warehouse. If the facts, dimensions and relationships underneath aren't clear, the semantic layer inherits the problem.
Is it a problem to build with AI coding assistants?
No, you should be using them. The problem is abdicating the thinking. Direct the assistant toward the rules and relationships you've decided on, rather than having it decide them for you.
Will a data model lock me into one warehouse or tool?
No. Your facts, dimensions and definitions live outside of any one tool. The same model can plug into Snowflake, BigQuery, dbt or whatever semantic layer you put on top.
Where this comes from
If data modeling is an area your team is struggling with, or you want a refresher on the concepts, I put together a free starter guide on data modeling.
And if you're leading a data team and you're being asked to deliver on AI, or you're looking at your architecture and want an outside opinion on whether it's ready, that's what I do. Get in touch here and we can talk through what you're working on.
Additional Free Resources
Starter Guides & Checklists
Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.