Data Architecture 101: The 5 Key Questions
Feb 25, 2026
Data architecture is a big topic, but you can get your arms around it by answering five questions. Where will we consolidate our data, how will it get there, how will we clean it, how will we analyze it, and how will we automate it. Each question maps directly to one component: a cloud database, data ingestion, data transformation, reporting, and version control. Answer the five and you have an end to end architecture, with everything else treated as an addition you make later if you need it. That is the sequence I follow with clients, because the foundation has to be solid before anything gets added on top.
Key takeaways
- Our job as data engineers is to give the business data it can make better decisions with. Everything else is extra.
- Five questions define the architecture: where data is consolidated, how it gets there, how it is cleaned, how it is analyzed, and how it is automated.
- Those five questions become five components: a cloud database, ingestion, transformation, reporting and version control.
- Default to one cloud database for storage. Adding cloud file storage in front is a common pattern, but it is an extra layer to maintain and one more place things break.
- Automation starts with version control and CI/CD, plus the scheduling already built into your loading and transformation tools. An orchestrator is a later addition.
- Separate the layers with databases or schemas: raw landing zone, staging, warehouse, marts. If you cannot query across databases, use one database and prefix the schemas.
- It is far easier to start simple and add functionality than to start with a pile of tools and unwind it later.
The five questions
These questions outline what data architecture is really for. Our job as data engineers is to provide the business with data so it can make better decisions and operate more efficiently.
1. Where will we consolidate our data?
Businesses run on many different systems. One of the most important things you do is bring them together, so the first decision is where that consolidation happens.
2. How will it get there?
Something has to connect to each source and move the data into that destination on a regular basis.
3. How will we clean it?
You are dealing with raw data from different systems, each with its own design and its own quirks. Something has to reshape all of that into a consistent model.
4. How will we analyze it?
Cleaned data is only worth the decisions it supports. This question is about how people actually see and use the result.
5. How will we automate it?
This applies to almost every part of a business, and as data engineers we know there are ways to do it. The question is which parts you automate first.
From questions to components
Each question becomes a box in the architecture. That is the whole translation step, and it is why the five questions are useful rather than philosophical.
- Where will we consolidate our data? A cloud database server.
- How will it get there? Data ingestion.
- How will we clean it? Data transformation.
- How will we analyze it? Reporting.
- How will we automate it? Version control.
Other components can be added later as you need them. A simple stack is built around these five because once they are in place you have a foundation you can scale, answer questions with, and serve the business from.
Two clarifications before you build
On storage
I recommend a cloud database. Many of them are designed for large volumes and for different types of data, so scale is unlikely to be your problem.
You can also use cloud file storage and then load into the database afterward. That is a very common approach, and it is the right call when you handle video, images or documents that do not fit as text.
Just know what it costs you. It adds a layer and a step, which means maintenance, team time, and more places where things can go wrong. Start with everything in the database and add the rest when you need it.
On automation
When I say version control for automation, I mean workflow automation: how you automate testing of your changes and how you deploy quickly.
You can run a lot of daily updates straight from a version control platform without any additional tooling. If you want something more robust later, an orchestrator is there.
Your loading and transformation tools also have their own scheduling built in. Use those first. Why add another tool before you need it.
On strategy over tools
There are endless methods, tools and languages that will get the job done. If you do not know where you are going and why, none of that matters.
The end to end architecture
Here is the full picture, lined up with those five components.
Step 1: the cloud database server
One server holds everything. I prefer a database for this because it removes unnecessary hops between systems, and most cloud databases handle far more data than we are likely to throw at them.
Step 2: data ingestion and the landing zone
Ingestion has two halves. On one side, source connectors reach out to your systems. On the other, the data lands in a target landing zone.
That landing zone is for raw data only: unformatted and straight from the source. Inside the server you have the database, and inside that, schemas that separate and isolate each source.
Step 3: data transformation
This is where raw data becomes cleaned data models. I keep it in a separate database so there is an actual process, instead of everything mixed together in one schema.
There are three levels, each one its own schema:
- Staging. The first pass over the raw data.
- Warehouse. The typical data warehouse with facts, dimensions and real data modeling.
- Marts. Often split by user group, so finance, operations and sales each get their own, or just one if that suits you.
The marts layer is what users interact with. Alongside all of this, each developer on the team gets a development environment, a safe space to work.
Step 4: version control and CI/CD
CI/CD means continuous integration and deployment, and it is how testing, validation and release happen. Version control also gives you a complete history of every change you and the team have made.
That comes with a test environment. Call it CI, UAT or test, the role is the same: a pre production environment that sits between development and production.
Step 5: reporting
This is how the company reports on the data. It could be a dedicated reporting tool, and it could be Excel.
Ideally it pulls only from the marts layer. That is what keeps the definitions consistent no matter who is asking.
When you cannot query across databases
Some platforms will not let you query easily across separate databases, and Postgres is a common case. The idea still holds.
Use a single analytical database and prefix the schemas instead, so your raw schemas stay isolated from staging, warehouse and marts without needing separate databases.
Key terms
Simple stack
An architecture built from only the five components the five questions produce, with anything else added later and only when the need is real.
Landing zone
An isolated area for raw data exactly as it came from the source, unformatted and separated by schema per source.
Source connector
The half of ingestion that authenticates to a source system and pulls its data, as distinct from the half that writes it into the landing zone.
CI/CD
Continuous integration and deployment: the automated testing, validation and release process that runs through a pre production environment before changes reach production.
Schema prefixing
Keeping raw, staging, warehouse and mart layers in one database by naming their schemas with a prefix, used when cross database queries are awkward.
Common questions
What are the five components of a data architecture?
A cloud database for storage, data ingestion to get data in, data transformation to clean it, reporting so people can use it, and version control to automate the workflow. Each maps to one of the five questions, which is what makes the set easy to remember and hard to leave gaps in.
Do I need a data lake or is a cloud database enough?
A cloud database is enough for most teams, and it avoids an extra hop. Add cloud file storage when you handle video, images or documents that cannot be stored as text. Otherwise you are maintaining a layer that is not earning its place.
Why is version control the answer to automation?
Because the automation that pays off first is workflow automation: testing changes and deploying them quickly. A version control platform handles that, and it can run scheduled updates too, so many teams never need a separate orchestrator early on.
How many environments do I need?
Three roles. A development environment per developer, a pre production environment for testing and validation, and production. They can live in the same platform, separated by database or schema.
When should I add an orchestrator?
Once the scheduling inside your loading and transformation tools stops being enough, usually when dependencies cross tools and need coordinating. Until then, an orchestrator is another system to run for no new capability.
Related reading
- A "Simple Stack" Data Architecture, Explained
- Data Architecture 101: The Modern Data Warehouse
- Data Architecture 101: Lambda Strategy (Stream + Batch)
- Data Architecture 101: Kappa (Real-Time Data)
Final takeaway
This is the exact process I walk through with the teams I work with, and the five questions are where every one of those engagements starts. Get the foundation strong, then add components when the need shows up rather than before.
Additional Free Resources
Starter Guides & Checklists
Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.