How to Build Modern Data Architectures Faster
Feb 26, 2025The fastest way I know to build a data architecture is to stop building it all at once. Pick one deliverable the business actually uses, plan the pipeline backwards from it, then build forwards from the source tables to that deliverable. You end up with one complete pipeline instead of twelve partial ones, and that pipeline carries your layers, your naming conventions, your testing and your core data model. I call this output led engineering, and the point is not speed for its own sake. Narrowing to a single output is what keeps the scope from running away while still giving the business something real early.
Key takeaways
- Building every pipeline in parallel produces activity, not progress. Do not mistake the count of models you shipped for value delivered.
- Start from the most used and impactful thing the business already relies on, and work backwards from it.
- One deliverable, built end to end, forces you to settle layers, naming conventions, testing, environments and the core data model once.
- That first pipeline becomes the template. After it exists, the team can work in parallel against a pattern everyone agreed on.
- Plan right to left, from the deliverable back to the sources. Build left to right, from the sources forward to the deliverable.
- Identify as few source tables as possible. Only what this deliverable needs, nothing speculative.
- The deliverable is a scoping device, not the goal. The goal is the pipeline and the process underneath it.
The wasted effort problem
A big moment for any company is deciding to really invest in a data architecture. Maybe it is greenfield, maybe there is something in place that needs an overhaul.
Either way it is exciting and a little overwhelming. There are dozens of sources to set up, expectations from stakeholders, and everything in the middle about how to model it and wire up an end to end pipeline.
Divide and conquer, and what it produces
The common response is to do it all at once. Split the sources across the team, build in parallel, and report progress as the number of models or pipelines built.
That is activity, not production. There is a John Wooden line I come back to here: never mistake activity for achievement.
Doing a lot of things does not mean the things are useful. With no consistent approach underneath, you mostly hurry up to make a mess.
The gap that opens up is between what you built and what is actually used. New stakeholder requests keep arriving, everyone keeps moving, and nobody has time to think about why any of it is being built.
Killing 2,000 reports
An example from my own experience, and the number is not a typo.
At the first company I worked at we had roughly 4,000 SSRS reports. SSRS is Microsoft's reporting tool, and like most reporting tools it makes creating a report very easy.
I could not believe we were using all of them, so I built a report on reports. It identified over 2,000 inactive SSRS reports.
Half of them or more were not used at all anymore, or used once a quarter, once a year, and a huge number were duplicates of each other. All I could think about was the time spent building them.
What we did instead
We decided to build a new data warehouse from scratch, to get away from the ad hoc, one off query habit that produced those 4,000 reports.
We began with the most used and impactful reports. We identified what the business was genuinely relying on and worked backwards from it, rather than aimlessly building from whatever source data happened to exist.
We did not do it all at once. We built the foundations, added on over time, and ended up with something robust that fed a lot of reports people actually opened.
The foundational pipeline
The alternative to scattering is to build one pipeline that goes all the way through, which is what I mean by a foundational pipeline and a minimum viable data model.
You still hit plenty of problems. You just hit them all in one place, and you solve them once:
- The layers of your pipeline and how they relate
- Automation and testing
- Naming conventions across the project
- The core data model: the keys, the relationships, the shape you want
It obviously cannot cover everything. It covers this scope, from beginning to end, which is far more than you get from poking at several pipelines at once.
Then you rinse and repeat. With one working example, the team can run in parallel against a pattern everyone has agreed on, and move much faster than they would have at the start.
Step 1: Narrow the scope to one deliverable
Output led data engineering is just a way to narrow scope by identifying a single deliverable that guides development.
It is not abandoning data modeling or process to rush out a result. The deliverable exists so the work does not get out of control in week one.
Deliverables that work well
- An existing report, dashboard or spreadsheet. The most commonly used one, like the report we started from at that first company.
- A grouping of similar metrics. With one client we picked three related metrics they were actively using and built around them.
- An overly manual process. Something the team does by hand today is a good candidate to pull into the new pipeline.
- A desired data set you do not have. Something the business wants that has to be built from scratch.
Do not overcommit
When you discuss this with stakeholders, resist the temptation to promise more than the one thing. Especially at the start.
The goal is to build the process and the pipeline. The report is how you scope that work, not the finish line.
Step 2: Reverse engineer the deliverable
Now plan right to left. Start at the deliverable on the right and walk backwards. Nothing is being built yet.
The mart
Ask what flat mart or presentation table would deliver that output. What columns do you want, and what are the ways you would want to slice it?
Try to shape it so it is not strictly one to one with this report. If the output is a sales report, what other reports or data sets could read from the same mart?
The fact
Next, what is the main fact behind that mart? By fact I mean the business action: the thing that actually happened.
This is where the star schema approach comes in, and the action you pick is the grain everything else hangs off.
The dimensions
Then the context around that action. Which dimensions do you need to slice the fact the way the mart requires?
Note the dimensions you can imagine adding later too. They are not needed for this deliverable, and seeing where they would plug in tells you the model is sound.
Staging and sources
Staging is one to one with the raw table, so these last two go together. Which source tables do you need to produce everything above?
Identify as few as possible. Only what this deliverable requires, which is the whole reason the scope stays manageable.
This is also where you set naming conventions and the landing zone separation, before anything loads. I cover that groundwork in Fix Your Data Pipeline From The Start.
A worked example
Here is the playground example I use when teaching this. The output is a daily NBA game summary report.
deliverable daily NBA game summary report
mart nba_games_detail -- also feeds other game level reports
fact fct_games -- the action: a game played
dims dim_teams, dim_date -- context around the action
staging stg_nba__games, stg_nba__teams
sources only the raw tables those staging models read
The staging convention there is stg, the source name, two underscores, the table name. It tells you immediately that it is a staging model, which source it came from and which table it maps to.
At the end of this step you have a game plan in front of you and you have not written a line of code.
Step 3: Implement left to right
Now reverse direction and build forward:
- Load only the source tables you identified into the landing zone, applying the naming conventions you chose.
- Build the staging layer on top of them.
- Set up the development environment and the workflow you want the team to follow.
- Create the warehouse tables, the fact and dimensions from your plan.
- Build the mart, then the final deliverable on top of it.
By the end you have a full pipeline and an example of more than the modeling. You have development environments, version control, automation and testing, all established by one deliverable.
Then you go back and do more with it. That is the minimum viable data model you rinse and repeat for everything that follows.
Key terms
Output led data engineering
Choosing one business deliverable and letting it define the scope of the pipeline you build, instead of starting from whatever source data you happen to have.
Minimum viable data model
The smallest complete model that supports a single deliverable end to end, including the keys and relationships the rest of the warehouse will reuse.
Foundational pipeline
The first pipeline built all the way through, which settles layers, conventions, testing and environments so later pipelines can copy it.
Presentation mart
The flat table that serves the deliverable, shaped so more than one report can read from it rather than being built one to one with a single report.
Right to left planning
Planning backwards from the deliverable through mart, fact, dimensions, staging and sources, then implementing in the opposite direction.
Common questions
How do I pick the first deliverable?
Pick what the business already leans on most, not what is easiest to build. An existing report with real usage, a small set of active metrics, or a painful manual process all work. Usage is the signal that the pipeline underneath it will get reused.
Is this the same as building a quick one off report?
No. The deliverable is a scoping constraint, and the actual output is a layered pipeline with conventions, tests and environments behind it. A one off report skips all of that, which is exactly how teams end up with thousands of reports and no model.
How long should the first pipeline take?
Long enough to get every layer right once, because you are deciding patterns the whole project will inherit. It will feel slow compared to parallel work, and the second and third pipelines are where that time comes back.
What if stakeholders want everything at once?
Show them one finished thing rather than progress on ten. Delivering a working output early buys more trust than a status report on partial pipelines, and it gives you a concrete reference for what the next request will take.
Does output led engineering work for a rebuild as well as greenfield?
Yes, and a rebuild often makes it easier. You already have reports in production, so usage data tells you which deliverable matters, and auditing what is actually used is usually the first worthwhile piece of work.
Related reading
- The First Step in a Data Architecture (Ingestion)
- How to Create a 3 Layer Data Model Pipeline
- The True Value of a Data Presentation Layer
- Managing the "End" of a Data Pipeline
Final takeaway
This is the approach I take with clients because it answers the two questions that stall a new architecture: where to start and when to stop. Choose one output the business already uses, plan backwards from it, build forwards to it, and let that pipeline become the template for the rest.
Additional Free Resources
Starter Guides & Checklists
Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.