#054: Why Data Migrations Go Wrong (3 reasons)

architecture Nov 08, 2023

Data migrations go wrong for reasons that have very little to do with the new tool. The three I run into most are copying bad code and bad practices into the new platform and expecting the platform to fix them, setting timelines so aggressive that developers take shortcuts to show progress, and staffing the project with people who have never used the tools before. Each one quietly cancels out the thing the migration was supposed to deliver, which is a more sustainable setup than the one you had. A migration is the scheduled opportunity to correct logic, restructure your model and agree on conventions. Treat it that way and the new stack really is an upgrade. Treat it as a lift and shift and you have paid to move the same problems into a nicer interface.

Key takeaways

  • Moving to a better tool does not fix poorly written queries, inconsistent formatting or contradictory logic. It relocates them.
  • The most common dbt migration mistake is pasting existing stored procedure logic into models and expecting modularity to appear on its own.
  • Unrealistic timelines do not produce speed. They produce shortcuts that get built on top of, and the finished migration is no better than what you started with.
  • No data project is ever truly done, so a migration planned as a one time sprint to a finish line starts from a false premise.
  • Being new to a tool means solving two problems at once: how the tool works and what the right logic is. Compressed deadlines bake the mistakes into the foundation.
  • You cannot replace experience with a quick training. One person who has done the migration before sets the practices everyone else follows.
  • Migrations are good projects to learn on. Just make sure learning is not the only thing holding the project up.

What usually goes wrong, and why it isn't the tool

A data migration is a project to move your existing pipelines, transformation logic and reporting onto a new platform. For most teams that means leaving a legacy or on premise stack for cloud tools.

Every year a new batch of teams starts one. I have been part of a handful in my career, and the disappointing ones rarely trace back to the technology.

They trace back to how the project was scoped, paced and staffed. Here are the three patterns, in the order they tend to show up.

Reason 1: Copying bad code and bad practices into a new tool

The assumption going in is usually that the new tool will sort things out. You move to something upgraded, and the old problems come out in the wash.

They don't. If your queries are poorly written, your formatting is inconsistent and your logic contradicts itself, none of that changes because it runs somewhere else now.

You might get short term relief from better performance or a nicer interface. Long term, you have the same problem in a new place.

The dbt version of this

I work on a lot of migrations to dbt. The team is coming from a graphical interface tool or from a pile of stored procedures, and the plan is to paste the existing logic into models.

That isn't a migration. That's relocating the problem.

dbt has its own way of building: modular models, clear layers, and references between models instead of hard coded table names. Pasted procedure logic ignores all of it.

What to do instead

Treat the migration as the one scheduled chance to fix things. You are already touching every piece of logic, so the extra cost of improving it will never be lower than it is right now.

  • Embrace how the new tool wants to be used. Break one long script into modular pieces rather than recreating the monolith.
  • Rethink the data model. Grain, naming and layering are all on the table while everything is being rebuilt anyway.
  • Fix the logic you already know is wrong. Every team has a few queries nobody trusts. Don't carry them across.

This takes longer than copy and paste. It is also the only version of the project that leaves you better off than when you started.

Reason 2: Overly aggressive timelines

Companies get excited about migrations. There are new tools, new processes, and stakeholders who want to see progress.

That excitement turns into a date. The date gets passed down to the developers actually doing the work.

What the pressure does to the code

None of this is an argument for moving slowly. Nobody should be taking forever, and being deliberate is not the same as being lazy.

The damage from an unrealistic deadline is specific. It shows up as shortcuts: skipped tests, pasted logic, naming nobody agreed on, structure decided in the moment.

Shortcuts compound. Each one is small, each one gets built on by the next task, and the finished migration ends up no better off than what it replaced.

Nothing is ever truly done

I have never seen a data project that was genuinely finished. There is always more work, and the platform keeps evolving underneath you.

A date that treats the migration as a single push across a finish line is working from a false premise. Plan it as a phase of work with room in it, not a sprint.

Pressuring people into closing individual tasks is the short sighted version. The long term question is what you want this system to look like in three years.

Reason 3: Lack of experience with the new tools

The third one is skill set, and I am not immune to it.

I was put on a project once where I had to build something from scratch in a tool that was new to me. There was a lot of expectation that I would simply be able to do it, and I struggled through it for a while before I got comfortable.

It had nothing to do with wanting to do a good job or understanding why we were migrating. It was a new tool, and that is just hard.

Two hard things at once

When the tool is unfamiliar you are solving two problems at the same time:

  • How does this thing work? Syntax, project structure, deployment, and all the quirks you only learn by hitting them.
  • What is the correct logic here? The modeling and business rules, which is the actual job.

Doing both under a compressed timeline is exactly how mistakes get built into the foundation of the new stack.

You can't shortcut experience

The failure mode I see most often is someone taking a quick training course and then being expected to know everything.

One person who has done this before changes the shape of the project. They put the right practices in at the start, and the rest of the team follows the pattern and runs with it long term.

If you are a team lead, that is the lever you control. Put people in a position to be successful instead of trying to save cost on experience and hoping it works out.

Migrations are still worth doing

Migrations are genuinely fun projects and a good place to learn on the job. Nothing will be perfect, and it doesn't need to be.

What matters is the mix. Sprinkle in people who have done it, give the work a realistic runway, and use the move as a reason to improve the code rather than relocate it.

Do those three things and the new stack is actually new. Skip them and you have paid for a change of address.

Key terms

Data migration

A project that moves existing pipelines, transformation logic and reporting from one platform to another, usually from a legacy stack to cloud tools.

Lift and shift

Moving existing code to a new platform unchanged, which carries the original design problems along with it.

Stored procedure

A sequence of SQL commands saved and executed inside the database itself, often on a schedule, and a common starting point for teams migrating to dbt.

Modular model

A transformation broken into small, referenceable pieces rather than one long script, which is how tools like dbt expect you to build.

Technical debt

Logic, naming and structure you know are wrong but keep working around, which a migration either clears out or copies forward.

Common questions

How long should a data migration take?

Longer than the first estimate, and it never fully ends. Plan it in phases with a realistic runway for each one rather than a single completion date, because a deadline that assumes a finish line pushes the team into shortcuts.

Should we rewrite our logic or copy it during a migration?

Rewrite the parts you know are wrong. You are touching every piece of logic anyway, so the extra cost of fixing it is lower now than it will ever be again. Copying it forward means paying for the move and keeping the problems.

Can we migrate to dbt without anyone on the team knowing dbt?

You can, but it is the most expensive way to learn. Bring in at least one person who has done it to set the project structure, naming and conventions at the start, then let the team take it from there.

Is a migration worth it if the current system still works?

Only if you plan to change how you build, not just where you build. The gain comes from modularity, version control and testing. If the plan is to recreate the same scripts in a new tool, the business case is thin.

What is the first thing to do when planning a migration?

Decide what you are going to fix. List the logic nobody trusts, the naming that is inconsistent, and the models at the wrong grain. That list is the real scope of the project, and it is what separates a migration from a copy and paste.

Related reading

Final takeaway

Every migration I have worked on that went well had the same thing in common: somebody treated it as a chance to rebuild rather than a chance to relocate. The tool you land on matters far less than whether the code, the timeline and the people on it were set up to produce something you would want to maintain.

 

Additional Free Resources

Starter Guides & Checklists

Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.

Browse Resources