#035: Why Data Teams Need Version Control

Mar 29, 2023

Version control is a system, usually git with a platform like GitHub or GitLab, that records every change to your code along with who made it and when. Data teams need it for three reasons. It keeps a full history, so you can see exactly what changed and revert when a deploy breaks something. It's the foundation for automation, because CI/CD workflows run from the repository and replace big release nights with small continuous releases. And it lets several people work on the same project without silently overwriting each other. Pushing changes straight to production feels fast in the short term, but a repository is the cheapest insurance a data team can buy.

Key takeaways

  • The fastest route to problems is pushing changes straight to production. It feels easy and it always costs you later.
  • Version control tracks every change at the line level, with the author and the timestamp. The change log comment block at the top of a stored procedure becomes unnecessary.
  • If a deploy breaks something, you revert to the previous version. Nothing is lost forever.
  • Code based tools are designed to be automated. GitHub Actions and GitLab pipelines run tests and deployments from a workflow file in the repository.
  • Continuous releases replace release nights. Stakeholders get their changes faster and in smaller pieces.
  • Two developers editing the same file get a conflict flagged immediately instead of overwriting each other's work.
  • Commit messages with a task ID give every change a consistent, traceable record.

The problem with going straight to production

I'm not above this. For years I worked on teams where saving a query meant updating production.

It seems easy in the short term. In the long term it leads to problems, and not ones you'll be proud to explain. Either a change gets lost and nobody can reconstruct it, or a small bug fix turns into a drawn out release night.

Version control is the way around it. Here are the three reasons your team should absolutely be using it.

Reason 1: Track every historical change

Version control records every individual step a piece of code has taken. That gives you transparency into what changed, who changed it, and when.

Most data tools today are open source and code based. dbt models, SQL files, Python scripts, YAML configs. Everything that matters is text, so everything can be tracked, and that's a feature you don't want to be missing.

Revert when something breaks

If you deploy something that breaks, it's not the end of the world. You go back to the previous version and pull it from the repository. You don't lose everything.

What a diff shows you

Open any commit and you get a diff: the exact lines removed and the exact lines added, side by side with the previous version. For a 400 line model, that's the difference between knowing one filter changed and rereading the whole file.

The comment block you can delete

One thing I still see teams doing: a comment block at the top of a stored procedure listing each change and who made it, in plain text.

-- CHANGE LOG
-- 2023-01-12  MK   Added region filter to the final select
-- 2023-02-03  JS   Fixed null handling on order_date
-- 2023-03-15  MK   Renamed cust_id to customer_id

Version control eliminates the need for this. It tracks who and when automatically, and it shows the specific lines and characters that were touched, not just a general description someone remembered to write.

So the first reason is simple. Track your history, revert when you need to, and never again save something and lose it forever.

Reason 2: Automate your workflow

Because the tools are code based, they're built to be automated. The common example is CI/CD, continuous integration and continuous deployment: changes are tested automatically when proposed and released automatically when approved.

Platforms like GitHub and GitLab have this built in. You create a workflow file inside your project, and the platform runs a script automatically once code is merged or checked in.

  • GitHub: GitHub Actions
  • GitLab: GitLab pipelines
  • Anything else: the same concept is a standard feature on most version control platforms

The end of release nights

If you're like me, you've spent a lot of days and nights staying up for a release, packaging all of the code into one big deployment.

With version control and these automation features, there's no reason for that anymore. Releases happen on a continuous, near real time basis, in small pieces, as changes are merged.

From a stakeholder's perspective, you're getting their changes to them much faster than before. That's the part they notice.

What a workflow looks like

A minimal GitHub Actions file that builds and tests a dbt project on every pull request looks like this:

name: CI

on:
  pull_request:
    branches:
      - main

jobs:
  build-and-test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v3
      - uses: actions/setup-python@v4
      - run: pip install dbt-snowflake
      - run: dbt build --target ci

That's the whole idea: the repository is the trigger. I walk through this file line by line in the post on CI/CD with a real life example.

Reason 3: Work better as a team

Most of us now work remote or hybrid, and that brings new challenges to work around as a development team.

Conflicts get called out

Picture two people working on the same file. Without version control, they both put their changes live, one overwrites the other, and nobody is sure whose version is running. A lot of time gets wasted figuring it out.

With version control, the conflict is flagged immediately so you can get to the root of it. Even if the two of you worked at completely different times, it still calls out the overlap. You never unknowingly override somebody's changes.

A review step for free

The platform adds a pull request on top: a colleague reads the change before it merges. Combined with the automated check from reason 2, that's two sets of eyes on every change, one human and one machine, before anything reaches production.

Consistent, traceable changes

Version control also lets you add some consistency and formality to how changes are described.

A common example is including the task number or ticket ID from your project tracker in every commit message:

DA-142: add fiscal quarter to dim_date
DA-143: fix null handling on fct_orders.order_date

Every commit then follows the same structure. The change is connected to the request that caused it, and anyone can tell exactly what it was for.

As a team, version control isn't only about the individual code. It helps you work together and avoid the unnecessary conflicts that lead to wasted time and unexpected results.

No excuse anymore

Hosted platforms are free for small teams, every modern data tool is built around a repository, and the three benefits compound: you track your history, you automate your workflow, and you work better together.

Nowadays there's really no excuse for a data team not to be using version control.

Key terms

Version control

A system that records every change to a set of files over time, with the author and timestamp, so any version can be viewed, compared or restored.

Commit

A saved snapshot of changes in version control, with a message describing what changed and why. Commits are the units of history you can inspect and revert.

Revert

Restoring code to a previous commit after a change breaks something, without losing the history of what was tried.

Merge conflict

What version control raises when two people change the same lines of the same file, so the overlap is resolved deliberately instead of one change silently replacing the other.

CI/CD

Continuous integration and continuous deployment: automation that tests proposed changes and releases approved ones, run by the version control platform from a workflow file in the repository.

Common questions

Do data analysts need version control, or just engineers?

Anyone who writes SQL that other people depend on benefits from it. Analysts who maintain dbt models, saved queries or dashboard logic lose work and overwrite each other just as easily as engineers do. The learning curve is a few commands, and most platforms have a visual interface for the rest.

Does version control work with SQL and stored procedures?

Yes. Version control tracks text files, and SQL is text. Keep each stored procedure, view or model in its own file in the repository and deploy from there. The change log comment block at the top of the procedure can go, because the history lives in the commits.

What's the difference between git and GitHub?

git is the version control software that runs on your machine and tracks the changes. GitHub, GitLab and Bitbucket are hosted platforms that store your git repository, add pull requests and permissions, and run automation like GitHub Actions or GitLab pipelines on top.

How does a data team get started with version control?

Create a repository, commit the SQL and config you already have, and agree on one rule: nothing changes in production unless it went through the repository first. Add a pull request step so a second person reviews changes. Automation and branching strategy can come later.

Can I still fix something quickly in production?

Yes, and faster than before. A hotfix is a small commit merged through the same workflow, and the automation deploys it in minutes. What you lose is the habit of editing production directly, which is exactly the habit that causes untraceable breakage.

Related reading

Final takeaway

When I start with a new client, the first thing I check is whether changes reach production through a repository or through someone's editor. Everything else I recommend, testing, automation, environments, is built on top of that one answer. Get the repository in place first.

 

Additional Free Resources

Starter Guides & Checklists

Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.

Browse Resources