#042: Data Automation (CI/CD) with a Real Life Example
May 17, 2023CI/CD stands for continuous integration and continuous deployment, and for a data team it means every change gets tested and released by an automated workflow instead of by hand. The practical version is a YAML file in your GitHub repository that runs on every pull request: it checks out the branch, installs dbt, builds the whole project into a CI schema and runs every test. If the build passes, the pull request gets a green check and you merge with confidence. If it fails, you find out before anything reaches production. This is a big part of why code based tools like dbt are worth the switch, and the whole setup is one file and a couple of secrets.
Key takeaways
- CI/CD is automated testing and deployment. For a data team it means a pull request builds and tests your models before anyone merges them.
- The workflow is one YAML file at
.github/workflows/ci.yml. GitHub Actions picks it up automatically. - Four sections do the work:
name,onfor the trigger,jobswith the runner and environment variables, andsteps. - Don't write every step yourself. Marketplace actions like
actions/checkoutandactions/setup-pythonhandle the setup. - Credentials come from GitHub secrets passed as environment variables, so nothing sensitive sits in the file.
- The build runs against a dedicated CI target and schema, separate from both development and production.
- Once it passes you get a check mark on the pull request. Every future pull request gets the same catch-all check with no extra effort.
What CI/CD looks like for a data team
One of the most fun parts of being a data engineer is building automations, and CI/CD is the one with the biggest payoff. Continuous integration means proposed changes are tested automatically. Continuous deployment means approved changes are released automatically.
The concept stays vague until you've seen it run. So this walkthrough builds a working CI workflow for a dbt project on Snowflake using GitHub Actions, the automation service built into GitHub.
The focus here is data quality checks: every pull request builds and tests the project before it can merge. The same pattern extends to deployments once you have it in place.
Create the workflow file
GitHub looks for workflows in one specific place. Create a folder called .github/workflows at the root of the repository. Every workflow you write lives there as its own file.
Workflows are YAML files. YAML is a plain text format for structured configuration, and here it's the set of instructions GitHub follows. Create one called ci.yml.
The goal for this file is simple. Every time a pull request is opened, read the branch that's trying to merge, build it, and test it.
Review the file layout
Here is the full workflow, cleaned up. The sections after it walk through each part.
name: CI
on:
pull_request:
branches:
- main
workflow_dispatch:
jobs:
build-and-deploy:
runs-on: ubuntu-latest
env:
DBT_USER: ${{ secrets.DBT_USER }}
DBT_PASSWORD: ${{ secrets.DBT_PASSWORD }}
steps:
- name: Check out branch
uses: actions/checkout@v3
- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: "3.10"
- name: Install dbt Snowflake
run: pip install dbt-snowflake
- name: Deploy and test models
run: dbt build --target ci
name and on
name is the label the workflow shows under the Actions tab once it's deployed. It doesn't have to match the file name.
on defines the trigger. Here it's a pull request, but only one targeting main. A pull request into any other branch won't fire it. GitHub supports plenty of other triggers, and the Actions documentation lists them.
workflow_dispatch adds a manual trigger. It puts a button in the UI so you can run the workflow on demand without opening a pull request.
jobs, runs-on and env
A workflow can have as many jobs as you want. This one has a single job called build-and-deploy, because there's no reason to overcomplicate it.
runs-on picks the virtual machine. ubuntu-latest is a Linux runner, the standard choice, and GitHub provides it free up to a usage limit. Windows and macOS runners exist too.
env sets environment variables for the job. The values come from GitHub secrets, encrypted values you store in the repository settings. The runner reads them at run time, so credentials never appear in the file.
steps
Steps are where it got confusing for me when I started. They run in order, top to bottom, and each one either uses a prebuilt action or runs a command.
- Check out branch pulls the branch under review onto the runner so the later steps have the code.
- Set up Python installs Python, which dbt needs to run.
- Install dbt Snowflake runs
pip install dbt-snowflake, exactly as you would in a terminal. - Deploy and test models runs
dbt build, which creates every model and runs every test in the project.
Use prebuilt actions
The first two steps have a uses line instead of a run line. That's how you pull in a prebuilt action from the GitHub Marketplace.
The Marketplace is a library of community and official steps that just work. Search Actions for "checkout" and you'll find actions/checkout, the one in this workflow. Its page explains what it does and lists every option.
actions/setup-python works the same way. You can pin a Python version and even run a script immediately after. Each action supports far more than the minimal version used here.
The point is that you don't have to write any of this yourself. Import the action, pass the inputs you need, and spend your time on the parts specific to your project.
Trigger the workflow
Push the branch, open a pull request against main, and the workflow starts on its own. Click Details on the check to watch it.
What the logs show
The run reads like a terminal session:
- A default setup step, then the checkout of your branch
- Python installed by the setup action
- dbt installed with pip
dbt buildrunning every model and test against the CI target
It behaves exactly like running the project locally, just on a hosted machine in real time. You'd think it would take a while. It's fairly quick.
The CI target
The build runs with --target ci, a dedicated target in profiles.yml that writes to its own schema, dbt_ci. Everything gets created and tested there, separate from each developer's schema and from production.
That separation matters. A failing test in dbt_ci costs nothing, while the same failure in production costs trust. If targets and environments are new to you, the post on dbt environments versus targets covers the setup.
The green check
When the run passes, the pull request shows a check mark and you're clear to merge. When it fails, the check goes red and the logs tell you which model or test broke.
From then on, every pull request gets the same check. Imagine changing a model or adding a new one: it gets tested before it ever reaches production, with nobody remembering to run anything.
That's CI/CD in practice. One file, a couple of secrets, and a catch-all check on every change your team proposes.
Key terms
CI/CD
Continuous integration and continuous deployment: automatically testing proposed changes and automatically releasing approved ones, driven by your version control platform.
GitHub Actions
GitHub's built in automation service. It runs workflow files from your repository on a hosted virtual machine when a trigger fires.
Workflow file
A YAML file in .github/workflows that defines a trigger, one or more jobs, the runner, environment variables and the ordered steps to execute.
Marketplace action
A prebuilt step published to the GitHub Marketplace, like actions/checkout, that you import with a uses line instead of writing the logic yourself.
CI target
A dbt target reserved for automated builds. It points at its own schema, such as dbt_ci, so test runs never touch development or production tables.
Common questions
What does CI/CD mean for a data team?
It means your transformation code is tested and deployed by an automated workflow rather than by someone running commands by hand. In practice: a pull request triggers a build of your dbt project, every test runs, and the result gates the merge. Deployment to production can run the same way once the merge happens.
Where does the GitHub Actions workflow file go?
In a folder called .github/workflows at the root of your repository. Each YAML file in that folder is a separate workflow. GitHub detects them automatically, so there's nothing to register or configure elsewhere.
How do I keep database credentials out of the workflow file?
Store them as repository secrets in GitHub settings and reference them in the env section with the secrets context. The runner injects them as environment variables at run time. Your profiles.yml then reads them with dbt's env_var function, so no password is ever committed.
Does the CI build run against production?
No. It should run against a dedicated CI target that writes to its own schema. That way a broken model or failing test shows up in an isolated place, and production stays untouched until the change is merged and deployed separately.
Does GitHub Actions cost money?
Public repositories get it free. Private repositories get a monthly allowance of free minutes on the standard Linux runner, which covers a small data team's pull request checks comfortably. Heavier use is billed per minute, and the warehouse compute for the build is billed by your warehouse as usual.
Should CI use dbt run or dbt build?
Use dbt build. It runs models, tests, seeds and snapshots in dependency order and stops downstream work when a test fails. dbt run only materializes models, so a CI job built on it would create tables without ever checking them.
Related reading
- Why Data Teams Need Version Control
- 3 Ways to Deploy Data Projects
- The 3-Environment Design for Your Database (DEV vs CI vs PROD)
- dbt Environments vs Targets: What's the Difference?
Final takeaway
Every team I've worked with that moved to a code based workflow eventually asked how to stop breaking production. This one file is the answer I give them: a pull request check that builds and tests every change before it merges. Set it up once and it covers every change after.
Additional Free Resources
Starter Guides & Checklists
Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.