5 Ways To Improve Your Data Workflow
Feb 18, 2026
Workflow can be a vague idea, so here are five specific things a small data team can add to its version control project to improve it. The first three are automations: a CI workflow that deploys and tests every pull request in a pre-production environment, a post-merge deploy that pushes merged code to production, and a daily refresh that rebuilds and tests the data on a schedule. The last two are files you write once: a pull request template that nudges everyone to document their changes, and a README that acts as the front page of the repository. Together they take a project with no process to one with automation, documentation and consistency built in. None of them take more than an afternoon.
Key takeaways
- A CI workflow runs when a pull request opens. It deploys the proposed code to a pre-production environment and runs tests, so nothing reaches production untested.
- A post-merge deploy is nearly the same workflow file with a different trigger: it runs on a merge to the main branch and targets production instead of CI.
- A daily refresh is the baseline expectation for any data team. It keeps data current and lets you catch breakages before a stakeholder opens a report.
- With those three automations in place, the whole development loop is covered: test before production, deploy after merge, refresh every day.
- A pull request template is one markdown file. It pops up every time someone opens a request and makes skipping documentation a deliberate choice instead of a default.
- The README is the most underused part of a data repo. Use it for what the project is, quick links, setup steps for new developers and who to contact.
What these five have in common
I talk a lot about workflow and why it matters so much for smaller data teams. The trouble is that "workflow" on its own doesn't tell you what to do on Monday.
So each of these five is a concrete thing you add to your version control project. You can check your own repo against the list.
The first three are automations that run on their own. The last two are files you write once and tweak as you go.
I show them with GitHub Actions because that's what I use most, but GitLab, Bitbucket and Azure DevOps all work the same way. If you host dbt in dbt Cloud, the same automations can be set up directly in there.
1. A CI workflow for every pull request
CI stands for continuous integration. In practice, it's the environment and the process that sit between development and production.
You have a development environment where you build, and production where the business reads from. CI is the middle step, a pre-production environment where new code gets deployed and tested before anyone merges it.
What it looks like in practice
When you open a pull request on your version control platform, an automation triggers. It checks out the code, deploys it to the CI environment in the database, and runs your tests.
In a dbt project the workflow file is roughly this:
name: ci
on:
pull_request:
branches: [main]
jobs:
deploy-and-test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: pip install -r requirements.txt
- run: dbt deps
- run: dbt build --target ci
That whole sequence, from the request opening to the tests finishing, is the CI workflow. It's one file in the repo, and it's something you can add to almost any project fairly easily.
The point is that nothing reaches production untested. You find out a change breaks a downstream model while it's still a pull request, not after it's live.
2. A post-merge deploy to production
The second automation is almost the same file with a different trigger.
Merging a pull request means moving code from your development branch into the main version of the project. When that happens, another workflow runs: deploy and test again, but this time to production.
What changes in the file
- The trigger. Instead of running when a pull request opens, it runs on a push to
main, which is what a merge produces. - The target. With dbt you point the run at the production environment instead of CI, so
--target prodinstead of--target ci.
Everything else stays the same. The recommendation is simply to have something fire right after the merge so you never deploy to production by hand.
It removes a manual step, and it keeps you in check. Production only changes through the process.
3. A daily refresh
The third automation is the simplest, and the one I still see missing most often: a scheduled run that refreshes and tests the data every day.
For most small teams once a day is enough. Two or three times a day if the business needs it. The frequency matters less than the fact that it happens without anyone pushing a button.
Why it matters even if it sounds obvious
Plenty of teams refresh manually, when it makes sense, or right after they change something. That means the data is only as current as the last time someone remembered.
A scheduled refresh gives you three things:
- You catch errors first. If a source changes and a model breaks, the morning run fails and the team finds out before a stakeholder opens a report.
- Tests run on a schedule. The job should include your tests, not just the build, so data quality is checked daily on your behalf.
- Data stays current. New rows land overnight and show up downstream without waiting on you.
I consider this the baseline expectation for any data team.
Three automations, most of the process
With those three in place, the development loop is covered:
- A pull request tests the change in pre-production
- A merge deploys it to production
- A schedule refreshes and tests everything daily
The remaining two recommendations aren't workflow files. They're template files, and they cover the human side of the process.
4. A pull request template
Every version control platform I've run into lets you add a template that pre-fills the description when someone opens a pull or merge request. It's usually one markdown file in the repo.
In GitHub it lives at .github/pull_request_template.md. Put in whatever headings you want people to fill out: what changed, why, how it was tested, anything downstream to watch.
Why a template changes behaviour
When a team is moving fast, documentation is the first thing to go. Notes get skipped, or they say nothing useful.
A template pops up automatically. You have to at least acknowledge it's there, and if you delete it to write nothing, that's a choice you made on purpose.
Most of the time people fill in at least a little, and every request ends up looking more or less the same. Set it up once and forget it unless you want to tweak it.
5. A real README
The README is the front page of the repository, and for most data teams it's either blank or the default project description. That's wasted space.
Like the template, you build it once and adjust as the project changes. It's not a big lift.
What to put in it
- What the project is. A high level description, the data sources it covers, the layers of the database it builds.
- Quick links. Style guide, data documentation, dashboards, wherever else context lives.
- Setup steps for new developers. For a dbt project, how to install, configure a profile and run the project locally.
- People. Who's on the team, who to contact, and credit to the people who've contributed the most.
It's the first thing anyone sees when they open the repo, whether they're on the team or not. New hires get oriented faster, and you spend less time explaining the same things.
Putting the five together
None of these is hard to do. Each is easier not to do, which is why so many projects skip them.
Combined, they take a project from no structure and no process to one that runs itself: automated testing before production, automated deployment after merge, a daily refresh, consistent pull requests and a repo that explains itself.
These are the recommendations I give small data teams most often, and the ones I'd start with on any new project.
Key terms
CI workflow
An automation that runs when a pull request opens, deploying the proposed code to a pre-production environment and running tests before it can be merged.
Pre-production environment
A database environment between development and production where changes are tested against real structures without touching what the business reads.
Post-merge deploy
An automation triggered by a merge to the main branch that deploys and tests the code in production, so nobody deploys by hand.
Daily refresh
A scheduled job that rebuilds and tests the data models once a day so data stays current and the team catches errors before stakeholders do.
Pull request template
A markdown file in the repo that pre-fills the description every time someone opens a pull request, prompting consistent documentation of changes.
Common questions
What is a CI workflow in a data project?
It's an automation that runs when a pull request is opened. It deploys the proposed changes to a pre-production copy of the database and runs tests. If the tests fail, you know before the code is merged. In GitHub it's a YAML file under .github/workflows.
Do I need CI/CD for a dbt project?
Once anyone depends on the output, yes. A CI job on pull requests and a deploy job on merge are two small workflow files, and they remove the risk of untested logic reaching production. dbt Cloud can run the same jobs if you'd rather not manage the workflow files yourself.
How often should a data pipeline refresh?
For most small teams, once a day. Go to two or three times a day if the business genuinely needs fresher numbers. What matters more than frequency is that the refresh is scheduled, includes tests, and doesn't depend on someone remembering to run it.
What should a pull request template include for a data team?
Keep it short: what changed, why, how it was tested, and anything downstream to watch. The goal is a consistent description on every request, not a long form. One markdown file in the repo does it.
What should go in a data project README?
A high level description of the project and its data sources, links to the style guide and data documentation, setup steps for a new developer, and who to contact. It's the front page of the repo, so write it for someone seeing the project for the first time.
Related reading
- 3 Simple Ways to Improve Your Data Workflow
- Don't Overlook Workflow (as a Small Data Team)
- The 3-Environment Design for Your Database (DEV vs CI vs PROD)
- The Art of Data Team Workflows
Final takeaway
Most of the small teams I work with have the code figured out and the process missing. Three automations and two files is the shortest list I've found that fixes that, and it's where I start with nearly every client repo.
Additional Free Resources
Starter Guides & Checklists
Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.