3 Simple Ways to Improve Your Data Workflow
Jun 26, 2025
The three simplest ways to improve a data team's workflow are version control, scheduled refreshes, and automated data quality checks with notifications. Version control means every change to your code goes through a platform like GitHub, GitLab or Bitbucket instead of being saved straight into the database or a reporting tool. Scheduled refreshes mean nobody runs the pipeline by hand each morning. Data quality checks mean tests run automatically in a pre-production environment, and a failure lands in Slack or email before a stakeholder sees it. None of the three require new tools. Together they matter more than the code itself once the business starts depending on your data.
Key takeaways
- As business impact grows, the workflow around how you develop becomes more important than the code. How changes reach production, how data stays updated, and how you hear about errors.
- Saving queries directly into the database or a reporting tool is a direct save to production with no backup. It's common, it feels fast, and it's where a lot of trouble starts.
- Version control is worth it even on a one-person team. You will get busy, forget why something changed, and eventually hand the work to someone else.
- Manual daily refreshes are a solved problem. Schedule ingestion and transformation in the tools you already have, with a gap between them, before reaching for an orchestrator.
- Automated checks in a pre-production environment plus an immediate notification are usually the first thing executives ask about, because they've been burned before.
- Errors happen. Trust erodes when they happen too often or when a stakeholder finds them first, and it's very hard to rebuild.
Why workflow matters more as you grow
For most data engineers, the first instinct when something comes up is to dive straight into the code. Early on that works fine.
As the business starts relying on what you build, the workflow around how you develop matters more than the code itself. Three questions tell you where you stand:
- How do changes actually get to production?
- How does the data stay updated across the board?
- How do you find out when something goes wrong?
Each of the three sections below answers one of them. Use them to look at what your team does today and see where there's a gap.
1. Version control
Most of the popular data tools today are code based, and even the ones that aren't usually ship with some form of version control. Despite that, plenty of teams still don't use it.
What no version control looks like
- Queries saved directly into the database as views or procedures, with no file anywhere else
- All the logic built straight into the reporting tool
- Every save is a direct save to production, with nothing backed up
It's how a lot of people start, and it feels fast because there's no process in the way. It's also a really common place for issues to begin, and at some point you're asking for trouble.
The fix
Move the code to a version control platform. GitHub, GitLab and Bitbucket are the common ones. Each gives you branches, commits, and a pull or merge request process for moving changes through.
The real benefit isn't the structure, it's what it prevents. A change that breaks something leads to someone making the wrong decision on bad data, and those issues snowball down the line. Doing things responsibly is your job, and this is one of the ways you do it.
Even on a one-person team
If nobody else reviews your code, I still recommend it. You'll get busy. You'll forget what you changed and why, or the steps that got you there. The history covers you.
And whether you want to admit it or not, at some point you won't be at that company. Someone will inherit your work, and a tracked history is how you leave the place better than you found it. It's looking out for your future self.
Version control is fundamental. Once you've worked this way, it's hard to go back to a team that doesn't.
2. Scheduled refreshes
Outside of development, the other core piece of workflow is how the data gets refreshed. A surprising number of teams still do it manually, every day.
Move away from that as quickly as you can. It's a solved problem, and the only thing manual refreshes buy you is lost time.
Scheduling without an orchestrator
Modern stacks split the work across tools: one for ingestion, a separate one for transformation. Almost all of them have scheduling built in.
The key is timing, so the jobs don't overlap. For example:
- Ingestion runs every day at 5:00 a.m.
- Transformation runs at 6:00 or 7:00 a.m.
If you know ingestion takes 30 to 60 minutes, that gap is enough. Both tools run on their own schedule, and 99 percent of the time it works. When it doesn't, the failure notifies you and you react.
When to add an orchestrator
The natural next step is an orchestration tool: software that sits on top of every tool in the stack and acts as a control plane, scheduling steps with real dependencies between them.
That's helpful as you scale. It's not required to get started, and it adds setup and maintenance. Basic schedules in the tools you already have go a long way, and you can layer orchestration on later as you see fit.
3. Data quality checks and notifications
With developers using version control and data refreshing on a schedule, the last piece is a process to keep an eye on quality.
This is where the architecture stops being a set of individual parts and becomes a system. The version control platform does more than track changes.
Automated checks before production
Set up checks that run through the version control platform when a change is proposed. With a command line tool like dbt, that means deploying the models to a pre-production environment and running dbt test automatically.
The new logic is validated in isolation before it ever reaches production. The platform checks your work on your behalf.
Get notified, immediately
Checking for errors isn't enough on its own. You need to hear about them right away.
Most tools, and certainly the version control platforms, can send an email or post into a messaging tool. A check fails, a message lands in your Slack channel, and you go fix it instead of keeping an eye on everything yourself.
Why executives ask about this first
This is usually the first thing leadership wants to know about an architecture. Whatever is being built, they want something in place to catch issues.
Often it's because they've been burned. An important stakeholder or an external customer found an error and reported it before the team caught it internally. That's never a good look.
The error itself isn't the problem. They happen, and you'll never be at 100 percent. The problem is when they happen too often, and the trust and confidence in the data starts to go. That is very hard to build back.
When something does slip through
Overcommunicate. Be transparent and get out in front of it as early as you can.
That shows people you're still on top of things. It also saves someone else from spending their time building proof that something is off, which is a bad use of their time and something nobody wants to do.
Workflow is part of the job
Writing code is the fun part, and it's easy to get lost in it. How you write it, how you move it to production, and how you keep it updated matter just as much, especially as things grow.
Look at what your team does today against these three. Most teams have at least one gap, and each one is fixable with the tools already in front of you.
Key terms
Version control platform
A service such as GitHub, GitLab or Bitbucket that stores code, tracks every change, and provides branches and pull or merge requests for moving changes through a process.
Direct save to production
Saving queries or logic straight into the database or reporting tool with no file backed up anywhere else, so every edit is live immediately.
Scheduled refresh
Running ingestion and transformation on a timed cadence inside each tool, with a gap between them, instead of triggering them by hand.
Orchestration tool
Software that sits above the whole stack as a control plane and schedules each step with dependencies on the ones before it.
Automated data quality check
Tests that run automatically in a pre-production environment when a change is proposed, with a notification sent to email or Slack when one fails.
Common questions
Does a one-person data team need version control?
Yes. Even without anyone reviewing your code, you'll forget what you changed and why. The history protects you, and it makes handing the work to someone else possible when you move on. It's also the foundation for the automated checks described above.
How do I schedule ingestion and transformation without an orchestrator?
Use the schedulers built into each tool and leave a gap. Run ingestion at 5:00 in the morning and transformation at 6:00 or 7:00 if ingestion takes under an hour. It works the vast majority of the time, and a failure will notify you so you can rerun.
When should a data team add an orchestration tool?
When simple timed schedules stop being enough: many steps, real dependencies between them, or runtimes that vary too much to leave a safe gap. Until then an orchestrator adds setup and maintenance without much benefit for a small team.
How do I get notified when a data pipeline fails?
Most ingestion, transformation and version control tools can send an email or post to Slack or Teams on failure. Turn it on for the scheduled refresh and for the automated checks on pull requests. The goal is to hear about a problem before a stakeholder does.
What should I do when a stakeholder finds a data error before the team does?
Overcommunicate. Acknowledge it, explain what happened and what you're doing about it, and get in front of it quickly. That shows you're on top of things and stops others from spending their time gathering proof. Then add a check so that class of error is caught automatically next time.
Related reading
- 5 Ways To Improve Your Data Workflow
- Don't Overlook Workflow (as a Small Data Team)
- What does data "Workflow" even mean? (3 examples)
- Why Data Teams Need Version Control
Final takeaway
When I review a small team's setup, these three are the first things I look for, and the data quality piece is the one leadership asks about before anything else. Get the code into version control, put the refresh on a schedule, and make sure a failure reaches you before it reaches them.
Additional Free Resources
Starter Guides & Checklists
Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.