AI For Data Teams (How to Be Ready)
May 27, 2026
Getting a data team ready for AI is less about picking a tool and more about the fundamentals you already know. When I say AI here I mean coding agents like Claude Code, Codex and Copilot, the ones that do real development work for you. The single most important thing you can do is give them good instructions, and good instructions come from naming conventions, project structure, a style guide, documentation and a clear workflow. Teams that have those written down move fast and stay in control. Teams that don't hand the steering wheel to the agent and end up responsible for something they can't explain.
Key takeaways
- When I say AI for data teams, I mean coding agents:
Claude Code,Codex,Copilotand the tools that do the heavy lifting of development. - Good instructions are the whole game. If you outsource the thinking, the agent takes control, because that is what it is programmed to do.
- There is a difference between moving really fast and getting nowhere, and moving really fast and being productive.
- The readiness work is boring and familiar: naming conventions, project structure, a style guide, documentation, a clear workflow.
- Treat an agent like onboarding an intern or a junior engineer. The difference is it keeps the context and picks up where you left off.
- It will confidently hallucinate, so the old guardrails matter more, not less: CI/CD, testing, environments, pipeline layers, modeling and tests on relationships.
- Being more efficient usually means being handed more work, not less. People who do a lot of work get more work to do.
What I mean by AI here
This topic comes up constantly with my clients and with other engineers, and plenty of us are tired of hearing about it. It still matters, because there are a lot of ways this can go and some of them are already showing up.
When I say AI in this context, I'm not talking about chatbots. I mean the coding agents and coding tools: Claude Code, Codex, Copilot, and whatever else is in that category by the time you read this.
These are the tools that expedite your development and genuinely do the heavy lifting of building.
Good inputs, and who stays in control
The biggest thing I'm learning, and seeing from other people, is that the most important thing you can do is give it good instructions. We already knew that from prompting. It matters more once development is involved.
When you outsource your thinking and your decisions, the agent takes control. That is what it's programmed to do, so it just moves.
It feels fast. You feel productive. But you're flying all over the place, and it's taking control of something you are then responsible for.
If instead you're clear about what you want, the expectations and the rules, you actually move faster. The agent does the grunt work and you steer the ship.
That combination is what teams are really after when they talk about AI. There's a difference between moving fast and getting nowhere, and moving fast and being productive.
It comes down to the fundamentals
So how do you become a team that uses this productively? Boringly enough, it comes down to having good quality conventions in place.
- Naming conventions. So the agent names things the way your team already names things.
- Project structure. So new files land where they belong instead of wherever the model guesses.
- A style guide. So the SQL that comes back looks like SQL your team wrote.
- Documentation. So the context isn't trapped in somebody's head.
- A clear workflow. So branching, testing and deploying happen the same way every time.
The difference that matters here is not letting the AI come up with all of that. It can support you and help your thinking, but you should be the one making those decisions.
A lot of teams already have most of this in place informally. The work is formalizing it, getting clear on what it is, and articulating it to the coding tool.
If you don't know what you want, that's a problem, and it's one I think we're starting to see play out.
Onboarding an agent is like training an employee
I like to think of this like training another employee. It isn't something you finish in a day unless you already have most of it in place.
Most of the time it takes ongoing reinforcement. You explain things over and over until the output is what you want, the same way you'd onboard an intern, an assistant or a junior engineer.
The difference is that once you train the agent, it remembers the context. There's no lag. It picks back up where you left off, keeps improving, and doesn't get tired.
There are elements of this work that I think will always be human. But the framing is useful: explain your process, explain your strategy, then say this is what we expect you to do.
Why learning the fundamentals still matters
It's easy to jump to the conclusion that nobody needs to understand development anymore. No reason to learn SQL or Python because the tool does it.
I think that's a critical mistake for a team to make, and if you're leading one I'd encourage you not to overlook it.
If you know what good looks like, you get there faster and more efficiently than if you let the AI make those decisions for you. I'm not saying that as a traditionalist. Knowing where you're trying to go is what lets the tool take you there quickly.
Efficiency tends to create more work
Here's a related thought from watching how careers actually go: people who do a lot of work get more work to do.
I have a hard time believing companies will get more efficient with AI and then say we're good, let's stop. There's going to be an appetite for more and more output, and probably more of a mess along with it.
No way to know for sure how that plays out. The point is that work doesn't become irrelevant. You may end up doing more of it because you're more efficient.
It will confidently hallucinate
Anyone who has used these tools for a meaningful amount of time knows they make mistakes, and that they will confidently hallucinate.
That's another reason to understand what's going on. You need to be able to review what's happening and catch it mid-stream, or have guardrails that catch it for you.
An agent may confidently build a project and deploy things, and something that should have failed may not. The pace at which that can happen is only going to increase.
The guardrails are the ones you already know
Oddly enough, the best protection is the stuff we've been doing for years:
- CI/CD, so changes run through an automated check before they land
- Testing, so a broken assumption fails loudly
- Separate environments, so development isn't production
- Layers in your pipeline, so you can isolate where something went wrong
- Good modeling and tests on your relationships
All of it already existed. The change is that you need to be more deliberate about it rather than loose, because looseness now leads to more accidental errors.
People make mistakes too. They push incorrect changes to production because something was overlooked and untested. The new wrinkle is that you now have to explain, on behalf of an agent, what was built to a stakeholder or to the rest of the business.
Finding the middle ground
I don't think this is black or white. It isn't outsource everything to AI and never look at it again, and it isn't refuse to use it because you don't trust it.
There has to be a middle ground. This is a tool and a technology, and it isn't going away.
We as data people still need to be in charge of the strategy, the conventions and the architecture. The tool expedites the process. Guardrails keep the speed from hurting you.
If you understand your naming conventions, your structure and your workflow, and you have documentation and a style guide, you can feed that into an AI tool and work with it. That combination is good for productivity, for getting work done and for consistency in your project.
Key terms
Coding agent
A tool like Claude Code, Codex or Copilot that does real development work on your project rather than returning a snippet for you to paste.
Naming conventions
The agreed rules for what things are called in your project, which is the first thing an agent needs so its output looks like the rest of your work.
Style guide
A written description of how your team writes code, from formatting to patterns, that you can hand to an agent as an instruction instead of correcting it every time.
Guardrails
The automated checks that catch a bad change before it reaches production: CI/CD, tests, separate environments and pipeline layers.
Hallucination
When an AI tool confidently produces something that looks right and isn't, which is why you need to be able to review what it did.
Common questions
What does it mean for a data team to be AI-ready?
It means your conventions, structure, style guide, documentation and workflow are written down well enough to hand to someone new. A coding agent is that someone new. If a new engineer could not get oriented from what you have, an agent can't either.
Which AI tools are we talking about for data engineering?
Coding agents that work inside your project, such as Claude Code, Codex and Copilot. The category is moving quickly, so the specific names matter less than the shift from chat prompts to a tool that reads and writes your codebase.
Do engineers still need to learn SQL and Python?
Yes. The value is in knowing what good looks like so you can direct and correct the tool. Without that you can't tell a correct result from a confident one, and you're the one who answers for it.
How do I stop an AI agent from breaking production?
The same way you stop a person: CI/CD, tests, separate environments and a pipeline with layers you can inspect. None of that is new, it just has to be deliberate rather than loose.
Will AI reduce how much work a data team has?
In my experience, people who do a lot of work get handed more work. I'd expect appetite for output to grow rather than demand to be satisfied, so plan for more throughput rather than less effort.
Related reading
- The AI Workflow for Data Engineering
- Why a Good Data Model Makes Your Team More AI-Ready
- How to Structure AI Projects for Data Engineering
- The Art of Data Team Workflows
Final takeaway
Nearly every small team I work with already has conventions, they just live in people's heads instead of in a file. Write them down before you point an agent at the project, and the tool makes you faster instead of making a mess you have to explain later.
Additional Free Resources
Starter Guides & Checklists
Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.