Create a dbt Project from Scratch w/ Claude Code

ai Apr 08, 2026

You can go from an empty directory to a working dbt project in a few minutes with Claude Code, as long as you give it something to copy first. The two inputs that matter are a CLAUDE.md file describing what the project is and how it should be built, and a sample dbt project that already follows the conventions you want. From there a single instruction covers the setup: create a virtual environment, install dbt for your warehouse, and lay the project out to match the sample. Claude Code writes a plan, you approve it, and it builds the directories, the profiles, the macros and the supporting files. You then add your own credentials in a .env file, confirm the connection, build one placeholder model to prove the loop works, and ask it to write rules files so the conventions carry into every future session. None of that removes the need to review the output, but it does take the grunt work out of getting from zero to one.

Key takeaways

  • Give Claude Code a sample dbt project before you prompt it. Describing a layout in prose is slower and less precise than handing it a project that already has the layout.
  • The CLAUDE.md file is the standing brief for the project. It gets read at the start of every session, so the setup instructions do not have to be retyped.
  • Claude Code writes a plan before it changes anything, and asks questions it cannot answer on its own, such as what to name the dbt project.
  • Keep credentials in a .env file that the gitignore excludes, and have the profiles read them through env_var. Nothing sensitive ends up in version control.
  • Run dbt debug before building anything. A passing connection is the line between a scaffold and a project.
  • Build one throwaway model with tests. It proves structure, credentials, run and test all work together before you commit to anything real.
  • Rules files per layer are what make the conventions stick. A lessons learned section in each one turns a single correction into a standing instruction.

Start with two files, not a prompt

The project I started from was a completely blank directory. Before typing anything into Claude Code, I dropped two files into it.

  • A CLAUDE.md file. A markdown file that says what this project is and how it should be built. Claude Code reads it at the start of every session.
  • A sample dbt project. A zipped copy of a project that already uses the conventions I want, so there is something concrete to point at.

The CLAUDE.md I used here is slimmed down compared to earlier versions of mine. It leans on the built-in conventions, rules and skills, and mostly says: this is a dbt project, here is what we install, here is how we work.

The sample project is the part people skip. It is the fastest way to nudge the output toward what you already know you want, instead of nitpicking your way there afterward.

The first prompt

One instruction covers the whole setup. I dictated this one with voice to text, which is worth trying, because the phrasing does not need to be tidy.

I want to start a new dbt project here and I want your help.

First, install Python and a virtual environment and make sure
those are all set up.

This is going to be a dbt project for Snowflake.

I've also uploaded a sample dbt project zip file. Look at that
and take inspiration for the structure, so the layout of our
dbt project aligns with it.

It plans before it builds

Claude Code read through the sample project, worked out the patterns in it, and then wrote a plan rather than starting to build. That plan is yours to approve, reject or adjust.

It also stopped to ask one question it could not answer for me: what should the dbt project be called. Every dbt project needs a name, so it asked instead of guessing. I called mine analytics.

What the plan covered

  • Create a virtual environment and install dbt-snowflake, which brings dbt Core with it, then freeze the dependencies.
  • Build the model layout from the sample: staging, warehouse broken into facts and dimensions, and marts.
  • Carry over the custom schema macro that overrides the schema name depending on environment.
  • Write a profiles file that reads credentials from environment variables.
  • Add the supporting files: a gitignore, a pull request template and an example env file.

Because I said Snowflake, it chose the Snowflake adapter. Name a different database and it infers that one instead.

I approved the plan with auto accept on and let it work. It created the virtual environment, installed dbt, wrote a requirements.txt listing every dependency, then made the directories and installed the dbt packages.

Check the choices it made

When it finished, two directories were not where the sample had them. Both .github and the project docs folder had landed outside the dbt project.

Rather than move them myself, I asked why.

I noticed that the project docs and the .github directories
are outside of the dbt project. Why was this done?

The answer was reasonable. GitHub only picks up pull request templates from a .github directory at the repository root, so that one has to sit there.

Project docs at the root was a judgment call. Its reasoning was that repo level documentation should stay generic, because ingestion, visualization or database setup code could live alongside dbt later.

That makes the repo a monorepo with dbt as one subdirectory of it. I had no problem with that, so I left both where they were.

analytics-repo/
├── .github/
│   └── pull_request_template.md
├── project-docs/
│   └── style_guide.md
└── dbt/
    ├── dbt_project.yml
    ├── profiles.yml
    ├── requirements.txt
    ├── .env.example
    ├── macros/
    └── models/
        ├── staging/
        ├── warehouse/
        │   ├── facts/
        │   └── dimensions/
        └── marts/

The point is not that every placement was right. It is that a one line question is cheaper than reorganizing the files yourself, and the reasoning is usually worth hearing.

Review it before you connect

Next I read through what it built. The dbt_project.yml already matched the sample: staging, warehouse and marts, with the configurations I would have set by hand.

The model directories were there too, with placeholder docs files showing how each one is meant to be used. Facts and dimensions had baseline files so nothing gets forgotten later.

The one gap was staging. A real project nests a directory per source under it, and there are no sources yet, so it left that empty and waiting.

Add your credentials in a .env file

The profiles file it generated reads everything from environment variables, which is exactly what you want. Nothing sits in plain text, nothing gets committed, and each person sets their own values locally.

analytics:
  target: dev
  outputs:
    dev:
      type: snowflake
      account: "{{ env_var('DBT_SNOWFLAKE_ACCOUNT') }}"
      user: "{{ env_var('DBT_SNOWFLAKE_USER') }}"
      role: "{{ env_var('DBT_SNOWFLAKE_ROLE') }}"
      warehouse: "{{ env_var('DBT_SNOWFLAKE_WAREHOUSE') }}"
      database: "{{ env_var('DBT_SNOWFLAKE_DATABASE') }}"
      schema: "{{ env_var('DBT_SNOWFLAKE_SCHEMA') }}"
      private_key_path: "{{ env_var('DBT_SNOWFLAKE_KEY_PATH') }}"

So I created my own .env file. It showed up greyed out straight away, because the gitignore it had written already excluded it.

What goes in it

Development credentials only: my user, my role, the warehouse to build with, the database to build in, and my own dev schema. I authenticate with a key pair, so I added the path to the key file as well.

The dev schema line matters more than it looks. Everyone building into their own isolated schema is one of those practices that costs nothing to set up and saves a lot of cleanup later.

Confirm the connection

Then a short prompt: I have updated this with my own .env file, confirm that you can connect to my database.

It proposed dbt debug, which is the right check. The connection passed, it reported how it had authenticated, and the setup was done in a handful of minutes.

Prove it with one model

Before trusting any of it, build something. I asked for a throwaway.

This looks good. Now create a sample placeholder model under
the staging directory, just to prove that this can work.

It created a directory under staging, as if that directory were a new source, then a model inside it following the naming convention from the sample: the stg prefix, the source name, then the model name.

It also added placeholder tests on the id column without being asked, then ran and tested the model.

dbt run --select stg_sample__placeholder
dbt test --select stg_sample__placeholder

All three tests passed, and the view appeared in my dev schema, created by the correct role. That is the whole loop working: structure, credentials, build, test.

Add rules so the conventions stick

One thing it had not done was set up the rules area, even though the CLAUDE.md calls for it. So I asked for it directly.

Everything looks good, but I noticed we didn't create anything
for the rules component based on the CLAUDE.md instructions.

Go ahead and create that, plus anything you think is important
to add to get us started.

It checked for a .claude directory, found none, and created one at the repository root alongside .github. Inside it wrote a markdown file of rules per layer.

  • Staging. One to one with the source tables, no joins.
  • Warehouse. Naming conventions for facts versus dimensions, and what belongs in each.
  • Marts. Wide denormalized tables, mart naming, joining dimensions to facts, unique and not null on the primary key.

Those came out of the CLAUDE.md and the style guide in project docs. That is why they read like my conventions rather than generic advice.

The lessons learned section

Each rules file also had a lessons learned section. When a staging model comes out wrong, you say so, and the correction gets logged there instead of being lost at the end of the session.

That is the part that compounds. A one off correction becomes a standing instruction, so the project gets more consistent the longer you work in it.

It is also where the data modeling concepts belong. What makes a fact, what makes a dimension, how wide a mart should be: write it down once and every model built after that inherits it.

Key terms

CLAUDE.md

A markdown file at the root of a project that describes what the project is and how it should be built, read by Claude Code at the start of every session.

Sample project

An existing project you hand to the AI purely as a structural reference, so the new project inherits conventions you already use instead of generic defaults.

Plan step

The written plan Claude Code produces before making changes, listing what it intends to create so you can approve or correct it first.

Environment variable credentials

Warehouse credentials stored in a local .env file and read into the profiles with env_var, so nothing sensitive is committed.

Rules files

Markdown files in the .claude directory that hold per layer conventions and a running log of corrections, so future work follows the same patterns.

Common questions

Do I need a sample dbt project to do this?

No, but the result is noticeably better with one. Without a sample you get a reasonable generic scaffold. With one you get your own layout, your own naming and your own macros, which is the difference between useful and something you then spend an hour rearranging.

How is this different from running dbt init?

The scaffold is only part of it. Here the virtual environment, the dependency install, the profiles wired to environment variables, the gitignore, the pull request template and the rules files all come together in one pass, shaped by conventions you supplied.

Where should the .github directory live?

At the root of the repository. GitHub only reads pull request templates and workflow files from that location, so it cannot sit inside the dbt subdirectory. If dbt is one of several projects in the repo, treat the repo root as the shared layer.

How do I keep warehouse credentials out of version control?

Have the profiles read every value through env_var, keep the real values in a local .env file, and make sure the gitignore excludes it. Commit a .env.example with the key names and no values so the next person knows what to fill in.

Do I still have to review the code?

Yes. Two directories landed somewhere I did not expect in this run, and the fix was asking about them rather than assuming. Treat the output the way you would treat a pull request from a new developer who is fast, consistent and occasionally wrong.

Related reading

Final takeaway

Every dbt project I set up with a client starts from the same conventions, and most of that first day used to be mechanical work. Handing over a sample project and a CLAUDE.md gets you to the same starting line in minutes, as long as you already know what good looks like well enough to check the result.

 

Additional Free Resources

Starter Guides & Checklists

Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.

Browse Resources