#034: Getting Started w/ Prefect (Task Orchestration)
Mar 15, 2023Prefect is an open source task orchestrator: a central place to coordinate, schedule and monitor the tools in your data stack. You write workflows in plain Python and add decorators, so a script you already have becomes something Prefect can track. The core building blocks are flows, which are the workflows themselves, tasks, which are units of work inside a flow, and deployments, which package a flow so it can be scheduled and triggered remotely. You can run the whole thing on your own server, or point the backend at Prefect Cloud, which has a free tier. Either way your code and data stay on your infrastructure, because Prefect only stores the metadata about where and how to run things.
Key takeaways
- Prefect coordinates the pieces of a data stack from one place: pipelines, custom extract and load scripts, schedules and alerting.
- It is built on Python, so there is no new DSL to learn. You decorate functions you already write.
- A flow is the workflow. A task is a unit of work inside it. A subflow is a flow called by another flow, and it gets its own entry in the UI.
- A deployment is a configuration file that tells Prefect where the code lives and how to run it. Build it, then apply it to push it to the server.
- Nothing runs until an agent is polling the work queue. The deployment puts a run in the queue, the agent picks it up and executes it on your infrastructure.
- Prefect's hybrid model keeps your code and data on your side. Only deployment metadata is registered with the backend.
- Blocks store credentials and configuration centrally, including GitHub storage so flows run from a repository instead of one person's laptop.
What Prefect is, and why teams use it
Modern data platforms are built from a lot of separate tools. That's fine until you need them to run in a particular order, with no central place to coordinate or monitor them.
A task orchestrator is that central place. Prefect is one, and these are the things that make it easy to start with:
- Open source and free to run yourself, with a cloud option that also has a free tier
- Built on Python, so basic Python is the only prerequisite
- A hybrid model, which separates where your code and data live from the Prefect backend
- A clean UI, which matters more than it sounds for getting a team to adopt it
How teams actually use it
- Coordination. Connecting the tools in a pipeline and monitoring the whole run in one view.
- Custom extract and load. When no prebuilt connector fits, you write the Python yourself and orchestrate it alongside everything else.
- Scheduling and alerting. Once runs are coordinated centrally, schedules and failure notifications come almost for free.
Install Prefect
Prefect installs with pip. As with any Python project, install it into a virtual environment, not globally.
From your project directory:
python3 -m venv prefect_env
source prefect_env/bin/activate
pip install prefect
prefect version
The activate line is the Mac and Linux form. If prefect version prints a version, you're installed.
The local server
Prefect ships with a UI you can run yourself:
prefect server start
That spins up the backend and the web interface on your machine. Every run you trigger from here on shows up there.
Turn a script into a flow
Start with an ordinary Python script. Mine was a file called hello.py that printed a message, nothing more.
Add the flow decorator
A flow is the most basic Prefect object. You import it, add the @flow decorator above your function, and that's it:
from prefect import flow
@flow
def hello_world():
print("Hello world")
hello_world()
Run it with python3 hello.py as usual. The Prefect engine kicks in and the run appears in the UI under a generated name.
Add a task
A task is a discrete unit of work inside a flow. Tasks aren't required, but they break a flow into separately monitored pieces:
from prefect import flow, task
@task
def create_message():
return "Hello world"
@flow
def hello_world():
message = create_message()
print(message)
hello_world()
Run it again and the UI shows the same flow run with a task run nested underneath, named after the function.
Add a subflow
A subflow is a flow called by another flow. The code looks almost identical to a task, just with a different decorator:
@flow
def print_result(message):
print(message)
The difference shows up in the UI and in what you can do with it. A task appears as a task run inside its parent. A subflow appears as a flow run of its own.
You can also create a deployment for a flow, which gives you control over infrastructure and scheduling. Think of flows as the bigger pieces you string together, tasks as units of work inside one.
Deployments
So far everything has been triggered by hand from one machine. Deployments let you schedule a flow, or let someone else run it without your laptop involved.
A deployment is an API representation of a flow. In practice it is a YAML file that packages your requirements, and the two settings that matter most are:
- Storage, meaning where the code lives: this machine, an S3 bucket, a GitHub repository
- Infrastructure, meaning how it runs: a local process, a Docker container, a VM
Build and apply
Creating a deployment is two steps. First build it, naming the entry point as the file, a colon, and the flow function:
prefect deployment build hello.py:hello_world -n demo-deployment
One adjustment is needed before this works properly. Wrap the call at the bottom of your script in the standard Python guard:
if __name__ == "__main__":
hello_world()
The build writes a YAML file holding the work queue, which defaults to default, the infrastructure, a local process, storage, null meaning this machine, and the entry point.
The file alone does nothing. Apply it to register it with the server:
prefect deployment apply hello_world-deployment.yaml
Refresh the UI and the deployment is there. Anyone with access can trigger it, you can attach a schedule, and the YAML goes into version control.
Agents and work queues
Trigger a run from the deployment and it sits there as late, doing nothing. The run went into a work queue, and nothing on your infrastructure knows it's waiting.
Start an agent
An agent is a lightweight polling service that checks the queue for work and runs it:
prefect agent start -q default
Start it and it picks up the queued run immediately, moving to pending and then running locally.
This is the hybrid model in practice. The backend holds metadata and passes along the fact that a run is due. Execution happens on your infrastructure, against your storage, with nothing else going back to Prefect.
Wherever you set up infrastructure, there needs to be an agent there too. Work pools group multiple agents so you can manage them together and set limits like concurrency.
Switch the backend to Prefect Cloud
Everything above ran against a server on my laptop. Moving the backend to Prefect Cloud takes one command, for three reasons:
- No server to host, and no overhead from running one
- Built in authentication, with permissions, service accounts and user management
- Team workflow, meaning workspaces, sharing, onboarding and automations
Connect your terminal to it
Sign up, create a workspace, then point your local CLI at it:
prefect cloud login
It offers a browser login or an API key. Choose the browser the first time and Prefect writes the key for you, visible afterwards under API keys in settings.
Now run python3 hello.py again and the run shows up in Prefect Cloud instead of your local server.
Re-apply your deployment
Your deployment definition was registered with the local backend, not the cloud. Run the apply command again and it lands in your cloud workspace.
Everything else behaves the same, agent included. Start one locally and the cloud triggered run executes on your machine.
Blocks
Blocks are how Prefect stores credentials and configuration, managed centrally through the UI. There are dozens of built in types: secrets, Slack credentials, GitHub, GitLab and more.
The integrations catalog adds more, including dbt, Fivetran, Snowflake, BigQuery and Redshift.
Using one in a flow
I created a simple string block to show the mechanics. Prefect gives you the snippet to paste into your script:
from prefect.blocks.system import String
string_block = String.load("future-nba-champs")
message = string_block.value
The .value is the part people miss: the load call returns the block, not what's inside it.
Edit the value in the UI, rerun the script unchanged, and the new value comes through. Credentials and config live in one place instead of scattered across scripts.
Store your code in GitHub
Local storage means a flow only runs on the machine holding the file. Pointing storage at a repository fixes that.
Create the GitHub block
Push your project to a repository first, with a .gitignore so the virtual environment and generated files stay out.
Then add a GitHub block: a name, the repository URL in HTTPS or SSH form, and a personal access token if the repo needs one. Tokens come from developer settings in your GitHub account.
Attach it to the deployment
Rebuild the deployment with the storage block flag, using the block type slug, a slash, and your block name:
prefect deployment build hello.py:hello_world -n test-deployment -sb github/prefect-repo
The YAML now points at the repository instead of a local path. Apply it and anything running that deployment pulls code from GitHub, not from my machine.
The same idea applies to infrastructure: swap the local process for a Docker or cloud run block and attach it the same way.
Automations and concurrency
Automations fire on state changes. When a flow enters a running or failed state, trigger a notification through a Slack or Teams block.
Concurrency limits cap how many tasks with a given tag run at once. That's how you stop an orchestrator from overwhelming a warehouse or an API.
One note on versions: the video predates several Prefect releases, and the agent and work queue commands in particular have evolved. The concepts hold, but check the current docs for exact command syntax.
Key terms
Flow
The basic Prefect object: a Python function wrapped in the @flow decorator that represents a workflow you can run, schedule and monitor.
Task
A discrete unit of work inside a flow, created with the @task decorator, used to break a workflow into pieces you can monitor separately.
Deployment
A YAML configuration that registers a flow with the Prefect backend, defining its entry point, work queue, storage and infrastructure.
Agent
A lightweight polling service running on your infrastructure that watches a work queue and executes the runs it finds there.
Block
A stored credential or configuration object, managed in the Prefect UI or in code, that flows and deployments load by name.
Common questions
What is Prefect used for?
Prefect is a task orchestrator for coordinating the tools in a data stack from one place. Teams use it to run pipeline steps in order, run custom Python extract and load code, schedule workflows and get alerted on failures.
Do I need to know Python to use Prefect?
You need the basics. Prefect workflows are Python functions with decorators added, so if you can write and run a script you can write a flow. There is no separate configuration language for the workflow logic itself.
What is the difference between a task and a subflow in Prefect?
A task is a small unit of work that appears nested inside its parent flow run. A subflow is a full flow called from another flow, so it gets its own flow run and can have its own deployment. Choose a subflow when the piece is big enough to deserve its own infrastructure or schedule.
Why is my Prefect flow run stuck as late?
Almost always because no agent is polling the work queue the deployment targets. The deployment puts a run in the queue, but something on your infrastructure has to pick it up. Start an agent against that queue name and the run moves to pending.
Is Prefect Cloud free?
There is a free tier that is enough to get started, with paid plans above it. The alternative is hosting the backend yourself, which is fully open source. The deciding factor is usually whether you want to maintain a server and build your own authentication.
Does Prefect see my data?
No. Under the hybrid model only deployment metadata is registered with the backend. Your code lives in your storage and your flows execute on your infrastructure, which is the main reason this design exists.
Related reading
- The Importance of Virtual Environments
- How to Create a Virtual Machine on GCP
- Data Automation (CI/CD) with a Real Life Example
- The 10 Key MDS Components: Part 1 (Essentials)
Final takeaway
On the consulting projects where I've used Prefect, the thing that made it stick wasn't the scheduler, it was that a flow is just the Python the team was already writing. Get one real script running as a flow with a deployment behind it, and the rest of the orchestration story becomes an incremental decision rather than a migration.
Additional Free Resources
Starter Guides & Checklists
Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.