The Biggest Cost of Owning Your Data (as a small team)

architecture leadership Jul 08, 2026

When data teams think about cost, they think about licenses and compute. For most small teams the biggest cost is neither. It is maintenance and the human time required to keep the system running, and the place it hides best is data ingestion, at the very start of the pipeline. Run the numbers on five sources with one mid level engineer and you land somewhere around $54,000 a year before any software bill, mostly in connector building, schema fixes and downtime. The real decision is not build versus buy, it is whether owning that work gives your business an advantage worth the hours.

Key takeaways

  • The largest line item on a small data team is usually engineering time, not the tools budget.
  • Ingestion sets the tone for everything downstream. Messy and hard to access at the start means a harder job at every later layer.
  • Reliable ingestion is how you build trust with stakeholders, and once that trust is gone it is very hard to earn back.
  • There are three common approaches: fully custom, an open source framework you host, or a managed cloud service.
  • A conservative total cost of ownership example with five sources came to roughly $54,000 a year, mostly in human time.
  • Opportunity cost is the other half of the number. Those hours could have gone to analytics, AI work or efficiency elsewhere.
  • Own the process yourself when it is a competitive advantage, when your company prefers open source, or when you genuinely have the capacity.

Why ingestion matters more than it looks

Picture the architecture left to right. Sources on the left, consumers on the right. Those consumers are no longer only internal stakeholders, they increasingly include AI agents and the tools built on them.

Whether or not you are doing that today, how you place and structure data at the start decides a lot. Three reasons stand out.

1. It sets the tone for the rest of the architecture

If data arrives disorganized and hard to access, every later layer is harder. You spend your time fighting APIs and fixing formats instead of modeling.

That hits small teams hardest, because the ambitions are big and the hours are not.

2. It builds trust

Stakeholders need to believe the data is accurate and current. If something breaks, you want to know before they do and have a way to act on it.

I have watched teams lose that trust and then spend all of their time fixing things instead of moving forward. Getting it back is extremely difficult.

3. It lets you scale

You move only as fast as your slowest component. If acquiring data is the bottleneck, nothing downstream gets better.

This is not about collecting more data for its own sake. A tuned ingestion process makes the next source cheap to add, which is what opens up new use cases.

The three common approaches

If you have been doing this a while you have probably used all three.

  • Do it yourself. Custom scripts in Python or whatever you know, connecting directly to databases and APIs, running on your own infrastructure. Sometimes written by you, sometimes inherited from someone who left ten years ago.
  • Open source. An established framework you run without a license fee. Better structure than homemade, plus a community, but you still own updates, networking and security.
  • Managed cloud service. The infrastructure is abstracted away. You configure a connector through an interface, set a schedule, and skip reinventing the wheel for every API. You pay for licenses and compute, and you give up some flexibility.

What total cost of ownership actually includes

To make this concrete I walked through a total cost of ownership calculator built for the ingestion component specifically. Here are the inputs I used.

  • Company size. 1 to 200 people, which covers a lot of small data teams.
  • Fully loaded salary. $120,000, including benefits, bonus and time off. A reasonable US benchmark for a mid level data engineer, which works out to about $62.50 per hour.
  • Number of sources. Five, which is typical for the teams I see.
  • Time to build each connector. Forty hours, halving the eighty hour default to stay conservative.
  • Schema change maintenance. Two hours a week fixing pipelines after schema or API changes. That is under an hour a day across the year.
  • Data downtime. Dropped to $2,500, for the weeks when pipelines stall and the business cannot decide anything.
  • Additional costs. Another $2,500 for self hosted infrastructure and manual handling.

The number

That comes out to about $54,000 a year. On conservative inputs, with the defaults it sits closer to six figures, and nudging the salary up pushes it past that.

Your numbers will be different. The point is that there is a very real price attached to your time and to these decisions.

The part the calculator misses

Beyond the direct cost there is opportunity cost. A few hundred hours a year is a meaningful amount of better analytics, new AI use cases, or efficiency work elsewhere in the business.

Everything here is a trade-off. There are no free lunches, and the mistake is letting this become a hidden expense you never priced.

What a managed platform looks like in practice

One of the platforms I have used many times and recommended to clients is Fivetran, which also sponsored the original video. Here is what the setup actually involves, loading a Postgres database into Snowflake.

  • Choose the connector. In this case Postgres, running on Google Cloud SQL.
  • Set a destination schema prefix, which groups each source schema in the target so nothing collides.
  • Enter connection details: host, user, password, authentication method. With no logical replication configured, it falls back to queries.
  • Whitelist the platform IP addresses on the Postgres server, then run the connection test.

The platform then reads the source schema. In this case it found one schema with ten tables, with options to filter rows and to hash sensitive columns so the column lands populated but without real values.

Schema changes and sync

Next it asks how to handle future schema changes. That is a one click decision here, which understates how tedious schema drift is to handle yourself.

The default sync frequency is every six hours and is adjustable. On the destination side you configure a dedicated user, a target database, a key pair, a role and a warehouse.

How the billing metric works

Monthly active rows is the billing unit, and it confused people for a while. A row counts as active only when it is new, modified or deleted in that month.

Initial historical syncs are not counted. Load three records and you have zero active rows. If the next day one value changes and one row is added, that is two active rows, not four.

Beyond loading

The same platform now covers more of the stack: transformations wired to a Git repository to run dbt models, and an activations section, which you may know as reverse ETL.

Reverse ETL means pushing modeled data back into the business applications people work in every day, rather than leaving it in a report. Fivetran and dbt Labs also merged into one company around the time of this recording.

How to decide

There is no universally correct answer. Three situations point toward owning the process yourself.

  • It is a competitive advantage. You have proprietary or very specific connector logic that no framework or cloud provider offers. For most teams I work with this is not the case, but it can be.
  • Your company prefers open source. Some teams value the control and flexibility strongly enough that hosting it themselves is the right call.
  • You genuinely have the capacity. A simple architecture where these responsibilities are not that time consuming.

If none of those describe you, a hosted platform gets you out of API changes, connector logic and codebase upkeep, with support and features you would otherwise build.

Key terms

Total cost of ownership

The full price of running a component, including engineering hours for building and maintaining it, downtime, and infrastructure, not just the software invoice.

Data ingestion

The process of acquiring data from sources and landing it in your analytics database, the first component of the architecture and the one most likely to hide cost.

Monthly active rows

A usage based billing metric that counts only the rows that were new, modified or deleted in a given month, rather than every row synced.

Reverse ETL

Taking modeled data out of the warehouse and pushing it back into the business applications stakeholders use daily, sometimes branded as activations.

Opportunity cost

The work your team did not get to because the hours went somewhere else, which on a small data team is usually the largest unmeasured cost.

Common questions

What is the real cost of building your own data connectors?

Budget the build and then the upkeep separately. At forty hours per connector and two hours a week of schema maintenance, five sources runs into tens of thousands of dollars a year in engineering time alone. The build is the part people estimate, and the maintenance is the part that actually accumulates.

Is open source ingestion cheaper than a managed platform?

There is no license fee, which is real. You still own updates, networking, security and the infrastructure, so the saving is a transfer from the software budget to your own hours. It works well when self hosting matches your team's skills.

When does building your own ingestion make sense?

When the connector logic is proprietary enough to be a competitive advantage, when your company has a strong open source preference, or when your architecture is simple enough that the upkeep is genuinely small. Those are the three cases worth it.

How do I estimate data downtime cost?

Estimate the hours per year your pipelines are stalled or delayed, then ask what decisions stop while the data is unavailable. The number varies enormously with how much the business leans on data, which is why it belongs in the model even as a rough figure.

Does an AI strategy change how I should handle ingestion?

It raises the bar. Agents and AI tools consume data on a schedule and without a human sanity checking it, so freshness and reliability matter more, not less. That generally argues for less bespoke plumbing.

Related reading

Final takeaway

On the small teams I work with, the hours spent keeping ingestion alive are almost never on anyone's budget, and they are almost always the biggest number. Price that time honestly, then decide whether owning the work is the best use of it.

 

Additional Free Resources

Starter Guides & Checklists

Explore additional free resources built on the same patterns I use with real clients so you can build your own with structure and confidence. Topics include data architecture, modeling and more specifically for small data teams.

Browse Resources