The founder of a wellness startup that had not yet launched was about to add two subscriptions: a CRM (the tool that tracks customers and sales conversations) and a customer messaging platform.
You could not blame them. The company already tracked employees, vendors, partners and customers across Notion, Shopify, an email waitlist tool, WhatsApp and a stack of spreadsheets. Five places, five partial views of the same handful of people. A CRM felt like the grown-up move: one more tool, this time to hold the customers properly.
We told them not to buy it. Not because CRMs are bad, but because the problem was not a missing tool. It was that every fact about the company lived in something the company rented.
That is the thesis of this piece, stated plainly: the smaller you are, the more you should centralise and own your data. Not later, when you are big enough to afford it. Now, while you have three systems instead of thirty, and owning them is weeks of work instead of a year.
Every tool you rent is a fact about your company you do not hold
The information existed. It just lived in so many places that no one person, including me, could see the company we were running.
Look at what a five-person company actually knows about itself. Who its customers are. Which partner introduced which deal. What each vendor was paid. Who on the team is working on what. Which prospect asked for a callback.
Now look at where that knowledge lives. Customers are in Shopify. Prospects are in the waitlist tool. Conversations are in WhatsApp. The team and vendor lists are in a spreadsheet. Context is in Notion. Each tool holds one slice, in its own shape, behind its own login, on someone else's servers, on a subscription you can cancel or that can cancel you.
That is not a data problem in the way people usually mean it. The data is not dirty. It is not missing. It is scattered, and it is not yours.
When I wrote up how we automated Edge8, the line I kept coming back to was this: "The information existed. It just lived in so many places that no one person, including me, could see the company we were running." I have been building companies for 20 years, and I could not answer basic questions about my own firm without opening six tabs. If that is true at our size, it is true at yours.
The founders I talk to already feel this. What they get wrong is the fix. They reach for one of two things: a big cleanup (organise the spreadsheets, dedupe the contacts, fix the tags) or a big cancellation (rip out three tools and consolidate onto one platform). Both start at the wrong end.
The order of operations, stated as a rule
Store everything, connect it, and let the models do the cleaning.

Three steps. The sequence matters more than any one of them.
First, build one home you own. One database, on infrastructure you control, where a customer is one record, a person is one record and a deal is one record. Nothing goes into it yet. It just exists, with a shape.
Second, integrate every tool into it. Shopify writes orders into it. The waitlist tool writes signups into it. The accounting system syncs invoices into it. The spreadsheets get loaded in as tables. You do not turn anything off. You point everything at the home.
Third, remove the applications you no longer need. Once a tool's data is landing in your home and the workflow that used the tool can run from your home instead, cancel the tool. One at a time, in the order the workflows ship.
Cleaning happens along the way, not up front. The Stanford Digital Economy Lab's Enterprise AI Playbook puts it more bluntly than I would have dared: "Store everything, connect it, and let the models do the cleaning." Duplicate contacts, inconsistent company names, three spellings of the same vendor: those are exactly the tasks a language model handles well once the records sit side by side. They are miserable to fix by hand across five tools that cannot see each other.
So: do not start by cleaning. Do not start by cancelling. Start by connecting into something that is yours.
Founders resist this because building a database sounds like a project and cancelling a tool sounds like a decision, and the decision feels faster. It is not. Cancel first and you lose the data before it has a home. Clean first and you polish records in a system you are about to leave. Connect first and every later step gets easier, because the facts are already where you can see them.
What it looks like when it is done: the Edge8 case

We ran the rule on ourselves before we ran it on anyone else.
Edge8 today runs on one database of 164 tables across nine entities. Every table has a named owner, one person accountable for what goes into it. As I write this, that home holds 925 people, 256 companies, 139 deals, 347 meetings and 300 tasks across 29 boards. Not big numbers. That is the point. A small company's entire operating reality fits in one place with room to spare.
What stayed. The accounting system. It is good at accounting and we had no reason to rebuild it. It syncs into our home weekly. The chat platform stayed too, and its meeting transcripts and calendar sync in nightly. Email goes out through a delivery API, a service that does nothing but send messages and report what happened to them.
What was integrated. All of the above, into the same tables. A meeting from the chat platform attaches to the person and the deal it belongs to. An invoice attaches to the company it was sent to. The record is the centre. The tools are inputs.
What was never bought. No CRM. No applicant tracking system. No HR platform. No project board. No survey tool. No coaching software. Every one of those functions runs from the same 164 tables, and we never signed up for a single one of those subscriptions.
The draft I quoted earlier describes the shape, and I have not found a better way to say it since: "Not a data warehouse project. Not an integration layer duct-taping forty tools together." What we wanted was "One home, where a customer is one record, a person is one record, a deal is one record, and everything that happens to them: meetings, invoices, applications, requests, attaches to that record."
Two things follow from owning the data that founders rarely anticipate.
The first is that the model underneath becomes swappable. Our brand voice, our editorial lenses and our colour palette each live in a database row, one per brand, not in a vendor's prompt box. Every agent we run reads from the same tables. If a better model ships next month, we change the model. Nothing else moves.
The second is security, and this one surprised me. Every server action in our system checks who is asking before it touches a row, and every write is logged with who made it. That is only possible because the data is ours. It is also what made it safe to put HR, payroll and client records through the same agents that handle everything else. The Stanford playbook has a line for this: "Security enables more than it blocks." Owning the home is what let us use it for the sensitive work, not just the convenient work.
The same rule at eleven systems
The obvious objection is that Edge8 is small and started clean. Fine. Here is the rule applied to a footwear retailer with eleven systems: an ERP (the system that runs inventory and orders), an e-commerce platform, a merchandise planning tool, payroll and rostering, accounting, a CRM and more.
This retailer had done something most companies never do. It had written down its workflows. Thirty-five of them, documented step by step. And it could not execute them.
Not because the processes were wrong. Because the data each workflow needed sat in different systems that did not talk to each other. Picture any process that needs a sales figure from one system, a stock level from a second and a plan from a third. On paper it is four steps. In practice it is three logins, two exports and a spreadsheet.
That is the sharper lesson from this engagement, and I would put it on a wall: a written process with disconnected data is a wish list. Documentation was never the blocker. Connection was.
So we applied the rule. Nothing was replaced on day one. One owned home was built next to the eleven systems, and the systems were integrated into it first. Along the way we found that orders from three of them already landed in the ERP, so we dropped a duplicate sync rather than building it. Integration is where you find those things. Removal comes after, one workflow at a time, as each of the 35 ships from the home instead of from the patchwork.
Eleven systems is a lot. The rule did not change. Home, integrate, remove. The retailer is simply further into the second step than a three-tool startup would be.
Three companies, one model

Between the retailer, the wellness startup and a PR advisory, here is where the rule landed.
| Company (industry) | Systems | Integrated first | Removed after | | --- | --- | --- | --- | | PR advisory | 3 | 3 systems | 8 client spreadsheets | | Footwear retailer | 11 | 11 systems, in progress | spreadsheets, as each of 35 workflows ships | | Wellness startup | 3 plus sheets | Shopify, waitlist tool | core spreadsheets; CRM and messaging SaaS never bought | | Total | 17 | | 50+ spreadsheets across the three |
The PR advisory's eight client spreadsheets, one per client, all live at once, became one table the firm owns. The wellness startup moved its core spreadsheets (employees, vendors, partners, contracts) into a home alongside Shopify and the waitlist tool, and bought nothing new. Because partners, waitlist and orders now sit together in something the company owns, it can personalise outreach per contact. No single rented tool could have offered that, because no single tool held all three.
More than 50 spreadsheets retired across the three companies. Every one of them after the data it held had a home. Not one before.
And here is the finding that should change how you think about the build: across a PR advisory, a footwear retailer and a wellness startup, the Edge8 core data model covered about 80% of what each company needed. The other 20% was their own tables on top.
Small companies are not special. They are the same model with fewer rows.
The model is a commodity. The owned layer is the asset.
The durable advantage is in the orchestration layer, not the foundation model.
Finding 11 of the Stanford playbook is the part every founder should read before their next AI purchase. "Model choice is a commodity for many use cases." More precisely: "For 42% of implementations, model choice was fully interchangeable." And the conclusion: "The durable advantage is in the orchestration layer, not the foundation model."
Orchestration layer needs unpacking. It means the part of your system that decides which data to pull, which steps to run in which order, and what to do with the result. It is the plumbing between your records and the model. And it can only be as good as the records it can reach.
Put the two halves together. The model you will use in two years does not exist yet, and it will not matter much which one it is. Whether the data it reads sits in one place you own or across tools you rent will matter enormously. Spend this year picking models and buying AI features inside your SaaS tools and you are investing in the interchangeable part. Spend it building the home and you are investing in the part that lasts.
This is why I do not think of the database as a technical project. It is the most business-critical thing a small company builds, because everything you later automate, analyse or hand to an agent runs off it.
The question is not what to build. It is who owns it.
If 80% of the model is the same across a PR firm, a shoe retailer and a wellness brand, then "what should our database look like" is mostly a solved question. The open question is who on your team owns it.
Not a consultant who builds it and leaves. Not a vendor whose roadmap decides what you can store. An engineer inside your company who owns the data layer and the workflow redesign that follows: someone who understands your ERP well enough to integrate it and your business well enough to know which of the eleven systems goes first. That is the engineer Edge8 places.
Your next step
Before you talk to us, do one thing. It takes an hour.
List every tool that holds a fact about your customers, your people or your money. Every one: the e-commerce platform, the waitlist tool, the payroll system, the WhatsApp group, the spreadsheet with the vendor contracts. Next to each, mark whether you own it or rent it.
For most small companies the second column is almost entirely "rent". That is not a failure. It is the starting position, and the shorter the list, the faster you can change it.
Then reply to Edge8 with that list. We will tell you what your one home should look like, which tool integrates first, and who should build it.
FAQ
How to centralize business data before buying another SaaS tool?
Follow three steps in order. First, build one database you own where a customer, a person and a deal are each one record. Second, integrate every existing tool into it, so Shopify, the waitlist tool, accounting and the spreadsheets all write into that home without turning anything off. Third, remove the applications you no longer need, one at a time, as each workflow runs from the home instead of from the tool.
What does the Stanford Enterprise AI Playbook say about model choice versus the data layer?
The Enterprise AI Playbook from the Stanford Digital Economy Lab, by Pereira, Graylin and Brynjolfsson, states that model choice is a commodity for many use cases and that for 42% of implementations model choice was fully interchangeable. Its conclusion is that the durable advantage sits in the orchestration layer, meaning the part of the system that decides which data to pull, which steps to run and what to do with the result. That layer can only be as good as the records it can reach, which is why owning the data matters more than picking the model.
Should a startup running Notion, Shopify and WhatsApp buy a CRM?
In the case Edge8 worked on, no. A pre-launch wellness startup already tracked employees, vendors, partners and customers across Notion, Shopify, an email waitlist tool, WhatsApp and spreadsheets, and was about to add a CRM and a messaging subscription. Instead it built one owned database, integrated Shopify and the waitlist tool into it, moved its core spreadsheets in and bought nothing new. Because partners, waitlist signups and orders now sit together in one place, it can personalise outreach per contact, which no single rented tool could offer.
How does Edge8 run without a CRM, HR platform or project board?
Edge8 runs on one owned database of 164 tables across nine entities, every table with a named owner, holding 925 people, 256 companies, 139 deals, 347 meetings and 300 tasks on 29 boards. The accounting system stayed and syncs in weekly, the chat platform stayed and its transcripts and calendar sync in nightly, and email sends through a delivery API. No CRM, applicant tracking system, HR platform, project board, survey tool or coaching software was ever bought, because every one of those functions runs from the same tables.
Should you clean your data before centralizing it?
No. Connect first, then clean along the way. The Stanford playbook puts it as store everything, connect it, and let the models do the cleaning, because duplicate contacts and inconsistent vendor names are exactly what a language model handles well once the records sit side by side. Cleaning first means polishing records in a system you are about to leave, and cancelling first means losing data before it has a home.
