AI ProgramsRetreat
AboutCareersBook a Conversation
The software was never the hard part

The software was never the hard part

You are about to buy another AI system. Maybe it is the second vendor, because the first one stalled in a pilot. Maybe it is a bigger model, or an agent platform that promises to finish what the chatbot started. Before you sign, look at one number, because it changes what you are actually buying.

In March, the Stanford Digital Economy Lab published "The Enterprise AI Playbook" (Pereira, Graylin and Brynjolfsson). The authors went to practitioners running real enterprise AI projects and asked a simple question: what was the hardest thing to fix? Not the most expensive invoice. Not the most embarrassing demo. The hardest thing. Their summary is six words long: "Technology is not the hardest part."

Then they put a number on it: "77% of the hardest challenges were invisible and intangible costs: change management, data quality, and process redesign." The system you are comparing vendors on is the other 23%.

I have been telling founders a version of this for years, usually after they have already bought the tool. What is new is that it is now a number. A number can go in a budget, and a budget is the one thing most AI projects never had.

Idea in Brief
The ProblemYou are about to buy another AI system after the first one stalled in pilot, and you are comparing vendors as if the model were the decision. Stanford's new enterprise study says the system is the 23%: the 77% that decides whether a project works is change management, data quality and process redesign, and none of it shows up on the invoice.
The InsightThe 77% is not a surprise cost, it is the budget you never wrote down. For every $1 of software companies spend up to $10 on intangibles, and the projects that scale are the ones where a person inside the company owns the process redesign and another owns the data layer.
The Way ForwardBefore you sign, list the three AI projects you have already paid for and write down what each one actually cost in people's time. Then reply to Edge8 with that list to talk about who inside your company should own the 77%, and whether that person exists yet.

What practitioners said was actually hard

61% of successful projects included at least one prior failure, whose costs never appear in the final ROI.

Horizontal bar chart of eight categories of enterprise AI challenges, with five invisible-cost categories totaling 77% led by change management and adoption at 33%, and three technical and visible categories totaling 23% ending with vendor and platform at 6%.

Exhibit 1. Change management and adoption alone is a third of what practitioners struggled with; the vendor and platform you spend months comparing is the smallest category at 6%. Source: Stanford Digital Economy Lab, The Enterprise AI Playbook (Pereira, Graylin and Brynjolfsson).

Here is the Stanford breakdown, invisible costs first.

  • Change management and adoption: 33%
  • Data quality and architecture: 17%
  • Process redesign: 10%
  • Quality and accuracy: 10%
  • ROI and business case: 7%

Those five add up to 77%. Stanford calls them invisible costs because none of them arrive as a line on a vendor's quote.

  • Technical and integration: 10%
  • Governance and compliance: 7%
  • Vendor and platform: 6%

Those three add up to 23%: the technical and visible costs, the ones a procurement process is designed to evaluate.

Look at the top and bottom of that list together. The single largest category, a third of everything practitioners struggled with, is change management and adoption. In plain language, that means getting people to work differently once the tool exists. The smallest category, at 6%, is the vendor and the platform. The thing you spend three months comparing is the thing least likely to be your hardest problem.

One more number from the same report, the one I would tape to the wall of the room where the vendor demo happens: "61% of successful projects included at least one prior failure, whose costs never appear in the final ROI." Six out of ten wins had a loss behind them that the spreadsheet forgot. When a vendor hands you a case study, you are very often reading the second attempt with the cost of the first attempt erased.

The 77% is not a surprise. It is the budget.

Timeline variance is organizational, not technical.

Comparison graphic showing one blue square for $1 of tangible tech investment beside ten mint squares for up to $10 of intangibles, above two bars showing 88% of organizations use AI in at least one function while only one-third have begun to scale at the enterprise level.

Exhibit 2. The software is the small square; the intangibles around it can run up to ten times larger, which is why 88% of organizations have started and only one-third have begun to scale. Source: Stanford Digital Economy Lab, The Enterprise AI Playbook (Pereira, Graylin and Brynjolfsson).

The Stanford authors connect their finding to something economists have measured before, the Productivity J-Curve. The name is literal: output dips before it climbs, tracing the shape of the letter J. Here is the line that matters: "Earlier research found that for every $1 of tangible tech investment, companies spend up to $10 on intangibles (process redesign, reskilling, organizational transformation), initially depressing productivity before gains are realized."

Read that as a founder, not as an economist. If you approved a budget for software and nothing for the people work around it, you did not underestimate the project. You left the largest line off the sheet entirely. Then the productivity dip showed up on schedule, nobody had planned for it, and the project got labeled a failure at exactly the point where the curve was about to turn.

That is my explanation for the most damning statistic in the report: "While 88% of organizations use AI in at least one function, only one-third have begun to scale their AI programs at the enterprise level." Two thirds of organizations remain in testing or proof of concept. Nearly everyone has started. Almost nobody has finished, because almost nobody funded the finishing.

Stanford is blunt about where the delay lives: "Timeline variance is organizational, not technical." And about what separated the companies that scaled from the companies that stalled: "The difference was executive sponsorship, existing organizational processes, and end user willingness." Three things on that list. Zero of them are a model.

The companies that get through the dip plan for it. The report cites Accenture research showing that successful AI scalers are 65% more likely to set one to two year timelines to move from pilot to scale. In the report's own words: "Intentional timelines beat 'move fast.'" The short pilot did not fail because the technology was slow. It failed because a few months was never a serious timeline for the 77%.

Data is the other half of the budget. "Data foundations are a major line item," the report says, and it shows what that line buys: strategic scalers are far more likely to possess a large, accurate data set, 61% versus 38% for non-scalers. Nobody scaled on top of a spreadsheet they did not trust.

Our own 77%: three failures that had nothing to do with the model

I do not write about this from the outside. Edge8 runs its own company on one database of 164 tables across nine entities. Twenty-five scheduled agents run on Vercel, nightly jobs run on a Mac mini, and every agent run is logged with the AI tokens it cost. As of this week that log holds 937 runs. I list that because we have broken this system repeatedly, and every break landed in the 77%.

Three examples, all real.

First, a coaching recap chain stopped. The model was fine. Production simply had no chat-platform credentials, so the recap had nowhere to go.

Second, a weekly sales scorecard task stopped. Again, the model was fine. A database password had gone stale and nothing told anyone.

Third, an intake survey started creating duplicate people records. The survey worked. Nobody had decided how the system should recognize that a person filling in a form was a person we already knew.

None of those are model problems. They are plumbing and process problems: who owns the handoff between environments, who owns credentials, who owns the definition of a person in our data. Each one was found and fixed by a human doing the 77%. If you had asked me to write the ROI of our automation program, none of those hours would have appeared, which is exactly how you get to 61% of wins hiding a prior failure.

So we made the hours visible. Edge8 measures the human side of every engagement as Human Tokens Tracked. The hours come from real working sessions under a per-person daily budget, never from a client's own estimate. Machines get their tokens logged on every run; people get theirs logged the same way. When the humans and the machines are both counted, the 77% stops being invisible and starts being a number you can manage.

I wrote this in a draft about automating our own company, and it is still the clearest way I can say it: "The information existed. It just lived in so many places that no one person, including me, could see the company we were running." The fix was not a better model. It was a person willing to pull the company into one place so that, as I put it in the same draft, "The humans walk in at the decision point, not the data-gathering point."

Eight spreadsheets and one head

It isn't a motivation problem. It's a translation problem.

Two-column table mapping a founder's stated problems, trapped knowledge with no playbook, handing work to senior hires, and three systems of record plus eight live spreadsheets, onto the Stanford categories of process redesign at 10%, change management and adoption at 33%, and data quality and architecture at 17%, with vendor and platform at 6% never coming up.

Exhibit 3. Every problem the founder named lands in the 77%; the vendor and platform category did not come up once. Source: Edge8 discovery call; Stanford Digital Economy Lab, The Enterprise AI Playbook.

Here is what the 77% looks like from the client side.

A PR and corporate communications advisory in Australia came to us running three systems of record plus one spreadsheet per client to track all PR activity. Eight of those spreadsheets were live. On the discovery call, the founder told us her biggest pain in her own words: the knowledge was trapped in her head, with no playbook and no workflow she could hand to senior hires.

Map that onto the Stanford list. Trapped knowledge and no playbook is process redesign. Handing work to senior hires who then have to be brought along is change management and adoption. Eight parallel sheets and three systems of record is data quality and architecture. Vendor and platform, the 6% category, did not come up once.

The system was the easy part. The work was writing the playbook down and retiring the eight sheets, and that work is done by a person inside the business, not by a license. As I wrote in that same draft: "It isn't a motivation problem. It's a translation problem." She was not short on will. She was short on someone whose job was to translate what she knew into a process a machine and a new hire could both follow.

You do not buy your way out of the 77%. You staff it.

This is the whole argument, so I will say it plainly. The buying decision in enterprise AI is not which model or which vendor. That is the 23%, and the market has made it a commodity. The buying decision is who inside your company does the 77%, and most companies have nobody.

Two roles cover most of it.

An AI officer owns the process redesign. In plain terms, that is the person who decides which workflows change, writes down how they change, secures executive sponsorship, and carries the end users through the dip. Stanford's three differentiators, sponsorship, process and willingness, all sit on this person's desk. The largest category on the list, the 33%, is their job.

An AI engineer owns the data layer. The data layer is everything that has to be true about your data before a model can act on it: where it lives, what a record means, who is allowed to touch it, and what happens when a password goes stale at two in the morning. The 17% on the list, and most of our own three failures, are their job. This is not a prompt engineer. It is someone who can look at 164 tables and tell you which ones lie.

Both of them belong inside the company, because the 77% does not end when a project ends. Consultants leave at the decision point. The people you need are still there the next morning when the scorecard breaks.

This is the first of three posts this week. On Wednesday I will get into the data side, the 17% that decides whether anything else works. On Friday, workflows, and what it actually takes to retire eight spreadsheets.

Your next step

Do not buy the next tool yet. Do this first. It takes an hour.

List the three AI projects you have already paid for. For each one, write down what it actually cost in people's time: the hours in meetings, the hours cleaning data, the hours spent chasing the vendor, the hours someone spent working around it. Be honest, because that number is your 77%, and I promise it is larger than the invoice.

Then reply to Edge8 with that list. We will talk about who inside your company should own the 77%, and whether that person exists yet.

FAQ

What does Stanford's Enterprise AI Playbook say about why AI pilots fail to scale?

The Enterprise AI Playbook, published in March 2026 by Pereira, Graylin and Brynjolfsson at the Stanford Digital Economy Lab, asked practitioners what was hardest to fix in real enterprise AI projects. It found that 77% of the hardest challenges were invisible and intangible costs: change management and adoption (33%), data quality and architecture (17%), process redesign (10%), quality and accuracy (10%) and ROI and business case (7%). Only 23% were technical and visible costs, and the vendor and platform category was the smallest at 6%. The report concludes that timeline variance is organizational, not technical.

What is the Productivity J-Curve and why does it matter for enterprise AI?

The Productivity J-Curve is a pattern economists have measured where output dips before it climbs, tracing the shape of the letter J. The Stanford authors cite earlier research showing that for every $1 of tangible technology investment, companies spend up to $10 on intangibles such as process redesign, reskilling and organizational transformation, which initially depresses productivity before gains are realized. If a company budgets only for the software, the dip arrives on schedule and the project gets labeled a failure right before the curve turns.

Why do only one third of organizations scale AI beyond pilots, according to Stanford?

The report states that while 88% of organizations use AI in at least one function, only one third have begun to scale their AI programs at the enterprise level, leaving two thirds in testing or proof of concept. What separated scalers from stalled companies was executive sponsorship, existing organizational processes and end user willingness, none of which is a model. Stanford also cites Accenture research showing successful scalers are 65% more likely to set one to two year timelines to move from pilot to scale, and strategic scalers are far more likely to hold a large, accurate data set (61% versus 38% for non-scalers).

What does it mean that 61% of successful AI projects included a prior failure?

Stanford found that 61% of successful projects included at least one prior failure whose costs never appear in the final ROI. In practice that means a vendor case study is very often the second attempt, with the cost of the first attempt erased from the spreadsheet. Those hidden costs are the hours spent in meetings, cleaning data, chasing vendors and working around broken tools, which is why Dave Hajdu recommends listing what past AI projects actually cost in people's time before buying the next one.

How does Edge8 measure the invisible 77% with Human Tokens Tracked?

Edge8 runs its own company on one database of 164 tables across nine entities, with 25 scheduled agents on Vercel and every agent run logged with the AI tokens it cost, 937 runs so far. Human Tokens Tracked applies the same logging to people: hours come from real working sessions under a per-person daily budget, never from a client's own estimate. Counting both machine and human tokens turns the 77% from an invisible cost into a number a founder can manage. Edge8's own failures, including a stale database password and duplicate people records from an intake survey, were all plumbing and process problems fixed by a person, not by a better model.

← All Posts