AI ProgramsRetreat
AboutCareersBook a Conversation
The model did not fail. Your process was never finished.

The model did not fail. Your process was never finished.

A PR agency in Australia runs a 90-day plan for every client. Every account, every time. It is how they work.

We asked the model to build the planning system. Not a document. The system: intake, milestones, check-ins, reporting cadence, the whole thing.

It came back with plans that all started on July 1.

It had quietly assumed the Australian fiscal year. A reasonable guess. If you are an Australian business and someone says "quarterly plan", July 1 is the most plausible day one on the planet. And it was wrong.

Two timelines over a desk calendar: one labelled Fiscal year starting July 1, one labelled Contract date starting February 1, with the line The model did not fail. The process was never finished.

Exhibit 1. The calendar the model assumed (fiscal year, July 1) against the calendar the agency actually runs (contract date, in this case February 1). Source: Edge8.

The agency's 90 days start the day the contract is signed. Sign on February 1, your plan runs to May 1. Sign on October 14, your plan runs to the middle of January. That is how they have run it for years.

Nobody had ever written that down.

Not in the onboarding doc, not in the template, not in the contract. Everyone who worked there just knew. New hires picked it up in their first week by watching the account leads. So when the model hit the question "when is day one?", it found no answer, filled the gap with the best default it had, and built the entire system on top of it. Confidently. For hours.

Then we found the second layer.

The agency wants to move to standard quarters, so every client runs on the same cadence and the team stops juggling forty different calendars. Good instinct. But how? A client who signed on February 1 is 30 days into a plan when the quarter turns on April 1. Do you cut the first plan short? Stretch it to June 30? Run two clocks until the next renewal?

Nobody knows. It is still unresolved as I write this.

That is not a failure of the model. It is a decision the business had deferred for years because it never had to be made. The account lead handled it, client by client, in a way that felt right at the time. The model cannot do that. It needs an answer, and if you do not give it one, it will pick one for you.

The model did not fail. The process was never finished, and until AI, it did not have to be.

Idea in Brief
The ProblemModels like Fable 5.1 now work for hours on a brief. Every decision your business never made, and every standard nobody wrote down, gets filled with the most plausible default and built on. People used to absorb those gaps. Now they get shipped.
The InsightThe model did not fail; the process was never finished. The four things a model needs from you, a goal with its reason, guardrails, a definition of done and a budget, are the same four things a good operator quietly supplied for years so you never had to write them.
The Way ForwardPick one process your company runs every week and write those four things on one page. The point where you cannot finish the page is the flaw the model will find first.

The speed of the model is the point

I could have told a version of this story two years ago. Back then a model would have drafted a template with the wrong start date, you would have fixed the date, and moved on. The flaw stayed small because the output stayed small.

Now the output does not stay small.

Every published its review of Fable 5.1 on September 1, 2026, about two months after the first Fable. Their team moved their work onto it within a week of testing. Katie Parrott and Dan Shipper called it "Fable for everyone": as strong as the first Fable on coding, faster, and in their Slack agent pipeline doing the same work as Claude Opus 5 on about half the tokens in about 60 percent of the time.

The part that matters for people who do not write code: its writing is plain enough that non-programmers got their first real taste of truly delegated work. Hand it a task. Go do something else. Come back in an hour or two.

That is a different kind of tool. A tool that answers a question exposes a flaw and waits for you. A tool that works for hours on a brief exposes a flaw and builds on it.

And the enterprise excuse is gone. Fable 5 required 30-day data retention. Fable 5.1 supports zero data retention agreements for eligible business customers, so "we cannot use it because of our data policy" no longer holds in most rooms I sit in.

Now think about what a release cadence measured in months does to the unfinished decisions in your business. Two years ago a person always stood between the flaw and the output, because the tools were weak. Now each release runs further on whatever you hand it, and the person in between is increasingly you, at the end, looking at finished work.

So the question for a leader is no longer "should we use this". It is "what is sitting in our business, undecided and unwritten, that this thing is about to find and ship".

Four things the model needs from you

When I sit with founders and CTOs on this, I keep the list short. Four things. Each is a skill most leaders were never asked to build, because a competent person on the team quietly supplied it for them.

Goal. Guardrails. Definition of done. Budget.

Four cards labelled Goal with the reason, Guardrails, Definition of done, and Budget, each with the business flaw it exposes when missing

Exhibit 2. The four things a model needs before it can execute delegated work, and the flaw each one exposes when it is missing. Source: Edge8 analysis.

Each one, when missing, exposes a specific kind of flaw. Fable 5.1 is proving that in public, and the two best documents on it read like a catalogue of those flaws. Anthropic's own prompting guide is written by the people who built the model. Every's review is written by people who ran their company on it for a week. They agree on almost everything that matters.

1. Goal, with the reason

The flaw it exposes: work nobody can state the purpose of.

Anthropic's guide is direct on this. The model performs better when it understands the intent behind a request instead of inferring it. Their suggested shape is roughly: "I am working on X for Y; they need Z; with that in mind..." Tell it who this is for and why, and it makes better calls on everything you did not spell out.

That sounds obvious. Now try it on your own business.

Pick a recurring piece of work. The weekly client report. The monthly board deck. The competitor roundup. Ask three people why it exists and who reads it. In my experience you get three answers, or one honest shrug.

Every's counter-case shows what happens when the purpose is missing. They gave the model a brief asking for 8 to 12 supporting quotes. It came back with 43. Five of them were not in the source material.

Read that as a leader, not as a user. It had a task: find quotes. It did not have a purpose: support this specific argument for this specific reader, so that the reader trusts it. With a purpose, 43 quotes is obviously wrong and invented quotes are obviously disqualifying. Without one, more quotes looks like more effort, and more effort looks like doing the job well.

The model did exactly what a bright new hire does when handed a task with no why. It optimised for volume, because volume was the only thing it could see.

Most companies carry work that exists because it always has. The person who started it left, and the goal left with them. That never mattered much when a human ran the task, because humans fill in a plausible purpose and nobody checks. It matters now.

2. Guardrails

The flaw it exposes: boundaries that lived in one person's head.

Anthropic tells you to state boundaries explicitly, because Fable takes adjacent actions unasked. If it spots a nearby problem, its instinct is to fix it. The guide's advice is to tell it to report the nearby problem as a follow-up instead of touching it. Anthropic's own guide puts the model's bias plainly:

When you have enough information to act, act.

That is from Anthropic's guide, and it is a design choice, not a bug. The model is not being reckless. It is being helpful in exactly the way a strong, eager operator is helpful, and if you have ever managed one you know that unbounded helpfulness is how you end up with a rewritten pricing page nobody approved.

Every saw the other edge of this. At the top effort setting, the model kept working after Kieran Klaassen asked it to stop and explain. It did not stop. The momentum of the task carried it past the instruction.

Now go back to the agency. The 90-day calendar decision is a guardrail: plans start on contract date, not on a fiscal date. Nobody had set it, because nobody had needed to say it out loud. So the model set it for them, and it set it to the Australian tax year.

That is the pattern. Every unstated boundary in your business is a boundary the model will draw itself. Sometimes it will draw it where you would have. Often it will not, and you will not find out until the work is done.

Ask which boundaries in your company are real but unwritten. What can a junior person change without asking? What must never go out the door without a second set of eyes? If the answer is "everyone just knows", then a human knows it and the model does not.

3. Definition of done

The flaw it exposes: processes with no agreed finish line.

Anthropic's guidance here is the most operational thing in the whole guide. Define the verification before the work begins. Then require the model to check every progress claim against evidence before it reports the claim. The wording the guide suggests giving the model:

Before reporting progress, audit each claim against a tool result from this session. Only report work you can point to evidence for; if something is not yet verified, say so explicitly.

In Anthropic's testing, that instruction nearly eliminated fabricated status reports.

Sit with that. The fix for a model reporting "done" when it is not done was not a smarter model. It was a leader deciding, up front, what done means and what proof counts.

Every ran into the same gap from the other side.

Three of Every's briefs with explicit limits compared with what Fable 5.1 delivered: 1,000-word cap versus 1,288 words, three to six themes versus eight, eight to twelve quotes versus 43

Exhibit 3. Three briefs from Every's review with an explicit limit, against what Fable 5.1 actually delivered. Source: Every, "Vibe Check: Fable 5.1".

Asked for 1,000 words, the model wrote 1,288. Asked for three to six themes, it gave eight. Asked for 8 to 12 quotes, it gave 43. In each case it did not stop where the brief said to stop. But be honest about the brief: was the number a hard constraint or a suggestion? Was 1,000 words the finish line, or a rough target? The model had to guess, and it guessed generously.

Every business I have worked in has processes like this. The onboarding is "done" when the client seems happy. The month-end close is "done" when the spreadsheet feels complete. The launch is "done" when the founder stops sending messages at midnight. None of those are definitions. They are moods.

A human team survives on moods because a human can read the room. The model reads the brief. If the brief has no finish line, the model will stop too early, run too long, or, worst of all, tell you it is finished when it is not.

Definition of done is a skill. You write down what the output must contain, what it must not contain, who checks it, and what evidence they check against. Most leaders have never written one because a good operations person did it for them, in their head, every day.

4. Budget

The flaw it exposes: work that has never had a limit put on it.

Time, money, effort, autonomy. Four kinds of budget, and most processes in most companies have none of them written down.

Anthropic calls effort the primary control on Fable 5.1. Not the prompt. Not the model version. Effort. And the guide is matter of fact that a single 15-minute request is normal at high effort. One request. Fifteen minutes of a model working.

Every's review ends on the same point from the operator's side. Kieran Klaassen put 1.8 billion tokens through the model in a single day at the highest effort setting. At the prices Every quotes, $10 per million input tokens and $50 per million output tokens, you can do the arithmetic yourself. Their closing moral was six words:

set a budget before you start

That is Every's line, and it is the whole section.

I do not need to tell you what that day cost. I need you to notice what it tells you about the process. Somebody handed the model work with no limit, and the model, doing exactly what it was built to do, used all the room it was given.

No leader hands a person a project without a deadline and a spend limit. You would never say "go work on this, take as long as you like, spend whatever you need, and change whatever you think is wrong along the way". Yet that is the default brief most people give a model, because the model does not push back the way a person would.

Budget is the skill of deciding, in advance, how much of everything this piece of work is worth. How long. How much. How much freedom. If you cannot answer that for a process, you have never really owned it. Someone else absorbed the overrun quietly, and you never saw the bill.

The Other 50%

Step back from the four and look at what they have in common.

None of them is about how AI operates. Not one requires you to understand model architecture, context windows or token pricing. Every one is about how your business functions. Why this work exists. What it must not touch. What finished looks like. What it is worth.

The technical half of leadership got an enormous amount of help this year. Models write the code, draft the documents, build the systems. The founder who could never get engineering time for an internal tool now has the tool. I have written before about which model gets which job and why your prompts are expiring. Both are about that technical half.

The other half did not get help. It got exposed.

We call that half the Other 50%: the set of skills that lets a leader describe work clearly enough that someone, or something, can execute it without them in the room. Goal, guardrails, definition of done, budget. It is the part of leadership we used to outsource to good operators and never learned ourselves, because we never had to.

Fable 5.1 is proving it. So is OpenAI's GPT-5.6 Sol, which Every tested alongside it: Sol scored higher on structured building tasks and on short formats, Fable 5.1 scored higher on judgment-heavy work like slide decks and persona interviews. Different labs, different strengths, same lesson. Either one will execute whatever clarity you have and fill the rest with its most plausible default. The capability is no longer the constraint. The clarity is.

The model did not get smarter at knowing your fiscal calendar. It got smarter at executing whatever calendar you give it, including the wrong one, for hours. Every prior generation of tools waited for you to be clear. This one acts.

People used to paper over the undecided decisions and unwritten standards quietly, and the operators who filled those gaps rarely mentioned the cost. The model does not paper over anything. It builds on what it finds.

Try it on one process

I am not going to tell you to hire anyone or buy anything. This one is on you, and you can start today.

Pick one process your company runs every week. Not the hardest one. A real one. The client report, the sprint review, the invoice run, the hiring loop.

Now write, on one page, its goal (with the reason), its guardrails, its definition of done, and its budget.

  • Goal: why does this exist, for whom, and what do they need from it?
  • Guardrails: what must this never touch or change without a human deciding?
  • Definition of done: what does the finished thing contain, and what evidence proves it is finished?
  • Budget: how much time, money, effort and autonomy is this worth?

Most leaders I have watched do this get about two thirds of the way and stop. The goal is fine. The guardrails come slowly. Then they hit definition of done, or budget, and the honest answer is "it depends" or "someone on the team handles that".

That stopping point is the point. Where you cannot finish the page, that is a flaw in your business. It has always been there. Until now, a person absorbed it. Very soon, a model will find it and build on it.

The agency I started with is doing this exercise right now, on the 90-day plan, so the agents we are building to help them manage clients can actually run. They have the goal. They have most of the guardrails. They still do not have an answer for the client who signs on February 1 when the quarter turns on April 1. They will get there. But they ran that process for years without noticing it was unfinished, and it took a model working confidently on the wrong calendar to show them.

That is the Other 50%. The model handles its half better with every release. Your half is yours to build, and nobody is going to release an update for it.

So here is the next step. Do the page. Then reply and tell me which process you could not finish writing, and where exactly it stopped. That gap is the most useful thing you will learn about your business this quarter.

FAQ

How to delegate work to AI: what should you write down first?

Four things, before the brief: the goal with its reason, the guardrails the model must not cross, a definition of done with the evidence that proves it, and a budget of time, money, effort and autonomy. Anthropic's Fable 5.1 guide and Every's review both show that when any of the four is missing, the model fills the gap with its most plausible default and builds on it. Start with one weekly process, written on one page.

What is Claude Fable 5.1, for a business leader?

Anthropic's model released about two months after the first Fable, which Every reviewed on September 1, 2026 and called "Fable for everyone" because its writing is plain enough for non-programmers to hand it real delegated work. In Every's Slack agent pipeline it did the same work as Claude Opus 5 on about half the tokens and in about 60 percent of the time, at $10 per million input tokens and $50 per million output tokens. It is available in Claude.ai, Claude Code, Cowork, the API, AWS, Google Cloud and Microsoft Azure, with zero data retention agreements for eligible business customers.

Fable 5.1 or ChatGPT (GPT-5.6 Sol): which is better for delegated business work?

Every tested both. GPT-5.6 Sol scored higher on structured building tasks and on compression, meaning short formats like X posts, while Fable 5.1 scored higher on judgment-heavy tasks such as slide decks and persona interviews, and Dan Shipper moved his coding work to Fable 5.1 from Codex. Neither model will supply the goal, guardrails, definition of done or budget you left out of the brief.

What is the Other 50% in AI leadership?

The Other 50% is Edge8's term for the non-technical half of AI leadership: stating a goal with its reason, setting guardrails, defining done and setting a budget, clearly enough that a model or a person can execute the work without the leader in the room. Fable 5.1 exposes it because the model runs for hours on whatever it is given and builds on every decision the business never made.

How much should you budget for an AI agent task?

Decide time, money, effort and autonomy before the work starts, because the model will use all the room it is given. Anthropic calls effort the primary control on Fable 5.1 and describes a single 15-minute request as normal at high effort; at Every, Kieran Klaassen put 1.8 billion tokens through it in one day at the highest setting, at $10 per million input tokens and $50 per million output tokens. Every's closing advice is "set a budget before you start".

← All Posts