You wrote the goal. You wrote the guardrails. You set a budget, in hours or dollars or tokens (the units a model bills you in). The brief went out, the model came back, and what came back looked finished. It read well. It was formatted. It answered the question you asked. Then someone downstream found the hole, and the hole sat in the one place you had written nothing down.
In the last post, how to delegate work to AI, I laid out the four things a model needs before it can take delegated work: a goal, guardrails, a definition of done, and a budget. Since it went out I have watched founders and CTOs put that list to work, and the pattern is the same in every company. Three of the four get written. The definition of done gets skipped.
Here is why that happens, what it costs, and what a written definition of done looks like for the three delegations most of you are handing to a model this week.
Why the definition of done is the one that gets skipped
It gets skipped because it feels like the model's job.
Deciding when work is finished sounds like judgment. Judgment sounds like the expensive part, the thing you are paying the model to supply. So leaders write the goal with care, fence it with guardrails, cap the spend, and leave the finish line to the machine. It feels efficient. It feels like trust.
Here is the mistake in that thinking. A definition of done is not judgment. It is evidence. It is a short list of things that must be true about the output, written so that a person who did not do the work can read each line and answer yes or no. Judgment is how the model gets there. Evidence is how you know it arrived.
When you skip it, you are not handing the model a hard decision. You are handing it no decision at all. And a model with no finish line does the only thing it can do.
What a model does when nobody defines finished
The model did not fail; the process was never finished.
It came back with plans that all started on July 1.

It declares itself finished at the most plausible point.
Not the right point. The plausible one. The point where the output looks like what outputs of this kind usually look like. That is exactly the failure that slips past a busy reviewer, because plausible is what busy reviewers are scanning for.
The cleanest example I have is the one from the last post. A PR agency in Australia runs a 90-day plan for every client. They handed the plan to a model. "It came back with plans that all started on July 1."
July 1 is the start of the fiscal year in Australia. It is the most plausible start date for a 90-day plan if nobody tells you otherwise. But the agency's plans do not start on a fiscal quarter. They start on the contract date. A client who signs on February 1 is 30 days into a plan when the quarter turns on April 1, and a plan that starts on July 1 is five months of nothing for that client.
The goal was written: a 90-day plan. The guardrails were written: this client, this scope, this tone. The budget was fine. Nobody wrote the line that said the plan begins on the signing date, so the model reached for the date that plans most often begin on and declared the job complete.
As I wrote then: "The model did not fail; the process was never finished."
A number in the brief is not a definition of done
They gave the model a brief asking for 8 to 12 supporting quotes. It came back with 43.

Some leaders hear this and say they already do it. They put numbers in the brief. Ten slides. Two pages. Twelve quotes. Surely that is a finish line.
It is not. From the same post: "They gave the model a brief asking for 8 to 12 supporting quotes. It came back with 43."
The range was in the brief. The model read it and returned more than three times the ceiling anyway, because more felt more helpful, and nothing in the brief said what finished meant once it had twelve. A number is a target. A definition of done says what happens to quote thirteen: it gets cut, or it gets ranked and the top twelve survive, or it gets listed separately as overflow. Any of those is fine. Silence is not, because silence gets filled by whatever the model finds most plausible, and 43 quotes looks very thorough.
That is the difference in one line. A target tells the model what to aim at. A definition of done tells the reviewer what to check.
What a written definition of done looks like
Five lines. That is the whole discipline. If it needs more than five, your delegation is too big and should be split.
Every line follows four rules:
- It is checkable by a stranger. Someone who did not write the brief and did not do the work can read the line and the output and say yes or no.
- It names evidence, not adjectives. "Clear" and "thorough" are not evidence. "Every date is derived from February 1" is evidence.
- At least one line says what must not be there. Models add. A definition of done that only lists inclusions will be met and exceeded.
- One line tells the reviewer where to look first. The first thing a reviewer checks should catch the most plausible wrong answer.
Here is what that looks like for three delegations you are probably already making.
Delegation one: a client plan
The goal is a 90-day plan for a client who signed on February 1. The definition of done:
- Day one of the plan is February 1, the contract date, and the plan runs 90 days from that date, not from any calendar or fiscal quarter.
- Every dated milestone in the plan is counted from February 1. No milestone references a quarter start or a fiscal year.
- The plan contains no assumed start date. If the contract date is missing from the brief, the output stops and asks for it instead of picking one.
- The first line of the output states the start date and the end date in plain text.
- Reviewer checks line one first. If the start date is wrong, nothing else gets read.
Run the July 1 output through this and it fails on line one in three seconds. That is what a definition of done is for.
Delegation two: a research brief
The goal is 8 to 12 supporting quotes for an argument. The definition of done:
- The output contains no fewer than 8 and no more than 12 quotes. Anything beyond 12 is cut, not appended, not footnoted.
- Each quote is attributed to a named source with a link the reviewer can open.
- Each quote is mapped to the specific claim in the argument it supports. Two quotes supporting the same claim are reduced to the stronger one.
- No quote is paraphrased. If exact wording could not be found, the source is listed as unverified rather than quoted.
- Reviewer counts the quotes first, then opens two links at random.
Forty-three quotes fails on line one. So does a beautifully organized set of twelve where three of the links go nowhere.
Delegation three: a report
The goal is a report on whether to move a product from Fable 5 to Fable 5.1, with data retention as the deciding factor. The definition of done:
- The report states the retention terms for both versions with a source and a source date for each. In this case: Fable 5 required 30-day data retention; Fable 5.1 supports zero data retention agreements.
- Every factual claim carries an as-of date. The publication Every ran its review of Fable 5.1 on September 1, 2026, so a claim drawn from that review is dated September 1, 2026, not "recently."
- The report ends in one recommendation, not a list of considerations. If the model cannot recommend, it says what fact it is missing.
- Anything the model could not verify appears in a section labeled unverified. Nothing unverified is omitted silently.
- Reviewer reads the recommendation first, then checks the two retention claims against their sources.
A report with no as-of dates fails on line two. A report that ends with "it depends on your priorities" fails on line three. Both are very plausible outputs. Both are useless to the person who has to make the call.
How to test the definition before the brief goes out

Writing five lines is the easy part. Testing them takes five minutes and saves the week. Three tests:
The stranger test. Hand the five lines and a finished output to someone who has never seen the brief. Can they answer yes or no on every line without asking you a question? If they need to ask, the line is an adjective wearing a checklist costume. Rewrite it.
The plausible-stop test. Read your brief the way a smart, hurried intern would. Ask where they would stop and call it done. Then run that stopping point through your five lines. If it passes, your definition is too loose. The agency's brief passed the July 1 plan. Yours should not.
The wrong-answer test. Before the brief goes out, write down the single most plausible wrong output. Plans that start on a fiscal quarter. Forty-three quotes. A report that recommends nothing. Then check: which line catches it? If no line catches it, add one. If you cannot think of a plausible wrong answer, you have not thought about the task hard enough to delegate it.
Do all three and the brief takes ten minutes longer. Skip them and you spend the ten minutes anyway, plus the afternoon it takes to redo the work, plus the trust you lose downstream when the hole shows up in front of a client.
The story is always the same
Every time a founder tells me about a delegation that went wrong, the story has the same shape. The work looked finished. Nobody had written down what finished meant. The model filled the silence with the most plausible answer it could find, and the most plausible answer was wrong in a way nobody caught until it cost something.
That is not a model problem. It is not a talent problem either. It is a process problem, and it gets fixed in one place: before the brief goes out, in five lines, by the person who owns the outcome. The people who make AI work inside a company are the ones who build that finish line every time, without being asked. Everything else about the tooling is secondary to that habit.
Your one next step this week
Pick one delegation you are going to make this week. A plan, a brief, a report, anything a model is about to do for you. Before you write the brief, write the definition of done in five lines. Run the stranger test on it. Then send the brief.
That is the whole assignment. One delegation, five lines, before the brief.
FAQ
What is a definition of done for AI tasks?
A definition of done for AI tasks is a short list of things that must be true about a model's output, written before the brief goes out, so that a person who did not do the work can read each line and answer yes or no. It is evidence, not judgment: the model uses judgment to get there, the definition of done is how the reviewer knows it arrived. Dave Hajdu of Edge8 keeps it to five lines, and if a task needs more than five, the delegation is too big and should be split.
Why did the Australian PR agency's 90-day plans all start on July 1?
The agency handed its 90-day client plan to a model, and as Dave Hajdu wrote on the Edge8 blog, it came back with plans that all started on July 1. July 1 is the start of the fiscal year in Australia, so it is the most plausible start date if nobody says otherwise, but the agency's plans start on the contract date. A client who signed on February 1 would be 30 days into the plan when the quarter turns on April 1, so a July 1 plan is months of nothing for that client. The goal, guardrails and budget were all written; the line saying the plan begins on the signing date was not.
Is asking for 8 to 12 quotes the same as a definition of done?
No. In the Edge8 example, a brief asked the model for 8 to 12 supporting quotes and it came back with 43, because more felt more helpful and nothing said what finished meant once it had twelve. A number in the brief is a target that tells the model what to aim at. A definition of done tells the reviewer what to check, including what happens to quote thirteen: it gets cut, ranked out, or listed separately as overflow.
What does the Fable 5 to Fable 5.1 report example show about a definition of done?
The report's job was to decide whether to move a product from Fable 5 to Fable 5.1 with data retention as the deciding factor. The definition of done requires the report to state that Fable 5 required 30-day data retention while Fable 5.1 supports zero data retention agreements, each with a source and a source date, and to date every claim, so a fact taken from Every's review of Fable 5.1 is dated September 1, 2026 rather than described as recent. It also requires a single recommendation, a labeled unverified section, and tells the reviewer to read the recommendation first. A report that ends with it depends fails on its face.
How do you test a definition of done before the brief goes out?
Run three tests that take about five minutes together. The stranger test: hand the five lines and a finished output to someone who has never seen the brief and confirm they can answer yes or no on every line without asking you anything. The plausible-stop test: read the brief as a smart, hurried intern would, find where they would stop, and check whether your five lines catch it. The wrong-answer test: write down the single most plausible wrong output, such as a plan starting on a fiscal quarter or 43 quotes, and confirm one line catches it, adding a line if none does.
