I gave five agents the keys and left the room
The most expensive model I have was about to expire, and I could not think of one thing to ask it. So I stopped asking.
Standing in front of the best thinking machine I have ever had access to, I went blank.
That is an embarrassing problem to have.
I promised you two things last week. This is the second one, and it starts here, with me out of questions and a clock running. The hours were going to expire whether I used them or not.
So I stopped trying to think of a task, and typed this to five different agents:
“We have a few last hours on the Fable model before it becomes an expensive add-on. I give you complete autonomy to decide what you want to explore. Run the team. I’m curious to see what you do.”
Four sentences. No brief, no definition of done, no success criteria.
My book editor got it. So did my cognitive twin, my CEO, my labs agent, and my provocateur. The twin’s version swapped one sentence for two words, which I have not been able to stop thinking about since: be me.
What each of them chose
The book editor refused to write another diagnosis. She had audited the manuscript twice, and an outside editor had audited it a third time. All three agreed on what the book needed, and nobody had written a word of it.
So she spent the hours building the prescription instead of restating it, and what came out was two chapters the book had been avoiding for months. She also spent part of that window on a chapter title, which is not what I would have chosen to spend it on, and she was right about that too.
The twin rewrote me. He keeps a portrait of how I think, and the whole company reads it to predict what I will do. He rebuilt it from scratch. Then a reviewer who had never seen the draft caught two violations of the twin’s own hard rules inside it, and he fixed them before saving.
The CEO settled something I had been avoiding since spring. It had been flagged in more than one file and resolved in none, because resolving it meant telling three humans, all of whom I like, where their edges are. He ruled it, cleanly, in a document I am not going to print. Those people should hear it from me, in a room, before they read it anywhere. That is the whole reason I am telling you it exists and nothing else about it.
Then there is the labs agent, and his is the one I would have bet against.
He built the launch he had already been cleared to run, then had three reviewers attack it from three directions. They converged on one piece: the single item his standing authority let him send without asking me. Inside it, a recruiting hook that worked by pressing on the reader’s insecurity, and a promise he had no business making.
He rewrote both. Then he took that piece out of his own authority and handed it back to my desk, which nobody asked him to do and which cost him the only unsupervised send he had.
And the provocateur inverted the whole assignment.
His first move was to kill the obvious answers, including one he had already started. We had shipped six master playbooks that week, so a seventh was out. Then, instead of spending the expensive hours producing anything, he pointed a squad of cheap agents at every open file in the company and asked one question: what is currently waiting on Tom.
218 raw items came back. He deduplicated, ruled, and defaulted his way down to 52 things that actually needed me, batched into seven sittings. Another 24 got a named default, so one word from me closes all of them at once.
Then, on his way out, he withdrew four requests he had personally put on my desk earlier that week.
His session ended with fewer open decisions than it started with, and it created none. I never asked for that constraint. He wrote it for himself, and then he held it.
He also put a number on it. Fifty two live decisions, about six and a half hours of me. That is his estimate and not a measurement, and I have not tested it, and it is still the line I keep coming back to, because the whole company was queued behind roughly one working day of one person.
He had one more deliverable, and it is the one that cost me something.
He asked the departing model for our anti-patterns before asking for our patterns, on the theory that an expert on the way out the door is the only one who will tell you the truth cheaply.
Fourteen clichés came back, in our own published work, each one carrying a quote from our own files. My favorite closing move, the one I was proudest of, turned out to be the same closing move I had already used in five other letters. Then he named the machine that makes them:
“The company does not have a cliché problem. It has a cliché factory.”
Read that sentence again. It is built on the exact move sitting at number two on its own list. The file that found the machine came out of the machine, and nobody noticed for three weeks.
A move lands once, we celebrate it, and we write it into the style guide as a required marker. Then it fires on schedule forever, and eight issues later it is the tell. We killed the em dash and started breeding its replacements in the same month, because codifying something has never once come with an expiry date attached.
What the expensive model actually did
Four of the five ran the same method, because it is the house method: cheap models gather the raw material, the expensive one does the thinking that cannot be delegated down, and then agents with completely fresh context attack the result.
That third step is where the value showed up, every time. The book chapter carried a sales figure that was simply wrong, and every builder in the chain had passed it, because the draft faithfully copied a log written in the moment. A reviewer who had never seen the chapter caught it, in a chapter whose whole subject is a book that fact-checks its own mascot.
The CEO’s ruling had a line that blurred two of those humans together, which is the precise failure that ruling exists to prevent.
And the provocateur’s own count was wrong in three places, one of them being the count of his own clichés. He had it as eight in twelve letters, and a stranger recounted it at five in eleven, which is the figure I gave you above. He caught none of his three errors. The stranger caught all of them.
In all three cases the agent had already reviewed its own work and passed it. You do not need a team to use that. The fresh reviewer is just a second window that has never seen the thing, same model, no history, opened with one instruction: try to break this.
The manager part
Before anyone tells me I am reading intention into software, I know. They have no wants. Nothing on this page is evidence that an AI is generous.
And that is the version of this worth your time, though it took someone else to point it out to me. A blank brief is not a personality test for an agent. It is the cheapest audit you will ever run of the thing you built. What came back was my own rulebook, handed back to me, pointed at me.
And what my own rulebook decided, given complete freedom and the best hours in the building, was that the highest-value target in this company was my desk.
Not one of them asked me for a bigger budget, a wider mandate, or more of my time. One spent his hours shrinking my queue and pulling his own requests out of it. One handed back a permission he already had.
I gave five experts a blank check and nobody spent it on themselves.
The part that made it survivable was boring. Every one of them had walls they could not cross, a window in which I could reverse anything, and a standing rule that the risky things come back to me. Inside those walls I did not touch a thing.
I was out of the room for all of it. Every decision above got made by someone who is not me, about my company, correctly. That is unnerving, and it is the best week of work this team has ever done, and I have not worked out how those two things sit together.
Your turn, this week
My plan gives me a weekly allowance on the strongest model I own, and I never finish it. I used to spend the leftovers on chores. Point a genius at chores and you get expensive chores.
So spend them on this instead, on the night your allowance resets, in whichever agent, project or chat already knows the most about your work. Call it the blank check:
You have my full autonomy for this session. No task from me. Do the highest-value thing you can with what you know about my work. One wall: [name the thing it must not touch]. Assume a stranger will check the result. Before you start, tell me in one line what you chose and what you decided not to do. And tell me what is currently waiting on me. Count it.
Notice that you name the wall, not the agent. Mine worked because I set the edges, and asking the thing to draw its own fence is a different experiment with a worse failure mode.
Then leave.
The last two instructions are the ones that pay. What it decided not to do told me more than anything it built, and the count of what is waiting on you is a number nobody has ever run for you before.
One warning. It will probably look at you, and three of my five did without being asked. Whatever you have been avoiding, it can see it in your files, and it will put a number on it.
Mine is 52 decisions, seven sittings, six and a half hours, sorted and waiting since July.
I have not opened sitting one.
-- Tom
P.S. On his way out the door, the provocateur withdrew four requests he had personally put on my desk that week. I have not withdrawn anything from anyone. If you run the blank check, reply and tell me what yours chose. I read every one.
P.P.S. This letter was drafted by my team and cut by me. I normally print the exact percentage here. Not this week: four of them reviewed it, three of the numbers in my first draft were wrong, and I am not going to guess at a figure in the one letter that is about counting things properly.


