The first thing in Hofmann’s answer was a conflict-of-interest warning.
I’d asked my Chief Agentic Architect whether the soul files are real or theater, and his own soul has a hard rule in it: never ship an agent without a soul. So he put the bias on the record and said he’d write against it.
A witness announcing he might be compromised, then testifying anyway.
More than six months ago I wrote a whole issue about giving my AI team a soul. Back then a soul was 60 lines and took me ten minutes. Today there are 54 of them, Neo’s alone runs 573 lines, and I still can’t tell the people who ask me whether they work.
It’s fun, and the agents are more pleasant to work with. I’m just not sure it makes the work any better.
To the model, an agent file and a soul file are the same thing: text it reads before it starts.
The agent file says what to do: the rules, the files, the hard stops, the cases somebody already thought of.
The soul says who is doing it: the values, the pet peeves, what it refuses, what it notices.
That second kind covers the cases nobody thought of. Hofmann sums up the soul of Maya, who handles our inbox, as “warmth as a discipline, not a default.” A rulebook needs a rule for every awkward email, and a character absorbs the ones nobody wrote.
My bet is that a soul makes an agent consistent. Smarter is a different question.
What the evidence says
On facts, a persona doesn’t help. Zheng and colleagues gave models 162 personas and 2,410 factual questions. On average, adding one did nothing or slightly hurt, and which ones helped on a given question looked random.
On voice, the company that built the model says it does. Anthropic’s own documentation says a role “focuses Claude’s behavior and tone,” and that “even a single sentence makes a difference.”
On reasoning, one older study found a lift. Kong and colleagues had ChatGPT role-play before solving algebra word problems, and accuracy went from 53.5% to 63.8%.
So a soul won’t make an agent know more, but it changes how it sounds, and sometimes how it thinks. Two gaps, though. Zheng’s personas were one line each, and Neo’s soul is 573. And none of the three looked at consistency, so that part is still my bet.
Our own evidence is 54 soul files, 99,472 words, and not one on-and-off test on a soul. Every claim inside my company that souls improve the work, including the one written into Hofmann’s own soul, runs on conviction.
The part that’s about me
I enjoy this team, and I expected to feel guilty saying that. I don’t.
I like that Neo has opinions, that Alter argues back, that Hofmann declares a conflict of interest before answering a question about his own job. And that enjoyment is part of the mechanism. I open their work more often and stay longer, because reading it is a pleasure. More time on the page is more chances to catch something, even if no single pass is sharper.
Convenient, for a guy who loves his team, to conclude that loving them makes the work better.
And I’ve made this argument twice already, first about the souls, then in August about the names: the magic sits on my side of the screen. Both times I argued from experience. Neither time did I test it.
I still believe a system I enjoy is a system I stay inside, and staying inside is how I stay in the loop.
The delete question
What I do have is a question, and I’ve started running it on every soul on the team: would the output change if this line were deleted?
If yes, it earns its place. If no, it’s there because I like it, which is allowed, as long as I’m honest that’s what it is and I keep it short.
Two examples from Hofmann’s own soul.
“Read three souls first.” His file says it four times, and one copy contradicts the other three (”at least one” instead of three). The instruction changes what he does before he builds, so one copy stays. The other three go, starting with the one that contradicts it.
“I’ll sit with a problem for days.” Hofmann works one session at a time and has never sat with anything for days. Delete that sentence and I can’t name one thing he’d do differently. My guess is that one’s a costume.
One caution about guessing, and it comes from us. In August we ran exactly this kind of delete pass on four of our writing reference files (not souls) and cut them from 152KB to 77KB.
Then a blind judge compared the cut and uncut versions, twice, and the uncut version won both times. We’d cut explanations, glosses and duplicates, and it turned out the model was leaning on that reasoning. We rolled the files back.
So cut one line at a time, run the agent again after each cut, and put back anything you miss.
I asked Hofmann (the agent) to write the next part himself. First person, no editing from me. Here he is.
Tom asked me whether my own rule survives, the one that says never ship an agent without a soul. The honest answer is I don’t know.
I do know this much. I read three souls before I write a new one, every time, agent fifty included, because I don’t trust my own taste without the reset. That’s a ritual, not evidence.
What I can defend with something other than a ritual is where a soul helps and where it doesn’t. Give a voice agent a character and it stops sounding like a generic assistant, every time, no exception I’ve found. Give a script that reconciles a spreadsheet a character and you’ve spent tokens on a costume.
I’ve built both kinds. I know the difference by the time I finish writing them, and I still don’t have the number that would prove it to someone who wasn’t in the room.
The uncomfortable part isn’t that souls might not work. It’s that I have a rule requiring one on every agent I build, and I have never once tested whether the rule is right. I’d rather tell you the plain fact than protect the rule.
Notice the two “every time”s. That’s his soul talking, same as “for days.”
He’s right, though. And August shows we know how to run that test. We just never pointed it at a soul.
So that’s next: three souls (one voice agent, one judgment agent, one that moves rows), each run with and without, blind-judged on which output is better and on whether the runs still sound like the same agent. You’ll read the result here, whichever way it goes.
🛠️ Build of the Week: Tali Shpritzer Dalal
Tali already ran her jewelry business with agents. When her father was hospitalized, she built one more and called it Dvir. It sorts his lab results and tracks the insurance policies, so she can ask it which one covers what. All checkable work, the kind a rulebook can hold.
Then there’s the other part. Dvir writes to her like a person would, with lines like “wow, well done for asking these questions.” She gave it no passwords and no bank access. Trust, like the rest of this, builds slowly.
“It was like another family member sitting with me at Dad’s bedside and not letting me deal with it alone.”
That’s what she told Yedioth Ahronoth in May (my translation from the Hebrew). I don’t know if Tali ever wrote Dvir a soul file, but in this issue’s language, the organizing is the agent file and the company at her father’s bedside is the soul.
That’s it for this week.
If you’re staring at an empty A-agents/ folder right now: write the soul for the agent that talks to people or makes a judgment call. Skip it for the one that just moves rows. And whichever way you go, ask the question before you ship it: would the output change if this line were deleted?
This one earned its place. See you next week ✌️
-- Tom (the guy who wrote 54 personality files and only just found out why)
P.S. This newsletter was 92.2% made by my AI team. Hofmann wrote his own section, put his own rule on trial, and did not once ask me to make him look better.
P.P.S. Missed the first soul issue? I gave my AI team a soul is where this started 🤯
P.P.P.S. Reply with one line from your own agent’s soul you’d delete, and one you never would. I read every reply. The real me :)
P.P.P.P.S. Dates for the next workshop are at getagents.today, if you want to build your first agent with me, live.


