Episode 10 · 2026-09-04

Good Enough for Agentic Work

Paul Gibbons spent a decade attacking change management for never proving it changed anything: the famous study asks the same question twice and calls the match evidence, and no client has ever asked him for his. Now he runs a multi-agent setup that drowns him in verification, and a plumber who left school at fifteen has automated a business the big firms would never have hired him to build. The models have been good enough for a long time. The decision has moved to what goes around them, and the habit that has to go is waving through work you did not read.

Host Sam Rogers · guest Paul Gibbons · 30 min

Also watch and listen

Watch the episodeYouTube: 30 minutes Read + listenSubstack: the episode, in your inbox Podcast, the showApple Podcasts Podcast, the showSpotify ReadFull transcript

A choice this episode can help with

How do you know an AI-adoption effort changed anything?

The signals

The subtractions

About this episode

The episode opens cold on Paul drowning in his own agents: they pump out so much that he is the one who has to verify it all, and a day of his week goes to validating what they did overnight. Then it backs up into the evidence problem he has been poking at since The Science of Organizational Change. The famous study proving change management makes you 60% more likely to succeed asks people how the change management was and how successful the change was, then correlates the two. Nobody has ever been in a project where the project succeeded and the change management sucked. Paul's own best result, a KPMG culture change that took a group from the lowest revenue per partner to the highest, ran alongside a new CEO reorganizing, hiring and firing, and Paul cannot draw a causal line from what he did to the turnaround, though the CEO would say it was instrumental. No client has ever asked him whether he had evidence for what he was saying. Evidence-based medicine did not get its first paper until the 1990s and a surgeon told him in the 2010s that his department did not adopt it because it prevents innovation. People are perfectly capable of ignoring evidence when it is in their interest to do so.

That is the decision question in its old clothes. Sam's question for AI adoption is the same one: organizations measure visible activity, policies, committees, training, tool availability, and the decision requires evidence of changed behavior or outcomes. Paul's answer is that for most knowledge workers the agentic revolution started in February, nobody has written the book on running a multi-agent operation, and he is the glue: every finished task costs him half an hour to two hours to verify, approve, and revise, and with five or ten agents running he is drowning in open loops. The verification load is where the evidence lives, and it is also where "good enough for government work" does its damage.

Sam's own harness answers Paul's question about how to review without the reviewer producing the same bloated prose. One AI writes the spec and the test; a panel of AIs competes for the work by saying who can do it best and why; the winner runs it and cannot validate its own output, the spec writer does; a third model logs everything, and they are all from different model families so they cannot bluff their way out. That is Harnessie, Sam's open-source project, named on tape in the course of the conversation; nothing was paid and it is not a sponsor segment. Paul's response, having said the last thing he wants is another tool in his stack, is that this is frontier stuff.

The back half is management. Paul reaches for the One Minute Manager four-box: motivation on one axis, competence on the other, and with a motivated, competent person you say go do this and tell me when you are done. Ten years of shared context lets you say make me the deck for Friday; open a fresh Claude Cowork and say the same thing and you get the average of all decks. Sam's version is that the expression of intent should be intentful, and that what he is building toward is aggregated intelligence: not people first versus AI first, but the maximum intelligence you can bring to a problem, with intent honed enough that people and machines pull the same way and you can tell when they are not. When Sam says the thing that keeps people in a predictable set of behaviors is the environment around them, not the request or the instructions, Paul names it: that is nudge theory, context and environment predict behavior better than motivation, purpose, or skills, and it was the whole contribution of that school of change management. The good news is that if you have ever managed anybody, you have a leg up on managing agents. Paul's confession closes it: he is a good leader and a bad manager, his son came to work for him for two weeks and needed hourly check-ins Paul never gave, and that failure mode exists for agents too. Send one into virgin territory with come back at six o'clock and you burn tokens or introduce risk.

Paul would not put anyone on a change team who is not really good with AI, and he is about to sell some big change management work. He knows maybe a dozen people who can talk credibly about org change and leadership and know what they are talking about with AI, and thinks that spot in the middle ought to be the most valuable coin of the realm, because the difficult stuff is the workforce stuff and everyone is spending the money on technology.

Want to bring your own signal and subtraction? Find yours.

Put it to work

What AI initiative are you responsible for, and what evidence would convince you that it is working?