Episode 10 · 2026-09-04
Good Enough for Agentic Work
Paul Gibbons spent a decade attacking change management for never proving it changed anything: the famous study asks the same question twice and calls the match evidence, and no client has ever asked him for his. Now he runs a multi-agent setup that drowns him in verification, and a plumber who left school at fifteen has automated a business the big firms would never have hired him to build. The models have been good enough for a long time. The decision has moved to what goes around them, and the habit that has to go is waving through work you did not read.
Host Sam Rogers · guest Paul Gibbons · 30 min
Also watch and listen
A choice this episode can help with
How do you know an AI-adoption effort changed anything?
The signals
- Paul: The plumber who left school at fifteen runs his business on AI. The speaker after Paul at a Denver talk dropped out of an American high school, joined the military, learned plumbing as a trade, and stood up a business whose entire back end runs on AI: the right part to the right plumber at the right time, invoices, cash flow, the whole hairy mess Paul says he is glad is not his job. If a firm with five thousand plumbers had paid someone five million dollars to build that, it would be a nice story. This man built it himself. Sam's point back: that firm would never have hired him. The bootstraps were built from the bottom up by someone motivated enough to make it happen and who did not know not to. If this lifts the bottom of the pyramid, the guy in Uganda spending weekends in Claude Code, then it is a technology that changes the human race, and that is the signal Paul says inspired him.
- Sam: Buyers are settling on good-enough models and going open. Paul's hypothesis first: for any knowledge work in business the models have been more than good enough for a long time, gains are at the margin, what matters is the rest, the harness. His Hermes setup has a model switcher and will route work to Qwen or Kimi at a tenth or a twentieth of the price. So how does anyone make money out of Fable 5.1 when most people's use cases could run on Kimi? There is no money in models. Sam's signal is that the revelation is landing: in discussions he is in, organizations are saying we have got the one that is good enough for us, and switching to an open-source model so they are not sharing all their information, because the data agreement on the frontier model is one many institutions cannot sign for legal reasons.
The subtractions
- Paul: Stop waving through work you did not read. Asked how a client would know AI adoption is going well, Paul says it is undiscovered country, then answers from his own practice: what are the behaviors, and how do you verify the work product AI generates? When he started, six pull requests waiting on GitHub got yeah, yeah, yeah, merge, merge, merge, and he ran without permissions because he has attention issues. Now he has to discipline himself to read code he does not fully understand, because he is using these tools so much he cannot give everything the pass. The tendency is human: that looks pretty good. A friend used to say "good enough for government work." Think about it hard and it is kind of disgusting: it means this is not important enough for me to devote time to verifying it. That thinking is terrible for the age of AI. If he writes something himself he knows there will be no ridiculous mistakes in it. He does not know that about AI, and he spent the better part of a day rewriting what Claude wrote from his own markdown.
- Sam: The model is no longer the decision. Some of us have been screaming about this for a long time: the model itself is now less important as a decision. Build the stack that is good enough for the work you are doing by how you divide up the work and, more importantly, the guardrails you place around it at the harness level. That is what constrains behavior and makes it trustworthy. It is not the model, it is what goes around the model. No different than with people, which is where last week's episode ended: managing agents is much more like managing people than anybody is comfortable with.
About this episode
The episode opens cold on Paul drowning in his own agents: they pump out so much that he is the one who has to verify it all, and a day of his week goes to validating what they did overnight. Then it backs up into the evidence problem he has been poking at since The Science of Organizational Change. The famous study proving change management makes you 60% more likely to succeed asks people how the change management was and how successful the change was, then correlates the two. Nobody has ever been in a project where the project succeeded and the change management sucked. Paul's own best result, a KPMG culture change that took a group from the lowest revenue per partner to the highest, ran alongside a new CEO reorganizing, hiring and firing, and Paul cannot draw a causal line from what he did to the turnaround, though the CEO would say it was instrumental. No client has ever asked him whether he had evidence for what he was saying. Evidence-based medicine did not get its first paper until the 1990s and a surgeon told him in the 2010s that his department did not adopt it because it prevents innovation. People are perfectly capable of ignoring evidence when it is in their interest to do so.
That is the decision question in its old clothes. Sam's question for AI adoption is the same one: organizations measure visible activity, policies, committees, training, tool availability, and the decision requires evidence of changed behavior or outcomes. Paul's answer is that for most knowledge workers the agentic revolution started in February, nobody has written the book on running a multi-agent operation, and he is the glue: every finished task costs him half an hour to two hours to verify, approve, and revise, and with five or ten agents running he is drowning in open loops. The verification load is where the evidence lives, and it is also where "good enough for government work" does its damage.
Sam's own harness answers Paul's question about how to review without the reviewer producing the same bloated prose. One AI writes the spec and the test; a panel of AIs competes for the work by saying who can do it best and why; the winner runs it and cannot validate its own output, the spec writer does; a third model logs everything, and they are all from different model families so they cannot bluff their way out. That is Harnessie, Sam's open-source project, named on tape in the course of the conversation; nothing was paid and it is not a sponsor segment. Paul's response, having said the last thing he wants is another tool in his stack, is that this is frontier stuff.
The back half is management. Paul reaches for the One Minute Manager four-box: motivation on one axis, competence on the other, and with a motivated, competent person you say go do this and tell me when you are done. Ten years of shared context lets you say make me the deck for Friday; open a fresh Claude Cowork and say the same thing and you get the average of all decks. Sam's version is that the expression of intent should be intentful, and that what he is building toward is aggregated intelligence: not people first versus AI first, but the maximum intelligence you can bring to a problem, with intent honed enough that people and machines pull the same way and you can tell when they are not. When Sam says the thing that keeps people in a predictable set of behaviors is the environment around them, not the request or the instructions, Paul names it: that is nudge theory, context and environment predict behavior better than motivation, purpose, or skills, and it was the whole contribution of that school of change management. The good news is that if you have ever managed anybody, you have a leg up on managing agents. Paul's confession closes it: he is a good leader and a bad manager, his son came to work for him for two weeks and needed hourly check-ins Paul never gave, and that failure mode exists for agents too. Send one into virgin territory with come back at six o'clock and you burn tokens or introduce risk.
Paul would not put anyone on a change team who is not really good with AI, and he is about to sell some big change management work. He knows maybe a dozen people who can talk credibly about org change and leadership and know what they are talking about with AI, and thinks that spot in the middle ought to be the most valuable coin of the realm, because the difficult stuff is the workforce stuff and everyone is spending the money on technology.
Want to bring your own signal and subtraction? Find yours.
Put it to work
What AI initiative are you responsible for, and what evidence would convince you that it is working?