Episode 13 · 2026-09-25
The Defaults Your Power Users Set
Markus Bernhardt debuts as cohost with a signal about who sets the terms. The more technical people usually bring an AI tool in for their own use, configure it for themselves, and those permissions stay exactly as set while the tool rolls out to everyone else. Nobody reviews them again, and underneath the hood the vendor keeps adding capabilities behind buttons that are already switched on. Sam adds the version with a number on it: a confidence score, where somebody decides where the line goes and everyone downstream treats it as a fact. Markus's subtraction is an insurance pilot that aimed for 15 to 20% automation, landed at 7 to 8%, and was celebrated as the success it was. Sam's is reaching for a large language model by default, when most of the job is sorting.
Host Sam Rogers · cohost Markus Bernhardt · 30 min
Also watch and listen
A choice this episode can help with
Who in your organization is actually setting the AI defaults everyone else inherits, and what would it take to move that decision somewhere else?
Run it on your own work: The Default Ledger, one page, any model.
The signals
- Markus: The people who bring the tool in set the terms everyone else lives under. Very often it's the more technical people who bring an AI tool into the organization, for their own use, and they set the terms it operates under. When the tool rolls out further, those permissions stay what they were at the start and nobody reviews them again. Two things follow. The organization lives with settings nobody chose for it. And underneath the hood the tool keeps changing: new capabilities arrive and hide under a button that's already switched on, which is a vendor changing how you work without asking. Office never worked this way, because every team used Outlook and Excel the same way. His sharpest example is recruitment: a technical team brings in an HR tool, signs the terms and conditions, and as the agentic piece improves they've effectively signed a new recruitment policy, one that updates itself in the background. The fix isn't new. It's decision rights for whoever signs things off, data contracts and a review cadence, permissions reviewed again whenever the tool reaches a new group, and the right two to four people in the room, legal included from the first meeting. Find them by following the tool's touch points through products, services, clients and ways of operating. Sales and marketing are the easy case; R&D, product, legal and HR are where it gets hot, and the sector changes the stakes.
- Sam: Confidence scores are defaults too. More of the AI being bought returns an answer and a number next to it, and somebody decides where the line goes. That line is a default of exactly Markus's kind: set once by whoever is most comfortable with the tool, and inherited downstream as if it were a fact. It's a policy. In Sam's own testing that week, a classifier returned a wrong answer at 95% confidence, so "trust only the confident ones" wouldn't have caught it. And once the results were in, moving the cutoff three points would have made every test pass, with no evidence for the move except that the tests passed, so it stayed where it was. The numbers are arriving faster than the people who should own them. Markus put it in human terms: a colleague you trust at 98% still gets checks and balances, and nobody rings the alarm when they slip a tenth of a point. A 62% win prediction at the US Open means that player loses 38 of 100, and a 95% system gets 5 of every 100 wrong while working exactly as designed.
The subtractions
- Markus: The overfocus on one number in an experimental environment. An insurance company set out to automate part of its claims under $150, the straightforward work where new hires start. In the kickoff, finger in the wind, the team said 15 to 20% automation after the pilot would be great. They documented the workflows and built a rules-based automation, not an LLM, and ran it for three weeks in parallel with the human team, suggesting what it would have done without being allowed to do anything. The comparison showed where the automation was unbelievably good and where it fell down, and also where the humans were less consistent than anyone had assumed. They landed, highly confident, at 7 to 8%. Then the old number resurfaced and someone clever pointed out it was half the goal. The team celebrated it anyway, correctly: they'd been upskilled, they understood what they'd gotten wrong in telling the automation how decisions are made, nobody was let go, and they were ready for the next 10%. A target in an experiment is a destination, not a pass or fail line.
- Sam: Reaching for a large language model by default. A lot of what gets handed to generative models isn't generation or even reasoning. It's sorting: which bucket does this go in, does this match that. That's classification. Sam tested TypeSafe's Jev, a small general-purpose classifier, on jobs he'd normally give a chat model. Matching plain-language questions to a fixed set of research categories, it got 15 of 15 where keyword search got 10, in about a quarter of a second each, and reviewing twelve of the show's own episodes cost less than a fifth of a cent. It also failed: asked whether a legal source actually supported a claim, it missed 3 of 22, which isn't good enough to trust without a person checking. That job looked like a sort and was a judgment. So the subtraction is the default, not the model. Before anything goes to a general model, ask which of three jobs it is. A rule goes to code, which always beats a confidence score. A sort goes to a classifier. A judgment goes to a general model or a person. And change it one call at a time: don't rewire the whole stack in an afternoon. Markus added that the three can be staged, with code first, then a classifier, then a language model working from better context.
About this episode
Sam's analogy for sequencing the three jobs: there's more than one right way to eat an orange, peel first, slice first or not at all, and the way you choose depends on who's eating it. What matters is that the team knows which way it's doing it.
Markus's larger point near the close: a clever idea used to save a team time, and now it mostly creates homework, because every idea has to be tested against the workflow. If everyone is using the same tools, the competitive edge is in the homework rather than the tool. Sam's version is that the bottleneck is the speed at which an organization learns, not the supply of ideas.
The episode's takeaway is The Default Ledger: one row per AI default in force, ending in the name of whoever set it. For the longer version of Markus's argument, including who holds authority behind a default nobody chose, he points to his brief The Decisions Your AI Is Making Without You at Endeavor Intelligence, and to the Endeavor Report, his non-sponsored publication of AI use cases told by the people who implemented them.
Mentioned on air, unpaid: TypeSafe's Jev, which Sam tested on his own work, and PAICE.work, whose error checking Sam described as an elaborate classifier. Disclosure: Sam founded PAICE.work. No money changed hands.
Want to bring your own signal and subtraction? Find yours.
Put it to work
Name one AI default in your org you did not choose and cannot change. Who set it?
The newsletter
Issue 70 · Wednesday 23 September 2026 · sent on Substack and LinkedIn.
One signal, one subtraction, one analogy, in five minutes. It draws on this episode but stands on its own. Every issue.
The takeaway
The method from this episode, handed over whole: one page with its copy-paste prompt and JSON schema. Every takeaway.