The Label Decision Card

What decision are you making from a label?

Two properties, named separately, so you can see they are not the same one.

Shipped with Episode 5

The Label Decision Card, filled in. A decision to reject a submission, resting on a likely AI-generated label. The label measures textual origin; the decision needs communicative quality. The two are marked as unequal and the verdict reads: wrong property.

The decision it forces

A label has appeared on something, and we are about to do something about it. Reject the submission. Flag the student. Decline the pitch. Discount the post.

Two questions, asked separately and in this order, because asked together they blur:

  1. What does the label actually measure?
  2. What does this decision actually need?

Most of the time those are two different properties, and once they are written on the same page a few inches apart, the gap is impossible to talk our way out of.

Three things collapsing into one

From issue 062 of the newsletter, which is where this argument is published in full:

1. Where the text came from. The button guesses. 2. How it was made. Only the author can say. 3. Whether it is any good. Nothing on the platform touches it.

A detector reaches for the first. Disclosure is the only thing that can speak to the second. The third, which is the one nearly every real decision is actually about, is untouched by either.

Sam's version of this on the episode is an apple:

But it doesn't tell you if it's a good apple. And the measures of good that we can create are far more nuanced and important, I think, than just slapping a label on it.

Provenance can be true and useful and still answer a different question than the one we are deciding. A local farm and no pesticides tells us something real about the apple. It does not tell us the apple is good.

And the positive half, from the same conversation:

When people ask me, "How was this article written?" I always want my answer to be, "Very well, thank you."

The five fields

Field What goes in it Who fills it
decision The action about to be taken. An action, not a feeling you
label_used What the label says, in its own words you
label_measures origin, disclosed process, quality, or unclear you
decision_needs The property that would actually justify the action you
gap Whether those two are the same property derived
better_evidence What would actually answer the decision you

There is no worked example on this page. A filled card would be a real decision about a real person's work, and publishing one would be doing the exact thing this card exists to slow down.

The gap is the finding

gap is derived, not entered. It is true whenever label_measures and decision_needs are not the same property.

In the overwhelming majority of real cases it comes back true, and that is the point rather than a defect in the method. A detector estimates where text came from. The decision resting on it is almost always about whether the work is any good, whether the person can be trusted, or whether a rule was broken. Those are three other properties, and none of them is the one that got measured.

A true label can support a false decision. The label being accurate is not a defence, because accuracy was never the thing in question.

The schema

Schema, label-decision/v0.1
{
  "schema": "label-decision/v0.1",
  "decision": "the action about to be taken",
  "label_used": "what the label says",
  "label_measures": "origin | disclosed process | quality | unclear",
  "decision_needs": "the property that would justify the action",
  "gap": "derived: true when label_measures and decision_needs differ",
  "better_evidence": null
}

better_evidence starts null on purpose. It is the field that turns a complaint into a piece of work.

Run it with any model

Nothing to install and no account. It interviews, and the one thing it is not allowed to do is decide for us what our decision needed.

Copy-paste prompt
Help me fill in a Label Decision Card. I am about to act on a label, and I want to see whether the label measures the thing my decision actually needs.

INTERVIEW ME. Ask at most three questions at a time, and work the fields in this order.

1. decision: the action I am about to take. Push me until it is an action with a consequence, not a feeling or an impression. "I would think less of it" is not a decision. "I would reject it" is.
2. label_used: what the label actually says, in its own words, not my summary of it.
3. label_measures: what property that label estimates. One of: origin, disclosed process, quality, unclear. If the label's own documentation does not say, the answer is unclear, and unclear is a real answer rather than a failure.
4. decision_needs: the property that would actually justify the action in field 1. Ask this WITHOUT showing me my answer to field 3 again, and do not remind me what I said there. The two answers have to be reached independently or the exercise is worthless.
5. better_evidence: what would actually answer field 4. Leave it empty if I cannot name anything.

RULES YOU DO NOT BREAK
1. Do not tell me whether the label is accurate, and do not speculate about how the detector works. Accuracy is not what this card is about, and a true label can still be the wrong evidence.
2. Do not fill decision_needs for me, and do not suggest candidates. If I answer it by repeating what the label measures, ask me once whether that is really the property the decision turns on, then record whatever I say.
3. Compare label_measures and decision_needs and tell me plainly whether they are the same property. That comparison is the only judgment you make here. Set gap accordingly.
4. If they differ, say so in plain words and do not soften it, do not tell me the label is still a useful signal, and do not suggest ways to justify the decision anyway.
5. Do not tell me what to do about the person or the work. This card is about the evidence, not the verdict.
6. Do not ask me for the name of the person, the publication, the student, or the submission. You do not need any of it, and I should not be typing it into a chat window.

OUTPUT
Valid JSON matching the schema label-decision/v0.1 as published at https://sigsub.show/takeaways/label-decision-card/, then the same as a short table, then one sentence naming what I would have to find out before the decision is actually supported.

The other route, which is not this one

Episode 5's guest, Limited Edition Jonathan, would not build this card. His position is that disclosure is the thing keeping the stigma alive, so the answer is to remove the signal rather than to strip its meaning. In his own words, on air:

I don't think people owe it to their audience, and I don't think the audience is owed, any kind of disclosure on how you do your writing.

And his last word on the episode:

I want people engaging with my ideas. I don't want people engaging with my process

He also put Sam's case better than this page does, unprompted:

Yeah, no, and I, and I can see, I, I can see your argument in the long term as well. Just steel man it. Revealing that you're using AI to help you with your writing is something that can normalize it over the long term.

Neither of them talked the other out of it, and the episode does not pretend otherwise. This card takes the route of keeping the signal and stripping the quality meaning out of it. That is a choice, not a settled question, and the other route is held by someone who has thought about it longer than most.

The part that's not optional

Fill decision_needs before looking back at label_measures. The whole method is two independent answers compared afterwards, and reading the first one before writing the second is how the gap quietly disappears.

Then sit with better_evidence being empty, if it is empty. An empty field there means the decision was never supported by anything we can name, and the label was doing work that nothing else was doing. That is worth knowing before we act on it, rather than after.