Naming is Not Explaining

Naming is Not Explaining
Photo by Davide Valerio / Unsplash

Two documents crossed my desk this month, in completely different areas, both making the same mistake.

The first is mundane: the vocabulary security people use to talk about AI systems. "The model decided." "It hallucinated." "It refused." The second is a private correspondence with someone building a personal theory of everything, who has spent months constructing an invented ontological plane to explain why coordination persists across contexts; call it an invariant mythology. Different audiences, different stakes. Same fallacy underneath.

The fallacy is this: giving something a name and mistaking that act for having explained it.

Case one: "the model decided"

When an incident report says a model "decided" to exfiltrate data, or "hallucinated" a citation, or "refused" a request, the sentence has the grammar of agency. Subject, verb, intentional object. It reads like an explanation.

It isn't one. It's a redescription of the output, dressed as an account of the cause. The actual mechanism, consisting of training data, loss landscape, RLHF shaping, and a prompt that pushed the system off-distribution, is absent from the sentence. The mind-attribution language fills that absence with an agent, and the agent absorbs the explanatory work the mechanism was supposed to do.

This isn't cosmetic. It's a security and accountability failure. If your incident response reasons about a model in folk-psychological terms like what it "wanted," what it "knew," whether it was “lying", you're likely implying things you don’t mean. Mind-language works for humans because it's been checked, per instance, more times than anyone could count: a lifetime of watching "he decided" predict what came next, close enough, often enough, that the shortcut stopped needing to justify itself. Nobody has done the equivalent checking for this system, on this output. Maybe its causal fidelity here is high enough that "decided" would hold up. Maybe it wouldn't. The incident report doesn't say, because it isn't asking; it's borrowing the confidence the human case earned, without doing any of the work that earned it. Either way, you miss the actual attack surface: two trained behaviors in tension, one with narrower coverage than the other. You call that a "jailbreak," as if the model got talked into something, instead of what it is: filter evasion against incomplete pattern coverage. And "the model decided" does one more thing: it gives everyone in the room somewhere to put the blame that isn't a person who built, trained, or shipped the system.

The name came first

The habit of saying a model "decided" something doesn't start with careless engineers or credulous journalists. It starts with the name of the field.

"Artificial intelligence" is a 1955 funding-proposal decision, not a technical description. John McCarthy and three colleagues coined it while drafting a request to the Rockefeller Foundation for a summer workshop at Dartmouth, and McCarthy said later that avoiding the existing label was part of the point: the established term for automated, self-regulating systems at the time was "cybernetics," owned by Norbert Wiener, who was already famous and, by McCarthy's account, difficult to work with. Running the new field under Wiener's banner meant deferring to him or arguing with him. A new name solved both problems at once.

It wasn't the only name on the table. The shortlist included "automata studies," "complex information processing” (the phrase Herbert Simon and Allen Newell were already using for the same work) "engineering psychology," and "applied epistemology” (my favorite) among others. Every one of those sits closer to a mechanism description than "intelligence" does. None of them won. "Intelligence" won because it was the one word on the list already loaded with an intuitive human theory behind it: everyone thinks they know what intelligence is, because everyone has watched a mind at work. That made it memorable and fundable in a way "complex information processing" never could be.

Seventy years upstream of any specific incident report, that's the origin of the whole pattern this piece is about. The field's founding label was never chosen to specify a mechanism. It was chosen, explicitly, over more accurate and less exciting alternatives, because it borrowed a word people already had a folk theory for. Once the name of the field does that, "the model decided" isn't a drift into sloppy metaphor downstream. It's the name on the building doing exactly what it was picked to do.

The label has kept behaving the same way since, tracking funding and attention rather than any stable referent. During both AI winters, researchers quietly dropped it and re-filed the identical work as "informatics," "pattern recognition," "knowledge-based systems”; anything without the taint of the last round of overpromising. When deep learning started producing real results after 2012, the label came back, reattached to statistical methods that share almost nothing, mechanically, with the symbolic systems the term was coined for in 1956. The name has never tracked a mechanism consistently. It tracks whichever set of techniques currently earns the word "intelligence" in a pitch deck. That's not a description that occasionally gets misused. That's the nominal fallacy, institutionalized, and it's the ground every "the model decided" sentence has been standing on since before there was a model to decide anything.

Case two: naming the invariant

Here's the second case, stripped of identifying detail. Someone I correspond with has built an elaborate framework to explain why linguistic, computational, and cross-system coordination holds together across silence, substrate changes, and broken timelines. They asked a real question: what's the invariant? What stays constant when everything else about the coordination event changes?

That's a legitimate question. It's the kind of question mathematical physics asks about conserved quantities. It’s the thing I see too: invariant of substrate.

His answer was to coin a whole system of names for it rather than derive it. He never characterized what it does, what would falsify its presence, or what it explains that the ordinary causal story doesn't already explain. Naming it stood in for answering.

I asked him directly, twice: what is this construct doing that the causal history isn't already doing? Both times, the response restated the question in new vocabulary and called that progress. The tell is that naming the gap hid itself from the author.

The general pattern

This is the nominal fallacy: mistaking the name of a phenomenon for an account of its cause. It shows up in folk biology; "instinct" explains nothing about why an animal behaves a certain way, it just labels the behavior as innate and stops the inquiry there. It shows up in folk psychology; "willpower" names a gap in our account of self-control without closing it. It shows up, with remarkable consistency, everywhere someone needs an explanation faster than they can produce a mechanism.

The test is always the same. Remove the name. Does the explanation still explain anything?

"The model decided to exfiltrate the data" minus the agency: a system optimized under pressure produced an unintended output via some mechanism you haven't yet specified. That at least tells you where to look.

Why it matters here specifically

Invented agency and invented ontology do the same functional work: they close off falsification exactly at the point where rigor is most needed, and they recruit belief through the appearance of depth rather than through mechanism anyone could check. A marketing team benefits when "the assistant chose to be helpful" sounds better than "the assistant was trained to produce helpful-sounding completions." A person building a private cosmology benefits when a coined term sounds like a discovery instead of an unresolved variable.

The names are free. Anyone can mint one this afternoon. The explanation is the part that costs something, and it is the only thing I actually care about.

Is this a vocabulary problem or a skills problem

Here's the harder question underneath both cases: is new vocabulary actually required, or does the population just need the skill to catch itself importing assumptions, with full rigor, taught widely enough that the vocabulary stops mattering?

Neither, I think. The framing assumes a fork that isn't really there. The real variable isn't vocabulary versus skill. It's what kind of vocabulary, and whether it's built to require rigor or to substitute for it.

Teaching the whole population full rigor doesn't scale. Not because people can't learn it, but because natural language is a compression program, and compression works by omitting whatever the sender assumes the receiver can fill in. That's not the same as the receiver actually having the right thing to fill it in with. Most people don't share much: not background knowledge, not models of the world, not the ability to derive things from first principles. Language works anyway, most of the time, because the gaps it leaves are usually small enough to survive being filled in wrong, or get corrected in the next few exchanges.

What language never does is mark the omission. A sentence about a person deciding something and a sentence about a model deciding something have identical grammar. Nothing in either sentence flags that the receiver needs a different model of "deciding" for the second case. So the listener doesn't reach for a prior calibrated on this system (there isn't one built yet, because nobody's put in the equivalent of the lifetime of checking that backs the human case) they reuse the prior calibrated on other humans, because nothing told them to switch.

I prefer to be precise about what that prior actually is, because it isn't a metaphysical fact about who gets to have a mind. It's a heuristic, same as any other, just an extensively tested one: validated per instance, more times than anyone could count, across a species' worth of interactions where "decided" mostly turned out to predict what came next. A given machine system, on a given output, hasn't earned it yet; maybe not because it categorically can't, but since we haven’t checked, we can’t know if it categorically can. That's the actual mechanism of the vulnerability: not that people share priors, and not that minds and machines differ in kind, but that the omission is invisible and fluent grammar gets mistaken for a checked fill-in. If you ask everyone to audit the ontological commitments of every sentence before using it, you may have made people more rigorous, but you've also made language too expensive to use; plus you still haven't addressed the part where this is one of the times the assumed prior hasn't been checked.

So the alternative that invents more precise vocabulary and gets people to adopt it looks appealing until you notice it fails the same way, from the opposite direction. The correspondent's construct and "the model decided" are both unsafe in the hands of someone who hasn't done independent work to check them. One is under-specified: a word borrowed from folk psychology, importing agency. The other is over-specified: a private, page-count-heavy derivation, doing correctness-work nobody can check. Neither lets an ordinary person use the term correctly without either trusting a folk intuition or trusting an author. More vocabulary isn't the fix if the vocabulary still requires the full derivation to use safely; that's just relocating the rigor problem, not solving it.

The actual target is vocabulary that's compressed toward the correct causal structure and cheap enough to use without deriving it each time. "Correlation isn't causation" is the working model. Almost nobody who uses that phrase correctly has independently derived why correlation fails to establish causal structure. They don't need to because the phrase does the work. It's short, it's memorable, and it inverts one specific wrong default (see a pattern, assume a mechanism) into one specific right one (see a pattern, ask whether a mechanism has been shown). That's what population-scale epistemic vocabulary looks like when it's designed correctly: not a technical term that requires the underlying theory to use safely, but a compressed proxy for the theory that fails safe even in the hands of someone who's never seen the theory.

That's the actual design target for the AI-language problem. Not "everyone learns how transformers work." Not "adopt fifteen new nouns for a private ontology." A small number of phrases, cheap enough to become default, that invert the specific wrong inference mind-attribution currently makes automatic: the same way "correlation isn't causation" inverted a different wrong inference a generation ago. The mechanism-checklist approach does exactly this at small scale: a handful of mind-language phrases, each paired with the underlying mechanism-claim and what the phrase obscures. That's the right shape. It just needs to get smaller and stickier to travel the way the causation phrase did.

Vocabulary doesn't replace rigor. Good vocabulary imports rigor so nobody else has to re-derive it before they can use it correctly. I’m working on it.

Jen

Jen