The Vacuum of Expertise

A few months ago in April, someone described a proposal floated in a government meeting: a ban on AI models above 100 billion parameters. The person relaying it wasn't a technologist. Neither, apparently, was whoever proposed it.

Parameter count is not a governance-relevant proxy. It never was. A model's parameter count tells you almost nothing about what it can do, how much it cost to train, or what risk it poses. Two models with the same parameter count can differ by orders of magnitude in capability, depending on architecture, training data, and how much compute went into them. Regulators figured this out years ago, which is why every serious framework including the EU AI Act and the US executive actions on frontier models uses compute thresholds, not parameter counts. FLOPs, not weights.

But FLOPs are themselves a proxy of a proxy. They measure operations performed, not intelligence produced or harm enabled. Watts might get you closer; actual energy expended is at least tethered to physical reality in a way that both parameter counts and FLOP counts are once removed from. None of these numbers is the thing itself. They're all approximations reaching for a target: capability, risk, or harm potential that we don't have a good way to measure directly. The question is which approximation is close enough to build policy on, and which is a category error masquerading as a number.

That distinction should be noted in any room where AI policy gets written, but so far, it isn't.

The instinctive defense of parameter count is environmental: bigger models generally do take more compute to train and run, so isn't parameter count at least a rough stand-in for energy use? The correlation is loose, and it breaks exactly where it matters. A sparse or distilled model can carry a large parameter count and a modest energy footprint. A smaller, denser model can burn more compute per parameter than an old-fashioned giant. If the actual goal is limiting energy use, the honest fix is to regulate energy use directly in measured FLOPs, measured watts, the same figures the EU AI Act already requires providers to disclose for high-risk systems. Reaching for parameter count instead doesn't get you closer to that goal. It gets you a rule a well-resourced lab can route around without changing its energy footprint at all, while a smaller, worse-optimized model with fewer parameters slides through untouched.

The EU AI Act's high-risk provisions take effect August 2, 2026. That's not a hypothetical future deadline: it's two weeks out as I write this. The thresholds attached to it are compute-based: 10^23 FLOPs for presumptive general-purpose AI status, 10^25 for systemic risk. Get the proxy wrong at this stage and you don't just embarrass yourself in a meeting; you write a rule that regulates the wrong thing, misses the models that matter, and catches the ones that don't.

The number that gets repeated in the room is usually the number people have heard, not the number that's right. Parameter count is the one that made it into headlines and product marketing for years ("175 billion parameters," "a trillion-parameter model”) so it's the one that sticks in a policymaker's mind when they're trying to sound informed. Compute thresholds are buried in technical appendices. Nobody puts them on a slide.

So the vacuum isn't really about parameters versus FLOPs. It's about who is in the room when the number gets chosen, and whether anyone there can tell the difference between a metric that tracks the thing you're trying to regulate and one that only sounds like it does. A serious proposal, one with legal teeth, one that will shape which models get built and where, should have someone present who can answer a simple question: proxy for what, exactly? 

This isn't a call for more credentials in the room, or even more ML practitioners. Credentials are their own proxy, and an increasingly bad one. Plus, in my experience, plenty of people who train and deploy these systems day to day don’t necessarily have the parameter/compute/energy distinction. Sitting with a model teaches you to build it, tune it, ship it. It doesn't automatically teach you which of its associated numbers are load-bearing and which are decoration once they leave the lab and become a legal threshold. That's a different skill, closer to measurement theory than engineering, and it's rare on both sides. The actual gap isn't technical versus non-technical. It's whether anyone in the room has thought carefully about what a number is supposed to be standing in for, and whether it still stands in for that thing once it's operationalized.

The cost of that gap doesn't show up immediately. It shows up eighteen months later, when a threshold set on the wrong metric turns out to regulate nothing, or regulates the wrong thing, or gets exploited by whoever noticed the mismatch first. By then the people who set it have moved on. The consequences land on whoever's left holding the framework.

There's an old name for the mechanism, if not the specific failure: Goodhart's Law, once a measure becomes a target, it stops measuring the thing it was meant to track. Parameter count didn't start as a target. It became one because it was the easiest number to repeat in a meeting, and the repeating is now doing damage that has nothing to do with what the number was ever supposed to indicate.

Fix the input, or keep paying for the output.

Jen

Jen