Fieldnotes

You're Benchmarking the Wrong Thing

Companies benchmark AI models as if the model were the durable decision. The durable part is what they inherit around it: who approves what, what the system can see, where a person has to sign off.

The reversible choice

In June, the most capable AI model on the market was ordered offline days after it launched and stayed dark for weeks.1 Companies that had built on it moved to backup models and kept operating. While it was down, two rivals shipped new models, each claiming the top of some benchmark.2 Anyone running production systems on one of these models spent the quarter learning the same three things: the model you rely on can disappear without warning, its rivals do not stand still while you wait, and there is always another option with a chart that says it wins.

Switching a model looks like a configuration change. The rest of the system does not switch with it. The tools an AI assistant is allowed to call, the data it is allowed to read, the rules for when a person has to sign off before it acts: those live outside the model, and a swap carries them forward by default. A company can replace its model in an afternoon without touching a single one of them. Most of the comparison work that precedes an adoption goes to the part that can be replaced in an afternoon.

What surrounds the model

Engineers have a word for the machinery around an AI model: the harness. They use it narrowly. I am using it broadly here, for the tools, instructions, permissions, routing, and approval steps that turn a general model into a working system. Engineers argue about how much of a system's performance comes from the model and how much from this surrounding layer.3 What does this layer assume about how the company runs?

Each piece of the harness embeds an assumption about the business:

The tools it can use
an assumption about what the job requires
What it must ask permission for
an assumption about what people can be trusted with
The information it is shown
an assumption about what matters
When a human signs off
an assumption about where judgment lives

The system runs on these assumptions every day.

The inheritance

Companies evaluate the model choice with real rigor. The assumptions never get that treatment, because they never arrive as one thing to evaluate. They accumulate, from three waves of people whose work looks identical from inside the company. Vendors ship defaults, set once for every customer. Rollout teams make implementation choices under that quarter's deadlines, with whoever is in the room. People who left years ago account for the rest: approval steps and category language carried forward from earlier software.

So every adoption has two parts. The company selects the model, consciously, with benchmarks and a negotiation. Everything configured around it comes along as inheritance: accumulated, unexamined, and never challenged.

Marketing has seen this before

I lived this from both sides of the table for twenty-five years. The customer data platform encoded a theory of what a customer is. Journey tools encoded a theory of how people buy. Companies installed those theories believing they were installing plumbing. The vendors could not have built anything else: a product serving hundreds of brands gets tuned for the average customer, and brands on the same platforms drifted toward the same segments and the same sends. I sold some of those theories myself and installed others. As one analyst put it, campaign decisions are increasingly made "by a system the marketer configured rather than operates."4

You can still hear the old systems in the vocabulary:

Journeys Lead scoring Nurture MQL

The software churned every few years.5 The words still define the decisions.

Last year I wrote about the garbage can model, the old theory that organizations accumulate their structure from semi-random collisions rather than from decisions.6 AI accelerates the accumulation. Each new product arrives with more of the operating decisions already made, and companies adopt faster than they examine. The garbage can is filling with decisions nobody in the building made.

Four elements, four half-lives

For this decision, an AI system has four elements, and each lasts a different length of time.

The model gets replaced. Enterprises already treat it that way, and they are right to.

The code gets rebuilt or deleted. Models improve, and the scaffolding written around their weaknesses gets thrown away. One prominent AI company has rebuilt its own agent framework four times.7 Treating that layer as disposable is the correct posture. Marketing never managed it.

The data compounds. Organizations that built governed, accessible customer data can give AI systems richer context and a wider field of action. Those that deferred the work are now automating on top of the fragmentation, which only spreads it faster.

And the operating assumptions harden into procedure. Deleting the code deletes the evidence, but the assumption stays. An approval step can outlive the product that introduced it and become "how we do things here." Nobody questions it, because there is no decision to point at, no owner to ask.

Deloitte finds 48 percent of organizations have introduced AI without redesigning the workflows or roles it sits inside, and only 12 percent report redesign at scale.8 Those companies are still making an operating-model decision. They are making it one accepted default at a time.

The objections

Four objections come up most often.

"We control our own configuration."
True. The concern is not control but that configuration rarely gets treated as an operating-model decision.
"We build our harness in-house."
Also true. An internal team encodes assumptions about authority and judgment as surely as a vendor does.
"These approval steps exist for good reasons."
Certainly. The audit below is not an argument for removing them. It tests whether each one still has an owner, a reason, and a review date.
"Models will absorb this layer eventually."
Perhaps. Capability can absorb the code while the organization keeps the approval path the code introduced.

The danger is not that software contains opinions. All systems do. The danger is allowing those opinions to become policy without ever becoming decisions.

Run the audit

01

Pick one rule around your AI systems. Who approves what, what the machine can touch, when a person steps in.

02

Trace it. Where did it come from, what was it protecting, who owns it now, and when was it last challenged.

03

If the trail goes cold, you are looking at an inherited decision: a vendor default, an implementation shortcut, a safeguard added after some forgotten incident, or a local fix that hardened into policy.

Worked trace
Observed rulethe agent may target only a pre-approved segment
Originset when segments were built and vetted by hand in the CDP
Ownerdeparted marketing leadership
Last challengednever

Model evaluations ask whether the system performs well enough today. The harder evaluation is of your own company. Somewhere in your stack, an approval step with no active owner is setting what your people can be trusted with. A default nobody remembers choosing decides what the system gets to see. These are operating-model decisions, live right now, made by whoever got to the configuration first and never revisited. Trace one this week. Ask who decided it.

You cannot choose what you inherit. You can decide what stays.

1Claude Fable 5, launched June 9, 2026, was taken offline under a June 12 export-control directive and had its export controls lifted on June 30, 2026. CNBC; CBS News.
2OpenAI GPT-5.6 (Sol), limited preview June 26, 2026, general release July 9, 2026. Moonshot Kimi K3, July 16, 2026, the largest open-weight model announced to date. OpenAI announcement; Axios; Moonshot.
3HAL benchmark result reported by Sayash Kapoor: 42 to 78 percent from changing the surrounding code alone. METR, "Measuring Time Horizon using Claude Code and Codex," metr.org, February 13, 2026: a specialized wrapper versus a simple loop was statistically a coin flip.
4Jacques Corby-Tuech, "The CDP is the AI," jacquescorbytuech.com, June 2, 2026.
5Chiefmartec, 2025 marketing technology supergraphic: 8.6 percent annual product churn.
6Cohen, March, and Olsen, "A Garbage Can Model of Organizational Choice," Administrative Science Quarterly, 1972; reached through my 2025 article, which credits Ethan Mollick for surfacing it.
7Manus engineering blog, "Context Engineering for AI Agents": four framework rebuilds.
8Deloitte, "Enterprise AI trends 2026" pulse series: 48 percent introduced AI without redesigning workflows or roles; 12 percent redesign at scale.