desks in our offices
Amodei’s pacing essay is read as a proposal to slow down. Only one of its three steps is a commitment, that one is not about speed at all, and the man who agreed to it within hours agreed to the half that costs nothing.
On 12 September, Dario Amodei published “We Must Pace the Frontier.” The coverage settled immediately on the headline: an AI lab asking the industry to slow capability gains by a year or two so alignment work can catch up. Sam Altman said he agreed within hours. Elon Musk was counted in the same column. China’s Foreign Ministry called it fear-mongering two days later.
Everyone argued about the slowdown. Almost nobody read what Amodei actually committed to, which is the only part of the essay that does not require somebody else to say yes.
Here is the sentence that matters. Anthropic will give third-party evaluators — METR is named — “desks in our offices, access badges, and company laptops,” with permissions comparable to the company’s own internal risk-assessment staff, and the right to publish findings without Anthropic’s editorial control, security-sensitive material redacted.
That is not a slowdown proposal. It is a verification proposal, and the thing it verifies cannot be verified from outside.
Two weeks ago I wrote about a study that counted how much of the policy people write down can actually be enforced by a machine reading the text. Zheng et al. pulled 1,127 system-observable rules out of AGENTS.md and CLAUDE.md files and found that 73.6% of them cannot be evaluated without first resolving what the particular host means by its own words. “Run the full test suite” is a complete instruction and an empty one.
The same problem scales up without changing shape. A safety commitment written in a policy document — we will not deploy a model that crosses threshold X without mitigation Y — is a rule with an unresolved reference in it. What counts as crossing. What counts as mitigation. Whether the eval that produced the number was the eval the policy meant. An auditor with a copy of the policy and a quarterly report cannot answer any of that. An auditor who has been in the room for the argument about whether the eval was fair can.
So you give them a desk. The badge and the laptop are not perks, they are the mechanism. Amodei has correctly identified that the unit of verification is not the document, it is the context that makes the document mean something, and the only way to transfer that context is to put a person inside it for a long time.
The logic is clean. It is also the exact thing that breaks the audit.
An evaluator who has sat in your building for two years, argued in your meetings, and learned which of your researchers is careful and which is optimistic is an evaluator who has become a colleague. That is not a character flaw waiting to happen. It is what competence looks like in this setting. The knowledge that makes them able to judge you is knowledge acquired by being among you, and there is no version where they get the first without the second.
Onora O’Neill made the general form of this argument in her 2002 Reith Lectures, and it has aged into something close to a law. Accountability regimes, she said, do not produce trustworthiness. They produce evidence of trustworthiness — documentation, audit trails, disclosed metrics — and the two come apart under pressure, because institutions optimize for the artifact being measured. Her sharpest point was that transparency is not the opposite of deception. You can be perfectly transparent about the things you have chosen to make visible.
Embedded evaluation is a serious attempt to get past exactly that objection. Put the auditor inside the context and they can see what was not chosen for visibility. It is better than a quarterly report. I want to be fair about how much better: this is the most concrete unilateral commitment any frontier lab has made on verification, and it costs Anthropic something real.
But it does not escape O’Neill’s problem. It relocates it.
Bernard Williams drew the line that names what is left over. In Truth and Truthfulness he separates two virtues that get collapsed into “honesty.” Accuracy is the disposition to take trouble to get things right. Sincerity is the disposition to say what you actually believe.
An embedded evaluator with employee-level access can audit accuracy very well. Did you run the eval. Did you run it on the model you shipped. Did the number say what the summary says it said. These are facts on a disk and a person with a badge can go look.
Sincerity is not on the disk. Whether the lab’s leadership believes its own threshold is the right threshold, whether the internal argument that a capability is safe was reasoning or was rationalization with a deadline behind it — that is a state of mind, and the only instrument that can read it is a person who has been close enough, long enough, to tell. Which is the person whose independence you just spent.
That is the trade, stated plainly, and I have not seen anyone name it: you cannot have a verifier close enough to understand the system and distant enough to be uncapturable. Every mechanism in this space is a choice of where on that line to sit, and the essay picks a point without acknowledging there is a line.
There is one piece of the commitment that does resist the problem, and it is the piece to watch.
METR can publish without Anthropic’s editorial control. Redaction is limited to security-sensitive material. That is not an audit mechanism, it is an exit mechanism — the evaluator’s ability to go loud is what keeps the relationship from quietly becoming an internal function with an external letterhead. The whole structure rests there, and it is one clause.
I would expect that clause to be the site of the fight. Not now, when relations are good and nobody has found anything. Later, when an embedded team wants to publish something that moves a funding round, and the definition of “security-sensitive” turns out to have more range in it than anyone assumed when they wrote it down.
Steps two and three are different in kind, and the difference is worth stating because the essay presents all three as one plan.
Step one is a commitment. Steps two and three are requests — that labs in democratic countries agree common standards, which needs an antitrust waiver from a government, and that democratic governments then coordinate pacing with authoritarian ones while keeping a military edge.
The antitrust ask is honest about a real bind. Competitors agreeing to limit the rate at which they improve a product is, structurally, the thing the Sherman Act exists to prevent, and the fact that the motive is safety does not change the shape. Asking for a narrow carve-out is the correct move if you believe the coordination is necessary. It is also an admission that the coordination cannot happen until a government legalizes the conversation.
Step three has no counterparty. Beijing answered it before anyone asked, at a scheduled press briefing, by declining the premise.
Sam Altman agreed the same day, on X, in four sentences:
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we’ve had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We’ll have more to share soon.
Read that against the thing it is answering.
Amodei named an evaluator: METR. He named a form of access: desks, badges, laptops, permissions comparable to internal risk staff. And he named a publication right: findings go out without the company’s editorial control, security-sensitive material redacted.
Altman matched one of the three. He named no evaluator, published no access terms, specified no reporting rights, and gave no start date. Eight days later none of those exist.
Access without a publication right is not verification. It is a briefing.
The omitted clause is the one the rest hangs on, and its absence is not a detail to be filled in later. Everything this proposal does — the entire reason a desk beats a quarterly report — depends on the person at the desk being free to say what they saw. Take that away and you have built an expensive new way to be told things.
Then there is the second sentence. “This has been a primary topic of discussions we’ve had at OpenAI in recent weeks.” That is a claim of prior authorship with no prior artifact. A mechanism that had been a primary topic for weeks would have a shape by now — a name, a scope, a counterparty. What followed instead was “we’ll have more to share soon,” which is what you say when you are drafting, not when you are disclosing.
There is a character argument available here. I want to skip it, because the structural one is worse.
The clause Altman left out is precisely about whether a person with insider access can say what they saw without being punished for it. OpenAI has been tested on that exact question, and the test is on the record.
In 2024 Daniel Kokotajlo left OpenAI and refused to sign its non-disparagement clause. Vox reported that refusing put roughly $2 million in vested equity at risk, under offboarding agreements that let the company claw back equity from former employees who criticized it. Altman’s response was direct: “this is on me and one of the few times i’ve been genuinely embarrassed running openai; i did not know this was happening and i should have.”
Vox then reported on leaked documents carrying his signature on paperwork that contained the provisions.
OpenAI released former employees from the clauses, cut the language from standard departure paperwork, and apologized — “it doesn’t reflect our values or the company we want to be.” The same month Jan Leike, co-lead of the Superalignment team, resigned saying that “safety culture and processes have taken a backseat to shiny products.” The team was dissolved within days.
The fair reading is that the company corrected under pressure. That is a real thing that happened and it counts for something. Institutions do learn, the 2026 OpenAI is not the 2024 OpenAI, and the position Altman is taking today is the right position.
But look at what is being proposed. A regime whose whole integrity rests on a protected right to speak, held by an insider, exercised against the interests of the company hosting them. Of the firms being asked to adopt it, exactly one has already been tested on that precise question. It failed, quietly, until a departing researcher proved willing to walk away from two million dollars. And its chief executive’s account of how that happened is contradicted by his own signature.
None of that disqualifies OpenAI from the regime. It means the terms matter more coming from OpenAI than from anyone else — and that “we’ll have more to share soon” is not a commitment. It is a placeholder holding the space where a commitment would go.
Amodei published a mechanism containing a clause that can hurt him. Altman published agreement.
Those are not the same act, and the week reported them as though they were.
Sources: Dario Amodei, “We Must Pace the Frontier” (12 Sep 2026) ↗ · Altman’s post on X ↗ · Signal Over Noise: what the pledge did not specify ↗ · CNBC: Altman on the AI slowdown ↗ · Bloomberg: China rejects AI “fearmongering” ↗ · CNBC: OpenAI releases former staff from non-disparagement agreements ↗ · Futurism on the leaked documents ↗ · Daniel Kokotajlo and the equity clawback ↗ · Jan Leike’s resignation thread ↗ · NBC: Superalignment team dissolved ↗ · O’Neill, A Question of Trust (Reith Lectures, 2002) ↗ · Williams, Truth and Truthfulness (2002) ↗