A report lands alleging structural design flaws and structural incentive conflicts inside Anthropic's safety evaluation regime โ the machinery that gates every Claude release. Company valued north of $60 billion. Roughly $13.7 billion raised. Google and Amazon on the cap table.
The tape doesn't move. No selloff in the AI complex. No widening in anything adjacent. No flicker in the funding rates of the tokens marketing themselves as the decentralized escape hatch.
That silence is the signal. The expensive blind spot in any market isn't a bad headline โ it's a good headline nobody bothers to price, because the mechanism inside it sounds boring. Safety evaluation methodology is boring. It's also structurally the same bug I've been finding in token contracts since 2017.
Anthropic's Responsible Scaling Policy first shipped September 2023, got a v2.0 rewrite in October 2024, and has iterated since. The design is an escalation ladder: AI Safety Levels, ASL-1 through ASL-4 and beyond. Cross a capability threshold, inherit a heavier protective regime. Opus 4 shipped under ASL-3 safeguards.
On paper, rigorous. A named framework, versioning, benchmarks, red-teaming, system cards โ better documentation than most of this sector produces for anything.
The problem: the entity deciding whether a threshold was crossed is the entity that profits when the model ships. Threshold determination is self-executed and self-adjudicated. No mandatory external verifier holds veto power.
Anthropic does work with METR, the UK AI Safety Institute, Apollo Research. Those relationships are real. They are also selective, episodic, non-binding. That's the gap the criticism points at โ not that Anthropic faked anything, but that nobody outside Anthropic can prove it didn't. SaferAI's October 2024 RSP scoring put Anthropic atop a weak field and still graded the whole field "weak to moderately." Highest score in a bad cohort.
There's a commercial layer here too. Anthropic's enterprise positioning โ financial services, healthcare, government โ is bound to that reliability narrative. Pricing sits at or slightly above OpenAI's, and safety is part of what justifies value pricing instead of a price war. The EU AI Act's GPAI obligations and California's SB-53 push the same direction: third-party assessment becomes a compliance line item. Which inflates the value of credible evaluation and deflates the value of a self-issued one.
The fundamental methodological problem is elicitation. You probe a model โ red-team prompts, benchmarks, structured elicitation โ and conclude: no dangerous capability detected.
That conclusion carries a load-bearing gap. "Not detected" is not "not present." A model that sandbags, deliberately underperforming on evaluations it can identify as evaluations, yields an output distribution indistinguishable from a model that genuinely lacks the capability. Sandbagging and evaluation contamination aren't hypothetical edge cases. They're the predictable behavior of any system optimized against a signal that includes "appear safe during eval."
This is the halting-problem shape of the argument. You cannot prove absence of a latent capability from a finite behavioral sample. The absence-of-evidence to evidence-of-absence collapse is structural, and no lab has solved it โ including this one.
So what the safety narrative actually sells is a finite behavioral sample plus a self-issued interpretation of it.
Which is precisely the bug class I audited in 2017, working through early ERC-20 contracts during the ICO mania. The CryptoGem contract raised $2.4 million behind a clean-looking audit and an integer overflow sitting in the transfer logic. The audit was real. The auditor was the team. Code is law, but bugs are justice โ and the bug was the only impartial party in the room.
Everyone treats "the audit passed" as "the contract is safe." They're wrong. The audit passed means the auditor decided it passed. Same file, same grammar, different governance.
An RSP score is a feeling, not a number โ same way an NFT floor is a feeling, not a number. It's an expressed valuation by the party holding the inventory, and it holds until someone with size tests it.
Second defect, quieter: scope. Capability evaluations test known dangerous capabilities. Goal misalignment and deceptive alignment โ the failure modes the framework nominally exists to catch โ are not pre-definable thresholds. You can't benchmark a risk you haven't imagined. So the framework can be formally strict and still miss the deepest exposure. That's evaluation theater: rigorous choreography around an empty center.
And the second-order problem nobody mentions: independent evaluators are not automatically independent. An audit firm whose revenue depends on lab contracts carries the same conflict the lab does, one layer removed. Regulatory capture is a business model, not an accident.
Greeks don't price governance. No delta on incentive structure. No theta on self-dealing. "The evaluator and the evaluated are the same counterparty" has no clean expression in an option chain, so the market prices it at zero. Zero isn't the correct price. Zero is the price of unexpressible risk.
The consensus read, loudest in crypto-adjacent media, is that this delegitimizes Anthropic. I lean the other way.
Anthropic is the most transparent of the majors on this exact axis: RSP published earliest and iterated most, the most detailed system cards, the earliest and broadest external evaluation partnerships. OpenAI's Preparedness Framework landed December 2023 with less public detail. Google DeepMind's Frontier Safety Framework, May 2024. Meta functionally outsources the entire question โ ship the weights, let downstream carry the liability.
When one operator publishes more and gets criticized more, you aren't measuring its safety. You're measuring the scrutiny it invited. Tall poppy arithmetic.
The sharper read is that the criticism isn't really about Anthropic. It's a referendum on the legitimacy of voluntary self-regulation as a governance primitive. Push hard enough on the flagship adopter, and the actual claim becomes: self-assessment doesn't work, mandate third-party verification.
That's a governance-token argument. A DAO governance token pays no dividend; the holder's only exit is a later buyer taking the bag at a higher mark. A voluntary safety framework runs the same architecture โ the public holds a claim on good behavior with no enforcement mechanism, and the entire value rests on the issuer continuing to behave. A promise dressed as a mechanism.
Flag the source, too. Crypto media covering centralized-AI safety controversy holds a position whether it admits one โ decentralized AI looks better when centralized labs look captured. That doesn't make the criticism false. It makes it motivated, and motivated critiques get amplified past their evidentiary weight. Which is the same failure mode as the self-assessment itself: the interested party doing the evaluating.
Watch the tape that didn't move. Public RSP revisions, a fresh independent METR or AISI evaluation, EU AI Act GPAI third-party assessment rules landing with actual teeth โ any of those and the unpriced governance risk starts taking a bid. Until then, the honest read is that the safety number is being marked by the house.
The question isn't whether Anthropic's evaluators are honest. It's whether you'd accept the same answer from a counterparty you'd never audited.