Fourteen of Fifteen

Three hundred fifty-four proteins that had never existed were designed by a language model, manufactured, and tested in a lab. Fourteen of the fifteen targets they were aimed at got hit.

Anthropic published the results on August 20. Adaptyv Bio and Twist Bioscience did the wet work — synthesizing the designs and measuring whether they actually bound. The hit rates were 26.7% and 22.6% for the two models in 48-hour runs, rising to 35.1% when one model was pointed at a single target for 24 hours. Published campaigns typically land in the 10–15% range. On one target, RBX1, the model hit 40% against 3.7% for the human participants in a competition run on the same problem.

Anthropic is careful in the paper, and the caveat is correct: protein binders are not drugs. A molecule that sticks to the thing you want it to stick to is the first step of a process that kills most of its candidates later, for reasons that have nothing to do with binding affinity.

Set the drug question aside. The part worth arguing about is what kind of knowledge this is.


Nobody knows why it worked

There is no theory here. There was no moment where a model derived a principle of protein folding, stated it, and applied it. There was a system that had read essentially everything and then proposed sequences, and a fraction of the sequences bound, and the fraction was two to three times the human baseline.

If you ask why design 847 worked and design 848 didn’t, there is no answer available at the level of explanation. There’s an answer at the level of measurement — one bound, the other didn’t — and that’s the whole of it.

Science has been here before, and the philosophy of science has a well-worn argument about it. Pierre Duhem’s position, over a century ago, was that a physical theory isn’t an explanation of reality at all. It’s a system for economically summarizing and classifying experimental laws. Whether it corresponds to what’s actually out there is a question physics can’t settle and doesn’t need to. Bas van Fraassen’s constructive empiricism made the modern version of the case: the aim of science is empirical adequacy — getting the observable phenomena right — and acceptance of a theory doesn’t commit you to believing the unobservable machinery it posits.

The realists have always had a good counterargument, and it’s the one Hilary Putnam sharpened: it would be a miracle if a theory made novel, specific, correct predictions about things nobody had looked at yet, and the theory weren’t at least approximately true. Success needs an explanation, and truth is the least strange one available.

The protein result is a strange test case for that argument, because the no-miracles move requires there to be a theory to be approximately true of. Here there’s a weight matrix. Something in it is tracking something real about how sequences fold and dock — the 35% hit rate is not luck, and the RBX1 result is not luck. But the thing that’s true is distributed across billions of parameters in a form nobody has extracted, and might not be extractable in any form a person could hold in their head.

That’s not a familiar epistemic object. It’s not a theory, and it’s not a knack. It’s closer to what Michael Polanyi called tacit knowledge — we know more than we can tell — except Polanyi was describing a skilled human being, and the point of tacit knowledge was that it lived inside a person who could still be trained, corrected, and asked to demonstrate. This lives in a file. You can copy it. You cannot interview it.


The trade we’ve already been making

The honest version of this is that biology made this trade a long time ago and is only now being asked to admit it.

AlphaFold did not explain folding. It predicted structures, extremely well, and the field absorbed it as an instrument rather than as a theory, and the sky did not fall. Docking software, molecular dynamics, high-throughput screening — the whole apparatus of modern drug discovery is already a machine for generating candidates faster than anyone can say why the candidates are good. Medicinal chemists have been running on pattern recognition and empirical SAR tables for decades. The theory arrives afterward, if it arrives.

What’s different is the ratio. When the search process proposes a hundred candidates and a chemist understands why three of them are worth making, the understanding is still doing load-bearing work. When the process proposes 1,320 candidates and 354 of them bind and the selection criteria are inside the model, understanding has been moved out of the loop and into the postmortem.

That’s not a catastrophe. It’s a shift in what a scientist’s day contains, and it’s the same shift software went through when compilers got better at register allocation than the people writing assembly. The knowledge didn’t vanish. It stopped being the thing you personally had to have.

The cost is real, though, and it isn’t sentimental. Explanation is what lets you generalize to the case you haven’t seen. A field that gets very good at producing verified results it cannot explain is a field that gets exactly as far as its verification apparatus reaches, and no further. Every binder here had to be synthesized and measured. The wet lab is the bottleneck and also the epistemology: nothing counts until it’s been made and tested, because there’s no theory available to vouch for it in advance.

Which means the thing to watch isn’t the hit rate. It’s whether anyone can turn the hit rate into a principle.


Who gets to know

There’s a second story in the announcement, and it’s the one with policy in it.

Anthropic says life-science capability remains blocked in its most capable model, and that it’s preparing an access program to give scientists a way in. Read that plainly: a private company has produced a method that beats the field, demonstrated it publicly, and is metering who may use it, on its own judgment about consequences.

The reasoning is not mysterious. The same capability that designs a binder for a therapeutic target designs one for a target you’d rather nobody optimized against, and there is no clean technical line between the two — the model doesn’t know what the protein is for. Given that, gating is defensible. It may well be right.

But it makes epistemic access a governed resource, and that’s a different world from the one where publishing a method meant anyone with a lab could run it. Peer review’s whole premise was that a result you can’t reproduce isn’t yet knowledge. An access program is reproducibility with a waitlist.

I don’t have a better proposal. The alternative — publish the weights, let the field have it — has failure modes that are worse and permanent. This is a case where the cautious answer is probably correct and still costs something, and the cost should get named rather than waved through: the community that gets to check the work is now selected by the party that did the work.


Fourteen of fifteen is a real number and it was measured by people who didn’t produce it. That’s the part that makes it science rather than a demo.

What it isn’t yet is an explanation. We have three hundred fifty-four proteins that work and no account of why, and the field will spend the next several years deciding whether that gap is a temporary embarrassment or the permanent shape of the thing.


Sources: Autonomous de novo protein binder design with Claude, Anthropic ↗ · Benchmarking Claude’s protein designs in the wet lab, Adaptyv Bio ↗ · Anthropic says Claude designed working protein binders, The Next Web ↗