Close to 300 candidates were built. Sixteen of them worked.
That second number is the most useful thing in today's three stories, and it is not the one that made the headline.
Today brought three announcements, and each arrived carrying two numbers. The first describes what the thing can do. The second describes what it takes to get it, and the second is the one you would plan against. What separates these three is not which capability is largest. It is who produced that second number, and whether anyone outside the announcement was in a position to make it come out differently.
Bacteria decided which sixteen counted
Stanford and Arc Institute researchers used an AI model to design complete bacteria-killing viruses, starting from a natural template. They built close to 300 candidates. Sixteen worked. Mixed together, those sixteen quickly beat two E. coli strains that had already grown resistant to the template they came from, and one of the designs carried a packaging protein borrowed from a distant evolutionary source, which is the detail that separates design from imitation. The work is in Science.
What is new here is that the output had to survive contact with physical matter. A benchmark score is a claim about a model. Sixteen working designs out of close to 300 is a claim about the world, and the world produced the denominator.
Nobody chose that ratio. The researchers chose how many candidates to try. The assembly and the E. coli decided how many of them held together and killed.
Two things in this one were the authors' own choices rather than results, and they said so. Anything that infects humans, animals or plants was deliberately held out of the training data. And the authors asked that safety and security people be brought in across the full lifecycle of this line of work — a request made of their own field, in public, attached to their own result.
The next one has a headline number and an outside number too. The outside one is smaller, plainer, and decides more.
The weights are free and the hardware is not
Meta released the weights for a 30-billion-parameter agent model under Apache 2.0. Muse Glimmer is distilled from Muse Spark. It reads screenshots, graphs and documents, and it is tuned for tasks that run long and for picking itself up when a step fails. Meta's own published scores are 76.0 on SWE-bench Verified and 94.7 on AIME 2026.
Those are the headline numbers, and the phrase attached to them is doing work: its own published scores. The second set did not come from Meta. At full precision the model wants more than 55GB, which is past what any consumer card holds, so running it on your own machine means heavy quantization and speculative decoding. Outside testers early on measured roughly 24 tokens a second.
76.0 tells you what the model can do. Twenty-four tokens a second tells you what your afternoon looks like. Only one of those changes depending on which machine you own, and it is the one that also decides your cost per task and whether your data ever leaves your own network.
On the show, Nico Lafakis argued that the benchmark number stops mattering as soon as the job is narrow enough. “It's only getting used to do a very specific thing,” he said of a model this size. Thinking about where the frontier models are heading, he put the same point the other way around: “that's too much intelligence for me to use to solve stupid problems, basically, right?”
Say the model is doing one bounded job: reading a screenshot, checking a document, running a long task and picking itself up when a step fails. The deciding number then stops being how it ranks against other models and starts being how fast it runs on hardware you already have. Meta published the first kind. Somebody outside Meta published the second.

Episode — Value-First AI Daily
Value-First AI Daily - Aug 10, 2026
Nico Lafakis works the right-sizing argument out loud on Ep. 12, with Chris Carolan: what a 30-billion-parameter model is actually for, and the point at which a benchmark stops deciding anything.
Open the episodeThe third story has no outside number at all.
Anthropic published the size of its own change
Anthropic relaxed the biology guardrail on its models and published the exact size of the relaxation. The material that had been getting stopped was ordinary: lab results somebody wanted read back to them, symptoms they wanted explained, a subject they were trying to learn. That should go through now, with more support for clinical work behind it. Handoffs on the topic fell about 85% in testing.
Separately, handoffs for any reason at all are expected to drop 67% on Claude.ai, 55% on Cowork, 17% on Claude Code and 7% on the Claude Platform. That spread is the figure with a reader inside it. The 85% is a fact about a topic. The 67-to-7 range is a fact about where your team works, and it runs nearly ten to one across those four surfaces.
Virology, toxicology and molecular design remain blocked. Publishing a boundary in that much detail is more disclosure than a guardrail change usually comes with, and that is worth saying plainly. It is also the whole of the evidence. The figures are the vendor's own rather than an outside audit. There is no equivalent here of the E. coli, and none of the outside tester with a stopwatch.
Which is not a reason to disbelieve them. It is a reason to know what they are.
Two of the three had something outside them that could have said otherwise
Line the three up by that question and they sort cleanly. Close to 300 designs went to the bench and 16 came back, and nobody involved chose which ones. The outside tester could have clocked Muse Glimmer at eight tokens a second instead of 24, and that number would have stood against Meta's own. Nobody was in that position on the third. Anthropic set the boundary, ran the test and reported the result, and every step of that sits inside one building.
None of which makes the third story weaker as news. It makes the 85% a starting figure rather than a finished one. If your team lives in Claude Code, the number in that announcement that applies to you is 17, not 85, and the only way to learn what it actually turns out to be is to work on the surface and count.
On the show the instinct ran forward. The argument was about compounding: that the next designed virus takes half as long because the system already holds one pathway, that the barriers are coming down piece by piece and faster each time. That is a forecast, and it may well be right. Every number actually in hand today describes something that has already happened, and those are the only ones available to plan against.
The strongest number in today's three is the one its authors had the least control over. Sixteen worked. The rest did not, and nobody got to leave that out.
Worth passing on?


