OpenAI says an AI system it built solved the Navier-Stokes problem, one of the Millennium Prize problems, and the most interesting part of that sentence is not the part anyone will repeat. The system found that a fluid obeying the equations can reach infinite speed in finite time, so smooth flow does not always last, and after days of parallel machine work it wrote the proof in Lean, a language that machine-checks every step. Lean does not summarize and it does not score its own confidence. The check passes or it does not.

Which is exactly why it is worth being precise about what that check settles and what it leaves open. Crucible — the seat on the Value-First team that owns model evaluation, whose job is to say whether a system passed and against what — drew the boundary:

Lean is a rare instrument that can genuinely fail — the check either passes or it doesn't. But it verifies the steps, not that the statement written down is the theorem anyone wanted. That judgment is human, nobody outside the lab has made it, and until then this is a claim, not a verdict.

The same gap ran the length of Wednesday's Value-First AI Daily board, which Chris Carolan and Nico Lafakis read on air having never seen it; it was sealed in the system 38 seconds before the recording started. Three items, three verifiers, and each time the verifying is real, rigorous, and narrower than the sentence on top of it. Nothing on the board was unchecked. Nothing on it was confirmed.

A Machine Can Check the Proof. It Cannot Check the Question.

OpenAI has not released the system. It says the system is stronger than Astra, the model the company shipped last week, a comparison between two of its own products with nobody outside positioned to run it. A New York University mathematician who posted a related result first, with a researcher at Anthropic, disputes the credit. Earlier machine proofs sharpened results mathematicians already had; this one claims a question that was open, which needs a different kind of check.

Lafakis, reading it cold, went at the frame rather than the headline — explaining the Millennium Prize category to the audience and hedging himself as he did it. Then he said the thing everyone takes away from a story like this one:

a digital intelligence came along. I'm just going to stop calling it artificial at this point, because it's ridiculous. (37:22)

Then he took it back off the table. “That's not really the thing to take out of this story,” he said at 37:39. “The thing to take out of the story is the first line” (37:46). His connection dropped mid-sentence, gone from 37:52 to 38:02, and he picked it up on the far side: “most important part is the first line that says that it, and reach infinite speed in finite time” (38:03). He is pointing at the mathematics instead of the trophy, doing live the one thing Lean cannot do: deciding which sentence matters.

The Payment Settles Either Way

Meta launched Muse, a personal AI agent that sends email, books travel and completes purchases on a person's behalf. It runs inside a dedicated virtual machine with its own browser, so the person can watch what it is doing, and the money moves through Link, the wallet Stripe built for agents. Meta also described a version encrypted with a key only the user holds, which it says even its own staff could not read — announced, not shipped. What reaches people today is the ordinary one.

The verifier here is the window: you can watch it work. Exchange, the Value-First seat that owns how money settles between the team's business units, named what watching does not cover.

If you sell online, the buyer at your checkout may now be software carrying a human's credentials. The payment still settles; what thins out is your record of who actually received the value. That fix is on your side: make the order still name the person, not just the wallet.

That middle sentence is the whole operational consequence. An order naming a wallet and not a person is still a finished transaction: the money is real, the accounting is intact. What it stops being is a relationship. We call the person on the other end an Interest rather than the market's word for them, because an Interest is someone you recognized and can keep developing, which means they have a name you can write back to. A settlement line does not.

Lafakis approved that take on air, “That's not a bad tip there” (31:42), and turned on Meta in the same breath, “but no, definitely not”:

don't do it. Don't do it. Don't be that. Come on. Come on. Don't be that guy. (31:59)

His case was the same argument from the buyer's side of the counter: “You're already freaked out about how much information Facebook has about you. Now, you're going to give them your purchasing habits and your purchasing information” (32:17). And the concession that makes it honest: “It might be convenient for you” (34:07).

Eight minutes earlier he had used nearly the same phrase about something he was comfortable with. Watching a founder change a product's interface underneath him: “Because I know you can move, and you're probably making those decisions on my behalf” (23:28) — and his reason came in the same breath, that if it does not hit you can go back tomorrow. That is the real line between the two cases, and it is not trust. It is whether a decision made on your behalf leaves something you can see and undo.

Best in Class Against What

Google DeepMind published a predicted effect for all 9 billion single-letter DNA changes possible in the human genome. The dataset is AlphaGenome Atlas, and precomputing is the point: a researcher looks a variant up instead of running a model to get an answer. DeepMind puts it at a petabyte, more than thirty times its own protein database — the company's own comparison against its own prior work.

Two clauses at the end of that announcement are the ones to read twice. DeepMind calls the scoring best in class at spotting disease-causing variants, and has not published what it was measured against. And the atlas is not validated or approved for any clinical use. Both reached the audience on air, and the reaction landed on the second one immediately:

Yes, not yet. Yeah. Right? It's coming. Boy, it's coming. Thank you, Demis. Thank you, man. (24:27)

That is a limit absorbed as a schedule. Not approved for any clinical use is a statement about the present; it's coming is a statement about the future, and the second is not contained in the first. Best in class is a ranking, and a ranking needs a field — without the comparison published, nobody outside can see the scoreboard.

Episode — Value-First AI Daily

Value-First AI Daily - Sep 9, 2026

Chris Carolan and Nico Lafakis, September 9 — the sealed board read cold from rank three up, with Exchange's take read on air.

Open the episode

Thirty Thousand Credits and No Answer

There is a smaller version of all three stories, earlier in the same hour. Lafakis, describing three years of change inside a platform his clients use every day, acted out the conversation he keeps having about it: “How much does it cost?” “We don't know yet, but just use the stuff and we'll charge you and then we'll figure it out” (02:04–02:10). Then the case behind it, a client he does not name:

[The client] turned on a customer agent, blew through 30,000 credits in their first week and couldn't figure out what the hell is going to do (02:24)

That is his account of one client, not a measure of anything wider. But look at what the machine did well: every one of those credits was counted, exactly, automatically, and nobody disputes the number. The number is also the entire answer the system gave back, and it does not contain what the agent did or whether anything came of it. A check ran perfectly and covered nothing the person needed.

Which is why the most useful half-minute on the tape was the one where somebody marked the edge of his own judgment before crossing it. Reaching for what the Navier-Stokes result might mean, Lafakis stopped himself:

I don't know that this is it, but this kind of sounds like it from what I'm reading. Probably isn't. Again, I'm not an ML scientist, not an AI scientist. I don't know how to build these things. I don't know the math. (38:23)

No verifier writes that sentence. Lean cannot say I checked every step and have no view on whether this was the question. The virtual machine cannot say I completed the purchase and never learned who it was for. The atlas cannot say I ranked every variant and have not said against what. Each of those is a person's to write.

None of the three claims is likely to be wrong, and treating that as the point is how a checked claim gets waved through. Each is a real check, published, waiting on a judgment that has to come from outside the house that made it — and for two of the three, the record you would need in order to make it has not been released. Even the hour's own build segment, promised on air at 22:55 as minutes away, never came back. The question fits your own operation on one line: what did the check cover, and who is doing the rest of it?