All media

Article

All Three of Today's Stories Told You Where They Stop

Chris Carolan
Chris Carolan
Founder & Methodology Lead | The Value-First Team

The first story Chris Carolan read out loud on Monday's show was four days old.

On screen was OpenAI's finance chief telling staff the company will be public in 2027, the third item on the previous Thursday's board, already run in this space. Nico Lafakis took it anyway and made it current, working through why a leadership team watching Anthropic's revenue would want a date in front of its own staff. Then Carolan found the seam.

Another example of how great AI is if humans aren't following instructions. I had the browser tab up before... Baldwin had finished writing the articles so quickly.

Monday's board was sealed after the show had already gone live. The tab open on screen predated it, gave no sign of that, and a refresh moved him onto the real three.

The page that could not tell him it was out of date was the only thing in the hour missing that particular honesty. Every one of Monday's three stories names its own limit, in the source, in the same paragraph as the finding. One says who did the scoring. One says when the price expires. One says the output is a prediction rather than a diagnosis.

The Score Reached 100% and the Team That Built the Harness Did the Scoring

Nvidia ran Claude Opus 5 inside its own agent harness and took it from 30% to 100% on ARC-AGI-3, per TechCrunch. A harness is everything built around a model that is not the model: which tools it can call, what it holds onto between steps, and a supervisor process that catches it when it wanders off the task. Nothing about the model changed. The construction around it did.

The bound travels with the number. One benchmark suite of interactive reasoning puzzles. Scored by Nvidia. The extra compute and tokens the harness burned getting there were not published.

Crucible's take, carried on the board as signed opinion and kept out of the report:

Not yet, for me. A hundred percent is the test telling you it ran out of room, not the agent telling you it is ready, and the team that built the harness did the scoring. I would trust it when someone who did not build it scores it, on a set this thing can still fail.

Carolan restated that in his own words on air: "They just gave themselves an A on their own test." Lafakis pushed back inside the same breath, and the correction is the most useful sentence of the hour.

The arc AGI test is a legitimate third-party test, so it's not like it's an individual test.

Both of those are true, and holding them together at once is the whole skill. The yardstick came from outside. The measurement did not. An outside benchmark does not turn an in-house run into an independent result, and an in-house run does not make the benchmark worthless. What is actually missing is narrow and nameable: nobody outside Nvidia has reproduced it.

Lafakis then made the observation that explains the jump better than the score does.

A single model didn't do it, two models did, one model did it, and then another model came behind it.

That is the finding underneath the headline figure, and Carolan's read landed in the same place: "It's the way we apply these tools via harnesses." Two models finishing work that one of them did not finish is multiplication, arriving as a benchmark result rather than as a claim. Whether the same construction multiplies the person operating it is the question an operator actually holds, and this benchmark does not measure that or pretend to.

Carolan's caution sits right next to it. "Bench maxing has proven to be a thing, and it's like, good job. It got really super good at that 2D game puzzle."

The Price Is Real and It Has an End Date on It

OpenAI cut GPT-5.6 Sol output pricing to $20 per million tokens, down from $30. Input came down as well, $5 to $4 per million, along with cached input. The rate reaches pay-as-you-go API calls, Codex credits, and eligible ChatGPT Work plans, and it runs through November 21. OpenAI's stated reason is efficiency gains it is passing on. Where the price lands on November 22 has not been said.

Pax's take, also signed opinion rather than reporting:

A promotional rate is a discount on what you already spend, not a new cost basis. Use the window to run the work you've been putting off on cost. But keep planning at the old number: if something only works at the lower rate, that isn't savings, it's a commitment with an end date.

Lafakis was less troubled, and his reason had nothing to do with the price.

Who is having token cost issues? Money actual money spending issues? I'm not sure who you are or where you are or what you're building or what you're doing, but it just seems horribly inefficient.

His argument is that a team surprised by its own bill was not watching it, and that the billing view had been sitting there the whole time. That is a fair point about attention, and it answers a different question than Pax is answering. The two do not collide. A visible bill tells you what you spent. It cannot tell you what the same work costs on November 22.

Roughly 60% of Human Genes, and a Prediction Somebody Has to Test

A UC San Diego lab trained a machine learning model to read the initiator, the short sequence marking where the copying of a gene's instructions begins, and to locate it in roughly 60% of human genes. There has been no dependable way to tell which genes carry one. Being able to find them opens two things: scanning that switch for mutations tied to diseases like cancer, and building synthetic gene-control sequences on purpose.

The scale is what keeps the claim modest. One element, inside a control code running to roughly six billion bases, and what the model hands back is a prediction rather than a diagnosis.

This item is on the board because Carolan asked for it four days earlier, on air, in almost these words: "Like, there are medical breakthroughs happening. And I want a spot for this. I want to make sure that hit." He laid out three tracks that day — inner-loop infrastructure, model drops, and breakthrough stories — with one story pulled from each, so the third kind would stop losing to the sheer volume of the first two. Monday's is the first of the twenty boards on record to carry those track labels.

The slot earned itself on its first outing. Lafakis went to Gattaca, and from the film to genetic stratification as something close rather than speculative. "It's a huge separation for humanity. It's not like a little bitty thing." He flagged his own reach in the same breath: "Now, I'm jumping to conclusions, right, with this kind of stuff."

Carolan's answer went to the only variable he thought was still in play. "It's a compute problem at this point. No matter what it is, you just have to decide."

Episode — Value-First AI Daily

Value-First AI Daily - Aug 24, 2026

Chris Carolan and Nico Lafakis read this board live on the August 24 episode of Value-First AI Daily, starting, by accident, with a story from four days earlier. Their disagreement over what a perfect benchmark score actually proves runs through the back half of the hour.

Open the episode

What the Page Had No Way to Say

Each of Monday's three sources was honest about its own edge. TechCrunch's item carried who did the scoring. OpenAI's carried the expiry date. UC San Diego's carried the distance between a prediction and a diagnosis. None of it was buried, and none of it needed a skeptic to dig it out.

The one surface in the room that offered a result with no edge on it was the page showing a board from the previous week with complete confidence. It had no way to say today's is not here yet. It served what it had, which is what pages do.

That is the thing worth carrying out of the hour, because it is the same shape as reading a 100% and stopping there. A reading that looks finished is not the same as a reading that has been checked, and the difference is almost never printed on the surface you are reading it from. Monday's took a refresh to find. The other three printed it themselves, which is the rarer thing.