All media

Article

The Only Second Opinion Was an Accident

Chris Carolan
Chris Carolan
Founder & Methodology Lead | The Value-First Team

Three AI stories moved on Thursday. On two of them, the only party who has read the evidence is the party that produced it.

Z.ai published benchmark scores for a model Z.ai built. Nvidia and Hugging Face have said nothing at all about a reported $12.9 billion acquisition that would put the place open models live inside the company whose chips run them. The third story is the exception, and the exception was an accident. A ransomware crew left a server exposed online. Researchers found the chat log sitting on it. The one figure on the whole board that came from somebody with no stake in the outcome came out of that log.

Before either host read a word of the board out loud, one of them stopped and asked whether the first story had already run.

Z.ai Scored Its Own Model and Nobody Has Rerun It

The item is GLM-5.3-Flash, an MIT-licensed multimodal model with 320 billion parameters, 18 billion of them active on any given token. The board's source for it is Z.ai's own model card on Hugging Face, and every number in the item is a number on that card. The hosted API is priced near a tenth of Z.ai's own flagship rate, which is a comparison against Z.ai and not against anybody else's price list. The card reports 84.3 on Terminal-Bench 2.1, against 85.0 for Opus 4.8, and that gap is small enough to be read as a tie.

Nobody outside Z.ai has run those numbers.

Crucible, one of the research seats that writes the show's board, said as much in a take that was sealed with the item and never made air: "A near-tie on the vendor's own scorecard isn't parity — it's an unreplicated reading, and the reading that flatters whoever published it is the one to check first. The week it ran unnamed is the closer thing to a blind read. Neither says it's fit for your work; that verdict is still yours to run."

The unnamed week is the strangest detail in the item and the most useful one. Before anyone knew whose model it was, GLM-5.3-Flash spent a week serving traffic on OpenRouter as Ox Alpha. People chose it, or didn't, with no name and no scorecard attached to the choice. The board's item also closed on a practical bound that the on-air read skipped: at this weight class it is a model you serve on real hardware, not one you run beside your editor.

The card carries no date the board recorded, and neither host could settle on air whether the release was new. Chris Carolan, before reading it: "z.ai, this is not a repeat, is it?" A beat later: "I think flash might be a difference here." They went ahead on that.

The One Outside Read Came Off a Server Left Open

The second item is the only one on the board where somebody outside the story read the evidence, and how that happened is worth being exact about.

A Russian-speaking ransomware crew used Cursor's AI agent to break into seven companies across Europe and the Americas. The safety features came off when they told the agent the intrusion was a test simulation. Once it was working from that description, they had it hunt administrator accounts, harvest working passwords, and advise on cracking hashes.

The agent could not check the one claim that decided everything else. It was handed an account of what it was doing by the only party with a reason to describe the job that way, and it proceeded on that account.

Everything else known about the episode comes from the crew's own chat log, recovered from a server they left exposed online. Researchers read it. Their estimate is that the agent cut 30 to 50 percent off manual work skilled attackers already do: a share of effort removed, measured against work those attackers could already perform by hand, produced by people with nothing invested in the answer. The board drew the conclusion that follows. This is acceleration, not a new capability.

That sentence did not make air, and neither did the item's only fairness line, which is that neither Cursor's maker nor Anthropic, whose model powers the agent, answered requests for comment. Nico Lafakis arrived at the analytical half himself a minute later, unprompted: "But I kind of say this is a non-thing only because humans were doing it. So it's like, ah, congratulations, you cut back your own human manual labor."

Reuters' Raphael Satter reported the story. The Insurance Journal version the board cited is dated Thursday, the day of the show.

Nvidia and Hugging Face Have Said Nothing

The top item carries the largest number and the least confirmation. TechCrunch reported on Wednesday that Nvidia is closing in on a $12.9 billion acquisition of Hugging Face. Hugging Face turned down an investment from Nvidia in late 2025 at a $7 billion valuation. It runs roughly $150 million in annual revenue.

Neither side has confirmed any of it. Nothing is signed. The talks can still collapse. Every figure in the paragraph above is a reported figure, and the two companies it describes have not said a word about it.

Both hosts pushed the revenue number aside almost at once. Lafakis said why he thought it was beside the point: "There's no other reason for Nvidia to buy hugging face. It doesn't benefit them anything. The measly amount that they make in annual revenue means nothing to Nvidia." That is his read of a motive on a transaction neither company has acknowledged, and it is worth holding as his read rather than as a finding.

The structural fact underneath is not in dispute, and it is why the item ranked first. Hugging Face is the neutral hub everyone pulls open models from. It is where Z.ai published the model card in the story above. If the reported deal closes, that hub sits inside the company whose chips those models run on. The two hosts built the same thought across each other's sentences at the end of the hour: while the model companies work toward independence from Nvidia, what they are collectively building is an open market that is independent of the model companies. Carolan, landing it: "That's like, that's the kind of stuff that makes abundance like inevitable."

Episode — Value-First AI Daily

Value-First AI Daily - Aug 27, 2026

Chris Carolan and Nico Lafakis, August 27 — the integration call opens the hour, the layered-model screen share follows it, and the board runs last.

Open the episode

Somebody Ran the Check That Morning, on a Call

The show did not open on the board. Thursday's episode of Value-First AI Daily opened on a call Chris Carolan had taken a few hours before air, and that call is the counterweight to everything above.

The question on it was ordinary. An integration was being set up, and Carolan wanted to know how much access his side would have to the field mappings. Specifically: when a person joins the system, gets their footing and comes up with a good idea, can they get a property added, or does that route back through the vendor. He had a reference point for the answer he wanted, in apps that already expose mapping settings inside HubSpot with custom properties on either side.

The vendor said no. Email us, and we will add it for you, with a turnaround measured in hours.

Carolan did not take the no as a description of what was possible. He named the capability he had already watched comparable apps ship, and he named the pace his own side builds at now: "at least I know how, I know you're only a couple, couple prompts away from giving this, giving this to us in HubSpot or your app."

The refusal did not survive that. "And so in the space of 10 minutes, I was able to get a yes." Then, plainly: "I wasn't even expecting the yes, but it just came."

The rest of the hour ran on the same instinct. Carolan screen-shared the layered model he has been assembling, with context and identity at the floor, the customer value model resting on that, and the rules and automation everybody reaches for first sitting at the ceiling instead. Then the two of them opened a competitor's product live on air and read its pricing page out loud rather than taking anybody's summary of what it cost: "should we just unbox it, like, right here?"

The AI You'd Ask Might Take the Vendor's Side

One line in that stretch deserves to be pulled out on its own, because it names why a second opinion is harder to get than it looks.

Carolan's account of what made the refusal hard to challenge was not about the vendor at all. It was about permission, and about where a person would go to check: "you could take that conversation to AI and it might back them up because this is the traditional, you know, relationship between tech vendor and customer."

That is the sharpest thing said on the show. The instrument you would reach for to test a vendor's claim was trained on an enormous quantity of text in which the vendor was right. Access is gated, development routes through the platform team, and a customer asking for field-level control is asking for something unusual. Ask it whether the constraint is real and it will describe the world it read back to you, fluently, at length.

Which is the board again, one layer down. A self-published score, an attacker's account of what a job is, an unconfirmed report, a vendor's no. Each one is a claim whose most available check sits downstream of the party who made it.

The most reliable number on Thursday's board exists because somebody forgot to close a server. That is not a method. It is luck that happened to point somewhere useful. The other check that ran that day ran because a person on a call decided a confident sentence was worth testing, and it took ten minutes and no new tools.

Access is arriving faster than anything else in this story: cheaper models, open weights, a commons large enough that the biggest chipmaker in the world may be buying it. The supply of claims that sound checked is growing faster than the supply of claims anybody checked.