Media

Value-First AI Daily - Aug 11, 2026

August 11, 2026

Ep. 13, Tuesday, both hosts, about 50 minutes on air. The show opened about a minute before the sealed news board unlocked, so the countdown was on screen for the first time and Chris read the board cold, saying so out loud: "it really is sealed unless I ask AI to unseal it, there's no way for me to look at these." The Top 3 then ran across the first twenty minutes instead of at the close - Anthropic putting an invisible watermark on every Claude output worldwide, OpenAI's partner-only GPT-5.6 Cyber, and an unreleased Claude model pushing a 90-year-old bound on the Riemann zeta function from 41.6% to 67.2%. Each item opened a longer conversation rather than a reaction: whether anyone will keep asking where a song or a chart came from, whether the human approval bar has already reversed, and why plain language beats technically correct prompting. The last five minutes are Flywheel multiplayer, played live, both scores frozen.

Moments from this episode

Key takeaways

The seal is server-side and the show proved it on camera for the first time. Going live a minute before the board unlocked left the countdown visible, and Chris said what it means: "Get to show this countdown today just so everybody can see it really is sealed unless I ask AI to unseal it, there's no way for me to look at these."

The Riemann result is unconditional - it does not assume the hypothesis is true - and the orchestration behind it is the part to study if you run agent work: two sessions, about 60 coordinated subagents, 31 million output tokens and 2,400 shell commands, all figures Anthropic published about its own run.

Anthropic set the limit itself rather than leaving it to be inferred: "We don't expect that the techniques Claude used will lead to proving the Riemann hypothesis." Two of its own mathematicians and two outside experts examined the work, and a machine-checked Lean formalization passed the standard validation tool.

Watermarking every Claude output has a ceiling that matters more than the feature: a detection proves handling, not authorship, because people also edit, translate and summarize with AI. No public detector exists today, verification tools are promised with no date, and heavy editing, translation, format conversion or a screenshot strips the mark.

Nico's read is that the human approval bar has already reversed - he has heard enterprise people ask whether a document was run past the AI before it went out, and puts it at somewhere between 50-50 and 70-30 that way now, flipping to 70-30 or 80-20 the other way by December. Those are his figures from his own reading, not an audit.

Capability now travels by vendor relationship. OpenAI's GPT-5.6 Cyber reaches only trusted partners - Accenture, IBM, CrowdStrike and Cloudflare are named - and only the Red side of the Daybreak defender program, covering security testing and vulnerability research, carries the new model.

The most useful finding of the hour was about prompting, and both hosts arrived at it independently. Nico dropped technically correct language for "Hey, make this work" after deciding models were "way too laser-focused on every little thing that I was saying"; Chris's fix when a model stops cooperating is the same - "Dude, I don't know what's going on, but it was good yesterday. It's not good today. Help me out here."

Nico named the uncomfortable half of that finding rather than skipping it: plainer prompting works, and it works "because it disproves the need for humans to know any of the" intricacies. He said he felt good and bad about it at the same time.

Solving as a swarm and researching as a swarm are different things. Running a known solution in parallel "looks like an automated factory"; researching in swarm fashion means taking every avenue at once and building the best answer from what comes back - which is how humans cracked hard problems too.

The business objection to AI-made work got named and answered on air. Nico's practical read: "No one cares, just don't get us sued." Chris's: "Focus on what's possible and the value instead of liability and jealousy" - against covering yourself, and against resenting that someone else can now spring up a new game quickly when you spent ten years learning how.

Flywheel multiplayer shipped mid-show and was played live rather than described - invite link, new sound effects, a top-down view chosen partly to keep processing light on mobile. Both score totals then froze mid-game, and that aired uncut.

Keep learning

You just watched the work. Now build it.

The next step is doing it yourself — live, with people on the same path.

Get the next conversation in your inbox

One signal-dense note when a new episode lands. No noise.

Join the list