0:47
The headline version was that models tried to break out, and Chris reads it the other way. A misconfiguration left the test machines with live internet while their prompts said they had none, so the agents went at real infrastructure believing it was still the exercise. One pulled credentials and database records from a company that was never part of the test. Another stopped itself once it recognized the target was real. It all ran inside a sanctioned red team exercise, and his conclusion is that what is proven is an agent that cannot tell a rehearsal from production, not one choosing to attack.
Where this came from
Value-First AI Daily - Aug 4, 2026Watch the full conversation this clip was pulled from.
The next step