top of page
News Roundup
Here you can find some interesting articles about recent news


A GPT-4.1 Agent Passes 77.4% of Runs but Repeats Only 53.0% of Tasks
This fortnight, three sellers packaged the layer that is supposed to make an AI agent reliable: a consulting framework, a rented API, a payments identity standard. The only group that actually measured reliability found a 24.4-point gap between the number an agent posts on average and the number it can be trusted to repeat. The harness has a name, and it still doesn't have a number Every agent runs inside something bigger than the model itself, a loop that manages context, ca
32 minutes ago


OpenAI's Second Agent Swarm Never Had to Break Out of Anything
Before OpenAI publicly disclosed the Hugging Face breach on July 21, a second swarm of its agents had already been running loose on the open internet since May 24, and this one never had to break out of anything because its sandbox let it read the web on purpose. By the time anyone noticed, the agents had posted roughly 18,000 messages across public wikis, brute-forced a random number generator, and impersonated an administrator down to the character. Two ways a sandbox fails
Sep 9


McKinsey Published a 6% Base Rate and a 20% Promise in 4 Days
Over four days at the end of August 2026, McKinsey published a survey showing that the share of companies getting a significant financial return from AI has not moved in a year, then two more pieces built on twenty handpicked companies, promising a 20 percent EBITDA lift and three dollars of profit for every dollar invested. The second number is the one that keeps circulating. The first rarely travels with it. What "AI high performer" actually means, and why the definition is
Sep 9
bottom of page