top of page
News Roundup
Here you can find some interesting articles about recent news


A GPT-4.1 Agent Passes 77.4% of Runs but Repeats Only 53.0% of Tasks
This fortnight, three sellers packaged the layer that is supposed to make an AI agent reliable: a consulting framework, a rented API, a payments identity standard. The only group that actually measured reliability found a 24.4-point gap between the number an agent posts on average and the number it can be trusted to repeat. The harness has a name, and it still doesn't have a number Every agent runs inside something bigger than the model itself, a loop that manages context, ca
33 minutes ago


McKinsey Published a 6% Base Rate and a 20% Promise in 4 Days
Over four days at the end of August 2026, McKinsey published a survey showing that the share of companies getting a significant financial return from AI has not moved in a year, then two more pieces built on twenty handpicked companies, promising a 20 percent EBITDA lift and three dollars of profit for every dollar invested. The second number is the one that keeps circulating. The first rarely travels with it. What "AI high performer" actually means, and why the definition is
Sep 9


Why AI Agents Pass Evals and Fail in Production
Enterprises are not short on AI ambition. They are short on proof that the AI they have deployed actually works once it leaves the sandbox. Four independent surveys published this month, alongside a major consulting report and a candid admission from the industry's own leading model maker, converge on the same uncomfortable finding: adoption has outrun the ability to measure, secure, and trust what has been adopted. What's actually at stake: software that acts, not software t
Jul 20
bottom of page