top of page
News Roundup
Here you can find some interesting articles about recent news


A GPT-4.1 Agent Passes 77.4% of Runs but Repeats Only 53.0% of Tasks
This fortnight, three sellers packaged the layer that is supposed to make an AI agent reliable: a consulting framework, a rented API, a payments identity standard. The only group that actually measured reliability found a 24.4-point gap between the number an agent posts on average and the number it can be trusted to repeat. The harness has a name, and it still doesn't have a number Every agent runs inside something bigger than the model itself, a loop that manages context, ca
32 minutes ago


McKinsey Published a 6% Base Rate and a 20% Promise in 4 Days
Over four days at the end of August 2026, McKinsey published a survey showing that the share of companies getting a significant financial return from AI has not moved in a year, then two more pieces built on twenty handpicked companies, promising a 20 percent EBITDA lift and three dollars of profit for every dollar invested. The second number is the one that keeps circulating. The first rarely travels with it. What "AI high performer" actually means, and why the definition is
Sep 9


The Harness Is Becoming the Product
This week, the agent harness stopped being plumbing and became a line item: DeepSeek open sourced one, Writer sells one, and NVIDIA routes through one. Four independent benchmarks published in the same window show why: swap the harness under a fixed model and the score moves 20 to 40 points, with almost no correlation between how models rank under one harness versus another. The Score Was Never Just the Model OpenAI's own developer guide for GPT-5.6 makes this argument first,
Aug 16
bottom of page