This week, the agent harness stopped being plumbing and became a line item: DeepSeek open sourced one, Writer sells one, and NVIDIA routes through one. Four independent benchmarks published in the same window show why: swap the harness under a fixed model and the score moves 20 to 40 points, with almost no correlation between how models rank under one harness versus another. The Score Was Never Just the Model OpenAI's own developer guide for GPT-5.6 makes this argument first,