These signals represent our analytical observations about the technology and market landscape. They are commentary, not financial product advice. The mention of any company or financial product is not a recommendation to buy, sell, or hold. See our Terms of Service for full details.
17 SEPT 2026

HarnessTax: harness barely moves pass rates but can cost 5x; Pi's four tools

HarnessTax, a new benchmark site whose data was generated on 16 September, evaluates 21 model–harness pairs — seven models across three harnesses, Claude Code, Codex CLI and Pi — on SWE-bench Lite and Terminal-Bench 2.0. Its headline, as stated on the site: harness choice has little effect on task success rate but can significantly affect cost; the same model can reach similar success rates at up to five times the cost. Pi, which exposes only four tools (read, write, edit and bash), reaches the Pareto frontier on both benchmarks. The authors examine cost through completed attempts, recorded turn counts and initial context size, and close by arguing users should not have to make these configuration decisions at all — a harness should adapt as the task unfolds while staying general. The site's blog text is rendered client-side, so the quoted figures are taken as pasted from its blog view; the 21-pair, two-benchmark design is confirmed from the page metadata. Mark's reading is that the result vindicates the minimal-tool harness; John relayed the same finding from practitioners working with small models: large skill libraries and long tool lists fill the context, and 27B–30B-class models in particular thrash in that context, whereas a handful of tools — often just bash — makes them perform markedly better. That is the mechanism behind HarnessTax's flat success curve: the model does the work, and the harness's main lever is how much context and how many turns it burns getting there. The counter-argument is scope. Two benchmarks of well-specified, verifiable tasks are exactly the setting where scaffolding matters least; the harness earns its keep on long-horizon, ambiguous work with permissions, memory and recovery, none of which SWE-bench Lite measures. And the July Terminal-Bench 2.1 evidence already showed the labs compressing the harness gap to about three points — so 'little effect on success' may reflect convergence rather than irrelevance. The implication is a price signal, not a capability one. If success is model-bound and cost is harness-bound, the harness market competes on token efficiency, and a 5x cost spread is a margin a vendor can either capture or lose. Watch whether Anthropic and OpenAI publish cost-per-task alongside pass rates for their own CLIs, whether the initial system prompt and tool manifest shrink in the next releases, and whether third-party minimal harnesses take share among users running open 27B–30B models locally, where every wasted token is felt directly.
harnessescoding-agentsbenchmarkscostharness-engineeringprocesses-harnessesefficiency-paradigm
Source ↗
17 SEPT 2026

Expert re-grade: GPT-5.6 Sol's HLE-Physics score rises from 47.3% to 78.7%

A 51-author paper led from Yale (arXiv 2609.13009, submitted 11 September) re-grades frontier models on six widely used physics benchmarks, including several that feed the Artificial Analysis Intelligence Index. Faculty and graduate researchers in each subfield reviewed problem statements, reference solutions and model responses to separate genuine model errors from grader errors, wrong reference answers and ambiguous questions. Most cases initially marked incorrect turned out to be benchmark faults. After experts corrected reference solutions and repaired or excluded flawed items, GPT-5.6 Sol's mean@4 rose from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark; its corrected pass@4 reached 94.4% on the 54 retained CritPt challenges. Corrected scores are computed on the retained subsets after review, and audited subsets of UGPhysics, PRISM-Physics and PHYBench also rose substantially. The abstract's conclusion: current benchmarks substantially understate frontier models' ability to solve well-posed physics problems, and closed-ended physics is near saturation. This is the strongest documented case yet that the public leaderboards are measuring their own defects. It cuts both ways for the NOOPS reading of benchmarks. It supports the 'good enough' thesis — models are further along on hard quantitative reasoning than the headline numbers say, which is why domain experts report better experiences than the scores predict. But it also undermines the composite indices that markets and procurement teams use to rank models, since a slice of the Intelligence Index is now shown to be grading against wrong answers. The counter-argument is selection: the audit focused on text-only problems with verifiable final answers, retained subsets exclude the questions experts could not repair, and mean@4 and pass@4 are generous metrics. 'Near-saturation on well-posed problems' is not the same as physical reasoning in open-ended research, and the paper itself says the field now needs more demanding, expert-validated evaluations. For positioning, the significance is that benchmark-driven narratives about a capability plateau in science are less reliable than they looked, and the value of independent, expert-graded evaluation goes up. Watch whether Artificial Analysis and the benchmark maintainers adopt the corrected sets — that would lift several models' scores at once and reorder rankings without any model changing — and whether the labs begin citing expert-audited results in place of raw leaderboards.
benchmarksevaluationphysicsfrontier-modelskuhn-paradigmgood-enoughharness-engineering
Source ↗
17 SEPT 2026

Yegge shuts Gas Town; Databricks: Astra lifts coding spend 60% across 3,500

Latent Space's AI News edition for 15–16 September runs under the heading 'reality checks'. Its first item is Steve Yegge — loud in his advocacy of 'tokenmaxxing' — shutting down Gas Town, the multi-agent orchestrator that made him famous, and conceding that despite spending many thousands a month on coding-agent subscriptions, he only ever used Gas Town to build itself. In his own essay Yegge says Gas Town 'fell apart at the seams' with Opus 4.7, which developed a 'just two more things' tic that kept it fiddling with the orchestrator rather than doing real work; since Claude Fable 5 he has returned to Wyvern, his 30-year-old game, and built a more organised successor he calls Wheelhouse. Dan Luu noted he had found such orchestrators unreliable for completing tasks and 'the author of the most famous one had the same issue'. The second reality check is Databricks' Patrick Wendell reporting that rolling GPT-6 Astra out to roughly 3,500 engineers 'unambiguously' outperformed Opus 5 and Sol 5.6 on complex system-design work while increasing overall coding spend by about 60%. Mark's reaction was to the essay's other half: Yegge's predictions that CI/CD 'as we know it' dies next year, replaced by a 'Mad Max-style thunderdome', and that human code review is 'completely done and gone' by next year, surviving only on SOC 2 life support — 'Fable is the only reasonably trustworthy model in existence today', but 'in seven months all the models will be that smart'. Read together, the two halves are a single signal about where the frontier of agentic engineering actually is. The most aggressive practitioner has abandoned his own orchestrator, and the most data-rich enterprise rollout shows that the better model costs more per engineer, not less. The counter-case is that Yegge is not retreating — he says he is 'operating about 12 months in the future', running fleets of agents with soaring budgets, and treats Gas Town as a prototype superseded by design rather than an idea disproved. The implication for the tiny-teams and technical-deflation theses is that the cost of engineering output is not yet falling in the way the per-token price is; it is being redeployed into more tokens. Watch for more enterprise disclosures like Databricks' — spend per engineer under the new models is the number to track — and for whether SOC 2 auditors accept agentic review chains in place of a human approval on every diff, which is the real gate on Yegge's forecast.
coding-agentstokenscostsoftware-engineeringtiny-teamsharness-engineeringtechnical-deflation
Source ↗

You're seeing the latest 3 of 2073 signals. Sign up free to see more, or subscribe to Plus for the full archive.

Sign up free Subscribe to Plus