Eleven cells moved. Here is what they mean for your business.
Reading the September 2026 State of AI verdict table: what improved, what declined, and what to do differently this quarter.
In June we published a table. Ten rows, one per AI capability, four columns each: what you can deploy today, where the demo outruns production, what is oversold, and what is underestimated. Forty judgements, every one dated and sourced.
This week we published the same table again, and for the first time we can say which cells moved.
Eleven of forty were amended; twenty-nine stand as written. That is the headline of the September 2026 edition, and it is worth pausing on before the detail, because it cuts against both stories you have been told about AI this winter. Read the paper’s register and the amendments sort further: four moved in your favour, five say more caution is now warranted, and two kept their verdict while the reason underneath changed. The field did not transform in a quarter. It also did not stall. It moved in specific, nameable places, and a business that knows which places can act on them.
This post walks the table. The paper behind it is long, deliberately so, because it is written to be quoted by AI search engines and checked by people who check things. You do not need to read it to use this. One tip before you go: the table has a toggle that swaps every cell to its June text. Use it once. The series only earns its keep if you can see the previous verdict beside the current one, and that contrast takes one click rather than two browser tabs.
Where the evidence moved in your favour
All four of these amendments are infrastructure. None of them is a smarter model.
Open-weight models got cheaper and longer. DeepSeek released V4-Flash on 31 July under an MIT licence with a one-million-token context window, priced by the vendor at US$0.14 per million input tokens. That is about 36 times below the input price of OpenAI’s flagship tier. Moonshot’s Kimi K3 followed with open weights. There are counter-signals worth knowing: Meta has left the open-weight frontier for a closed, paid line, and Alibaba’s newest Qwen shipped under a tighter licence than Apache 2.0. The practical reading is unchanged from June and stronger: for bounded tasks, extraction, classification, drafting against a template, an open model running under your own control is now the price comparison every quote should be tested against. Read the licence first.
The integration standard grew up. The Model Context Protocol, which is how AI systems connect to your tools, published its 2026-07-28 specification with something it lacked in June: a lifecycle policy. Features now get at least twelve months’ notice before removal, there is a conformance suite, and there are proper extensions for long-running tasks and structured skills. For a business that has held off connecting AI to its systems because the plumbing looked like it might change under them, that objection is materially weaker than it was.
The payment rails went live. Visa’s agent-payment programme moved from pilot to production in Europe on 2 July with more than 30 issuing banks. No transaction volumes have been published, which matters, and the September edition is careful about that. But “live with named banks and no volume” is a different state from “pilot”, and the US Treasury’s stablecoin rulemaking (18 August) means the settlement layer beneath it is getting a rulebook. For most readers this is still a watch category. It is a watch category that moved.
Australian regulators kept writing. The Fair Work Commission published AI-disclosure requirements on 26 August that take effect on 20 October: if you use generative AI in a document you file, you say so, you verify the citations, and false evidence carries up to twelve months’ imprisonment. The TGA listed software as a medical device among its twelve compliance priorities for 2026 and 2027, which every practice using an AI scribe should read as a signal. The OAIC updated its facial-recognition guidance after the Bunnings decision. June’s verdict was that Australia’s rules of the road were “documented, achievable, mostly ignored”. September’s is that there are more of them, they are dated, and they are still mostly ignored.
Where more caution is now warranted
All five are about trust: in measurement, and in products.
The benchmark in the sales deck is now disputed by the labs themselves. On 8 July, the day before it released GPT-5.6, OpenAI published an analysis estimating that about 30 per cent of the tasks in SWE-bench Pro, the coding benchmark we recommended in June as the contamination-resistant one, are broken. The day before that, a competitor had posted a better score on it. Whatever you make of the timing, the effect for a buyer is simple. June’s rule was to ask which benchmark, measured by whom, with what scaffolding, and how many runs. September adds a fifth question: on which subset. The public leaderboard for that benchmark tops out around 62 per cent; the labs quote 65 and 80 on a subset you cannot see. Two cells declined on this evidence, one in language models and one in code generation, and they declined together because it is the same lesson.
Grounded AI has more ways to fail than have been studied. A peer-reviewed taxonomy published on 4 July catalogues 33 distinct failure modes in retrieval-augmented systems, the “ask questions of your own documents” pattern that most business knowledge tools use. Twelve of the 33 have never been empirically studied, and that includes all eight that involve agents, which the paper calls the fastest-growing deployment pattern with the least scrutiny. The June verdict was that source-grounding does not mean accuracy. The September verdict is the same, with the addition that we understand the failure surface less well than the products imply. Grounded Q&A over a curated corpus with citations checked remains the right pattern. “Ask anything about your business” remains a pitch.
A fourth flagship product was shut down. OpenAI closed its Atlas browser on 9 August, less than ten months after launch. Google retired Veo 2.0 and Veo 3.0 on 30 June and Imagen 4 on 17 August. Sora’s API sunsets on 24 September, and the leading automated video pipeline has already removed it from its roster. In June we counted three flagship shutdowns from the best-resourced labs inside nine months. It is now four inside about a year, plus five model retirements from Google in one quarter. The verdict is unchanged and the evidence for it is heavier: build on capabilities and standards, rent products, and assume any single video model you depend on will be replaced within the year.
The one machine-payment rail with real numbers lost most of them. Daily settlement on x402, the rail we cited in June as “real and tiny”, fell 93 per cent year to date, from a late-2025 peak that approached US$800,000 a day to roughly US$28,000 to 42,000 by August. “The agent economy is here” was oversold in June. It is more oversold now.
Where the verdict stands but the reason changed
A US appeals court vacated the injunction against Perplexity’s Comet browser on 4 August, holding that when an AI assistant acts on a website, it is the user who accesses the site, and the assistant is “a tool, not a person for statutory purposes”. Read that carefully before you read it as a green light. It narrows one legal theory against the companies that make agents. It moves the exposure toward the business that runs the agent under its own account. And it says nothing at all about whether the agent works, which on the only independent office-task evidence available is still about 30 per cent of the time.
The second reframing is the payments one above: agent-initiated card payments are live but unmeasured, which is a different gap from the one June described.
The twenty-nine that stand
Half of the unchanged cells were confirmed by new evidence. The other half had no new evidence of the required standard, and I want to be direct about what that means, because it is easy to read an uncoloured cell as “nothing happened”.
In voice agents, for example, we could not find a single independent evaluation of containment or accuracy in the two quarters this series has looked. Bland raised a Series C and reports 3.5 million calls a week; ElevenLabs reports US$500 million in annual revenue. Both are vendor figures about scale, not evidence about reliability. The planning assumption from June, roughly half of routine calls fully contained, stands because nothing has tested it, not because something has confirmed it. The December edition will treat a third consecutive quarter of silence in any cell as a finding about that category, and voice is the likeliest candidate.
What I would do differently this quarter
The adoption posture from June holds. Deploy the bounded, verifiable work with a human check on low-confidence cases. Pilot the plausible things against a metric you set before you start. Wait on unsupervised agents for anything consequential. Walk away from “hallucination-free” and from any ROI figure nobody audited.
Three things are new. Re-price every bounded task you run on a closed model against an open-weight one; the gap widened this quarter and it is now large enough to change decisions. If you have been holding off connecting AI to your systems because the standard looked unstable, the twelve-month deprecation guarantee is the answer to that objection. And if you file anything with the Fair Work Commission, or run a clinical scribe, or use facial recognition in a retail setting, there is a dated rule with your name on it that did not exist in June.
The full edition, with every source, the scorecard of what June got wrong, and the watchlist of thirty claims we still cannot verify, is at the State of AI hub. The June edition keeps its own permanent link, unchanged, because a series is only worth quoting if it scores itself.
If you want to work through what this quarter’s movement means for your own operation, start with a conversation.