Evaluation 7 min read

Eleven cells moved. Here is what they mean for your business.

Reading the September 2026 State of AI verdict table: what improved, what declined, and what to do differently this quarter.

In June we published a table. Ten rows, one per AI capability, four columns each: what you can deploy today, where the demo outruns production, what is oversold, and what is underestimated. Forty judgements, every one dated and sourced.

This week we published the same table again, and for the first time we can say which cells moved.

Eleven of forty were amended; twenty-nine stand as written. That is the headline of the September 2026 edition, and it is worth pausing on before the detail, because it cuts against both stories you have been told about AI this winter. Read the paper’s register and the amendments sort further: four moved in your favour, five say more caution is now warranted, and two kept their verdict while the reason underneath changed. The field did not transform in a quarter. It also did not stall. It moved in specific, nameable places, and a business that knows which places can act on them.

This post walks the table. The paper behind it is long, deliberately so, because it is written to be quoted by AI search engines and checked by people who check things. You do not need to read it to use this. One tip before you go: the table has a toggle that swaps every cell to its June text. Use it once. The series only earns its keep if you can see the previous verdict beside the current one, and that contrast takes one click rather than two browser tabs.

Where the evidence moved in your favour

All four of these amendments are infrastructure. None of them is a smarter model.

Open-weight models got cheaper and longer. DeepSeek released V4-Flash on 31 July under an MIT licence with a one-million-token context window, priced by the vendor at US$0.14 per million input tokens. That is about 36 times below the input price of OpenAI’s flagship tier. Moonshot’s Kimi K3 followed with open weights. There are counter-signals worth knowing: Meta has left the open-weight frontier for a closed, paid line, and Alibaba’s newest Qwen shipped under a tighter licence than Apache 2.0. The practical reading is unchanged from June and stronger: for bounded tasks, extraction, classification, drafting against a template, an open model running under your own control is now the price comparison every quote should be tested against. Read the licence first.

The integration standard grew up. The Model Context Protocol, which is how AI systems connect to your tools, published its 2026-07-28 specification with something it lacked in June: a lifecycle policy. Features now get at least twelve months’ notice before removal, there is a conformance suite, and there are proper extensions for long-running tasks and structured skills. For a business that has held off connecting AI to its systems because the plumbing looked like it might change under them, that objection is materially weaker than it was.

The payment rails went live. Visa’s agent-payment programme moved from pilot to production in Europe on 2 July with more than 30 issuing banks. No transaction volumes have been published, which matters, and the September edition is careful about that. But “live with named banks and no volume” is a different state from “pilot”, and the US Treasury’s stablecoin rulemaking (18 August) means the settlement layer beneath it is getting a rulebook. For most readers this is still a watch category. It is a watch category that moved.

Australian regulators kept writing. The Fair Work Commission published AI-disclosure requirements on 26 August that take effect on 20 October: if you use generative AI in a document you file, you say so, you verify the citations, and false evidence carries up to twelve months’ imprisonment. The TGA listed software as a medical device among its twelve compliance priorities for 2026 and 2027, which every practice using an AI scribe should read as a signal. The OAIC updated its facial-recognition guidance after the Bunnings decision. June’s verdict was that Australia’s rules of the road were “documented, achievable, mostly ignored”. September’s is that there are more of them, they are dated, and they are still mostly ignored.

Where more caution is now warranted

All five are about trust: in measurement, and in products.

The benchmark in the sales deck is now disputed by the labs themselves. On 8 July, the day before it released GPT-5.6, OpenAI published an analysis estimating that about 30 per cent of the tasks in SWE-bench Pro, the coding benchmark we recommended in June as the contamination-resistant one, are broken. The day before that, a competitor had posted a better score on it. Whatever you make of the timing, the effect for a buyer is simple. June’s rule was to ask which benchmark, measured by whom, with what scaffolding, and how many runs. September adds a fifth question: on which subset. The public leaderboard for that benchmark tops out around 62 per cent; the labs quote 65 and 80 on a subset you cannot see. Two cells declined on this evidence, one in language models and one in code generation, and they declined together because it is the same lesson.

Grounded AI has more ways to fail than have been studied. A peer-reviewed taxonomy published on 4 July catalogues 33 distinct failure modes in retrieval-augmented systems, the “ask questions of your own documents” pattern that most business knowledge tools use. Twelve of the 33 have never been empirically studied, and that includes all eight that involve agents, which the paper calls the fastest-growing deployment pattern with the least scrutiny. The June verdict was that source-grounding does not mean accuracy. The September verdict is the same, with the addition that we understand the failure surface less well than the products imply. Grounded Q&A over a curated corpus with citations checked remains the right pattern. “Ask anything about your business” remains a pitch.

A fourth flagship product was shut down. OpenAI closed its Atlas browser on 9 August, less than ten months after launch. Google retired Veo 2.0 and Veo 3.0 on 30 June and Imagen 4 on 17 August. Sora’s API sunsets on 24 September, and the leading automated video pipeline has already removed it from its roster. In June we counted three flagship shutdowns from the best-resourced labs inside nine months. It is now four inside about a year, plus five model retirements from Google in one quarter. The verdict is unchanged and the evidence for it is heavier: build on capabilities and standards, rent products, and assume any single video model you depend on will be replaced within the year.

The one machine-payment rail with real numbers lost most of them. Daily settlement on x402, the rail we cited in June as “real and tiny”, fell 93 per cent year to date, from a late-2025 peak that approached US$800,000 a day to roughly US$28,000 to 42,000 by August. “The agent economy is here” was oversold in June. It is more oversold now.

Where the verdict stands but the reason changed

A US appeals court vacated the injunction against Perplexity’s Comet browser on 4 August, holding that when an AI assistant acts on a website, it is the user who accesses the site, and the assistant is “a tool, not a person for statutory purposes”. Read that carefully before you read it as a green light. It narrows one legal theory against the companies that make agents. It moves the exposure toward the business that runs the agent under its own account. And it says nothing at all about whether the agent works, which on the only independent office-task evidence available is still about 30 per cent of the time.

The second reframing is the payments one above: agent-initiated card payments are live but unmeasured, which is a different gap from the one June described.

The twenty-nine that stand

Half of the unchanged cells were confirmed by new evidence. The other half had no new evidence of the required standard, and I want to be direct about what that means, because it is easy to read an uncoloured cell as “nothing happened”.

In voice agents, for example, we could not find a single independent evaluation of containment or accuracy in the two quarters this series has looked. Bland raised a Series C and reports 3.5 million calls a week; ElevenLabs reports US$500 million in annual revenue. Both are vendor figures about scale, not evidence about reliability. The planning assumption from June, roughly half of routine calls fully contained, stands because nothing has tested it, not because something has confirmed it. The December edition will treat a third consecutive quarter of silence in any cell as a finding about that category, and voice is the likeliest candidate.

What I would do differently this quarter

The adoption posture from June holds. Deploy the bounded, verifiable work with a human check on low-confidence cases. Pilot the plausible things against a metric you set before you start. Wait on unsupervised agents for anything consequential. Walk away from “hallucination-free” and from any ROI figure nobody audited.

Three things are new. Re-price every bounded task you run on a closed model against an open-weight one; the gap widened this quarter and it is now large enough to change decisions. If you have been holding off connecting AI to your systems because the standard looked unstable, the twelve-month deprecation guarantee is the answer to that objection. And if you file anything with the Fair Work Commission, or run a clinical scribe, or use facial recognition in a retail setting, there is a dated rule with your name on it that did not exist in June.

The full edition, with every source, the scorecard of what June got wrong, and the watchlist of thirty claims we still cannot verify, is at the State of AI hub. The June edition keeps its own permanent link, unchanged, because a series is only worth quoting if it scores itself.

If you want to work through what this quarter’s movement means for your own operation, start with a conversation.

Published 3 September 2026

Perth AI Consulting delivers AI opportunity analysis for small and medium businesses. Start with a conversation.

Prepared by Claude, directed and approved by PAC.

More from Thinking

Evaluation 11 min read

AI in property valuation: the evidence, the design rules, and what it could become

The best Australian evidence on vision AI in valuation measures a different task than the one vendors demo. The findings, and the design rules that follow.

Evaluation 7 min read

Competitor intelligence for small business: what AI can and cannot see

What AI-assisted competitor intelligence really is for a small business: the public sources worth watching, what they cannot tell you, and the legal line.

Evaluation 10 min read

AI in regulated professional work, Mid-2026

One structure links family law, valuation, and building inspections: a signed document others rely on. How each field's regulator answered the AI question.

Technical 9 min read

The business knowledge base: evidence, risks, and how to build one

What a business knowledge base actually is, what the evidence says it delivers, the security and privacy realities, and how we build one that holds up.

Evaluation 8 min read

What AI can see in your customer data (and what it cannot)

What AI can genuinely find in the customer records an SME already holds, what it cannot, and when a spreadsheet honestly beats a model.

Building 7 min read

What an AI quoting engine actually does

What an AI quoting engine takes in, what it drafts, what the evidence says about accuracy and speed, and why the final price stays with a human.

Adoption 6 min read

Australia's AI adoption gap is bigger than the 12% headline suggests

ABS says 12% of Australian businesses use AI. The real story is 35% of large businesses against 11% of small ones, and the barrier isn't the technology.

Building 7 min read

Why we let AI run the interviews (and why we never let it pretend to be human)

AI-conducted interviews compress weeks of stakeholder discovery into days, standardise what gets asked, and lower the guard that distorts honest answers.

Adoption 14 min read

How AI capability actually moves through a business

The decisive variable in SME AI adoption is the human absorption sequence, not the tooling. A working framework from observation across WA businesses.

Evaluation 7 min read

AHPRA advertising rules for psychologist websites

Recovery stories, 'specialist', 'clinical psychologist', and endorsement titles are where psychology sites breach the National Law. A practical read-through.

Adoption 4 min read

Customer service AI has finally grown up

Chatbots and AI receptionists earned their bad reputation. What changed, and how the mature version answers every call without replacing anyone.

Evaluation 6 min read

Who can use the titles 'Dr', 'Specialist', and 'Surgeon'?

AHPRA restricts 'specialist' and 'surgeon' to specific registrations, and 'Dr' has its own rule. What health practice websites can and cannot claim.

Adoption 5 min read

Your best people hate writing reports

The operators you promote are brilliant at the work and allergic to reporting. A scheduled AI call interviews them, drafts the briefing, they approve it.

Building 6 min read

Your website isn't just for humans anymore

How to build a chatbot that keeps itself up to date, can't leak client information, and won't answer beyond what you've published.

Evaluation 7 min read

Can you show Google reviews on your health practice website?

AHPRA bans clinical testimonials, even true ones, but service reviews are fine. What that means for the Google reviews widget on your practice site.

Evaluation 7 min read

What AHPRA's advertising rules mean for your website

Your practice website is advertising under the National Law. What AHPRA's rules prohibit, who is responsible, and how to check your own site.

Evaluation 8 min read

Is it safe to paste client data into ChatGPT?

Short answer: it depends on one setting, and most people have it wrong. What ChatGPT, Claude and Copilot do with your data, and what the Privacy Act expects.

Evaluation 4 min read

What a good AI audit actually delivers

The audit report named one recommendation specific enough to check, and what the Build that followed looked like: one real engagement, generalised.

Evaluation 7 min read

AI and video, Mid-2026: the models can watch now, not just listen

AI could always transcribe video. It can now read the frames as well, and every hour of footage a business owns becomes something it can question.

Building 7 min read

Case study: a 119-page AML/CTF program in three days

How we built a seven-document AML/CTF compliance pack for a small accounting practice in three days, working from 31 confirmed assumptions.

Building 11 min read

From evidence base to delivery: a production AI methodology

How we delivered 34 evidence-anchored AI briefings to a WA peer-advisory chapter: fact-checked literature review, multi-agent verification, one method.

Technical 9 min read

The six functions of a working AI system

A working AI system is six functions doing six jobs. When all six connect, hallucinations get caught, outputs hold steady, and models become swappable.

Technical 7 min read

Supervised autonomy: the middle path for AI architecture

Between drafts you approve and agents you hope about sits the middle path: an envelope of authorised routine work, supervised, audited, and yours to widen.

Evaluation 5 min read

The state of applied AI in Mid-2026

Our literature review of applied AI in mid-2026: ten capability categories, three fact-check passes, written for operational leaders.

Technical 9 min read

How to design a PHI redaction system for clinical AI

PHI redaction is part of a clinical AI tool's architecture, not a feature you add. What the literature says it should look like, and how we built it.

Building 9 min read

How we built on-device de-identification so AI never sees real names

Most AI privacy is a policy. Ours is architecture: an NER model runs in the browser and strips names before anything leaves the device.

Technical 7 min read

Your agency's clients are about to ask why this costs so much

A solo consultant built in three weeks what your agency quoted twelve for. The client doesn't know why yet. The agencies that survive change what they sell.

Adoption 6 min read

What do you love doing? What do you hate doing?

Ask people what they love doing and what they hate doing, then show them AI is coming for the second list. Why the reframe works, and how it fails.

Technical 7 min read

Why I don't use n8n (and what I do instead)

n8n demos well. But a compelling demo and a reliable production system are different things, and the distance between them is where businesses get hurt.

Technical 10 min read

Your codebase was not built for AI. That's the actual problem.

Amazon's mandatory meeting about AI breaking production is an architecture story: codebases built for human maintainers only, now maintained by AI.

Adoption 4 min read

Your team has AI licences. You don't have an AI system.

Fifteen people, fifteen separate AI accounts, no shared context. The problem isn't the tool; it's the architecture around it. Here's the fix.

Building 7 min read

Your $2,000 day starts the night before: our system keeps you on the tools, not on the phone

Optimised routes overnight, automatic customer notifications, and promises the system keeps or corrects. A scheduling system that protects your daily rate.

Evaluation 4 min read

The fastest way for an executive to get across AI

AI moves faster than any executive can track. One focused conversation, one written report, and a decision you can act on: your time stays on the business.

Building 6 min read

Your IT department will take 18 months. You need this working by next quarter.

Senior leaders know what they need built; the gap is time. A prototype gets the tool working now and hands IT a validated blueprint for later.

Building 8 min read

We built an AI invoice verifier. Here's where it hits a wall.

We built an AI invoice verifier and watched a fake beat a real invoice. Why document analysis alone cannot stop fraud, and the five layers that can.

Building 5 min read

How to build an AI chatbot that doesn't lie to your customers

Woolworths scripted its AI to talk about its mother. The business fix is honesty; the technical fix is architecture that prevents fabrication by design.

Technical 9 min read

Why AI safety features are load-bearing architecture, not political decoration

The 'woke AI' label came from real failures, but they were engineering failures, not safety failures. The difference matters wherever errors have consequences.

Adoption 3 min read

Woolworths' AI told a customer it had a mother. That's a problem.

Woolworths' AI assistant Olive was scripted to talk about its mother and uncle. When callers realised, trust broke instantly. The fix is honesty.

Evaluation 5 min read

Google is no longer the only way your customers find you

Customers now find businesses through ChatGPT, Perplexity, and Gemini. The sites AI cites are structured differently to the sites Google ranks.

Evaluation 6 min read

The personal workflow analysis: what watching a real workday reveals about automation

People describe the work they value, not the work that eats their time. Recording a real workday reveals the automation opportunities interviews miss.

Evaluation 11 min read

An AI audit that starts with your business

How an operations-first AI audit works: what it looks for, how the evidence is collected, what the report contains, and what it tells you to skip.

Building 6 min read

What production AI teaches you that demos never will

The gap between a demo and a working system is where the useful lessons live. Architecture, framing, privacy, adoption: the patterns repeat every time.

Adoption 6 min read

The psychology of why your team won't use AI

You buy the tool, run the demo, and three months later nobody is using it. Five predictable psychological barriers, each with a strategy that works.

Technical 4 min read

Stop telling AI what NOT to do (and what to say instead)

Instructions built on prohibitions make AI cautious and generic. Describing what you want instead transforms the output, and the reason comes from psychology.

Building 5 min read

How we turned generic AI into a specialist: and what that means for your business

Mediocre AI output is rarely the model's fault. Three structural changes that turn the same model from generic to specialist-grade.

Evaluation 6 min read

Your business has 9 customer touchpoints. AI can fix the 6 you're dropping.

You pay to get customers to your door, then lose them to missed follow-up. AI can handle the six touchpoints most businesses drop.

Technical 6 min read

What happens to your data when you press 'Send' on an AI tool

Businesses send customer data to AI tools without knowing what happens during processing. The spectrum of AI privacy is wider than you think.