Quarterly evidence review · Edition 2

The State of AI: September 2026 Update

What AI can actually do right now, what it can't, and which way each verdict moved since June, so you can plan and invest with confidence.

Evidence as at 2 September 2026 · v1.0 Next edition: December 2026

This page always serves the current edition. Past quarters keep their own permanent links in the archive; this edition's permalink is /resources/state-of-ai/september-2026/.

Perth AI Consulting · Verdict table

The State of AI: September 2026 Update

Ten capability sections, four judgements each, every cell complete as it stands. Eleven of the forty judgements were amended this quarter: amended cells are tinted, and the toggle reads the superseded June text in place. Every row links into the published review.

Evidence as at 2 September 2026 · v1.0
Deploy today Text and document work with human review, structured extraction with validation, coding assistance; now materially cheaper on open-weight models for bounded tasks.
Demo vs production Amended Anything quoted in benchmark points, and any SWE-bench figure without its subset, scaffold, and run count.
Oversold “Hallucination-free” claims in any domain, and headline benchmark margins between frontier models, which the labs themselves now dispute.
Underestimated Amended Open-weight models: for many bounded tasks, closed-frontier quality at a fraction of the cost, under the business’s own control. Read the licence.
§3.2 Agentic systems and tool use September 2026 · v1.0
Deploy today Coding agents, research agents with human review, scoped internal automations with approval gates, MCP-connected tools built on the current specification.
Demo vs production End-to-end “digital employee” demos still showcase the roughly 30 per cent of attempts that succeed on the only independent office-task evidence (2025); METR finds autonomous research agents have minimal effect.
Oversold Amended Autonomous agents for consequential unsupervised work, pervasive “agent washing”, and any reading of the Ninth Circuit’s ruling as a green light: it decides who accesses a computer, not whether the agent works.
Underestimated Amended The compounding value of the MCP standard, now with a twelve-month deprecation guarantee, and supervised agents as a labour multiplier for staff who delegate and verify.
§3.3 Vision and multimodal models September 2026 · v1.0
Deploy today Document extraction with confidence-gated human review, construction progress capture, regulator-cleared medical tools inside their cleared indication.
Demo vs production Handwriting and degraded documents (clean-benchmark 95 per cent falls to roughly 75 on real material, last measured 2025 to 2026). Medical tools moved across sites.
Oversold Vision-against-standards products quoting no independent accuracy data, for a second consecutive quarter; consumer-grade diagnostic apps.
Underestimated Ordinary document work when the workflow includes verification: for many SMBs still the single highest-ROI AI capability.
§3.4 Video and image generation September 2026 · v1.0
Deploy today Image generation for marketing and mock-ups, short-form video with human selection from multiple generations, AI-assisted pipelines built to swap models.
Demo vs production Shot-to-shot continuity, on-screen text, precise brand fidelity: the showreel is the best of dozens of attempts, and a 30-second single pass is a duration, not a continuity guarantee.
Oversold Amended “Automated content engine” revenue claims, “studio quality” as a general proposition, any single product’s durability (Sora’s API sunsets 24 September; Google shut down Veo 2.0, Veo 3.0, and Imagen 4 this quarter).
Underestimated How cheap and fast competent short-form visual content has become for businesses that previously could not afford it.
Deploy today Voice agents for bounded after-hours and overflow handling with human escalation, transcription with review, TTS for IVR, content, accessibility.
Demo vs production Clean-audio accuracy claims; vendor containment rates roughly 1.4 to 2 times delivered rates, on evidence not independently refreshed since 2025.
Oversold “Indistinguishable from human” as a blanket claim, fully autonomous phone-based sales, most published voice-agent ROI statistics (vendor case studies).
Underestimated Missed-call response and after-hours intake (narrow, cheap, capture lost revenue), and the cloning-fraud exposure Australian regulators are now naming in scam reporting.
Deploy today Grounded Q&A and synthesis over curated, current document sets with citations checked, meeting synthesis, audio digestion.
Demo vs production Enterprise search over an ungoverned file estate: the demo corpus is clean, yours is not.
Oversold Amended “Ask anything about your business” positioning, and the implication that source-grounding means accuracy; the July 2026 literature catalogues 33 failure modes and has studied none of the eight involving agents.
Underestimated NotebookLM-class tools for regulated professionals working against bounded authoritative texts, provided the verification habit holds.
Deploy today AI-assisted development for professional teams (code review and tests non-negotiable), non-developer internal tools holding no sensitive data.
Demo vs production The app that works in the demo has no auth, no edge cases, and no attackers; the developer who feels faster has not, on the last randomised evidence, been measured as faster.
Oversold Amended “Anyone can ship production software”, and benchmark-derived claims: the labs now dispute the flagship coding benchmark among themselves.
Underestimated Coding agents cut custom internal software cost: processes too small in 2023 now justify software, provided someone competent owns security; this quarter’s price cuts push further.
Deploy today Missed-call text-back and after-hours intake, FAQ-grade deflection on a well-maintained knowledge base, workflow automation with human approval steps.
Demo vs production Resolution rates quoted from simple-traffic mixes, “AI employee” demos concealing throttles and the escalation tail, vendor revenue growth quoted without customer counts.
Oversold Enterprise agent suites as plug-and-play (the data engineering is the project), virtually all published ROI numbers, now including a US$1.5 billion ARR figure that says nothing about delivered resolution.
Underestimated The boring automations (where the dependable money is), and outcome-based pricing: charging per resolution is a falsifiable claim.
§3.9 Crypto and AI convergence September 2026 · v1.0
Deploy today For almost all readers, nothing: a watch category (stablecoin settlement for international payments is the relevant adjacent capability, now with a US federal rulebook in draft).
Demo vs production Amended Agent-initiated card payments are live in Europe with named banks and no published volume; agent-to-agent commerce demos still stand against no consumer-scale deployment that has survived contact with merchants.
Oversold Amended Token projects citing volume that is actually speculation or subsidy; “the agent economy is here”, against a machine-payment rail whose daily volume fell 93 per cent this year.
Underestimated Amended The rails themselves: incumbents moved from pilot to production inside a quarter, and confidential compute for AI workloads now has a first measured, small, real adoption figure.
§3.10 Vertical AI applications September 2026 · v1.0
Deploy today AI scribes under AHPRA/RACGP-compliant consent and review on ARTG-included products, legal research with mandatory citation verification (a Fair Work Commission disclosure obligation from 20 October), confidence-gated bookkeeping, trade-platform quote drafting.
Demo vs production Time-saved claims: vendor surveys say 40 to 60 per cent, controlled studies (none refreshed this quarter) say minutes per day, still worth having.
Oversold “Hallucination-free” professional tools, against an incident record past 2,000 cases with Australia second; autonomous compliance documents of any kind (SOAs, building reports, HR policies, tribunal filings).
Underestimated Amended Compounding value of modest validated savings in high-frequency workflows, and Australian regulators’ already-published rules, which gained three dated moves this quarter and remain mostly ignored.
The four judgements are a set; reading one without the other three misreads the evidence. Movement classifications are claims and carry citations in the paper's register. Executive one-pager →
PERTH AI CONSULTING EVIDENCE REVIEW · EXECUTIVE PAGE

The State of AI: September 2026 Update

What the evidence supports, on one page

Evidence as at 2 September 2026 · Version 1.0 · Edition 2 · Ten capability sections

The September 2026 edition holds the Mid-2026 frame fixed and records how each of forty judgements moved. Eleven moved, nine directionally. Four improved, all in infrastructure and regulation: open-weight models, the MCP standard, agent-payment rails, and Australia's regulatory floor. Five declined, all in measurement and durability: the flagship coding benchmark is disputed by its own users, grounded systems have more failure modes than studied, a fourth flagship product was shut down, and the one measured machine-payment rail lost most of its volume.

Every claim in this review carries a date, every vendor-originated number is labelled a vendor claim, every carried-forward judgement is cited to the June edition, and this page should be read as a photograph of 2 September 2026, not a standing description. Check the edition and version before relying on any figure.

What to do with this

Deploy now, with verification built in
Assisted drafting, transcription with review, grounded Q&A over curated corpora, extraction from consistent formats, confidence-gated bookkeeping; re-price bounded tasks against open-weight models.
Pilot deliberately, with falsifiable metrics
Narrow agent workflows, document triage, coding assistance beyond boilerplate, and, new this quarter, bank-offered agent payments with spending limits. Set the exit metric before the pilot starts.
Wait
Multi-step autonomous agents, touchless document processing, broadcast-grade video, any pitch resting on a benchmark chart (the flagship coding benchmark is now disputed by the labs themselves). The date on this page tells you when the evidence was last checked.
Walk away
Any deployment without a human checkpoint, any claim that cannot name its evidence, and any benchmark chart that does not name its subset.
The June edition's thirty unverifiable claims were revisited one by one: two are now verified, one gained a complication, two are closed as moot, twenty-five are unchanged. Eight new items were opened. The June edition's own predictions are scored item by item in the paper's front matter.
Read the full review → perthaiconsulting.com.au/resources/state-of-ai/

How the series works

Each edition keeps the same ten sections and four judgements so that movement can be measured rather than re-described. Every cell is classified against the previous edition (improved, declined, reframed, or unchanged), the previous edition's predictions are scored item by item, and its unverifiable claims are carried forward on a standing watchlist until they are verified, refuted, or retired. Where no new evidence of the required standard appears, the previous judgement is restated in full and cited, so each edition stands on its own.

Use the toggle on the verdict table to read the June 2026 text in place. The previous edition remains at its own permanent link: The State of AI in Mid-2026.

How this review is checked

Seven parallel research streams gathered dated, source-labelled evidence for the window 14 June to 2 September 2026. The draft then went through an independent fact-check by a separate model pass that had not seen the drafting and that traced every cited source: 118 findings, of which 13 were refuted and all corrected. The corrections log is published as Appendix C, the queries behind every "no new evidence" classification are listed in Appendix A, and the standing watchlist of unverifiable claims is Appendix B.

The same pipeline delivers our client work: the production methodology is documented in From evidence base to delivery.

Read the full review

About 23,000 words including the movement register, the scorecard of the June edition's predictions, the standing watchlist, and the corrections log; 80 plus cited sources.

Questions about this review

How often is the State of AI review updated?

Quarterly. This page always serves the current edition; each edition keeps a permanent dated permalink in the archive. The current edition is September 2026 (evidence as at 2 September 2026); the next edition is planned for December 2026.

What does 'vendor claim' mean in this review?

A number that originates from a vendor's own marketing or case studies rather than independent measurement. The review labels every such figure as a vendor claim so readers can weight it accordingly.

Can I rely on these verdicts after the quarter ends?

Treat each edition as a photograph of its stated date, not a standing description. The evidence moves; check the edition date and version before relying on any figure, and read the current edition at this page.

Editions

Every edition keeps a permanent link. Editions are superseded, never edited.

  • September 2026 Evidence as at 2 September 2026 · v1.0 · current edition · PDF
  • Mid-2026 Evidence as at 13 June 2026 · v1.1 · superseded 2 September 2026 · PDF