Quarterly evidence review · Edition permalink · Archive

The State of AI in Mid-2026

What AI could do as at June 2026, and what it couldn't. The first edition of the series.

Evidence as at June 2026 · v1.1 Superseded 2 September 2026

This is the permanent link for the Mid-2026 edition, kept unchanged for citation. The current edition, which classifies every one of these forty judgements against this one, lives at /resources/state-of-ai/.

Perth AI Consulting · Verdict table

The State of AI in Mid-2026

Ten capability sections, four judgements each. The set travels together; every row links into the published review.

Evidence as at June 2026 · v1.1
Deploy todayText and document work with human review, structured extraction with validation, coding assistance.
Demo vs productionAnything quoted in benchmark points.
Oversold“Hallucination-free” claims in any domain, and headline benchmark margins between frontier models, now within measurement noise.
UnderestimatedOpen-weight models: for many bounded tasks, closed-frontier quality at a fraction of the cost, under the business’s own control.
Deploy todayCoding agents, research agents with human review, tightly scoped internal automations with approval gates.
Demo vs productionEnd-to-end “digital employee” demos showcase the 30 per cent of attempts that succeed.
OversoldAutonomous agents for consequential unsupervised work, and pervasive “agent washing” of ordinary automation.
UnderestimatedThe compounding value of the MCP standard, and supervised agents as a labour multiplier for staff who delegate and verify.
Deploy todayDocument extraction with confidence-gated human review, construction progress capture, regulator-cleared medical tools inside their cleared indication.
Demo vs productionHandwriting and degraded documents (clean-benchmark 95 per cent falls to roughly 75 on real material). Medical tools moved across sites.
OversoldVision-against-standards products quoting no independent accuracy data, consumer-grade diagnostic apps.
UnderestimatedOrdinary document work when the workflow includes verification: for many SMBs the single highest-ROI AI capability in 2026.
Deploy todayImage generation for marketing and mock-ups, short-form video with human selection from multiple generations, AI-assisted (not AI-run) pipelines.
Demo vs productionShot-to-shot continuity, on-screen text, precise brand fidelity: the showreel is the best of dozens of attempts.
Oversold“Automated content engine” revenue claims, “studio quality” as a general proposition, any single product’s durability (Sora’s seven-month arc).
UnderestimatedHow cheap and fast competent short-form visual content has become for businesses that previously could not afford it.
Deploy todayVoice agents for bounded after-hours and overflow handling with human escalation, transcription with review, TTS for IVR, content, accessibility.
Demo vs productionClean-audio accuracy claims, vendor containment rates roughly 1.4 to 2 times delivered rates.
Oversold“Indistinguishable from human” as a blanket claim, fully autonomous phone-based sales, most published voice-agent ROI statistics (vendor case studies).
UnderestimatedMissed-call response and after-hours intake (narrow, cheap, capture previously lost revenue), and the cloning-fraud exposure most SMBs have not addressed.
Deploy todayGrounded Q&A and synthesis over curated, current document sets with citations checked, meeting synthesis, audio digestion.
Demo vs productionEnterprise search over an ungoverned file estate: the demo corpus is clean, yours is not.
Oversold“Ask anything about your business” positioning, and the implication that source-grounding means accuracy.
UnderestimatedNotebookLM-class tools for regulated professionals working against bounded authoritative texts, provided the verification habit holds.
Deploy todayAI-assisted development for professional teams (code review and tests non-negotiable), non-developer internal tools holding no sensitive data.
Demo vs productionThe app that works in the demo has no auth, no edge cases, and no attackers.
Oversold“Anyone can ship production software”, benchmark-derived capability claims.
UnderestimatedCoding agents cut custom internal software cost: processes too small in 2023 now justify software, provided someone competent owns security.
Deploy todayMissed-call text-back and after-hours intake, FAQ-grade deflection on a well-maintained knowledge base, workflow automation with human approval steps.
Demo vs productionResolution rates quoted from simple-traffic mixes, “AI employee” demos concealing throttles, configuration burden, and the escalation tail.
OversoldEnterprise agent suites as plug-and-play (the data engineering is the project), virtually all published ROI numbers (none independently audited).
UnderestimatedThe boring automations (where the dependable money is), and outcome-based pricing: charging per resolution is a falsifiable claim.
Deploy todayFor almost all readers, nothing: a watch category (stablecoin settlement for international payments is the relevant adjacent capability).
Demo vs productionAgent-to-agent commerce demos, against no consumer-scale deployment that has survived contact with merchants.
OversoldToken projects citing volume that is actually speculation or subsidy, “the agent economy is here”.
UnderestimatedThe rails themselves: incumbents’ speed standardising agent-payment authorisation suggests they expect demand, and confidential compute addresses a real regulated-firm privacy problem.
Deploy todayAI scribes under AHPRA/RACGP-compliant consent and review, legal research with mandatory citation verification, confidence-gated bookkeeping, trade-platform quote drafting.
Demo vs productionTime-saved claims: vendor surveys say 40 to 60 per cent, controlled studies say minutes per day (still worth having).
Oversold“Hallucination-free” professional tools, autonomous compliance documents of any kind (SOAs, building reports, HR policies).
UnderestimatedCompounding value of modest validated savings in high-frequency workflows, and Australian regulators’ already-published rules: documented, achievable, mostly ignored.
The four judgements are a set; reading one without the other three misreads the evidence. Executive one-pager →
PERTH AI CONSULTING LITERATURE REVIEW · EXECUTIVE PAGE

The State of AI in Mid-2026

What the evidence supports, on one page

Evidence as at June 2026 · Version 1.1 · Ten capability sections

The field in mid-2026 supports neither the enthusiast's story nor the sceptic's. Across ten capability categories the evidence shows one consistent shape: assistive uses with human review are reliably in production, while autonomous end-to-end claims remain ahead of the data. AI agents completed roughly 30 per cent of simulated office tasks (Carnegie Mellon's TheAgentCompany; best-agent result reported by The Register, 2025), and grounded question answering over curated, current corpora is judged deployable today.

Every claim in this review carries a date, every vendor-originated number is labelled a vendor claim, and this page should be read as a photograph of June 2026, not a standing description. Check the version before relying on any figure.

What to do with this

Deploy now, with verification built in
Assisted drafting, transcription with review, grounded Q&A over curated corpora, extraction from consistent formats.
Pilot deliberately, with falsifiable metrics
Narrow agent workflows, document triage, coding assistance beyond boilerplate. Set the exit metric before the pilot starts.
Wait
Multi-step autonomous agents, touchless document processing, broadcast-grade video. The evidence is moving; the date on this page tells you when it was last checked.
Walk away
Any deployment without a human checkpoint, and any claim that cannot name its evidence.
Thirty circulating claims we could not verify are listed in Appendix B, each with the specific reason. Several widely repeated facts about AI in 2026 live there rather than in the body, which is itself a finding.
Read the full review → perthaiconsulting.com.au/resources/state-of-ai/

How this review is checked

The review went through a three-pass independent fact-check before publication. Across the passes, 135 individual fact-check findings were logged and resolved, and the full corrections log is published as an appendix in the document itself. Thirty circulating claims that could not be verified are listed with the specific reason each failed checking, rather than silently dropped.

The same pipeline delivers our client work: the production methodology is documented in From evidence base to delivery.

Read the full review

About 13,000 words, 100 plus cited sources, corrections log included.