Evaluation 7 min read

AI and video, Mid-2026: the models can watch now, not just listen

AI could always transcribe video. It can now read the frames as well, and every hour of footage a business owns becomes something it can question.

For most of the last few years, “AI and video” meant one thing: transcription. Feed in a recording, get back the words. Useful, and completely deaf to everything that was not spoken. The model was listening. It could not see.

That has changed, quietly, over the past year. The same models that read documents can now look at what is on the screen: the slide, the diagram, the product, the thing someone pointed at and said “this bit here.” It is a smaller technical leap than it sounds and a much bigger practical one. This is a note for any operator sitting on a library of recorded footage: production shops, marketing teams, practices full of recorded meetings, trades with site walkthroughs on a phone. What has actually shifted, where it lands, and where it does not yet.

What’s stable

The craft is untouched. A model can log footage; it cannot judge which moment carries a piece, feel the beat that stops a viewer scrolling, or make the creative call a client pays a named person for. Whatever changes below, the judgement layer of video work stays human, and businesses whose value is the visual craft keep that value.

The economics of the archive have also been stable for years, in the bad sense. Most of working life is now recorded and never looked at again: every video call, every webinar, every training session someone captured “so we have it.” That footage sits on drives as dead weight, because nobody has three hours to scrub back through a two-hour call for the one exchange that mattered. The recorded hour has long been the most expensive thing to get value back out of. That is the fact the current shift attacks.

What’s changed

The words were already solved. Transcription is mature, cheap, and mostly free: platforms hand over caption tracks, and speech models cover anything without one. This is the part every tool already had, and it describes the audio only.

The pictures are the new part, and the engineering detail matters. To let a model see a video, the video has to become still frames, because that is what a vision model reads. Tools that grab a frame on a timer, one every thirty seconds or so, miss almost everything on edited footage; a slide that changed four times in that window reads as one slide. The stronger approach is scene-change detection: capture a frame every time the shot actually changes, so the sequence of frames matches the sequence of decisions someone made in the edit. The model then reads the video roughly the way an editor sees it, as deliberate moments rather than a slideshow on a metronome. That one architectural choice separates a tool that understands footage from one that guesses, and it is a useful test to put to any vendor in this category.

Frames and transcript read together are the unlock. Pair the scene frames with the words and hand both to a model that reasons over text and images jointly, and the questions change from “what was said” to “what was shown, when, and does it match the claim.” Neither stream alone can answer that.

The capability is assembled from open parts. This is not a product category a business has to wait for. The pieces are ordinary and public: a downloader for the video and captions, scene detection for the frames, a speech model as fallback, a vision-capable model to read it all. The open-source skill that demonstrated the pattern (claude-video, by Brad Bonanno, since extended by others) runs inside tools like Claude from a pasted link. For a business already running AI in its workflow, this is a small setup, not a project.

Where it lands

The narrow use is the obvious one: a production shop letting the machine do the logging, captioning, and platform versioning that consumes the hours between shoots. Real, and the smallest of the applications.

The broader use is the library. A marketing team sitting on a decade of shoots and event footage can finally find the usable thirty seconds without opening every file. A recorded onboarding session becomes a searchable set of steps. The quarterly all-hands becomes something staff can ask, rather than re-watch. A builder’s site walkthrough becomes a dated, described record that can be queried when there is a dispute. A researcher or lawyer reviewing hours of interview or camera footage gets a first pass that flags where to actually look. The pattern underneath is one idea: any pile of video a business owns is now a pile it can question, and an archive that was dead weight starts behaving like inventory.

Where it doesn’t yet land

It sees stills, not motion. Even with good scene detection, the model reads snapshots and infers what happened between them. Continuous motion, fast physical action, a gradual change with no clean cut: these fall through the gaps. It is watching a flip-book, not the film.

Scene detection is a heuristic, not comprehension. A camera pan or lighting shift can spawn frames nobody needs, while a long static talking-head with critical detail on a shared screen might barely trigger at all. Both directions are the tool guessing at where the meaning is.

It misreads with confidence. A vision model can misread text on a slide or describe something that is not quite there, fluently. The failure mode is not a blank “I don’t know”; it is a plausible wrong sentence. On anything that matters, the output gets verified against the source, which is the same verification rule that applies to every AI capability worth using.

Long footage costs real time and money, caption quality varies, and the speech-model fallback adds its own errors on top.

And there is a privacy decision underneath all of it. Turning footage into frames and sending them to a model means client faces, shared screens, and whatever was visible in the background now travel to wherever that model runs. For a public webinar, that is nothing. For a recorded client session or an internal call, it is a decision to make before the upload, not after. The considerations are the same as for any data pressed into an AI tool.

What this means for the operator’s posture

The sensible first move is small and low-stakes: point the capability at footage that could not hurt anyone, a recorded webinar, an old training video, a public talk, and ask it something specific. Ten minutes of that teaches more about where it is strong and where it is soft than any amount of reading. From there, the question is not “which video tool should we buy” but the same question as everywhere else in applied AI: which pile of recorded material does this business own, what would it be worth if it could be questioned, and what has to be true about privacy and verification before that is allowed to happen.

The interesting part was never the tool. It is that the recorded hour, which used to be the most expensive thing to get value back out of, is quietly becoming one of the cheapest.


If you are working out where a capability like this fits in how your business actually runs, start with a conversation.

Published 6 July 2026

Perth AI Consulting delivers AI opportunity analysis for small and medium businesses. Start with a conversation.

Prepared by Claude, directed and approved by PAC.

More from Thinking

Evaluation 11 min read

AI in property valuation: the evidence, the design rules, and what it could become

The best Australian evidence on vision AI in valuation measures a different task than the one vendors demo. The findings, and the design rules that follow.

Evaluation 7 min read

Eleven cells moved. Here is what they mean for your business.

Reading the September 2026 State of AI verdict table: what improved, what declined, and what to do differently this quarter.

Evaluation 7 min read

Competitor intelligence for small business: what AI can and cannot see

What AI-assisted competitor intelligence really is for a small business: the public sources worth watching, what they cannot tell you, and the legal line.

Evaluation 10 min read

AI in regulated professional work, Mid-2026

One structure links family law, valuation, and building inspections: a signed document others rely on. How each field's regulator answered the AI question.

Technical 9 min read

The business knowledge base: evidence, risks, and how to build one

What a business knowledge base actually is, what the evidence says it delivers, the security and privacy realities, and how we build one that holds up.

Evaluation 8 min read

What AI can see in your customer data (and what it cannot)

What AI can genuinely find in the customer records an SME already holds, what it cannot, and when a spreadsheet honestly beats a model.

Building 7 min read

What an AI quoting engine actually does

What an AI quoting engine takes in, what it drafts, what the evidence says about accuracy and speed, and why the final price stays with a human.

Adoption 6 min read

Australia's AI adoption gap is bigger than the 12% headline suggests

ABS says 12% of Australian businesses use AI. The real story is 35% of large businesses against 11% of small ones, and the barrier isn't the technology.

Building 7 min read

Why we let AI run the interviews (and why we never let it pretend to be human)

AI-conducted interviews compress weeks of stakeholder discovery into days, standardise what gets asked, and lower the guard that distorts honest answers.

Adoption 14 min read

How AI capability actually moves through a business

The decisive variable in SME AI adoption is the human absorption sequence, not the tooling. A working framework from observation across WA businesses.

Evaluation 7 min read

AHPRA advertising rules for psychologist websites

Recovery stories, 'specialist', 'clinical psychologist', and endorsement titles are where psychology sites breach the National Law. A practical read-through.

Adoption 4 min read

Customer service AI has finally grown up

Chatbots and AI receptionists earned their bad reputation. What changed, and how the mature version answers every call without replacing anyone.

Evaluation 6 min read

Who can use the titles 'Dr', 'Specialist', and 'Surgeon'?

AHPRA restricts 'specialist' and 'surgeon' to specific registrations, and 'Dr' has its own rule. What health practice websites can and cannot claim.

Adoption 5 min read

Your best people hate writing reports

The operators you promote are brilliant at the work and allergic to reporting. A scheduled AI call interviews them, drafts the briefing, they approve it.

Building 6 min read

Your website isn't just for humans anymore

How to build a chatbot that keeps itself up to date, can't leak client information, and won't answer beyond what you've published.

Evaluation 7 min read

Can you show Google reviews on your health practice website?

AHPRA bans clinical testimonials, even true ones, but service reviews are fine. What that means for the Google reviews widget on your practice site.

Evaluation 7 min read

What AHPRA's advertising rules mean for your website

Your practice website is advertising under the National Law. What AHPRA's rules prohibit, who is responsible, and how to check your own site.

Evaluation 8 min read

Is it safe to paste client data into ChatGPT?

Short answer: it depends on one setting, and most people have it wrong. What ChatGPT, Claude and Copilot do with your data, and what the Privacy Act expects.

Evaluation 4 min read

What a good AI audit actually delivers

The audit report named one recommendation specific enough to check, and what the Build that followed looked like: one real engagement, generalised.

Building 7 min read

Case study: a 119-page AML/CTF program in three days

How we built a seven-document AML/CTF compliance pack for a small accounting practice in three days, working from 31 confirmed assumptions.

Building 11 min read

From evidence base to delivery: a production AI methodology

How we delivered 34 evidence-anchored AI briefings to a WA peer-advisory chapter: fact-checked literature review, multi-agent verification, one method.

Technical 9 min read

The six functions of a working AI system

A working AI system is six functions doing six jobs. When all six connect, hallucinations get caught, outputs hold steady, and models become swappable.

Technical 7 min read

Supervised autonomy: the middle path for AI architecture

Between drafts you approve and agents you hope about sits the middle path: an envelope of authorised routine work, supervised, audited, and yours to widen.

Evaluation 5 min read

The state of applied AI in Mid-2026

Our literature review of applied AI in mid-2026: ten capability categories, three fact-check passes, written for operational leaders.

Technical 9 min read

How to design a PHI redaction system for clinical AI

PHI redaction is part of a clinical AI tool's architecture, not a feature you add. What the literature says it should look like, and how we built it.

Building 9 min read

How we built on-device de-identification so AI never sees real names

Most AI privacy is a policy. Ours is architecture: an NER model runs in the browser and strips names before anything leaves the device.

Technical 7 min read

Your agency's clients are about to ask why this costs so much

A solo consultant built in three weeks what your agency quoted twelve for. The client doesn't know why yet. The agencies that survive change what they sell.

Adoption 6 min read

What do you love doing? What do you hate doing?

Ask people what they love doing and what they hate doing, then show them AI is coming for the second list. Why the reframe works, and how it fails.

Technical 7 min read

Why I don't use n8n (and what I do instead)

n8n demos well. But a compelling demo and a reliable production system are different things, and the distance between them is where businesses get hurt.

Technical 10 min read

Your codebase was not built for AI. That's the actual problem.

Amazon's mandatory meeting about AI breaking production is an architecture story: codebases built for human maintainers only, now maintained by AI.

Adoption 4 min read

Your team has AI licences. You don't have an AI system.

Fifteen people, fifteen separate AI accounts, no shared context. The problem isn't the tool; it's the architecture around it. Here's the fix.

Building 7 min read

Your $2,000 day starts the night before: our system keeps you on the tools, not on the phone

Optimised routes overnight, automatic customer notifications, and promises the system keeps or corrects. A scheduling system that protects your daily rate.

Evaluation 4 min read

The fastest way for an executive to get across AI

AI moves faster than any executive can track. One focused conversation, one written report, and a decision you can act on: your time stays on the business.

Building 6 min read

Your IT department will take 18 months. You need this working by next quarter.

Senior leaders know what they need built; the gap is time. A prototype gets the tool working now and hands IT a validated blueprint for later.

Building 8 min read

We built an AI invoice verifier. Here's where it hits a wall.

We built an AI invoice verifier and watched a fake beat a real invoice. Why document analysis alone cannot stop fraud, and the five layers that can.

Building 5 min read

How to build an AI chatbot that doesn't lie to your customers

Woolworths scripted its AI to talk about its mother. The business fix is honesty; the technical fix is architecture that prevents fabrication by design.

Technical 9 min read

Why AI safety features are load-bearing architecture, not political decoration

The 'woke AI' label came from real failures, but they were engineering failures, not safety failures. The difference matters wherever errors have consequences.

Adoption 3 min read

Woolworths' AI told a customer it had a mother. That's a problem.

Woolworths' AI assistant Olive was scripted to talk about its mother and uncle. When callers realised, trust broke instantly. The fix is honesty.

Evaluation 5 min read

Google is no longer the only way your customers find you

Customers now find businesses through ChatGPT, Perplexity, and Gemini. The sites AI cites are structured differently to the sites Google ranks.

Evaluation 6 min read

The personal workflow analysis: what watching a real workday reveals about automation

People describe the work they value, not the work that eats their time. Recording a real workday reveals the automation opportunities interviews miss.

Evaluation 11 min read

An AI audit that starts with your business

How an operations-first AI audit works: what it looks for, how the evidence is collected, what the report contains, and what it tells you to skip.

Building 6 min read

What production AI teaches you that demos never will

The gap between a demo and a working system is where the useful lessons live. Architecture, framing, privacy, adoption: the patterns repeat every time.

Adoption 6 min read

The psychology of why your team won't use AI

You buy the tool, run the demo, and three months later nobody is using it. Five predictable psychological barriers, each with a strategy that works.

Technical 4 min read

Stop telling AI what NOT to do (and what to say instead)

Instructions built on prohibitions make AI cautious and generic. Describing what you want instead transforms the output, and the reason comes from psychology.

Building 5 min read

How we turned generic AI into a specialist: and what that means for your business

Mediocre AI output is rarely the model's fault. Three structural changes that turn the same model from generic to specialist-grade.

Evaluation 6 min read

Your business has 9 customer touchpoints. AI can fix the 6 you're dropping.

You pay to get customers to your door, then lose them to missed follow-up. AI can handle the six touchpoints most businesses drop.

Technical 6 min read

What happens to your data when you press 'Send' on an AI tool

Businesses send customer data to AI tools without knowing what happens during processing. The spectrum of AI privacy is wider than you think.