AI and video, Mid-2026: the models can watch now, not just listen
AI could always transcribe video. It can now read the frames as well, and every hour of footage a business owns becomes something it can question.
For most of the last few years, “AI and video” meant one thing: transcription. Feed in a recording, get back the words. Useful, and completely deaf to everything that was not spoken. The model was listening. It could not see.
That has changed, quietly, over the past year. The same models that read documents can now look at what is on the screen: the slide, the diagram, the product, the thing someone pointed at and said “this bit here.” It is a smaller technical leap than it sounds and a much bigger practical one. This is a note for any operator sitting on a library of recorded footage: production shops, marketing teams, practices full of recorded meetings, trades with site walkthroughs on a phone. What has actually shifted, where it lands, and where it does not yet.
What’s stable
The craft is untouched. A model can log footage; it cannot judge which moment carries a piece, feel the beat that stops a viewer scrolling, or make the creative call a client pays a named person for. Whatever changes below, the judgement layer of video work stays human, and businesses whose value is the visual craft keep that value.
The economics of the archive have also been stable for years, in the bad sense. Most of working life is now recorded and never looked at again: every video call, every webinar, every training session someone captured “so we have it.” That footage sits on drives as dead weight, because nobody has three hours to scrub back through a two-hour call for the one exchange that mattered. The recorded hour has long been the most expensive thing to get value back out of. That is the fact the current shift attacks.
What’s changed
The words were already solved. Transcription is mature, cheap, and mostly free: platforms hand over caption tracks, and speech models cover anything without one. This is the part every tool already had, and it describes the audio only.
The pictures are the new part, and the engineering detail matters. To let a model see a video, the video has to become still frames, because that is what a vision model reads. Tools that grab a frame on a timer, one every thirty seconds or so, miss almost everything on edited footage; a slide that changed four times in that window reads as one slide. The stronger approach is scene-change detection: capture a frame every time the shot actually changes, so the sequence of frames matches the sequence of decisions someone made in the edit. The model then reads the video roughly the way an editor sees it, as deliberate moments rather than a slideshow on a metronome. That one architectural choice separates a tool that understands footage from one that guesses, and it is a useful test to put to any vendor in this category.
Frames and transcript read together are the unlock. Pair the scene frames with the words and hand both to a model that reasons over text and images jointly, and the questions change from “what was said” to “what was shown, when, and does it match the claim.” Neither stream alone can answer that.
The capability is assembled from open parts. This is not a product category a business has to wait for. The pieces are ordinary and public: a downloader for the video and captions, scene detection for the frames, a speech model as fallback, a vision-capable model to read it all. The open-source skill that demonstrated the pattern (claude-video, by Brad Bonanno, since extended by others) runs inside tools like Claude from a pasted link. For a business already running AI in its workflow, this is a small setup, not a project.
Where it lands
The narrow use is the obvious one: a production shop letting the machine do the logging, captioning, and platform versioning that consumes the hours between shoots. Real, and the smallest of the applications.
The broader use is the library. A marketing team sitting on a decade of shoots and event footage can finally find the usable thirty seconds without opening every file. A recorded onboarding session becomes a searchable set of steps. The quarterly all-hands becomes something staff can ask, rather than re-watch. A builder’s site walkthrough becomes a dated, described record that can be queried when there is a dispute. A researcher or lawyer reviewing hours of interview or camera footage gets a first pass that flags where to actually look. The pattern underneath is one idea: any pile of video a business owns is now a pile it can question, and an archive that was dead weight starts behaving like inventory.
Where it doesn’t yet land
It sees stills, not motion. Even with good scene detection, the model reads snapshots and infers what happened between them. Continuous motion, fast physical action, a gradual change with no clean cut: these fall through the gaps. It is watching a flip-book, not the film.
Scene detection is a heuristic, not comprehension. A camera pan or lighting shift can spawn frames nobody needs, while a long static talking-head with critical detail on a shared screen might barely trigger at all. Both directions are the tool guessing at where the meaning is.
It misreads with confidence. A vision model can misread text on a slide or describe something that is not quite there, fluently. The failure mode is not a blank “I don’t know”; it is a plausible wrong sentence. On anything that matters, the output gets verified against the source, which is the same verification rule that applies to every AI capability worth using.
Long footage costs real time and money, caption quality varies, and the speech-model fallback adds its own errors on top.
And there is a privacy decision underneath all of it. Turning footage into frames and sending them to a model means client faces, shared screens, and whatever was visible in the background now travel to wherever that model runs. For a public webinar, that is nothing. For a recorded client session or an internal call, it is a decision to make before the upload, not after. The considerations are the same as for any data pressed into an AI tool.
What this means for the operator’s posture
The sensible first move is small and low-stakes: point the capability at footage that could not hurt anyone, a recorded webinar, an old training video, a public talk, and ask it something specific. Ten minutes of that teaches more about where it is strong and where it is soft than any amount of reading. From there, the question is not “which video tool should we buy” but the same question as everywhere else in applied AI: which pile of recorded material does this business own, what would it be worth if it could be questioned, and what has to be true about privacy and verification before that is allowed to happen.
The interesting part was never the tool. It is that the recorded hour, which used to be the most expensive thing to get value back out of, is quietly becoming one of the cheapest.
If you are working out where a capability like this fits in how your business actually runs, start with a conversation.