What AI Still Can't Do in Video
Every claim here comes from a model's own documentation, a vendor's terms of service, or published research — sources with no incentive to understate what these tools can do.
Generative video has had two extraordinary years, and the demonstrations keep getting better. What follows is not scepticism for its own sake: every claim below comes from a model's own technical documentation, from a vendor's terms of service, or from published research. In other words, from sources with no incentive to understate what these tools can do.
The gap between a curated showreel and a working day is still wide, and knowing exactly where it sits is what separates a usable tool from a disappointment.
The eight-second shot
The first limit is the easiest to verify because the vendors publish it themselves. Veo 3.1, the current version of Google DeepMind's model since November 2025, generates shots of four, six or eight seconds at 24 frames per second, in two aspect ratios only. Runway Gen-4.5, announced in December 2025, produces two to ten seconds at 720p. Kling 3.0 reaches fifteen seconds of continuous video at 1080p.
What the industry calls an AI-generated film is therefore an assembly of fragments, none of which reaches the length of an ordinary cinema shot. Editing, continuity and direction remain entirely human, and that is where most of the working time goes.
The resolution figures are worth reading closely too, because they contradict the marketing language around "4K cinema". The model Runway positioned at the top of comparative rankings at the end of 2025 outputs 720p. Veo 3.1 does offer 4K, but its fast variant stops at 1080p and the reference images you feed it are capped at 1080p.
The most telling event of the year, though, is a withdrawal. OpenAI shut down the Sora application on 26 April 2026 and is retiring the corresponding API on 24 September, stating that it wants to concentrate its compute and its research team elsewhere, specifically on world simulation for robotics. The most discussed video product of the period lasted under eighteen months in consumer form.
Only one vendor publishes its failures
Of the major generators, exactly one documents what its model gets wrong. Runway publishes a section headed "Limitations" on its research page, and it is more severe than an outside observer would be.
Three defects are named. Causal reasoning: effects sometimes precede causes, a door opening before the handle is pressed. Object permanence: objects disappear or appear unexpectedly between frames, a cup vanishing after being occluded. And success bias: actions succeed disproportionately, a poorly aimed kick still scoring the goal.
That third item deserves more attention than it gets. Generated video rarely shows a fumble, a hesitation or a failure, because the training material is full of things going right. The consequence is not merely aesthetic — it rules out an entire category of shots, the ones where something does not work. For documentary, instructional or product-honest advertising, that is a hard boundary.
Runway adds that these limitations are common to AI video generation models. It is the vendor generalising, not a critic, and it is what makes the page valuable: everyone else documents capability and stays silent on failure, which means the knowledge of limits comes from academic research rather than from manufacturers.
What the research measures
Evaluation has changed character over the last two years. The first generation of benchmarks scored image quality, motion smoothness and subject consistency within a single shot. Those dimensions are now largely saturated, and the field has moved to what the VBench 2.0 project calls intrinsic faithfulness: not whether the image is plausible, but whether it is physically and logically coherent.
The dimensions added in that second generation read like a list of what still resists — human anatomy, the persistence of a character's identity and clothing over time, the ordering of movements, interaction between people, mechanical and thermal state changes, geometric consistency across viewpoints, and the plain rationality of motion. Scores on complex plots and on dynamic spatial relationships remain very low across every model evaluated.
A study published in June 2026, covering twenty-three models and several thousand human-annotated videos, ran a deceptively simple protocol: the camera turns away, then comes back. Has the object continued to evolve while unobserved? Usually not. It stays frozen in the state it was left in, disappears, reappears in the wrong place, or regresses to an earlier condition. The authors note that the failure crosses every model family and every parameter scale, and that adding parameters does not fix it — in several cases the larger model performs worse than the smaller one.
That result lines up precisely with the object permanence limitation Runway acknowledges, and together they state the underlying difference between these tools and what people expect of them: a generator produces sequences of plausible images, it does not maintain a state of the world.
One caveat on scope. The models in these studies are mostly open-weight systems plus a few commercial ones; neither Veo 3.1 nor Gen-4.5 appears in most of the protocols. It would be wrong to derive a failure rate for frontier models from them.
On-screen text and non-English dialogue
Legible text inside the frame remains a documented weakness. A benchmark devoted to that single question, published in May 2025 and based on human evaluation of ten open and commercial systems, concludes that most models struggle to produce text that is both readable and stable from frame to frame. For an advertisement where a brand name must stay identical for eight seconds, that is disqualifying, and the workaround is to composite the text afterwards.
Lip synchronisation outside English calls for more care, because the sources are thin and frequently misread. What is established is narrow. Kling documents five languages for its native audio — Chinese, English, Japanese, Korean and Spanish. Google's technical sheet for Veo 3.1 lists English only, but as the prompt language rather than the language of generated dialogue, and the distinction matters. A benchmark on multi-speaker dialogue published in February 2026 identifies lip-sync errors as the most frequent failure type in its corpus, which is entirely English.
What does not exist is any public evaluation of lip synchronisation in French, German or Italian on commercial generators. It would be tempting to conclude that non-English languages are structurally handicapped, and a research paper from October 2025 covering twelve languages shows the opposite: an architecture aligning phonemes to visemes delivers stable performance across all of them. The problem is not intractable. It is simply not solved in shipping products.
Who owns the output
The question looks secondary until you invoice a client, and it has different answers depending on the vendor. The gap between marketing language and contract language is measurable.
Runway is the clearest and the most permissive: users retain ownership of what they upload and generate, commercial use is explicitly allowed, from monetised video to advertising, and no credit is required.
Kling is far more restrictive than most secondary guides claim. Its terms of service, updated in April 2026, do recognise the user's intellectual property in inputs and outputs — and then state that without written permission the user may not use, reproduce, distribute or create derivative works from that output for commercial purposes. The company also reserves a non-exclusive, royalty-free licence over the content, and requires attribution on public distribution. You own the video without being able to exploit it.
For Google and OpenAI the position is less documented than one might expect, since applicable terms depend on the access channel, and we found no public clause squarely settling ownership of Veo outputs. Sora's shutdown raises a question nobody has answered: what becomes of the rights in, and the usability of, videos generated on a platform that closes and deletes its data.
Marking has been mandatory in the EU since August
The transparency obligations of the European AI Regulation have applied since 2 August 2026, which makes them an operational constraint rather than a forecast. Article 50 requires providers to mark synthetic content in a machine-readable way, and deployers to clearly label content that resembles real people, objects, places or events in a way that would falsely pass as authentic.
There is relief for work that is manifestly artistic, creative, satirical or fictional, where the obligation narrows to disclosing the existence of generated content in a way that does not impair the enjoyment of the work. Manifestly impossible content falls outside the deepfake definition altogether. An advertisement showing a real product in a credible situation sits squarely inside the scope.
Two technical schemes carry the marking, and both have limits worth knowing. Content Credentials, promoted by the C2PA coalition, wrap signed metadata into the file; adoption is broad, but ordinary web pipelines strip metadata or recompress files, which destroys the container, and competing versions of the standard create incompatibilities between manufacturers.
The deeper problem was demonstrated in August 2026, when a researcher published a plainly fabricated photograph carrying cryptographically valid Content Credentials — a certificate chain resolving to the coalition's own trust list, issued by Google for Pixel cameras, with a signed timestamp. Google closed the report as "Won't Fix (Infeasible)" and classed it as not security-bulletin material. The case concerned a still image rather than video, which is worth stating precisely, but the principle carries: a cryptographic signature attests which device signed a file, not that its contents are true.
SynthID, Google DeepMind's invisible watermark, is designed to survive cropping, frame rate changes and lossy compression, and has been applied to more than ten billion images. Its detector, however, is still not publicly available — access runs through a waiting list. Google states in its own paper that the technique will not solve the problem alone and remains vulnerable to regeneration attacks, and the published robustness figures cover images rather than video.
What this changes in practice
None of this argues for avoiding these tools. What it argues for is knowing the perimeter.
Short, visually convincing shots on subjects that do not require long-range coherence are well within reach: an atmosphere, a camera move across a landscape, an illustrative sequence. Holding an identical character across a dozen shots, producing legible on-screen text, showing a gesture that fails, or guaranteeing that an object leaving frame returns in the state it left — none of those are solved. Editing, art direction and verification remain human, and they account for most of the schedule.
Frequently asked questions
Can you keep a character consistent from shot to shot?
Partially, through reference-image techniques several models offer. But no public benchmark currently measures character consistency across distinct shots of the same edit, which happens to be the central problem in professional use. The absence of a standard evaluation on that point is itself informative.
Do these models improve simply by getting bigger?
Not on the dimensions that cause trouble. Several independent studies report that physical coherence and state persistence do not improve with parameter count, and occasionally regress. The authors read this as requiring an architectural change rather than more scale.
Must AI-generated advertising be labelled in the EU?
Since 2 August 2026 the transparency obligations of the AI Regulation apply, and an advertisement featuring people, places or events of apparently real character falls under the labelling requirement. The relief for manifestly creative work does not cover content designed to look authentic.
Does an invisible watermark prove a video is genuine?
No, and the confusion is common. A cryptographic mark or watermark attests to the technical origin of a file, not to the truth of what it shows — a point demonstrated in August 2026 by a fabricated image carrying a valid camera signature. Google's watermark detector is also not open to the public, which leaves ordinary verifiers without a means of checking.
Further reading
- LED Wall Virtual Production: The Tech Behind The Mandalorian — how in-camera VFX on an LED volume actually works
- Green Screen and Chroma Key: How to Get a Clean Key — the four decisions that make or break a chroma key
- Why Google Isn't Indexing Your Videos — the rule that decides whether Google indexes your video