- What's happening when a model makes a video
- The places the weirdness hides
- A real example from a client shoot that never happened
- Step by step: how I check a batch of AI clips before anyone else sees them
- The uncomfortable bit nobody selling you a tool wants to say
- What's improving, and what isn't
- What to do about it right now
- Frequently asked questions
- Official documentation
The short version: AI video models predict pixels frame by frame based on patterns in training data, they don't understand physics, gravity, or object permanence, which is why hands warp, blinks stutter, and backgrounds drift. Some of this is improving fast, some of it is a structural limit of how these models work and won't be "fixed" by the next update. The fix for most people isn't waiting for better AI, it's knowing where the flaws hide and editing around them.
What's happening when a model makes a video
An AI video tool isn't filming anything. It's generating a sequence of images that are statistically likely to follow one another, based on billions of frames of real video it was trained on. There's no camera, no set, no actor who understands that a coffee cup has to stay the same size when you pick it up. The model is guessing, frame by frame, what pixels come next.
Most tools, including Runway Gen-3, Kling 1.6, and Luma Dream Machine, generate at a lower internal frame rate (often somewhere around 8 to 16 frames per second) and then interpolate up to a smooth 24 or 30fps for the final export. That interpolation is exactly why motion sometimes has a faint gliding quality, like the footage is floating half a second off from where your eye expects it. You're not imagining it. You're watching a model fill in gaps it never generated.
The places the weirdness hides
After running hundreds of generations for client work, the flaws cluster in the same few spots almost every time:
- Hands and fingers. Still the weakest point across every major model. Six fingers, fused knuckles, a thumb that vanishes mid-gesture.
- Teeth and eyes. Blinking that doesn't sync with speech, teeth that shift number or shape between frames.
- Object permanence. A logo on a mug that changes text halfway through a shot. A background chair that quietly disappears.
- Physics. Liquid pouring wrong, cloth that moves like it's underwater, shadows pointing two directions at once.
- Audio and lip sync. Mouths that move slightly out of time, or a voice that has that flat, slightly reverby quality even in a supposedly "quiet room."
If you want the deeper mechanics on why the sound side is often the giveaway before the visuals even are, I go into that in what makes AI videos sound weird and how to fix it, because audio artefacts are often easier to catch on a phone speaker than the visuals are on a laptop screen.
A real example from a client shoot that never happened
Last year I was helping a small coffee roastery client build a product teaser without booking a studio day. We generated 40 separate clips of "a hand pouring coffee beans into a grinder" using Runway. Out of 40, three were usable without obvious flaws. Two more were usable with careful cropping. The rest had at least one of: a hand with the wrong number of fingers, beans that fell upward for a frame, or a grinder that changed shape between the start and end of the clip.
That's roughly a 12% clean-usable rate on a fairly simple, low-motion prompt. Anything with faster movement, multiple people, or dialogue drops that rate further, sometimes to one usable clip in twenty. The client still saved money versus a studio shoot, but the time spent generating, reviewing, and discarding clips ate up almost as many hours as a short real shoot would have. That's the part nobody puts in the sales pitch for these tools.
Step by step: how I check a batch of AI clips before anyone else sees them
This is the exact process I run through before a client ever sees a generated clip:
- Watch it once at full speed, no sound. Your eye catches motion errors faster without dialogue distracting you.
- Watch it a second time at half speed on hands, mouths, and any text or logos. This is where 80% of the visible glitches show up.
- Check the first and last frame side by side. Anything that changed shape, colour, or position between the start and end of the clip is a red flag.
- Play the audio alone with your eyes closed. If it sounds flat or slightly robotic on its own, it'll sound worse once viewers are watching lips move too.
- Ask someone who hasn't seen the prompt to watch it cold. They'll spot the "something's off" feeling faster than you will, because you already know what to expect.
If a clip fails any of the first three checks, I don't try to fix it in the edit. I regenerate it. Trying to patch a warped hand in post costs more time than a fresh generation almost every time.
The uncomfortable bit nobody selling you a tool wants to say
Here's the part that gets glossed over: this isn't purely a "the tech is young" problem that a future update quietly solves. Models like Sora, Kling, and Runway are trained to predict plausible pixels, not to simulate a physical world. Making them understand that a hand has five fingers and that liquid falls downward requires either a totally different architecture or an enormous jump in compute cost that most companies aren't rushing to fund, because the current output is already good enough to sell subscriptions.
The honest expectation to hold isn't "this will be fixed by 2027." It's "some of this will keep improving steadily, and some of it is a permanent tell, the same way CGI in a 2004 film still looks like CGI even though the technology moved on." Planning your content around that reality saves you a lot of wasted regeneration credits.
What's improving, and what isn't
Hand rendering has gotten noticeably better over the past 18 months, closer prompts with fewer than five fingers showing up are rarer than they were in early 2025. Camera movement and consistency across longer clips (past 10 seconds) has also improved, models handle a static or slow pan far better than fast motion. What hasn't moved much: multi-person scenes with interaction (a handshake, two people passing an object), fine facial expression during speech, and anything involving text that needs to stay legible and consistent across frames.
If your content leans into a slightly surreal, floaty quality rather than fighting it, some brands have leaned into that look deliberately, there's a good breakdown of why that dreamlike effect happens and when it works in your favour in why AI generated videos sometimes look like dreams.
Want AI doing the heavy lifting in your marketing?
I build the systems that handle the boring 80 percent, so you get your week back. Done properly, with the human kept in.
What to do about it right now
If you're publishing AI generated video for a business, three practical habits cut the "weird" feeling down a lot:
- Keep shots under 6 seconds where possible, longer clips give the model more chance to drift.
- Avoid close-ups on hands, mouths, and text unless you've generated enough variations to be picky.
- Add real audio (voiceover, music, sound effects) rather than relying on generated speech, since that's usually the fastest tell.
If you're already editing the footage afterward, there's a proper walkthrough of the enhancement side, colour grading, stabilising, and layering fixes without making the final cut look more obviously fake, in how to enhance a video using AI tools without making it look fake. And if you're weighing up whether the visible flaws matter for where you're planning to publish, I've also covered how platforms treat this content in whether uploading AI videos to YouTube affects monetisation, because the rules there have shifted more than most creators realise.
For a full technical rundown on the visual and audio side together, my longer piece on why AI generated videos can look and sound off goes into the render pipeline in more depth than I've covered here.
Related reading: instagram why are some videos brighter.
Frequently asked questions
Will AI generated videos stop looking weird eventually?
Some flaws, like hand rendering and short-clip consistency, are improving steadily year over year. Others, like true physical understanding and multi-person interaction, are tied to how these models fundamentally work and are unlikely to disappear soon, so expect gradual improvement rather than a sudden fix.
Why do AI video hands still look wrong in 2026?
Hands have many small moving joints that change shape constantly, and the model has to predict each finger's position frame by frame with no understanding of anatomy. It's improved noticeably since early 2025 but remains the single weakest point across every major tool.
Is it worth using AI video for a small business instead of filming real footage?
For simple, short, low-motion shots (product close-ups, background b-roll, social teasers) it can save real money and time. For anything involving hands, dialogue, or multiple people interacting, expect to generate far more clips than you keep, sometimes as few as one usable clip in every eight to twenty attempts.
What's the fastest way to spot a fake AI clip before publishing it?
Watch it at half speed with the sound off, focusing on hands, mouths, and any text or logos, then compare the first and last frame side by side. Most visible errors show up in one of those two checks before a viewer ever notices them at normal speed.