Asset 20 8 2
Does AI recommend your business? Run the free check →

Join 15,000 business owners, marketers and entrepreneurs. The Sunday newsletter you'll be annoyed only arrives once a week.

Article

Why AI Generated Videos Can Look and Sound Off (And What’s Really Going Wrong)

The short version: AI generated videos look and sound off because the model is predicting pixels and sound waves frame by frame with no underlying understanding of anatomy, physics, or speech, so tiny inconsistencies (a blink rate that’s too low, a mouth shape that doesn’t match the sound, a shadow that moves the wrong way) stack up across a clip until your brain flags it as wrong even if you can’t say why. The fix isn’t a magic prompt. It’s knowing exactly where these tools fail and checking those specific spots before you hit publish.

The blink test I use on every AI clip

Last March I built a 14-second product demo for a client pitch using one of the newer text-to-video tools. Small thing, a founder “talking” to camera about a new feature. I watched it four times before I worked out what was bothering me. She wasn’t blinking enough.

Humans blink between 15 and 20 times a minute on average, sometimes more when we’re talking. The AI-generated founder blinked twice in fourteen seconds. That’s roughly a third of the normal rate, and it’s exactly the kind of gap your eyes catch before your brain can explain it. I’ve since made this the first thing I check on any AI video before it goes anywhere near a client: count the blinks, roughly, in the first ten seconds. If it’s under four or five, something’s wrong with the source clip and no amount of colour grading fixes it.

This is a known limitation of diffusion-based video models. They’re trained on huge volumes of footage and they’ve learned that faces mostly look a certain way, but blinking is a fast, irregular motion that gets smoothed out or dropped entirely because the model is optimising for visual coherence between frames, not biological accuracy.

Why the eyes and mouth never quite agree

The next thing that gives AI video away is lip sync, and it’s a slightly different problem to the blinking one. Most tools generate the video and audio as separate passes, or generate video first and then map audio onto it (or vice versa with tools like HeyGen and Synthesia, which build talking avatars from a photo plus a voice track). That means the model is guessing which mouth shape goes with which sound, using phoneme prediction, and it gets close but rarely exact, especially on hard consonants like “b,” “p,” and “f” where your lips need to close.

I tested this for a piece I wrote on what makes AI videos sound weird and how to fix it, and the pattern held across every tool I tried: wide vowel sounds sync fine, closed-mouth consonants lag by a few frames. Your ear hears “impossible” and your eye sees a mouth that’s still open. That mismatch, even at 100 milliseconds, is enough for viewers to describe the video as “creepy” without knowing why.

The audio problem is a separate fight entirely

People assume the voice is the easy part because text-to-speech has existed for decades. It’s not the easy part anymore, it’s just a differently broken part. AI voices, even the good ones from ElevenLabs or Play.ht, tend to get three things wrong: breathing, pacing, and emotional consistency.

Real human speech has micro-pauses for breath roughly every 8 to 12 words. AI voice generation often skips these entirely or places them at grammatically odd points, because the model is predicting the next likely sound, not managing lung capacity. The result is a voice that sounds slightly too smooth, like someone reading without ever needing to breathe, which is unsettling in a way most people can’t name.

Pacing is the second issue. AI voiceover tends to keep an almost identical rhythm across a whole script, sentence after sentence at the same clip, where a human narrator naturally speeds up on simple points and slows down on important ones. And emotional consistency breaks down over longer scripts, a voice that sounds warm in the first ten seconds can flatten into something closer to a weather report by second forty, because the model isn’t tracking emotional intent across the whole piece, it’s generating in chunks.

Hands, teeth, and the small objects nobody notices until they do

Faces get most of the attention because they’re what we look at first, but hands are where AI video still falls apart most obviously. Fingers, in particular. I’ve had clips where a hand reaching for a coffee cup grew a sixth finger mid-motion, and clips where a handshake simply merged two hands into one shape for three frames before separating again.

This happens because hands have far more possible positions than faces do, and far less consistent training data showing hands doing everyday things at the angles video needs. Teeth have a similar problem, they tend to look slightly too uniform, too white, or shift subtly in number and shape as a mouth opens and closes across a talking sequence.

Background objects cause quieter trouble. Text on a sign in the background of an AI clip is very often garbled, because the model has learned what text generally looks like (letters, spacing, that a sign has words on it) without learning what any specific word says. Clocks are a good tell too. If there’s a clock in an AI-generated scene, check whether the hands move in a way that makes sense, and check the reflections in windows and mirrors, they very often don’t match what should be reflected.

The bit most people writing about this skip

Here’s the part that’s uncomfortable to say if you sell AI video tools for a living: better prompting does not fix most of this. I see a lot of advice telling people to add words like “cinematic,” “natural blinking,” or “realistic hand movement” into their prompts, as if the model has a switch marked “make it more human” that’s just been left off. It doesn’t work that way. These models don’t understand blinking or hands or breath as concepts you can dial up, they’re predicting the statistically likely next frame based on training data, and the flaws in that data (rare hand angles, inconsistent blink patterns, imperfect audio-visual alignment) are baked into every generation regardless of how the prompt is worded.

What happens when you add those words to a prompt is you shift the output slightly, sometimes for the better, often at random, and you can convince yourself it’s working because you’re comparing two flawed clips rather than a flawed clip against a realistic benchmark. I’ve watched marketing agencies charge clients extra for “prompt engineering” that, when I tested it side by side with plain prompts across a batch of forty clips for a workshop I ran, produced no measurable difference in blink rate, lip sync accuracy, or hand consistency. The honest fix is choosing shorter clips, tighter framing, and accepting the current limits, not paying someone to whisper better instructions at the model.

A five-step check before you publish anything AI-made

This is the checklist I run through now, takes about ninety seconds per clip:

  • Watch the first ten seconds and count blinks. Fewer than four or five and the clip will read as unnatural to most viewers.
  • Mute the video and watch only the mouth on hard consonants like “b” and “p.” If the lips never fully close, the sync is off.
  • Watch the hands specifically, in slow motion if your player allows it, especially any moment someone picks something up or gestures.
  • Listen for breath sounds and natural pauses. If the voice sounds like it’s reading without stopping for air, viewers will pick up on it even if they don’t consciously notice why.
  • Check any background text, clocks, mirrors, or reflective surfaces. These are the easiest tells and the ones people forget to look at.

If a clip fails two or more of these, I don’t publish it as is. I either regenerate with a tighter shot (close crops on the face hide hand and background issues), trim it shorter, or swap to real footage for that section.

What helps right now

Some practical things do reduce the “off” feeling, even though they don’t eliminate it. Shorter clips help enormously, because most of these flaws compound over time; a five-second clip has far less chance to drift than a thirty-second one. Tighter framing on the face, cutting out hands and busy backgrounds, removes two of the biggest tell categories entirely. Using a real recorded voice for anything client-facing and saving AI voice for internal drafts or lower-stakes content is a trade worth making, at least for now. And picking the right tool for the job matters more than most people think, which is why I wrote up where to make AI videos as a small business owner after testing the main options myself rather than going off marketing claims.

If budget is the constraint, there are usable free options too, though the trade-offs are different, longer render queues, watermarks, tighter length caps, and I’ve laid those out in where to create AI videos for free in 2026. Worth knowing before you commit hours to a script that a free tier won’t let you export cleanly.

It’s not just video, and that matters for how you use it

The same underlying issue, models predicting the next likely token or pixel rather than understanding what’s true, shows up in AI writing too. If you’ve ever read an AI-written paragraph that felt slightly hollow despite perfect grammar, that’s the written version of the blink problem. I go into the specific tells in what AI writing looks like compared to human writing, and the patterns rhyme with what’s happening in video: technically plausible, statistically likely, and missing the small irregularities that real human output has by default.

There’s also a resource cost to all this generating and regenerating that most people don’t think about until someone asks. If you’re running dozens of clips through a tool trying to get one usable version, which is common, it’s worth knowing what that draws in energy and water, covered with real figures in does making AI videos use a lot of energy and water.

Where this is heading

These tools are improving fast, month to month rather than year to year, and some of what I’ve described here (the blink rate gap especially) will likely close within the next twelve to eighteen months as training data and model architecture improve. But the fundamental issue, a model that predicts plausible frames rather than understanding a face, a hand, or a breath, isn’t going away with the next update. It’s a structural feature of how these systems work, not a bug waiting for a patch. Which is exactly why the checklist matters more than chasing the newest tool. Learn to spot the five tells, and you’ll catch the off feeling before your audience does, whichever model you’re using in six months’ time.

Related reading: ai videos look like dreams.

Frequently asked questions

Why do AI generated faces look almost right but still feel wrong?

Because the model has learned the general shape and structure of a face extremely well but hasn’t learned the small irregular behaviours, blink timing, breathing, subtle muscle movement, that make a face read as alive rather than rendered. Those tiny gaps are what your brain flags even when nothing looks obviously broken.

Can better prompts fix AI video lip sync issues?

Not reliably. Prompt wording shifts the output slightly but doesn’t correct the underlying gap between how the audio and video are generated, which is usually two separate prediction processes being mapped together. Shorter clips and tighter face framing help more than prompt tweaks do.

Why do AI voices sound slightly off even when the words are correct?

Mostly because of missing or misplaced breath pauses and unnaturally even pacing across a script. Real speech varies rhythm and includes breathing roughly every 8 to 12 words. AI-generated voice tends to skip or misplace these, which reads as smooth but slightly inhuman.

What’s the fastest way to check if an AI video will look off to viewers?

Watch the first ten seconds and count blinks (fewer than four or five is a red flag), then mute it and watch the mouth on hard consonants, then check the hands and any background text or reflections. That ninety-second check catches most of the common tells before you publish.

Official documentation

Published and maintained by the Lilach Bullock team, covering marketing, AI and business growth.
Your buyers are asking AI who to use. Does it say you?

See for free whether ChatGPT, Claude, Perplexity, Gemini and Google name you, and get the plan to become the answer.

Check my AI visibility →
Sundays only

Get the Sunday newsletter.

One email a week. AI experiments, marketing tactics, and the workflows Lilach is building right now in her own business.

Subscribe free

Let’s get your marketing running on AI.

Book a free 30-minute call

We figure out what you need, where AI fits in, and what working together would look like.

Book the call →

Or take the 30-second calculator

You’ll see the hours and the money quietly leaking out of your week, and the three workflows worth building first.

Take the calculator →

Or grab the free AI resource library

Prompt packs, templates, checklists, and swipe files. The exact tools I build for paying clients. Yours, free.

Get the library →
Keep reading

More from the blog.