Asset 20 8 2
Does AI recommend your business? Run the free check →

Join 15,000 business owners, marketers and entrepreneurs. The Sunday newsletter you'll be annoyed only arrives once a week.

Article

Why Do AI Generated Videos Sometimes Look Like Dreams?

Straight answer: AI generated videos look dreamlike because the models making them have no real memory of what a hand, a dog, or a candle flame is supposed to do from one frame to the next. They are guessing pixel patterns from a compressed version of millions of videos rather than running an actual physical world underneath. That gap between “looks right for half a second” and “behaves like a real continuous object” is precisely the mechanism that makes our own dreams feel off in the same way.

There is no skeleton under the pixels

When you watch a normal film, every object in it obeys physics because it was, at some point, a real thing in front of a real camera. A dog’s legs stay attached to the dog. A wine glass keeps its stem. AI video tools like Runway, Luma Dream Machine, Kling, and Pika don’t start from a real dog or a real wine glass. They start from noise, and then a diffusion model reworks that noise into something that statistically resembles the training data it learned from, over roughly 25 to 50 denoising steps per clip, before decoding it back into pixels you can watch.

There is no bone structure holding the dog together between frame 14 and frame 15. There is only a prediction, and predictions drift. A slightly different prediction between two consecutive frames is what your eye reads as a leg briefly disappearing, a face reshaping, or text on a sign warping into nonsense. The model isn’t animating an object, it’s re-imagining the whole scene many times a second and hoping the re-imaginings line up closely enough that you don’t notice.

What happened when I tested four tools on one simple prompt

I ran a small test in January for a client who wanted a faceless walking-dog clip for an Instagram Reel, the kind of content I’ve written about before when covering faceless social media video production. I put the exact same prompt, a golden retriever trotting across a park at golden hour, into Runway Gen-3, Luma Dream Machine, Kling, and Pika, and generated five clips on each, twenty clips total.

Three of the four tools produced at least one clip where the dog briefly had five legs mid-stride before correcting itself two frames later. One clip from Luma turned the dog’s tail into a second, smaller dog for about half a second, then merged it back. In another, the grass texture in the background pulsed slightly, like it was breathing. None of this was a glitch in the technical sense, nothing crashed and no error appeared. The model was doing exactly what it’s built to do, generating its best guess frame by frame, and the guesses simply didn’t stay locked to one consistent object the whole way through.

The clip we ended up using was the one with the least motion, a slow pan across the dog sitting rather than walking. Static and near-static shots hold up far better than anything with full-body movement, because there’s less for the model to keep track of between frames.

Why this specifically reads as dreamlike rather than just broken

A jammed video codec looks broken, blocky, obviously wrong. AI video artifacts don’t look broken, they look wrong in a way that feels familiar, and that’s the interesting part. In an actual dream, a room can quietly become a different room, a person’s face can shift without you registering the moment it happened, and objects can behave with slightly wrong gravity, and your dreaming brain accepts all of it without alarm bells going off. Sleep researchers have described this quality of dream cognition, the smooth continuity despite logical impossibility, for decades, and you can read a good overview of the underlying research on Wikipedia’s dream entry.

AI video does the same thing to your visual system. The lighting is consistent, the colour grading is consistent, the general composition is consistent, so your brain classifies the whole clip as “real footage” and only afterwards flags the impossible bit, the extra finger, the morphing tail, the sign that says three different words in three consecutive shots. That delay between acceptance and doubt is the exact texture of a dream.

The part nobody selling these tools wants to say out loud

Here’s the uncomfortable bit. The companies building these video generators are not racing to eliminate the dreamlike wobble, because a good chunk of what makes AI video get watched, shared, and talked about is that wobble. A physically perfect two-second clip of a dog walking is unremarkable, people scroll past it in half a second. A dog whose legs briefly multiply and then correct themselves is the sort of thing people stop for, screenshot, and send to a friend with “look at this.” I have watched a client choose the glitchier of two versions on purpose because it pulled three times the shares on Reels compared with the cleaner, more stable take.

That’s uncomfortable if you’re trying to sell AI video as a mature, reliable production tool, because it means some of what people find compelling about it right now is the very thing the labs are technically trying to fix. Fixing it fully might make the output more useful for serious commercial work and, at the same time, quietly less interesting to watch. Nobody wants to put that in a product launch email.

Step by step: reducing the dream effect in your own clips

You can’t remove this entirely with today’s tools, but you can push a clip a long way toward the stable end of the spectrum. This is what I do now before I generate anything for a client:

  • Keep single generations to 3 to 5 seconds. Longer clips give the model more frames to drift across, and drift compounds.
  • Favour close, static, or slow-panning shots over full-body walking, running, or hand movement, which are the hardest thing for these models to hold together.
  • Use start-frame and end-frame image anchoring where the tool supports it, so the model is solving for two fixed points instead of guessing freely across the whole clip.
  • Generate 4 or 5 variations of the same prompt and pick the most stable one rather than accepting the first result, which is standard practice on every shoot I run now.
  • Add explicit negative prompting, phrases like “no extra limbs, consistent object count, no morphing,” which several tools now honour.
  • Cut every 2 to 3 seconds in the edit rather than letting one AI clip run long, because a cut hides drift far better than any prompt trick does.

If you’re weighing up which tool to build a workflow around, I broke down the strengths of the current crop in more detail in my piece on AI video tools you can use to make high-converting videos, which covers where each one is strongest and where each one still slips.

Use the dreaminess, or fight it, depending on the goal

Not every use case needs this fixed. A brand mood film, a music video, an art piece, a fashion reel, all of those can lean into the surreal quality and it reads as stylistic choice rather than error. What you should not do is put a wobbly, morphing AI clip on a landing page or a checkout flow and expect it to build trust. I talked through this exact tension between what gets attention and what moves people to buy in my conversion rate optimisation webinar for Simplilearn, and the principle holds here too: a clip that gets stares is not the same as a clip that gets conversions, and mixing the two up costs businesses money.

This is the same trap I’ve raised before around vanity metrics generally. Views and shares on a strange AI video feel like proof it worked, but I said on the C-Sweet Talks episode on engagement not equalling ROI that attention without intent behind it rarely turns into revenue. A dreamy, glitchy clip might be exactly right for top-of-funnel brand awareness on TikTok and completely wrong for a product demo further down the funnel where someone needs to trust what they’re seeing before they hand over card details.

Where this gets better, and how slowly

The labs building the next generation of these models, OpenAI’s Sora, Google DeepMind’s Veo line, and the teams behind Kling and Runway, are all working on what’s loosely called world modelling, giving the system a rougher sense of physical consistency, object permanence, and cause and effect, rather than pure frame-by-frame guessing. Google’s own Veo 2 research claimed noticeably better physical consistency across 8-second clips compared with its predecessor, and it is a real improvement. I still watched a wine glass lose its stem mid-pour in one of my own tests with it a few weeks ago.

My honest read, after two years of testing this stuff for client work, is that the dreamlike quality will fade for short, simple, low-motion clips well before it disappears for anything long, complex, or full of moving limbs. Budget for another 18 to 24 months before AI video handles a five-second clip of two people shaking hands without one of the hands briefly doing something a hand can’t do. If your business is building a proper video pipeline around this and you don’t want to burn a quarter figuring it out through trial and error the way I did, that’s exactly the kind of practical build-out I help clients with through hands-on AI implementation coaching.

I talked more broadly about where AI is pulling marketing and sales next on the UNmiss podcast episode on the future of marketing and sales, and video is the clearest example of a tool that’s brilliant for some jobs and unfinished for others right now. Knowing which job you’re pointing it at is most of the battle.

Frequently asked questions

Why does AI video look weird when the still frames look fine?

Because each frame is generated as its own best guess rather than as a locked continuation of the last one, so a hand or a face can look perfectly normal in isolation while shifting slightly between frames in a way that only becomes visible once the frames play in sequence.

Can you fix the morphing and glitching with better prompts?

You can reduce it a lot with shorter clips, less full-body motion, start and end frame anchoring, and negative prompts telling the model to avoid extra limbs or object duplication, but you can’t remove it fully with the current generation of tools.

Is the dreamlike look ever a good thing to keep?

Yes, for brand mood pieces, art content, and music videos it often reads as an intentional stylistic choice, but for product demos, ads, or anything that needs to build trust in a viewer, it usually undermines the message rather than helping it.

Which AI video tool currently produces the most stable, least dreamlike footage?

In my own testing, close static or slow-pan shots on Runway Gen-3 and Google’s Veo 2 held up the most consistently, while full-body movement caused visible drift on every tool I tried, including those two.

Published and maintained by the Lilach Bullock team, covering marketing, AI and business growth.
Your buyers are asking AI who to use. Does it say you?

See for free whether ChatGPT, Claude, Perplexity, Gemini and Google name you, and get the plan to become the answer.

Check my AI visibility →
Sundays only

Get the Sunday newsletter.

One email a week. AI experiments, marketing tactics, and the workflows Lilach is building right now in her own business.

Subscribe free

Let’s get your marketing running on AI.

Book a free 30-minute call

We figure out what you need, where AI fits in, and what working together would look like.

Book the call →

Or take the 30-second calculator

You’ll see the hours and the money quietly leaking out of your week, and the three workflows worth building first.

Take the calculator →

Or grab the free AI resource library

Prompt packs, templates, checklists, and swipe files. The exact tools I build for paying clients. Yours, free.

Get the library →
Keep reading

More from the blog.