Straight answer: yes, generating an AI video uses noticeably more energy and water than sending a text prompt, and the difference isn’t small. Some researchers put a single short AI-generated video clip in the same energy ballpark as running a microwave for over an hour, though the exact number depends on the model, the length, and how many times you hit “regenerate.” The bit almost nobody tells you is that your drafts, not your final video, are where most of that cost sits.
Why video costs more than a chat message
I’ve spent a fair chunk of the last year making short AI videos for client campaigns, so this isn’t theoretical for me. Text generation is quick because the model is predicting the next word, one token at a time, using a relatively small amount of compute per response. Video generation is a different beast entirely. The model has to predict thousands of frames, keep objects consistent across those frames, render motion, lighting and sound, and do it all through multiple passes of a much larger model. That takes far more GPU time, and GPU time is what turns into electricity, and electricity is what turns into heat, and heat is what the data centre then has to cool down, often using water.
Think of it like the difference between writing a shopping list and filming a thirty-second advert. Both use paper and a pen, technically. One takes ten seconds. The other needs lighting, a crew, several takes, and a lot more electricity running through the building.
The actual numbers people quote
Nobody outside the big AI labs has the exact per-video figures, because OpenAI, Google and Runway don’t publish them. What we have are independent estimates and studies that give a rough shape of the problem.
- MIT Technology Review’s 2025 energy investigation found that generating a five-second video with a leading video model could use roughly as much electricity as running a microwave for more than an hour, somewhere in the region of 1 to 1.5 kilowatt-hours depending on the model and resolution.
- A single text prompt to a large language model, by comparison, has been estimated at somewhere between 0.0003 and 0.01 kilowatt-hours, depending on the model size, which is why I’ve written before about whether ChatGPT uses water in everyday business use as a separate, much smaller question.
- Google’s own 2024 environmental report showed data centre water withdrawal rising by around 17 percent year on year, and the company directly attributed part of that increase to AI workloads, including generative video and image tools.
- Microsoft reported a jump of roughly 34 percent in water consumption between 2021 and 2022, the years its AI investment accelerated hardest, according to its own sustainability disclosures.
None of these numbers are exact, and I want to be honest about that rather than pretend I have a perfect figure for you. Energy use varies by which model you use, which data centre it runs in, what the local electricity grid looks like, and what cooling system that particular facility uses. A data centre in Iowa running on wind power is a very different story to one in a hot, dry region cooling itself with fresh water. The honest position is: it’s meaningfully more than text, it’s real, and it’s currently under-reported.
What about the water specifically
Water comes into this because most large data centres use evaporative cooling. Hot server racks need to be cooled, and one of the cheapest ways to do that at scale is to run water through cooling towers where some of it evaporates, the same basic principle as sweating. That water is gone from the local supply, not returned, which is why water-stressed regions like parts of Arizona and the American Southwest have started pushing back on new data centre builds.
A widely cited 2023 University of California Riverside study estimated that training GPT-3 alone consumed around 700,000 litres of fresh water, mostly for cooling. That’s a training cost, a one-off, not a per-video cost, but it tells you the infrastructure behind these tools is thirsty before a single customer has typed a prompt. Every video you generate afterwards adds a smaller, ongoing draw on top of that.
The client video that made me stop assuming
Here’s my actual experience with this, because I think it matters more than the abstract numbers. I was putting together a set of short product intro videos for a client’s launch using an AI video tool, testing different openings to see which one hit hardest. I generated something like thirty separate clips before we landed on the final three. Thirty renders, most of them ten seconds long, most of them thrown away.
Nobody on that call thought about energy use once. Why would we, we were focused on the brand, the pacing, whether the product looked right in the light. But sitting down to write this post made me go back and count, and thirty renders of a video that reportedly costs somewhere near a kilowatt-hour each is not nothing. That’s potentially close to 30 kilowatt-hours for one afternoon of iterating on a single client project, before we’d even picked the winner. Multiply that across every agency, every solo creator, every brand doing the same thing this week, and you start to see where the real cost sits. It isn’t the finished video. It’s the pile of drafts nobody ever watches.
Where the footprint piles up
This is the part that gets skipped in most articles on this topic, and I think it’s worth saying plainly. The environmental conversation around AI video tends to focus on the individual user feeling guilty about making one video, when the actual scale problem is corporate and structural. It’s the labs racing to release bigger, more capable video models every few months without slowing down to make them efficient. It’s the marketing departments generating fifty variations of an ad because storage and compute feel invisible and free. It’s the “sustainability” pages from these same companies that talk about carbon offsets and renewable energy targets while quietly not publishing the actual per-generation numbers a journalist or researcher could check.
Worrying intensely about your own three-video social post while a company runs data centres on diesel backup generators in a drought region is a bit like feeling guilty about a plastic straw while a container ship idles offshore. Your usage matters, but it is tiny next to the industrial scale of the problem, and the companies selling you the tool would rather you focus on your own guilt than their infrastructure choices.
How to make AI videos with a lighter footprint
You don’t need to give up the tool. You need to be less wasteful with it, the same way you’d be less wasteful with printer paper or petrol. Here’s what cuts the number of renders, based on what changed my own workflow after that client job.
- Write a fuller prompt the first time. Vague prompts get vague results, which means more regenerations. Specify camera angle, lighting, pacing, and mood in one go rather than nudging it across five attempts.
- Use image-to-video instead of text-to-video where you can. Starting from a still image you already like gives the model less to guess at, which typically means fewer wasted takes to get something usable.
- Batch your prompts and review before generating the next round. Generate three variations, watch all three, decide, rather than generating one, watching it, tweaking, generating another, one at a time.
- Repurpose existing footage instead of creating new clips from scratch. If you’ve already filmed something, cutting it down is far cheaper, computationally and financially, than generating new video. This is exactly the gap that AI video clipping tools built to turn long content into short clips are designed to fill, and they use a fraction of the compute a full generation model does.
- Keep clips short. Duration scales the cost roughly in line with frame count, so a three-second clip is meaningfully cheaper to generate than a fifteen-second one.
- Download and archive the final version once. Don’t keep regenerating a clip because you lost the file. Save it the first time, the same instinct that applies when you’re figuring out how to grab a video off social media rather than trying to recreate it, similar to the steps people search for when they want to know how to download an Instagram Story video to their phone instead of screen-recording it badly five times.
Should this stop you using AI video for your business
No, and I’ll say that plainly. Used with a bit of discipline, AI video is still far cheaper in time, money and often carbon than flying a crew somewhere, hiring equipment, and reshooting until you get it right. The comparison people rarely make is against the alternative, not against doing nothing at all. A studio shoot has its own footprint too, travel, lighting rigs, air conditioning, catering, all of it.
What I’d push back on is the guilt spiral some business owners fall into, where they read one scary headline about AI and water and quietly stop experimenting altogether. That’s not sustainability, that’s fear dressed up as principle, and it usually costs them the growth they were chasing in the first place. If this is the kind of thing that sends you spiralling into self-doubt about every tool decision you make, it might be worth reading through the common ways entrepreneurs beat themselves up over choices like this, because most of the guilt isn’t proportionate to the actual harm. Use the tool with intention, cut your waste where you can, and keep moving. If you need a nudge to stay consistent with new tools rather than abandoning them after one uncomfortable headline, my guide to staying motivated every day has more on pushing through that exact kind of hesitation.
Frequently asked questions
Does making one AI video really matter environmentally?
One video, on its own, is a rounding error against the scale of global data centre use. What matters is the pattern, how many drafts you generate, how often you regenerate rather than reuse, and multiplied across millions of users the small individual number becomes a large collective one.
Which uses more energy, an AI video or an AI image?
Video uses considerably more, because it requires generating and keeping consistent thousands of individual frames rather than one static image. Estimates generally put video generation at ten to fifty times the energy cost of a single image generation, depending on length and resolution.
Do shorter AI videos use less energy than longer ones?
Broadly yes. Energy use scales roughly with frame count, so a three-second clip costs meaningfully less to generate than a fifteen-second one from the same model, which is one of the simplest ways to cut your own footprint.
Can I find out exactly how much water or energy my AI video tool used?
Not currently, no. Almost no consumer-facing AI video company publishes per-generation energy or water figures, so the numbers available are independent estimates from researchers and journalists rather than official data from the companies themselves.