The Neogen Brief
AI Video Production

AI Video Production in 2026: What It Costs, Where It Fails and When to Shoot Live

What AI video production really costs, what it still can't do, how we keep a character on-model across a whole series, and the four kinds of shot we still send to a camera crew.

Rehdhil Siyad
Rehdhil Siyad
Founder · Neogen Media
29 September 2026
17 min read
Matte black cinema camera on a red lacquer plinth beside three chrome lenses, one lens being mounted

AI video production in 2026 is good enough to replace a camera for B-roll, explainers, product-in-scene shots, stylised ads and a recurring animated series. It is still not good enough for on-screen text, hands working small objects, long unbroken dialogue or anything where a real face is the point. We run Seedance 2.5, Kling 3.0 and Veo 3.1 on client work, and we pick a model per shot, not per project.

The short version: Seedance 2.5 carries anything with a recurring character, because it takes the character sheets and a voice sample on every generation. Kling 3.0 wins on calm camera work and costs the least. Veo 3.1 is kept for photoreal close-ups and short synced dialogue. None of them can spell, and Seedance will add music you told it not to.

Below is what AI video production costs per second on our own plan, where each model fails, the rules we use to keep a character identical across an episode, and the four kinds of shot we still send to a camera crew. The model-by-model render logs are in our companion post on AI video generation.

What is AI video production, and how is it different from using an AI video generator?

AI video production is the full pipeline around the model: script, a locked reference library for characters, props and voices, shot-by-shot model selection, generation, then edit, colour, sound and typography in post. An AI video generator is one step inside it. Most of the quality comes from the steps around the model.

That distinction matters because the top search results for this keyword are mostly tool lists and courses. A tool list tells you which generator exists. It does not tell you why the same character has a different jawline in shot 4, or why the logo on the product melted in shot 7. Those are production problems, and they get solved before and after generation, not by switching tools.

Here is the pipeline we run on every job:

  • Script and shot breakdown, with each take's length timed from its dialogue and beats before anything is generated. A take padded out to the model's maximum length comes back slow and drifting.
  • A reference library: character sheets, wardrobe, props, environments and a clean voice sample for every speaking character. Props are locked in one plain, inert state, and the video model animates any glow or change from that.
  • Model selection per shot, recorded against the shot list with its cost tier.
  • Generation with the character sheets attached to the video job itself. We only make start frames for locations and props. Passing a character through an image model first redraws it every time, and on one project that detour cost us 403 credits and produced nothing usable.
  • Edit, grade, sound design and voice-over in DaVinci Resolve, with ElevenLabs or a human voice. All typography is added here.

A 30-second spot typically uses two or three models. You can see how we scope that on our AI video production service page.

Which AI video model is best for each kind of shot in 2026?

There is no single best model, and the rankings move every few months. In September 2026, MiniMax H3, Google's Gemini Omni Flash, Wan 3.0 and Seedance 2.5 led Arena's blind-vote image-to-video leaderboard within 18 points of each other, while Kling 3.0 and Veo 3.1 had fallen out of the top ten (Arena leaderboard). We build on what we have proven on client work: Seedance 2.5 for recurring characters and long takes with dialogue, Kling 3.0 for cheap calm camera work, and Veo 3.1 for the odd photoreal close-up. On a real job, the shot decides the model. Our AI video generation comparison covers the newer models we are testing.

The spec differences that actually change our decisions:

  • Clip length: Seedance 2.5 generates up to 30 seconds in one pass. Kling 3.0 generates any length from 3 to 15 seconds. Veo 3.1 only generates 4, 6 or 8 seconds.
  • Resolution: ByteDance's own API lists Seedance 2.5 at 480p and 720p only. Some platforms, including the one we render on, sell a 1080p tier at 71% more credits per second, and independent comparisons describe that 1080p as a provider-side addition to a 480p and 720p model (CellCog pricing comparison). Kling Standard is 720p and Pro is 1080p. Veo 3.1 goes to 1080p and 4K, but only on 8-second clips.
  • References: the platform we use accepts up to 50 reference files per Seedance 2.5 job, but characters start dropping out above about seven. Veo 3.1's Ingredients mode takes up to three reference images.
  • Cuts inside one generation: Seedance handles several hard cuts in a single render. Kling 3.0's multi-shot mode handles up to six shots. Veo needs stitching.
  • Audio: all three generate native audio. Seedance 2.5 also takes a voice sample as an audio reference, so a recurring character keeps the same voice from episode to episode. Veo's lip-sync on short lines is the cleanest of the three.
  • Frame rate: Veo and Seedance output a fixed 24fps. If the delivery spec says 30fps, you conform in post.

Where Seedance 2.5 wins, and where it fails

Seedance 2.5 is our default for any character who has to stay on-model. We attach the character sheets, a location sheet and a voice sample to every job, and a 30-second take comes back with picture, dialogue and ambience together. In one week of September 2026, a single account on one animated series logged 32 Seedance 2.5 renders and 6,204 credits.

It fails in predictable places. Past about seven references, characters dissolve or disappear from the frame. It adds background music even when the prompt says no music. It fills whatever duration you give it, so an over-long take turns into slow motion. Its close-up micro-expressions are weaker than Veo's, so we frame emotional beats a little wider.

Where Kling 3.0 wins, and where it fails

Kling 3.0 is no longer a top-ranked model, and we do not use it for anything that needs one. It is our default for B-roll, establishing shots and product animation. It follows camera directions well (orbit, dolly, tracking, speed ramp) and it can hold several art styles in one frame without blending them. Because length is free-form between 3 and 15 seconds, you generate exactly the duration the edit needs instead of paying for 8 seconds to use 5.

It fails in predictable places. Faces lose detail in wide shots, so we keep characters in mid or close framing. Lip-sync drifts after about 10 seconds, so dialogue goes in the first 10. Cramming several actions into one long take produces invented movement. The fix is shorter shots, not better prompts. When it has silence to fill, it sometimes fills it with invented background dialogue, so our prompts say "no background dialogue" explicitly.

Where Veo 3.1 wins, and where it fails

Google has already moved its own users on: Gemini Omni replaced Veo as the default video engine in the Gemini app after Google I/O in May 2026, and Veo 3.1 now lives on through the Gemini API and Vertex AI. We still use it for photoreal hero shots: skin, fabric and lighting are the most convincing of the three, and it reads cinematography language literally. Write "slow dolly-in about 10% over 4 seconds" and that is roughly what you get. Its Ingredients mode, fed three angles of a product or a face, is a reliable way to keep one identity consistent across separate generations.

It is also expensive per second, capped at 8 seconds per clip, and weaker with more than one person in frame. Scene Extension can chain clips toward two and a half minutes, but extended clips drop to 720p, so we render hero shots standalone at 1080p or 4K and only extend where resolution matters less.

How much does AI video production cost per second?

Generation is the smaller part of the budget, but re-rolls multiply it. Google's published API price for Veo 3.1 Standard is $0.40 per second at 720p or 1080p and $0.60 at 4K. Veo 3.1 Fast is $0.10 to $0.12 per second, and you are only charged for videos that generate successfully, according to the Gemini API pricing page.

On our own subscription ($237 a month for 5,400 credits, about $0.044 a credit), one second of generated footage costs:

Model and tierCredits per secondAbout $ per secondOne 10-second take
Seedance 2.5, 1080p tier12$0.53$5.30
Seedance 2.5, 720p7$0.31$3.10
Seedance 2.5, 480p draft3$0.13$1.30
Kling 3.0 Pro2$0.09$0.90
Kling 3.0 Standard1.75$0.08$0.80

Now the math on a 30-second spot. Say it is ten shots averaging 3 seconds, with a one-second handle on each side for the edit: about 50 seconds of footage you actually keep. Assume a planning ratio of five takes per kept shot, which is conservative for character work. That is 250 generated seconds.

  • All on Seedance 2.5 at the 1080p tier: 250 x $0.53 = about $133
  • All on Veo 3.1 Standard: 250 x $0.40 = $100
  • All on Veo 3.1 Fast at 1080p: 250 x $0.12 = $30
  • All on Kling 3.0 Pro: 250 x $0.09 = about $22

For one spot, even the top figure is small next to a crew day. A series changes the maths. A failed 30-second Seedance take at the 1080p tier costs about $16, and our first episode needed a 40% re-roll rate. The saving is in not re-rolling: when one cut inside a take fails, we regenerate only that moment as a short insert, anchored to a still from a take we kept. On one scene that cost 372 credits against 528 for re-rolling the full takes.

The rest of the cost is human time: the script, building the reference library, choosing and rejecting takes, fixing continuity, the grade, the sound mix and the typography. That is the position we would defend against anyone quoting AI video as "cheap because the model is cheap". A studio that underprices AI video is usually skipping one of those steps, and it shows in shot 7.

Budget pressure is real on the buyer side too. Wistia's 2026 State of Video report found that almost 40% of companies spent under $5,000 producing videos last year. More than a third already use AI in their video workflow, and nearly a quarter more plan to start. Wistia's CEO Chris Savage put the direction plainly in an interview with Sacra: "The cost per minute of creation of these things will be driven down over time, and the quality will be pushed up."

What can AI video production do better than a live shoot today?

AI video now beats a shoot on B-roll, explainer visuals, product shown inside an environment, stylised or impossible locations, pre-visualisation, recurring animated characters and social cut-downs. In each of these the value is in the image and the camera move, not in a real person's face or a legible label.

  • B-roll and establishing shots. A drone-style pass over a skyline at dusk, a slow push through an empty office, rain on a shopfront. Kling handles these in one or two takes. We never leave B-roll static: every insert gets a camera move, because a locked-off AI shot reads as a still image.
  • Explainer visuals. Abstract processes, data flows and before-and-after states that would otherwise be motion graphics. Where the explainer is mostly type and diagrams, our motion graphics team still does it better, because type is the one thing AI cannot render.
  • Product in scene. A bottle on a café table, a device on a desk at golden hour. Veo's Ingredients mode with three product angles keeps the product recognisable across the campaign.
  • Locations you cannot get to: period sets, underwater, space, or simply a location that would cost a travel day.
  • Pre-visualisation before a real shoot, so blocking and lighting are agreed with the client before anyone books a crew.
  • Recurring characters. We produce an animated children's series where the same characters appear across episodes. Each has a locked character sheet and a voice sample attached to every Seedance 2.5 generation, so neither the face nor the voice is left for the model to invent.

If you are deciding how video fits into a wider brand system, our branding playbook for Indian businesses covers where AI video sits next to identity and static design.

Where does AI video still fail?

AI video still fails at on-screen text, hands handling small objects, continuity of an object's state between renders, long unbroken dialogue, real likenesses and "no music" instructions. These are not prompt problems. They are model limits in 2026, and a good production plan routes around them instead of re-rolling.

Text and logos

All three models garble text beyond two or three words. Signs, labels, captions and supers come out as plausible nonsense. Our rule is strict: no typography is ever rendered in-model. Every word on screen is added in post, which also means the client can change a line without regenerating footage.

Hands and small objects

Pouring, folding fabric and using small tools still look uncanny. We anchor a hand to a single object, avoid two-handed interactions, and cut away before the hard part of the action.

State continuity

This one surprises clients. Reference images carry a character's identity from one render to the next. They do not carry the state of things. A bitten popsicle comes back whole in the next batch, a smile resets to neutral, a door that was open is closed. We handle it by extracting the last frame of one batch and using it as the start frame of the next, and by planning cuts so state changes happen on screen rather than between renders. Effects follow the same rule. A picture of a glowing portal attached as a reference got pasted into the scene as a flat card, so we now describe effects in words and carry them forward from the last frame.

Pictures overrule the prompt

Attached images outweigh the words in the prompt. Attach a product sheet for a prop that should be small in frame and it comes back at hero size, whatever the prompt says. One of our rejected takes rendered a magic portal as a still photograph of a parked train, because the train's exterior was attached as a reference. We leave the sheet off when a prop must be small, and every prompt carries a fixed scale block that says how big each character and prop is in frame.

Music you did not ask for

Seedance adds background music even when the prompt opens with a no-music rule. On our pilot, one take had the picture approved and the audio rejected for exactly that, and an automated silence check did not catch it. We now listen to every take on its own before it goes into the cut, and rebuild the sound from clean effects when a take leaks music.

Dialogue and real faces

Lip-sync holds for short lines. Past about ten seconds it drifts on Kling and Veo, so long speeches get split across clips and covered with cutaways. Seedance 2.5 holds longer takes, but only when the words fit the seconds: eleven lines packed into 30 seconds came back as babble on our pilot. Real-person likenesses are filtered heavily, and we would not use one without written permission anyway.

How do you keep an AI character consistent across a whole series?

Attach the character's sheets to every video generation, keep each job to seven references or fewer, lock props in one plain state, and repair broken cuts instead of re-rolling whole takes. Consistency comes from what the model is given, not from how the prompt is worded. These are the rules we run on our animated series:

  • Character sheets go on the video job itself. Never route a character through an image model first, because it redraws the character every time.
  • Seven references at most. Above that, characters dissolve. When a job is over the cap, drop hero poses and prop sheets first, never a character.
  • Establish every character in the first cut of a take, so the model has each identity fixed before the action starts.
  • Lock props static. The reference shows the prop plain and inert, and the video model animates any glow or change.
  • Time each take from the script: spoken words divided by 2.2, plus 1.5 seconds per reaction beat, plus 1 second per cut.
  • Split a long scene only on a cut to a different subject, never in the middle of a shot, so the join is invisible.
  • Carry state forward with the last frame of the previous take, not with a new picture.
  • Repair, don't re-roll: regenerate only the broken cut as a short insert, anchored to a still from a take you kept.

When should you still book a real shoot instead of AI video?

Book a camera when authenticity is the message, when a recognisable person is the creative anchor, when a macro product shot is the whole spot, or when the video is mostly someone talking for minutes. In those four cases, AI either cannot deliver or delivers something that erodes trust.

  • Customer testimonials and founder stories. A generated customer is a lie, and viewers who notice will not trust the rest of the ad.
  • Star talent or a known public figure.
  • Hero product macro: jewellery, a beverage pour, automotive paint. Live photography still wins on finish.
  • Long-form talking content: webinars, interviews, training where one person speaks for several minutes.

The middle ground is a hybrid: shoot the founder or the product for real, and use AI for the inserts, environments and B-roll around it. On most of our brand films that is the actual answer, and it is the first thing we work out on a discovery call.

How do we choose the model for each shot?

We decide per shot from three questions: does a recurring character appear, is there a photoreal face in close-up, and how hard does the camera or subject move. The answers route it to Seedance, Veo or Kling, and we record the choice and its cost tier before generating.

  • A recurring character, or a long take with dialogue: Seedance 2.5, with the character sheets and voice sample attached.
  • Photoreal close-up, emotional beat or a short synced line: Veo 3.1.
  • Fast camera move, action, dance, or several cuts in one beat: Seedance.
  • Calm move, slow drift, locked composition, B-roll, product animation: Kling 3.0. Spending Seedance or Veo money on a slow drift wastes budget.
  • A scout pass to test framing and timing: the cheapest tier available (Seedance 2.5 at 480p, Veo 3.1 Lite or Fast), then re-render the winner at full quality.

That routing is also why we do not sell AI video by model. Nobody should pay Veo rates for a shot Kling does just as well, and no single model covers a whole 30-second spot well.

Frequently asked questions

Who owns the rights to AI-generated video?

On the paid commercial tiers of Kling, Veo and Seedance, which is what we render on, commercial use of the output is permitted and the delivered film is yours to run on paid media, TV, web and social. For regulated categories like health and finance, we keep the prompt, model and licence chain on file so your legal team has an audit trail.

How long does an AI video take to produce?

A single 30-second spot with script, reference library, generation, edit, grade, sound and two revision rounds takes us 10 to 14 days. A campaign of one hero spot plus six to ten social cut-downs takes three to four weeks. Once your reference library exists, simple single-scene social spots can turn around in three to five days, because the locked assets get reused.

Can you put our founder or staff into an AI video?

Only with written consent, and even then the models filter real likenesses heavily, especially Seedance. For founder content we usually recommend filming the person for real and building AI environments and B-roll around that footage. It is faster to approve and it keeps the part of the video that earns trust real.

What do we need to provide to start?

A script or a clear brief, your logo and brand colours, product photos from at least three angles if a product appears on screen, and a short, clean voice recording with no music for any character who speaks across several videos. Three angles is what Veo's Ingredients mode needs to hold a product consistently. If you have none of this yet, building the reference library becomes the first week of the project.

Should we learn the tools and do this in-house?

For simple social B-roll, yes: a tool like Kling on a monthly plan is enough, and there are good courses. The in-house gap shows up on character continuity across a campaign, per-shot model routing, and the post-production finish. Those are the parts that separate a clip from a film.

If you have a script or a concept and want to know which parts should be AI, which should be shot, and what that costs, talk to our team. We will tell you honestly which lane your project is in.

Rehdhil Siyad
Rehdhil SiyadFounder · Neogen Media

Founder and Director at Neogen Media. Writing field notes on AI automation, growth systems, and the integrated playbook we ship for Indian SMBs. Based in Kochi.

Follow on LinkedIn
Next Step

Want a system like this shipped for you?

If the playbook above maps to your stack and you'd rather we implement it than read about it, book a 30-minute strategy call. We'll map the priorities, tell you what's actually worth building, and leave you with a plan either way.

Book a Strategy Call
30 MINFREE AUDITNO DECKNO OBLIGATION
Or send us a WhatsApp
// What You Walk Away With
  • 01

    A map of every manual task worth automating

  • 02

    Ballpark ROI on your top 3 automation opportunities

  • 03

    Honest read on whether we are a fit — or who is

Usually responds within 24 hours