· 11 min read

Google Quadrupled the Clip Length and Left the Price at Ten Cents a Second. Budget by Attempts, Not by Seconds.

Google DeepMind released Gemini Omni 1.1 Flash on August 27. The headline number is that generated sequences go from 10 seconds to a cumulative 40. The number nobody put in a headline is that the price did not move: roughly $0.10 per second of 720p, the same rate since the developer API launched on June 30.

Four times the length at the same unit price sounds like the moment a video pipeline becomes affordable for one person. I spent an hour with the pricing page and a calculator, and I do not think it is, yet. Here is the arithmetic.

What actually changed

A single generation still stops at 10 seconds. What is new is that you can chain 10-second segments through scene extension until the sequence reaches 40 seconds cumulative, up from a 10-second ceiling in the previous release.

The continuity mechanism improved along with it. The model now reads up to 10 seconds of prior footage before continuing a shot, rather than only the final second. If you have tried to extend an AI clip and watched the character's jacket change colour at the seam, that one-second context window was why. Reading ten seconds of prior footage is a much better shot at holding a subject together.

You can also now specify the first and last frames of a shot and let the model generate the footage between them. That is aimed at camera orbits, zoom transitions, and loops, and it is the sort of control that turns a slot machine into a tool. Video input can include up to three seconds of reference footage to hold a character or a visual style across shots.

The pricing, spelled out

Billing runs on tokens, and the per-second figure is derived rather than listed:

  • 5,792 output tokens for each second of 720p video
  • $17.50 per 1M video output tokens
  • $1.50 per 1M input tokens

So one second of 720p is 5,792 x $17.50 / 1,000,000 = $0.101. The widely quoted "about ten cents a second" is exactly right, and it is nice when the published numbers actually reconcile.

Which makes a finished 40-second 720p sequence roughly $4.06 of video output, before input tokens.

Four dollars for forty seconds of video is genuinely cheap against any live-action alternative. That is not the number that matters.

The number that matters

Higgsfield's Hell Grind is the most useful data point I have seen on this, and it is not a Google number. It is a 95-minute AI-generated science fiction film that screened around Cannes in May 2026. It was made in roughly two weeks for about $500,000, of which about $400,000 went to AI compute. It still needed a 15-person team.

The detail that should reset your budgeting: the first 25 minutes alone took more than 16,000 initial generations, cut down to 253 final shots.

That is about 63 generations per shot that survived.

Now redo the arithmetic. If your ratio is anything like theirs, a 40-second sequence you actually ship is not $4. It is $4 times some attempt multiplier, and the multiplier is the entire cost model. At 20 attempts per keeper you are at $80 for forty usable seconds. At Hell Grind's 63 you are past $250.

I am not claiming your ratio will be 63. A feature film chasing narrative and character consistency across 95 minutes is close to the worst case, and a solo operator making a product demo loop with fixed first and last frames is close to the best. But whatever your number is, it is not 1, and every "AI video is ten cents a second" post I have read this week quietly assumes it is.

Budget by attempts and draft at 360p

This is what the 360p draft mode is for, and it is the most useful thing in the release for anyone cost-sensitive.

360p draft runs up to 60% faster than standard 720p and is priced at one third. The speed figure is Google's own system throughput comparison as of August 27, so treat it as a vendor number. Finished output exports at 1080p or 4K.

The workflow that follows is obvious once you see the attempt ratio: iterate on composition, motion, and prompt at 360p where a throwaway costs a third as much, and only render at 720p or above once the shot is decided. If you are burning 20 attempts to land one shot, doing 19 of them at a third the price is not a micro-optimisation, it is most of your bill.

The gap in the pricing page

Google publishes no per-second rate for 360p, 1080p, or 4K. The pricing page lists the 720p token rate and nothing else.

The 360p "one third" figure comes from Google's announcement rather than the pricing table. The 4K export price is simply undisclosed.

I would not build a product with a 4K output tier on an undisclosed rate. That is not a hypothetical caution: 4K is roughly nine times the pixels of 720p, and if the token cost tracks pixel count even loosely, a 40-second 4K export is a very different line item from $4. Until Google publishes it, treat 4K as "ask before you promise it to a customer."

There is also no free tier. Omni 1.1 is listed on the paid tier only, which is unusual across the Gemini lineup and means the first clip you generate while evaluating costs money. Budget for the evaluation, not just the production.

Two constraints that will decide whether this fits your use case

Every generated clip carries a SynthID watermark. For most solo operator use, social content, product demos, b-roll, that is a non-issue. If your client contract says no watermarking of any kind, read it again before you promise delivery.

Character consistency across edits and accurate text rendering are both still listed as open problems on the model card. Google saying so itself is worth more than any benchmark. If your concept requires the same person recognisable across six shots, or a legible logo or price on screen, this is the wrong tool today and the 40-second ceiling is not your limiting factor.

Worth knowing on competitive positioning: ByteDance's Seedance 2.5, announced June 23, does clips up to 30 seconds in a single generation, exports 4K, and accepts up to 50 simultaneous reference inputs. Omni's 40 seconds is cumulative across chained 10-second generations, which is a different and in some ways weaker guarantee. Meanwhile OpenAI's Sora 2 models and Videos API are deprecated and shut down on September 24, 2026, so if you built on those you have under a month.

What I would actually do

If you are evaluating this for a content pipeline, run a real test before you plan around it. Pick one shot you actually need. Generate it at 360p until you have something you would ship, and count the attempts. That single number, your attempt ratio for your kind of shot, is worth more than every pricing analysis including this one, because it is the only variable that is about you.

Then multiply: attempts at draft price, plus one final render at 720p or 1080p, times the number of shots in a piece. That is your real cost per finished minute. My guess is that most solo operators will find it lands somewhere between "cheap enough to be interesting" and "cheaper to license stock footage," and which side depends almost entirely on how specific your shot list is.

Where I could be wrong

The Hell Grind ratio may be a badly chosen anchor. A narrative film with recurring characters is the hardest case for a model whose own card lists character consistency as unsolved. A solo operator generating abstract b-roll or a looping product shot with locked first and last frames is doing something categorically easier, and the new keyframe control exists precisely to cut down the number of attempts. It is plausible the honest solo-operator ratio is 3 to 5, not 63, in which case the economics are much better than I have made them sound and I am scaring people off a tool that would work fine for them.

The counterargument I find harder to dismiss is simpler: this is the worst and most expensive that AI video will ever be. The June 30 to August 27 stretch delivered four times the sequence length at an unchanged price, which is a real deflation in cost per finished second even with no headline price cut. Planning around today's ratio may just mean planning around a number that is wrong by December.

Author

Sources

Stay in the Loop

Get new posts delivered to your inbox. No spam, unsubscribe anytime.

Newsletter coming soon. Set PUBLIC_CONVERTKIT_FORM_ID in .env to activate.

Related Posts