saturday, september 5, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

One model for the picture, the motion and the sound

Price a finished clip in seconds rather than files, keep a look consistent across shots, and stay on the right side of the new labelling rules.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

For years, making a short piece of visual content with AI meant running a relay race. One tool made the still. A second tool animated it and quietly changed the product’s colour. A third tool made a voiceover that did not match the mouth. A fourth tool stitched the pieces together while you fixed the drift by hand. Every handoff was a place where the thing you had approved turned into something slightly different, and most of your time went into the seams rather than the work.

That relay is being collapsed into one model. Black Forest Labs describes FLUX 3 as “one multimodal model for Image, Video, Audio and Action-Prediction,” producing clips up to 20 seconds in a single generation, with optional native audio covering multilingual dialogue, synchronised speech, sound effects and environmental ambience, at no extra charge on top of the video [1]. OpenAI and Google sell comparable generation through their own APIs [3][4]. This guide is for a solo operator, freelancer or small team who needs a 15-second product clip or a set of consistent stills and does not have an editor, a colourist or a sound designer on call. If you already have those people, the seams were never your bottleneck and this guide will not tell you much.

The seam you used to pay for has moved

The practical change is not that the output looks better. It is that a single generation now returns something you could ship. FLUX 3 renders in HD at up to 1 megapixel per frame or FHD at up to 2 megapixels per frame, and the audio arrives inside the same clip rather than as a separate track you have to line up [1]. When one model generates the step and the sound of the step together, they land on the same frame, because nothing had to be synchronised after the fact.

What that removes is a category of small, boring failure: the voice that drifts out of sync, the product that changes shade between the hero still and the motion shot, the ambience that stops dead at a cut. What it does not remove is the need to decide what the shot is. The work that used to go into repairing handoffs now goes into the brief, the reference frame and the review. That is a better trade, but it is a trade, not a saving.

You are buying seconds, not files

Stop thinking in clips and start thinking in seconds of finished footage, because that is the unit every vendor bills in. FLUX 3 text-to-video and image-to-video cost $0.06 per second in draft, $0.17 per second at HD and $0.29 per second at FHD [2]. OpenAI charges $0.10 per second for sora-2 at 720p, rising to $0.30 per second for sora-2-pro at 720p and $0.70 per second for sora-2-pro at 1080p, with batch pricing roughly half of each [3]. Google’s Gemini Omni Flash works out at approximately $0.10 per second for 720p video [4].

Stills are a different order of magnitude. FLUX.2 [klein] 4B starts at $0.014 per image, FLUX.2 [pro] at $0.03 and FLUX.2 [max] at $0.07 [2]. Gemini 3.1 Flash Image comes to $0.045 per 0.5K image and $0.151 per 4K image [4]. An image costs about as much as a fifth of a second of video. That ratio is the single most useful number in this guide, and it should change the order in which you work.

The list price is never what you pay, though, because you will not get the shot on the first attempt. A 20-second clip at HD costs $3.40 to render once [2]. Nobody renders once. Your real cost is the render price multiplied by however many times you go around, and the way to control it is to go around cheaply.

Resolution is the other lever, and it is worth being honest about which end of your work needs it. The jump from HD to FHD on FLUX 3 is $0.17 to $0.29 per second, a 70% increase [2]. On OpenAI’s side, sora-2-pro goes from $0.30 per second at 720p to $0.70 at 1080p [3]. A clip that will be watched inside a social feed on a phone does not need the top tier, and the money is better spent on more attempts at a lower resolution. Where a delivery genuinely needs the higher tier, OpenAI’s batch rates for Sora run at half the standard price, $0.05 per second for sora-2 at 720p and $0.35 for sora-2-pro at 1080p [3], and Google halves its image output rate under batch as well [4]. Batch suits an overnight run of variants far better than it suits a client sitting next to you, so plan the queue before the meeting rather than during it.

calculator
What one finished clip actually costs
$ per finished clip

FLUX 3 draft HD at $0.06/s and standard HD at $0.17/s [2]. Computed in the page; nothing is sent anywhere.

Draft mode is not a preview, it is the workflow

At $0.06 per second against $0.17 for the same shot at HD, draft costs roughly a third of a finished render [2]. That gap is large enough to change your habits. The temptation with a good model is to render at quality every time because the quality is right there. Resist it. Run every idea in draft until the composition, the pacing and the audio are what you wanted, and spend the HD budget once, at the end, on the version you have already approved.

This inverts how most people plan visual work. Planning used to be cheap and production expensive, so you storyboarded carefully and shot once. Now trying is cheap and deciding is expensive, so the productive move is to generate 6 mediocre versions of a shot in draft and use them to work out what you actually want, rather than to reason about it in a document first. Write the shot brief in Claude, ChatGPT or Gemini if it helps you be specific, then let cheap renders do the arguing.

Consistency comes from what you feed it, not what you ask for

The hardest thing in generated video for a small business is not making one good shot. It is making the second shot look like it came from the same afternoon. Asking politely in the prompt for “the same product, same lighting” does not reliably work, and each retry burns seconds you are paying for.

Feed the consistency in instead. FLUX 3 supports image-to-video with multiple frames and video continuation as first-class modes [1]. Given the price ratio, the sequence that costs least is: generate stills until you have one that is genuinely right, at a few cents each [2][4]; use that still as the input frame for the video; and extend with continuation rather than generating a fresh clip that has to rediscover your product from a text description. You are spending image money to avoid spending video money, and image money is roughly a fifth of a second of video per attempt [2][4].

The same logic applies to a set of stills for a site or a deck. Settle one reference image first, then derive the rest from it through editing endpoints rather than generating each one from scratch and hoping they converge. On FLUX.2, editing starts at $0.014 per image on [klein] 4B and $0.045 on [pro], against $0.014 and $0.03 to generate the same image from text [2]. At the [pro] tier you pay a premium of about half again to edit rather than generate, and what you get for it is that the shot starts from something you have already approved instead of from a description you hope still matches.

Keep the inputs as project files, not as chat history. The reference frame, the prompt text and the settings that produced an approved shot are what let you make the eighth clip in a campaign match the first one three weeks later, and none of that survives in a browser tab. This is the least glamorous habit in the guide and the one that saves the most money.

Read what the vendor may do with your client’s material

If you are generating for yourself, skip this. If you are putting a client’s unreleased product, unreleased campaign or staff photographs into a generation API, the terms matter more than the price.

Black Forest Labs’ FLUX API service terms have the developer grant the company “a fully paid, royalty-free, perpetual, irrevocable, worldwide, non-exclusive, and fully sublicensable right and license to use, sub-license, distribute, reproduce, modify, adapt, publicly perform, and publicly display Developer’s Input and Output,” including to “train and improve its artificial intelligence models, algorithms, and related technology, products, and services” [5]. The service terms as published set out no opt-out from that grant [5]. Compare OpenAI, which states that data from the API Platform after 1 March 2023 “isn’t used for training our models, unless you have explicitly opted in,” and retains API inputs and outputs for up to 30 days to provide the service and identify abuse [6].

That is not a verdict on either company. It is a question you now have to be able to answer when a client asks it, and a clause you should read before you paste anything covered by an NDA. Different vendors also apply different terms in different regions, so check the version that applies where you are [5].

The practical response is not to avoid the cheaper vendor. It is to sort your inputs. Public material, your own brand assets and generic reference images can go anywhere. Unreleased products, identifiable staff, client photography and anything under a confidentiality clause should only go to a service whose terms you have read and could quote back. Doing that sorting once, at the start of a project, costs an hour. Doing it after a client asks costs the project.

The same terms also constrain what you can build on top. Black Forest Labs’ service terms bar a developer from hosting an API endpoint to the FLUX models for third parties [5], which matters if your plan was to wrap generation in a tool you resell rather than to deliver finished work.

Labelling is an obligation now, not a courtesy

Since 2 August 2026, the EU’s transparency rules for AI systems apply [8]. Providers of generative systems must apply a machine-readable mark to synthetic content across text, images, video and audio, with an exception for assistive editing that does not substantially alter the content [8]. Deployers, which is you when you publish the clip, must clearly label deepfakes, meaning AI-generated or manipulated media that falsely appears authentic, and AI-generated text on matters of public interest published without human editorial review [8]. Penalties reach €15 million or 3% of annual worldwide turnover for companies [8].

There is a carve-out worth knowing: the marking obligation is exempted until December 2026 for generative AI systems placed on the market before 2 August 2026 [8]. That means some of the models you use today may not be marking their output yet, and the labelling duty on you as publisher does not wait for them.

Vendor policy points the same way. The FLUX usage policy prohibits users from acting to “circumvent, remove, alter, suppress, or otherwise interfere with any C2PA Credentials, digital watermarks, or other content provenance signals,” and from realistically depicting a real person in a “sexual, intimate, degrading, defamatory, fraudulent, misleading, or otherwise abusive manner without their verified, documented, and informed consent” [7]. It also prohibits content that “falsely represents events, things or people in a manner likely to deceive a reasonable viewer without appropriate or lawful disclosures” [7]. In practice: keep the metadata your export produces, do not put a real person’s face or voice in a clip without written permission, and say on the post that it is generated when a viewer could otherwise mistake it for a recording.

checklist
Before you render a client clip
0 of 8 · saved in this browser only

What still goes wrong

Twenty seconds is a shot, not a film [1]. Anything with a narrative still has to be assembled from pieces, and continuation keeps the look stable without keeping the story coherent, so the editing judgment you were hoping to skip comes back at the sequence level. Native audio covers multilingual dialogue [1], which means someone has to listen to every second of every take before delivery, sometimes in a language nobody on the job speaks. That review time is real and the cost calculator does not show it.

The provenance picture is genuinely unfinished. The marking exemption running until December 2026 for systems placed on the market before 2 August 2026 means you cannot assume a file carries a machine-readable mark just because a model made it [8], and the fact that vendor policy has to forbid removing provenance signals tells you they are removable [7]. Treat any mark as a helpful signal, not proof, and keep your own record of what you generated and when.

The costs above are also list prices for the render alone [2][3][4]. They do not include the hour you spend reviewing, the retries you did not count, or the two days you lose when a client’s legal team reads the same service terms you read and says no. Price that in before you quote.

sources
  1. 01Black Forest Labs — FLUX 3 model pagebfl.ai
  2. 02Black Forest Labs — API pricingdocs.bfl.ml
  3. 03OpenAI — API pricingdevelopers.openai.com
  4. 04Google — Gemini API pricingai.google.dev
  5. 05Black Forest Labs — FLUX API service termsbfl.ai
  6. 06OpenAI — Enterprise privacyopenai.com
  7. 07Black Forest Labs — FLUX usage policybfl.ai
  8. 08European Commission — Quick facts: transparency rules for AI systemsdigital-strategy.ec.europa.eu
next guide
How to use a model when you don't know who built it
10 min · verified 2026-09-05
related guides