A 27 Credit AI Product Ad: Price Before You Generate, Prompt Positively, Animate Last
A four cut vertical product ad, built on a 100 credit Higgsfield budget for 26.97. The pricing grid that has no pattern in it, why negative phrasing made every output worse, where brand consistency actually comes from, and the rule for when a video model is worth 5 credits and when ffmpeg will do.
TL;DR
- One five second clip costs about forty stills on the same platform, so find the frame with stills and animate only a frame that is already right.
- Pay a video model when the content of the frame must change (liquid, steam, people, light). When only the camera moves, ffmpeg does it for nothing.
- Negative phrasing made things worse: banning text on a product doubled the invented text and broke an unrelated clause that had been working.
- A hex code in a prompt is not a colour instruction. Soul 2.0 rendered
#111111as the characters1014printed on the can and left the colour alone. - Consistency comes from passing an approved image as a reference. In one job a referenced product matched exactly while a described navy came back 51 of 255 off.
One five second clip costs as much as forty stills
I gave myself 100 credits on Higgsfield (its prepaid unit of generation, charged per job and quoted before the job runs) and a brief: a complete vertical ad for a fictional cold brew brand, plus the product stills, wordmark, palette and social assets around it. The finished piece is four cuts, 20 seconds, 704x1280. It cost 26.97 credits across eighteen jobs, and every charge matched its quote exactly. The stills, the wordmark and the finished ad are on the project page.
The ratio that shapes every decision: a still from Higgsfield Soul 2.0 is 0.12 credits, and a five second 720p clip from Grok Imagine 1.5 Lite is 5. Roughly forty stills per clip, so the expensive thing is deciding. Do it where it costs pennies: read the model's parameters, price the exact call, iterate on stills until a frame is right, animate that frame once, then cut, zoom and stitch locally.
Read the model, then price the call
Two CLI commands before any spend, both free:
higgsfield model get text2image_soul_v2 # accepted params, REQUIRED and CONSTRAINTS
higgsfield generate cost text2image_soul_v2 --prompt "..." # the exact credits for the exact callThe first is the one people skip. It told me that Soul 2.0 has no text parameter because it cannot set type on request, and that it accepts a style_id (borrow a look) but refuses to combine it with image_references (reproduce a subject). Both facts changed the plan.
The second is where the surprises live. Resolution pricing follows no pattern between models. Measured October 2026:
| Model | Cheap | Expensive |
|---|---|---|
| Soul 2.0 | 1.5k: 0.12 | 2k: 0.12, same price, take the 2k |
| GPT Image 2.5 | 1k/low: 0.25 | 2k/high: 2.75 |
| Recraft v4.1 | 1k: 1.25 | 2k: 8, a 6.4x cliff |
| Nano Banana Pro | 1k and 2k both: 2 | 4k: 4 |
| Grok Imagine 1.5 Lite | 480p and 720p both: 5 | 1080p: 20, four times |
Recraft and Grok hide a cliff one step up. Soul 2.0 and Nano Banana Pro give a tier away. GPT Image 2.5 goes up elevenfold when resolution and quality rise together. Nobody reasons their way to that table: price the grid, which costs nothing.
Keep a ledger, one row per job: id, model, credits, prompt, verdict, and why you rejected it. The rejected jobs taught me the most. The API is asynchronous: you submit a request to a model endpoint, then poll for the result or receive a webhook when processing finishes. Download the output in the same step that logs the job.
Negatives attract what they name
The first asset was supposed to be easy: a clean, unbranded matte black can. Three attempts on Soul 2.0:
| Attempt | Phrasing | Invented text on the can | Unwanted vapour |
|---|---|---|---|
| 1 | unbranded | 2 bold lines | clean |
| 2 | no text, no lettering, no logo, no typography | 4 lines | plume returned |
| 3 | smooth featureless surface, bare unprinted metal, clean blank cylinder | 2 faint, one barely readable | clean |
Banning text doubled the text. A negation still puts the noun in the prompt, and the noun wins.
The vapour column surprised me. Attempts 1 and 2 both carried no steam, no hot vapour. Attempt 1 stayed clear; attempt 2, with the text negatives added, brought the steam back. Piling on negations broke one that had been working, which I can report but not explain. Attempt 3 dropped every negation and said clear dry air. The air cleared, the text shrank to two faint lines, and the slate, condensation and rim light landed for the first time.
So: name what you want in the frame, never what you want kept out of it.
The limits. Three attempts on one model is a prior to test on yours, not a law. A dedicated negative prompt parameter is a different mechanism, and neither Soul 2.0 nor GPT Image 2.5 has one. And no wording produced a truly blank can on Soul 2.0: that is a model limit, so stop paying to fight it and switch models.
Keep hex codes out of the prompt
I asked Soul 2.0 for a matte black #111111 aluminium can. It printed 1014 on the can in large type, and the colour sampled as a warm brown. A hex code in prose is a string the model may decide to render.
Hex belongs in a real colors parameter, which Recraft has and Soul 2.0 and GPT Image 2.5 do not. For those, generate the photographic part and add exact-colour elements locally: my end card composites a vector wordmark into space the prompt was told to leave empty. Exact colour is a compositing problem, not a prompting problem.
Consistency comes from a reference image, not your brand description
One job showed both halves at once. I generated a thumbnail on Nano Banana Pro with the approved hero still passed as image_references, and the brand palette described in the prompt. Drift is the largest difference on any one RGB channel, out of 255:
| Asked for | How | Result | Drift |
|---|---|---|---|
| the product | image reference | same label, type, placement, condensation | 0 |
pale ice blue #BFD8EA | words | #C7E4F2 | 12 |
deep navy #2E4A6B | words | #071A38 | 51 |
| background "on near black" | words | came back white | instruction dropped |
The referenced pixels came back faithfully; every described colour drifted, by different amounts. For a consistent brand across assets, pass the approved image.
Two corollaries. First, type needs a model that can set type: of the image models I tried, only GPT Image 2.5 put exact words on a product, for 0.25, right first time. That still became the hero frame.
Second, baked text survives image to video. The brand name stayed legible at 0, 2.5 and 5 seconds of the first cut, through a hard camera push that ends tight on the label. The route to text in a video ad is a text capable still, then animate that still. Never ask a video model for words.
Pay the video model only when the frame's content changes
A video model is worth 5 credits when the pixels have to become something they are not in the still: liquid pouring, steam, cloth, people, a lighting change, real parallax. When only the camera moves over a static subject, ffmpeg does it for nothing and more predictably.
Graded against that rule, my own four cuts:
| Cut | What moves | Verdict |
|---|---|---|
| 1, hero push in | camera only, but ends on a macro with a real focus shift | borderline, and I paid 5 |
| 2, the pour | liquid | worth it |
| 3, rooftop | ambient motion and light | worth it |
| 4, end card | camera only, over a logo | should never have been a clip |
So cut 4 is a slow ffmpeg zoom on the approved still, 0.25 for the still and nothing for the motion:
ffmpeg -loop 1 -i endcard.png -f lavfi -i anullsrc=channel_layout=stereo:sample_rate=48000 \
-t 5 -filter_complex "[0:v]scale=2112:3840,zoompan=z='min(zoom+0.0004,1.12)':d=120:\
x='iw/2-(iw/zoom/2)':y='ih/2-(ih/zoom/2)':s=704x1280:fps=24,setsar=1[v]" \
-map "[v]" -map 1:a -c:v libx264 -pix_fmt yuv420p -c:a aac -shortest 04.mp4Two details are load bearing. The 3x upscale gives zoompan spare pixels, so the end card stays sharp as the zoom tightens. The silent anullsrc track is there because the generated clips carry audio and the concat filter requires every segment to carry the same number of streams of each type. Drop it and the filtergraph fails outright, which is at least loud.
The concat is the quiet failure. Grok Lite did not return a consistent frame width: identical parameters (9:16, 720p, 5 seconds) gave 704x1280 twice and 720x1280 once. Joining the clips with -c copy exits 0 and yields 20 seconds and 480 frames. But the container declares 704x1280 while a frame at 12 seconds decodes at 720x1280, and the audio reports non-monotonic timestamps at every join. The exit code will not tell you.
The FFmpeg filter documentation says why: all corresponding streams must have the same parameters in every segment, and while the filtering system picks a common pixel format for you, other settings such as resolution must be converted explicitly. Normalise, then join. Scaling to cover and cropping trims 8 pixels from each side of the wider clip rather than squashing it:
ffmpeg -i 01.mp4 -i 02.mp4 -i 03.mp4 -i 04.mp4 -filter_complex \
"[0:v]scale=704:1280:force_original_aspect_ratio=increase,crop=704:1280,setsar=1,fps=24[v0];\
[1:v]scale=704:1280:force_original_aspect_ratio=increase,crop=704:1280,setsar=1,fps=24[v1];\
[2:v]scale=704:1280:force_original_aspect_ratio=increase,crop=704:1280,setsar=1,fps=24[v2];\
[3:v]scale=704:1280:force_original_aspect_ratio=increase,crop=704:1280,setsar=1,fps=24[v3];\
[v0][0:a][v1][1:a][v2][2:a][v3][3:a]concat=n=4:v=1:a=1[v][a]" \
-map "[v]" -map "[a]" -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a aac -b:a 192k ad.mp4Probe every clip's width and height with ffprobe as it arrives: the same aspect ratio and resolution settings did not give the same frame size.
How to start on a hundred credits
On a small budget the order matters more than the prompts:
- Run
model geton each model, then price its resolution grid. Both are free. - Iterate on the cheapest model that can render your subject. Describe what is in the frame: no negations, no hex codes.
- Make one approved still the reference for everything else, and get type from a model that can set it.
- Animate once, from that frame, only where its content has to change.
- Log every job with its verdict, and download as you log.
Almost every lesson here came from a rejected generation I had written a reason down for.
The prompts, to try yourself
These are the exact prompts I ran, stills first and then the video prompt that animated each one. The first three are the negation test, so run them as a set to see the effect.
Attempt 1 on Soul 2.0, 9:16, 2k:
matte black aluminium can, unbranded, on a wet slate counter at dawn, condensation beads on the metal, warm rim light from the left, shallow depth of field, no steam, no hot vapourAttempt 2, the negations that doubled the text and brought the steam back:
matte black aluminium can on a wet slate counter at dawn, no text, no lettering, no logo, no typography, plain unmarked surface, condensation beads on the metal, warm rim light from the left, shallow depth of field, no steam, no hot vapourAttempt 3, the same scene described positively:
matte black aluminium can with a smooth featureless surface, bare unprinted metal, clean blank cylinder, on a wet slate counter at dawn, condensation beads on the metal, warm rim light from the left, shallow depth of field, clear dry airThe hero still, cut 1, on GPT Image 2.5 at 1k, low quality, 9:16:
matte black aluminium can on a wet slate counter at dawn, the label reads exactly NORTHWIND COLD BREW in a bold geometric sans serif, below it in small type SLOW-STEEPED, condensation beads on the metal, warm rim light from the left, clear dry airThen animated on Grok Imagine 1.5 Lite, 9:16, 720p, 5 seconds, with that still as the start image:
slow push in on the can, condensation beads catching the light, the label stays locked and legible, clear dry airThe pour, cut 2, on Soul 2.0:
dark cold brew coffee pouring over clear ice cubes in a tall glass, slow motion, backlit at dawn, droplets suspended in the air, wet slate counter, warm rim light from the left, shallow depth of field, clear dry airAnimated the same way:
cold brew pouring in slow motion over the ice, the stream catching the backlight, droplets settling, the glass filling, clear dry airThe rooftop, cut 3, on GPT Image 2.5 with the hero still passed as the reference image:
a person on a city rooftop at sunrise holding the same matte black can from the reference, backlit silhouette, the skyline soft behind, the label reads exactly NORTHWIND COLD BREW, handheld framing, available light, clear dry airAnimated the same way:
slow dolly in on the person at the rooftop railing, the can held steady, the skyline soft behind, the label stays locked and legible, clear dry airThe end card still, cut 4, on GPT Image 2.5 with the same reference. The empty space above the can is where the wordmark goes. This one has no video prompt: ffmpeg zooms it, not a video model.
the same matte black can from the reference centered low on a near black background, soft overhead spotlight, condensation beads, generous empty space above the can, the label reads exactly NORTHWIND COLD BREW, studio product photography, clear dry airAnd the two that failed in an instructive way. The hex code that printed as text on Soul 2.0:
a hand holding a matte black #111111 aluminium can on a city rooftop at sunrise, backlit, the skyline soft behind, handheld framing, available light, slightly imperfect composition, clear dry airThe thumbnail, with the hero still as the reference, where the product held and every described colour drifted:
YouTube thumbnail, the same matte black cold brew can from the reference filling the left third, condensation beads, high contrast, on the right the words WAKE UP COLD in very large bold condensed type, deep navy and pale ice blue palette on near black, crisp edges, legible at small sizeSources
- Higgsfield API - Higgsfield API Docs (opens in a new tab), Higgsfield API Docs
- FFmpeg Filters Documentation (opens in a new tab), ffmpeg.org
generative-aiimage-generationvideo-generationpromptingffmpegcost