The short answer: one scene, five parts
Gemini Omni Flash turns one described scene into a 4 to 10 second clip. It does best when the prompt reads like a single line from a shot list: who or what is on screen, what happens, how the camera behaves, what the light looks like and which visual style to aim for. Prompts that try to tell a whole story, list several locations or change subject halfway through give the model too many targets and usually produce a clip that matches none of them well.
So the practical rule is simple. Describe one moment, in the present tense, with concrete nouns and verbs. If you need a second moment, generate a second clip and cut them together. This is also the cheapest way to work, because each attempt stays short and you can see exactly which sentence caused a change.
The five-part prompt formula
Write the five parts in this order. The order matters less to the model than it does to you: a fixed order makes it easy to change one part at a time and compare results.
- Subject: the main thing on screen, with two or three identifying details. "A ceramic coffee mug with a blue glaze" beats "a mug".
- Action: one verb phrase that fits inside a few seconds. "Steam rises and the mug slowly rotates" fits; "the barista makes coffee and serves it" does not.
- Camera: shot size and movement. Close-up, medium shot or wide shot, plus static, slow push-in, orbit, tracking or handheld.
- Light: time of day or light source and its quality. Soft window light, golden hour, hard noon sun, neon at night.
- Style: the look you want, described in plain words. Clean product commercial, documentary, film grain, pastel animation.
12 copy-ready prompt examples
Each example follows the formula and is sized for a single 4 to 10 second clip. Paste one into the generator above, keep the default 6-second 1080p setting for a first look, then change one part at a time.
| Use case | Prompt |
|---|---|
| Product hero | A matte black wireless earbud case on a wet stone slab, the lid opens slowly, macro close-up with a slow push-in, soft studio light with a cool rim light, clean premium commercial look |
| Food | A bowl of ramen with a soft-boiled egg, chopsticks lift noodles and steam curls upward, overhead close-up, warm tungsten light, shallow depth of field, food magazine style |
| Skincare | A glass serum bottle on pale marble, a single drop falls from the dropper, extreme close-up, bright diffused morning light, minimal beauty advertisement |
| Fashion | A model in a long camel coat walks toward the camera on an empty city crosswalk, medium shot tracking backward, overcast daylight, editorial film look |
| Portrait | An elderly fisherman mending a net on a wooden pier, he looks up and smiles, medium close-up, golden hour side light, documentary style with natural colors |
| Travel | A narrow street in Lisbon with a yellow tram climbing the hill, wide shot from a low angle, slow pan right, late afternoon sun, warm travel vlog look |
| Nature | Fog drifting through a pine forest at dawn, sunbeams break through the trees, slow aerial push forward, cool blue light turning warm, calm cinematic look |
| Real estate | A bright minimalist living room with floor-to-ceiling windows, sheer curtains move in the breeze, slow gimbal walk-through at eye level, soft daylight, architectural photography style |
| App promo | A hand holds a smartphone showing a colorful dashboard, the thumb scrolls once, close-up, clean white background with soft shadows, modern tech advertisement |
| Pet | A golden retriever puppy chases a red ball across a sunny lawn and tumbles, low tracking shot at dog height, bright summer light, playful home video feel |
| Vertical social | A barista pours latte art into a cup, the heart shape forms, vertical 9:16 close-up from above, warm café light, satisfying social media style |
| Event teaser | Confetti falls over an empty stage as spotlights sweep the crowd area, wide shot, slow push-in, colorful concert lighting with haze, energetic teaser look |
When to add reference images
Words are good at motion, camera and mood. They are bad at exact appearance. If a logo, a product shape, a face or a fabric pattern has to look a specific way, attach it as a reference image instead of describing it. The generator accepts up to 7 PNG, JPG or WebP images of up to 10 MB each and uses them as visual anchors.
When you attach references, shorten the subject part of the prompt and spend the words on action, camera and light. Repeating in text what the image already shows adds nothing and can pull the result away from your reference. Only upload images you own or have permission to use.
Five mistakes that waste credits
Most failed attempts come from the same few habits. Fixing them is cheaper than rerolling the same prompt and hoping.
- Several scenes in one prompt. Split them into separate clips.
- Abstract adjectives only. "Epic, stunning, cinematic" gives the model nothing to draw; name the light and the shot instead.
- Changing everything between attempts. Change one of the five parts per attempt so you learn what works.
- Rendering the final length first. Draft at 4 seconds, then render the length you need once the draft looks right.
- Choosing the aspect ratio last. 9:16 and 16:9 frame the subject differently, and switching means generating again.
A cheap iteration loop
Start with a 4-second draft at 90 credits and judge only composition and motion. If the framing is wrong, rewrite the camera part. If the subject looks off, add a reference image. If the mood is wrong, rewrite the light and style parts. When the draft is right, render the final clip; a default 6-second 1080p clip costs 120 credits and 720p costs the same as 1080p, so there is no reason to finish at 720p.
Failed generations are refunded automatically, so a technical error never costs credits. A clip that simply is not what you wanted still uses them, which is why the short draft step pays for itself.