What video to video means here
Video to video is the opposite trade to text to video. Instead of describing a scene from nothing, you supply footage and ask for a change. The clip you upload fixes the structure — the timing, the motion, the camera path — and the prompt decides what the footage should look like afterwards.
That makes it the practical choice whenever the movement matters more than the look. If you already have a shot with the right pacing and framing, there is no reason to gamble on generating that motion again from words. Keep the take you filmed and re-render its appearance instead.
How to run a video to video render
Uploading a video switches the workspace into its editing path automatically — there is no mode to select.
- Sign in with an email address and open the workspace.
- Attach the source clip, up to 100 MB. Keep the revision close to the original length.
- Describe the change in one sentence: the new style, lighting, setting or mood.
- Pick 16:9 or 9:16, then 720p or 1080p or 4K, then a length between 4 and 10 seconds.
- Generate, review and download. Change one clause and re-run if the look is close but not right.
What is kept and what changes
The motion comes from your footage: the model follows the movement it sees rather than inventing one. What changes is the surface — the grade, the weather, the time of day, the environment behind the subject, the overall style. That is what makes the technique useful for reusing one take in several different settings.
Some things survive a re-render less well than others. Fine text and logos drawn into the frame can warp, because they are being redrawn rather than copied. Faces stay recognisable as faces, but not necessarily as photographs of a specific person, so do not rely on it for identity-accurate material. And if the source clip is much longer than 10 seconds, the render covers a short span of it; trim the take to the moment you care about before uploading.
- Kept: timing, motion, camera path, general composition.
- Changed: colour grade, lighting, weather, setting, style and mood.
- Unreliable: on-screen text, logos, branding and exact facial likeness.
What one render costs
Renders with a source video are billed per render at a flat rate, so a 4-second re-render and a 10-second one cost the same. 720p and 1080p share one price and 4K is higher, which means the length and resolution of the output are quality choices rather than pure budget ones. A new account gets 300 credits without a card, and failed generations are refunded.
| Output resolution | Credits |
|---|---|
| 720p or 1080p | 240 credits |
| 4K | 360 credits |
Where this fits in a real workflow
The most common use is versioning. One shot of a product on a tabletop can become a summer version, a night version and a rainy-day version without reshooting, because the movement of the hand and the timing of the pan stay the same across all three. That consistency is hard to get from separate generations and easy to get from a single source clip.
The second use is salvage. A take with good movement but flat light can be relit; a background that was fine on the day but wrong for the campaign can be replaced. Where a conventional edit would need masks, tracking and a lot of patience, a sentence describing the target look can be enough — with the honest caveat that it will not be frame-exact.
Common mistakes
The biggest one is treating the prompt like a video-editing instruction list. "Cut at 00:04 and speed up the middle" describes operations this tool does not perform; the render works on appearance, not on the edit decisions. Describe the picture you want at the end, not the steps you would take in a timeline.
The next is uploading an over-long clip and hoping the interesting moment survives. The output is a short render, so trim to the span you want first. The last is judging the result at full 4K: test at 1080p, where the render is cheaper and faster, and only pay for 4K when the look is already right.