1. Choose a Generation Mode
Use Text to Video when the scene starts as an idea. Use Image to Video when the opening composition must match an image. Use Reference to Video when identity, motion, audio, style or written creative direction comes from several sources.
2. Add Frames or References
Image mode requires a first frame and accepts an optional last frame. Reference mode supports images, short video, audio, one document and one public link.
Use files you own or have permission to use. Keep each reference relevant to the same creative direction.
3. Write the Prompt
Describe the subject, action, setting, camera, lighting, style and sound. For a 20 or 30 second video, explain the action in a short sequence of beats.
The prompt field supports detailed production direction, but clarity matters more than length.
4. Select Output Settings
Choose Adaptive when the model should follow the reference composition, or select a fixed ratio. Choose 480P for quick iteration, 720P for balanced output or 1080P for final detail.
Set a duration from 2 to 30 seconds and leave native audio enabled when the final scene should include synchronized sound.
5. Generate the Video
Sign in, review the estimated credit cost and submit the task. The interface uploads references first, then creates a Kie task and polls until the provider returns the video or an error.
Keep the tab open while the result is processing. The task ID is retained for the active generation.
6. Review and Iterate
If the subject or movement is wrong, change only the relevant part of the prompt. If the look is wrong, update lighting, color or reference images without rewriting the action.
Validate difficult scenes at shorter duration before spending credits on a longer 1080P version.