Video models
Klipse lets you choose from a range of video generation models depending on what you are making. Models differ in start and end frames, reference images, video input, clip length, aspect ratio, resolution, and how they handle audio, so it matters that you pick the one that matches how you want to work.
What are the differences?
Video models differ in more than output quality — they differ in how the video is made. You can generate from a prompt alone, control the scene with start and end frames, or feed in a reference image or an existing video.
I want prompt-driven video in a range of aspect ratios.
Generating video while setting length and aspect ratio yourself across a wide range
Seedance 2.0 SeriesI want to set the first and last scene myself.
Using a start frame and an end frame to control where the scene goes
Kling 2.6 / 3.0 · Veo 3.1I want to make full use of reference images or existing video.
Using image and video input together to give the model more material to work from
Kling Omni · Gemini Omni FlashI want the audio generated from the prompt along with the video.
Video audio generated automatically from the prompt instead of a separate sound on/off switch
Veo 3.1 Series · Gemini Omni FlashModels at a glance
Compare the main input methods and generation options supported by the Klipse video node.
| Model | Specialty | Start | End | Reference | Video input | Length | Ratio | Resolution | Audio |
|---|---|---|---|---|---|---|---|---|---|
| Seedance 2.0 Mini | Fast, lightweight generation | - | - | - | - | 4–15 sec | 6 ratios | 480p, 720p | Supported |
| Seedance 2.0 Fast | Fast repeated generation | - | - | - | - | 4–15 sec | 6 ratios | 480p, 720p | Supported |
| Seedance 2.0 | High-resolution general generation | - | - | - | - | 4–15 sec | 6 ratios | 480p–4K | Supported |
| Kling 2.6 | Start and end frame control | O | O | - | - | 5–15 sec | Follows start | 720p, 1080p | Supported |
| Kling 3.0 | Frame control + high resolution | O | O | - | - | 3–15 sec | Follows start | 720p–4K | Supported |
| Kling Omni | Generation from multiple inputs | O | - | O | O | 3–15 sec | 9:16, 16:9, 1:1 | 720p–4K | Supported |
| Gemini Omni Flash | Image and video reference + automatic audio | O | - | O | O | 10 sec | 9:16, 16:9 | No information available yet | Automatic, from the prompt |
| Veo 3.1 Pro | Frame and reference based + automatic audio | O | O | O | - | 4–8 sec | 9:16, 16:9 | 720p, 1080p | Automatic, from the prompt |
| Veo 3.1 Fast | Fast frame and reference based generation | O | O | O | - | 4–8 sec | 9:16, 16:9 | 720p, 1080p | Automatic, from the prompt |
| Veo 3.1 Lite | Simple start and end frame generation | O | O | - | - | 4–8 sec | 9:16, 16:9 | 720p, 1080p | Automatic, from the prompt |
Seedance 2.0 Series
Suited to generating video from the prompt alone, without start and end frames or reference images, while setting length and aspect ratio yourself across a wide range.
Seedance 2.0 Mini
A lightweight setup supporting 480p and 720p. Good for quick tests or low-resolution drafts.
Seedance 2.0 Fast
Supports 480p and 720p. A good option for repeated generation and a fast production flow.
Seedance 2.0
Supports 480p through 4K, so it suits cases where the same generation method needs a higher output resolution.
Kling 2.6 / Kling 3.0
A setup built for specifying the first and last scene of a video directly as images. Neither model has a separate aspect ratio setting — both follow the ratio of the start frame image.
Video with a set first and last scene
Supports 5–15 sec lengths and 720p or 1080p. It suits designing how the scene changes using a start frame and an end frame.
Frame control + up to 4K
Supports 3–15 sec lengths and up to 4K. It suits wanting both start and end scene control and a high output resolution.
Kling Omni
Along with a start frame, it can take reference images and existing video as input, so it is built for video generation that draws on several pieces of media.
Video generation from multiple inputs
Because it can use a start image, a reference image, and video input together, it suits cases where you need to give the model more production information than a single image can carry.
Gemini Omni Flash
Suited to video production that uses a start frame, reference images, and video input together while also generating the audio automatically from the prompt.
Multiple inputs + automatic audio
It generates 10-second video and supports the 9:16 and 16:9 ratios. Sound is not a separate on/off setting — it is generated automatically from the prompt you enter.
Veo 3.1 Series
Suited to specifying start and end frames and generating the audio from the prompt at the same time. Pro and Fast can also use reference images, while Lite offers a simpler input structure.
Frames + reference image + automatic audio
Suits video production that has to use start and end frames together with a reference image.
Fast frame and reference based generation
A good choice when you want the same input structure as Pro but are working toward fast repeated production.
Simple start and end frame generation
Suits making video from start and end frames and the prompt, without a reference image.
Common to all: 4–8 sec · 9:16 / 16:9 · 720p / 1080p · audio generated automatically from the prompt
Recommendations by task
I need a range of ratios — vertical, horizontal, square.
Six ratios are supported, for when each content channel needs a different frame
Seedance 2.0 SeriesI want the first scene to change into a specific last scene.
For when the start frame and end frame have to be specified clearly
Kling 2.6 / 3.0I want the model to reference both images and existing video.
For when a reference image and video input have to be used together
Kling Omni · Gemini Omni FlashI want the sound made along with the video in one go.
For when you want the video and audio built together from the prompt
Veo 3.1 · Gemini Omni FlashI need resolution up to 4K.
Final deliverables that need a high output resolution
Seedance 2.0 · Kling 3.0 · Kling OmniI need a relatively long clip, up to 15 seconds.
For when you want a clip up to 15 seconds long from a single generation
Seedance · Kling SeriesChoosing a model in the video node
The actual video generation happens in the video node in Node Canvas. When you select a model, the input slots, length, ratio, resolution, and sound options it supports change automatically.
What the video node and AI Models each cover
The video node page explains how to connect inputs and generate video in practice, while AI Models · video models compares what each model specializes in and how to choose.
Frequently asked questions
Which model is best?
No single model is best for every task. Choose based on the input method, video length, aspect ratio, resolution, and audio generation you need.
Which model should I choose to use both a start frame and an end frame?
Kling 2.6, Kling 3.0, Veo 3.1 Pro, Veo 3.1 Fast, and Veo 3.1 Lite support start and end frames.
Can I use a reference image and a video together?
Kling Omni and Gemini Omni Flash let you use a reference image and video input together.
Which models support 4K?
In Klipse today, Seedance 2.0, Kling 3.0, and Kling Omni support resolutions up to 4K.
Can the sound be generated automatically along with the video?
Gemini Omni Flash and the Veo 3.1 family generate audio automatically from the prompt.