DocsAI ModelsVideo models

Video models

Klipse lets you choose from a range of video generation models depending on what you are making. Models differ in start and end frames, reference images, video input, clip length, aspect ratio, resolution, and how they handle audio, so it matters that you pick the one that matches how you want to work.

What are the differences?

Video models differ in more than output quality — they differ in how the video is made. You can generate from a prompt alone, control the scene with start and end frames, or feed in a reference image or an existing video.

I want prompt-driven video in a range of aspect ratios.

Generating video while setting length and aspect ratio yourself across a wide range

Seedance 2.0 Series

I want to set the first and last scene myself.

Using a start frame and an end frame to control where the scene goes

Kling 2.6 / 3.0 · Veo 3.1

I want to make full use of reference images or existing video.

Using image and video input together to give the model more material to work from

Kling Omni · Gemini Omni Flash

I want the audio generated from the prompt along with the video.

Video audio generated automatically from the prompt instead of a separate sound on/off switch

Veo 3.1 Series · Gemini Omni Flash

Models at a glance

Compare the main input methods and generation options supported by the Klipse video node.

ModelSpecialtyStartEndReferenceVideo inputLengthRatioResolutionAudio
Seedance 2.0 MiniFast, lightweight generation----4–15 sec6 ratios480p, 720pSupported
Seedance 2.0 FastFast repeated generation----4–15 sec6 ratios480p, 720pSupported
Seedance 2.0High-resolution general generation----4–15 sec6 ratios480p–4KSupported
Kling 2.6Start and end frame controlOO--5–15 secFollows start720p, 1080pSupported
Kling 3.0Frame control + high resolutionOO--3–15 secFollows start720p–4KSupported
Kling OmniGeneration from multiple inputsO-OO3–15 sec9:16, 16:9, 1:1720p–4KSupported
Gemini Omni FlashImage and video reference + automatic audioO-OO10 sec9:16, 16:9No information available yetAutomatic, from the prompt
Veo 3.1 ProFrame and reference based + automatic audioOOO-4–8 sec9:16, 16:9720p, 1080pAutomatic, from the prompt
Veo 3.1 FastFast frame and reference based generationOOO-4–8 sec9:16, 16:9720p, 1080pAutomatic, from the prompt
Veo 3.1 LiteSimple start and end frame generationOO--4–8 sec9:16, 16:9720p, 1080pAutomatic, from the prompt

Seedance 2.0 Series

Suited to generating video from the prompt alone, without start and end frames or reference images, while setting length and aspect ratio yourself across a wide range.

4–15 sec9:1616:91:121:94:33:4

Seedance 2.0 Mini

A lightweight setup supporting 480p and 720p. Good for quick tests or low-resolution drafts.

Seedance 2.0 Fast

Supports 480p and 720p. A good option for repeated generation and a fast production flow.

Seedance 2.0

Supports 480p through 4K, so it suits cases where the same generation method needs a higher output resolution.

Kling 2.6 / Kling 3.0

A setup built for specifying the first and last scene of a video directly as images. Neither model has a separate aspect ratio setting — both follow the ratio of the start frame image.

Kling 2.6

Video with a set first and last scene

Supports 5–15 sec lengths and 720p or 1080p. It suits designing how the scene changes using a start frame and an end frame.

5–15 sec
Kling 3.0

Frame control + up to 4K

Supports 3–15 sec lengths and up to 4K. It suits wanting both start and end scene control and a high output resolution.

3–15 sec4K

Kling Omni

Along with a start frame, it can take reference images and existing video as input, so it is built for video generation that draws on several pieces of media.

Video generation from multiple inputs

Because it can use a start image, a reference image, and video input together, it suits cases where you need to give the model more production information than a single image can carry.

Video Input3–15 sec9:16 / 16:9 / 1:1Up to 4K

Gemini Omni Flash

Suited to video production that uses a start frame, reference images, and video input together while also generating the audio automatically from the prompt.

Multiple inputs + automatic audio

It generates 10-second video and supports the 9:16 and 16:9 ratios. Sound is not a separate on/off setting — it is generated automatically from the prompt you enter.

Video Input10 sec9:16 / 16:9Audio Auto

Veo 3.1 Series

Suited to specifying start and end frames and generating the audio from the prompt at the same time. Pro and Fast can also use reference images, while Lite offers a simpler input structure.

Veo 3.1 Pro

Frames + reference image + automatic audio

Suits video production that has to use start and end frames together with a reference image.

Veo 3.1 Fast

Fast frame and reference based generation

A good choice when you want the same input structure as Pro but are working toward fast repeated production.

Veo 3.1 Lite

Simple start and end frame generation

Suits making video from start and end frames and the prompt, without a reference image.

Common to all: 4–8 sec · 9:16 / 16:9 · 720p / 1080p · audio generated automatically from the prompt

Recommendations by task

I need a range of ratios — vertical, horizontal, square.

Six ratios are supported, for when each content channel needs a different frame

Seedance 2.0 Series

I want the first scene to change into a specific last scene.

For when the start frame and end frame have to be specified clearly

Kling 2.6 / 3.0

I want the model to reference both images and existing video.

For when a reference image and video input have to be used together

Kling Omni · Gemini Omni Flash

I want the sound made along with the video in one go.

For when you want the video and audio built together from the prompt

Veo 3.1 · Gemini Omni Flash

I need resolution up to 4K.

Final deliverables that need a high output resolution

Seedance 2.0 · Kling 3.0 · Kling Omni

I need a relatively long clip, up to 15 seconds.

For when you want a clip up to 15 seconds long from a single generation

Seedance · Kling Series

Choosing a model in the video node

The actual video generation happens in the video node in Node Canvas. When you select a model, the input slots, length, ratio, resolution, and sound options it supports change automatically.

What the video node and AI Models each cover

The video node page explains how to connect inputs and generate video in practice, while AI Models · video models compares what each model specializes in and how to choose.

Frequently asked questions

Which model is best?

No single model is best for every task. Choose based on the input method, video length, aspect ratio, resolution, and audio generation you need.

Which model should I choose to use both a start frame and an end frame?

Kling 2.6, Kling 3.0, Veo 3.1 Pro, Veo 3.1 Fast, and Veo 3.1 Lite support start and end frames.

Can I use a reference image and a video together?

Kling Omni and Gemini Omni Flash let you use a reference image and video input together.

Which models support 4K?

In Klipse today, Seedance 2.0, Kling 3.0, and Kling Omni support resolutions up to 4K.

Can the sound be generated automatically along with the video?

Gemini Omni Flash and the Veo 3.1 family generate audio automatically from the prompt.