Video Nodes
Video Nodes generate video from text and reference images, and edit, extend, lipsync, and upscale existing clips.
Video Nodes generate short animated clips directly from text prompts and optional reference images. They are ideal for creating cinematic previews, marketing videos, quick scene animations, or visual storytelling elements within Atlas workflows.
Two nodes are available: a full-featured version and a simplified fast-generation version.
When to use video nodes
Reach for video nodes when a static image isn't enough and a fully-rigged animation pipeline is overkill. Common use cases:
Marketing and trailers. Generate cinematic clips from concept art for pitch decks, store listings, social media, or community announcements.
Cutscene prototyping. Roughly visualize a cinematic before committing animator time to a polished version.
Animated moodboards. Turn a single style reference into a short looping clip that conveys mood and motion direction to the team.
NPC and character animation prototyping. Use Lipsync to put a generated voice on a static character portrait for dialogue review, or Reference to Video to test motion choreography against a character reference.
Video edits and continuations. Use Video Edit to restyle existing footage and Video Extend to grow short clips into longer sequences without re-rendering.
Text + Image -> Video
Generates a video from a text prompt, an optional continuation prompt, and up to three reference images.

Inputs
Prompt — main instruction for video content
Continuation Prompt (optional) — describes how the animation should evolve
Reference Images — up to 3 images to control style, subject, or composition
Video Settings
Duration: 4, 6, 8, 16, 24, 32, or 40 seconds
Resolution: 720p or 1080p
Aspect Ratio: Landscape or Portrait
Output
A rendered video clip in the chosen format
This node is suited for more controlled, style-specific video generation, especially when reference images are important.
Example Usecase
Simple Text + Image -> Video
A streamlined version optimized for fast, lightweight video generation.

Inputs
Prompt — primary description
Input Image (optional) — style or subject reference
Video Settings
Duration: 5, 10, or 12 seconds
Resolution: 420p, 720p, or 1080p
Aspect Ratio: Landscape, Portrait, Standard, or Square
Fixed Camera Position: enable or disable
Seed: control variation (
-1= random)
Output
A quick-rendered video clip
This version is ideal for rapid prototyping or generating simple animated assets for marketing or social media.

Example Usecase
End Image input — provides control over both the start and end frames of the generated video clip.

Use Cases
Marketing videos from a single concept image
Animated moodboards
Scene previews for game or environment design
Quick animations for pitch decks or client presentations
Stylized loops for social media
Video Nodes provide a fast way to bring static concepts to life using text prompts and reference imagery.
Video Edit
Transforms an existing video by applying a new creative direction or environment using a reference image and text prompt.

Inputs
Source Video — the original video clip to transform
Reference Image — visual guide for the target style, environment, or look
Prompt — text description of the desired edit
Parameters
Backend selector — choose motion-path generation method (some backends use prompt-based motion, others use reference-driven paths)
Seed — control variation (
-1= random)
Output
Edited video clip matching the reference image style and prompt direction
Useful for: changing the setting or atmosphere of placeholder footage, creating environmental variations of cutscenes, or adapting generic vehicle or character animations into themed game contexts (expedition tours, combat zones, fantasy landscapes).
Accepts input videos from 3 to 60 seconds in length
Supports up to 5 reference images
Output resolution selectable as 720p or 1080p
Audio handling mode: automatic or original (preserve source audio)
Some backends support instruction-based edits with optional style reference image
Supports style transfer driven by a reference video in addition to reference images

Supports placing synchronized sound effects onto a video
Video Extend
Extends an existing video clip forward in time by generating additional frames based on a text prompt and the final frames of the input.

Inputs
Source Video — the video clip to continue
Prompt — text description guiding the extended footage
Backend selector — choose generation engine (different backends produce varying motion styles and continuation approaches)
Parameters
Duration — length of the extension (available durations depend on the selected backend)
Resolution — output resolution (options vary by backend)
Seed — control variation (
-1= random)
Output
Extended video clip appended to the original
Useful for: creating longer cutscene sequences from short generated clips, looping environmental footage, or prototyping extended NPC actions and vehicle animations without re-rendering the entire scene.
Seedance Reference to Video
Generates video content by combining multiple reference inputs—images, video clips, and audio—with a text prompt to produce a cohesive animated result.

Inputs
Prompt — text description guiding the generation
Reference Images — up to 9 still images for style, character, or environment guidance
Reference Videos — up to 3 video clips (e.g., motion choreography, background loops, camera movement)
Reference Audio — up to 3 audio clips to influence pacing, rhythm, or mood
Parameters
Seed — control variation (
-1= random)
Output
Generated video clip synthesizing all provided references

Useful for: creating NPC dance sequences synced to in-game music, generating character performances driven by reference choreography, or producing cutscene animations that blend concept art, motion samples, and soundtrack cues.
Lipsync
Synchronizes a character's mouth movements to match an audio track, producing a video of the character speaking the provided dialogue or narration.

Video Input — the character or face video to animate
Audio Input — the speech or dialogue track to sync
Backend Selection — four selectable backends, each with different processing characteristics
Output — a lip-synced video clip with matching mouth movements
Requires a clearly visible speaker facing the camera during the audio timeframe
Best used with close-up or medium shots where facial features are distinct

Upscale Video
Increases the resolution of an input video to produce sharper, higher-fidelity output from lower-resolution source clips.

Video input. Accepts an existing clip as the source for upscaling.
Video output. Returns the upscaled clip at increased resolution.
Resolution ceiling. Upscales output up to 1080p at 30fps.
Preserves timing. Frame count and duration of the source are retained; only spatial resolution is increased.
Typical use. Finalize prototyped or generated footage for trailers, store listings, and cutscene review at higher fidelity.
Useful for: promoting rough or low-resolution generated clips to presentation quality without re-rendering the original video.
Common pitfalls
Mismatched aspect ratio across reference images. When supplying multiple reference images to the Text + Image -> Video node, references with very different aspect ratios produce inconsistent framing. Pre-crop references to a consistent shape via 2D Post Processing before generation.
Expecting backends to share durations and resolutions. Different backends in the Video Extend node support different duration and resolution options, and backend-driven motion in the Video Edit node varies similarly. Check the dropdown per node rather than assuming parameter parity.
Feeding Lipsync unclear audio. The Lipsync node produces visibly worse mouth movement from background noise, music behind dialogue, or low-quality recordings. Clean or generate the speech track via Audio Nodes first.
Feeding Video Edit a source outside its length range. The Video Edit node accepts input videos from 3 to 60 seconds; clips outside that range cannot be transformed.
Expecting Upscale Video to add motion or length. The Upscale Video node increases spatial resolution only and retains the source frame count and duration; it does not regenerate content or extend length. Use the Video Extend node for additional duration.
Related nodes
Input Nodes — Input Image and Input Images supply reference content for video generation; Input Video supplies source clips for the Video Edit and Video Extend nodes.
Image Nodes — generate or refine reference frames before producing video.
Audio Nodes — generate voice tracks for the Lipsync node and audio clips used as reference audio in the Seedance Reference to Video node.
Utility Nodes — combine multiple references, prompts, or extracted document content into video-ready inputs.
Mesh Nodes — provide rigging and animation workflows for real-time character animation, an alternative to pre-rendered video clips.
Frequently asked questions
What is the difference between the Text + Image -> Video and Simple Text + Image -> Video nodes?
The Text + Image -> Video node accepts more parameters—a continuation prompt, up to three reference images, and longer durations—for more controlled, style-specific output. The Simple Text + Image -> Video node is a streamlined text to video AI node with one optional input image and fewer settings, optimized for fast, lightweight generation. Use the Simple node for rapid prototyping and the full node when reference images and control matter.
How long can generated videos be?
Duration depends on the node. The Text + Image -> Video node supports 4, 6, 8, 16, 24, 32, or 40 seconds. The Simple Text + Image -> Video node supports 5, 10, or 12 seconds. The Video Extend node appends additional duration to an existing clip, with available durations depending on the selected backend. For longer sequences, chain Video Extend calls.
Can I generate video with sound?
The video generation nodes can produce visual output with baked-in audio according to the backend. For audio, you can also generate it separately via Audio Nodes and combine it in post-production, or use the Lipsync node to synchronize generated dialogue onto a character video directly.
Why does my video look stylistically different from my reference image?
Some video generation backends prioritize motion fidelity over strict style adherence. If style consistency matters, use the Seedance Reference to Video node, which accepts up to nine reference images, three reference videos, and three reference audio clips, or the Text + Image -> Video node rather than the Simple Text + Image -> Video node.
Can I use video output as a real-time engine asset?
Generally no. Video outputs are pre-rendered clips, not interactive content. Use them for cinematics, marketing material, in-engine playback such as UI screens and in-world displays, or as reference for traditional animation. For real-time character animation, use the rigging workflows in Mesh Nodes instead.
What is the recommended workflow for animated NPC dialogue?
For lipsync generation, generate the speech track via Audio Nodes text-to-speech, then feed it into the Lipsync node along with a character video. The Lipsync node offers four selectable backends and synchronizes mouth movement to the audio; it requires a clearly visible speaker facing the camera and works best with close-up or medium shots. For broader body animation, use the rigging and animation pipeline in Mesh Nodes.
How do I set the start and end frames of a generated clip?
The Simple Text + Image -> Video node accepts an optional Input Image and an End Image input, providing control over both the start and end frames of the generated video clip.
How do I change the setting or style of existing footage?
Use the Video Edit node, which transforms a source video using a reference image and a text prompt. It suits changing the atmosphere of placeholder footage, creating environmental variations of cutscenes, or adapting generic animations into themed game contexts. A backend selector chooses the motion-path generation method.
How do I improve the resolution of a generated clip?
Use the Upscale Video node. It increases the resolution of a lower-resolution source up to 1080p at 30fps while retaining the original frame count and duration, which finalizes prototyped or generated footage for trailers, store listings, and cutscene review without re-rendering.
What is the difference between Seedance 2.0 Reference to Video and Text + Image → Video?
Seedance 2.0 Reference to Video: A highly multimodal node. It lets you link multiple images (up to 9), videos (up to 3), and audio clips (up to 3) together in your prompt (e.g., matching a face to a voice reference). Text + Image → Video: A multi-engine playground. Best for high-fidelity animations, smooth morphing from a start image to an end image, or writing complex multi-shot timelines.
Use Seedance 2.0 Reference to Video when: You need to synthesize video, motion, and audio (like matching a character portrait to a specific voiceover/sound reference).
Use Text + Image → Video when: You want maximum cinematic quality, need to animate a single image, or want to morph a start image into a different end image.
Last updated