A ComfyUI custom node package for streamlined media loading and video pipeline assembly. Provides intuitive nodes that simplify media resource editing and loading with user-friendly parameters, making it easier to build and configure video processing workflows.
Important
It is strongly recommended that before installing this node package, you first ensure that FFmpeg has already been installed in your system environment
cd Your_ComfyUI_Path/custom_nodes
git clone https://github.com/yolain/ComfyUI-Easy-Media.gitAfter installing, open ComfyUI and find the bundled example workflows in the Templates panel on the left sidebar — look for entries under ComfyUI-Easy-Media.
The package includes the easy-media-multitrack-workflow Skill. Give Codex your assets and requirements in natural language to create a workflow from the bundled template or modify an existing one, automatically configuring the timeline, prompts, continuity modes, and sampling settings.
Use $easy-media-multitrack-workflow to create three 5-second tasks from these reference images and prompts.
Use Shot for the first segment and Context for the next two, then generate a new workflow JSON.
By default, the Skill only generates JSON. When explicitly asked to run the workflow, it can upload assets and submit it to ComfyUI. See the Skill documentation for installation and detailed usage.
Introduced in v1.3.0, the MultiTrack Project Pipeline connects timeline arrangement, segment generation, context continuity, and final video assembly. MultiTrack Project currently supports MiniMax H3; MultiTrack Editor remains independent of any model and can still feed other model workflows.
Use the workflow generation Skill to configure this pipeline through natural language, either from a blank template or by modifying an existing workflow.
MultiTrack Editor ── TRACKS_INFO ──→ MultiTrack Project ── PROJECT_NAME ──→ MultiTrack Project Video Combine ── VIDEO ──→ SaveVideo
↑
model_loader
(optional model_loader_2nd)
Important
v1.3.0 loads media on demand. When a timeline has no Slot references, the editor no longer loads all images, audio, and video in advance. Its IMAGES, AUDIO, and VIDEO outputs are None by design. The project pipeline only needs TRACKS_INFO; older workflows that consume media directly should use MultiTrack Task Output.
Select the MiniMax format, set the target dimensions and frame rate, enter prompts for each task segment, and add reference images, video, or audio as needed. See the MultiTrack Editor overview below for tracks, task modes, and media editing.
The new continuity mode determines how a task follows the previous segment. It is configured separately from task modes such as text-to-video, image-to-video, and reference-to-video:
| Continuity Mode | Generation Behavior | Use Cases |
|---|---|---|
Shot (shot) |
Generates independently, without inheriting motion or audio context from the previous segment | New shots, scene changes, and deliberate cuts |
Context (context) |
Uses the tail of the previous result's audio/video latent to continue motion and sound | Continuous action, long takes, and ongoing audio |
Drift Control Context (context_drift) |
In the first pass, applies Drift-Control to the copied video prefix: a denoise mask recalculated for each sampling step allows more change away from the seam and tapers to a fixed boundary. This limits visual drift while carrying motion forward; audio transitions softly. The second pass uses the standard high-resolution anchor without Drift-Control. | Change a subject or appearance while retaining the previous motion |
The first segment starts in Shot mode. Set subsequent segments individually or select multiple tasks to change them together. For example, “Shot → Context → Context → Shot” creates three connected segments followed by a new shot. This setting affects generation, rather than adding a crossfade during assembly. Context does not guarantee seamless continuity across arbitrary scene or prompt changes.
Connect the editor's TRACKS_INFO to tracks_info, then connect a model_loader containing the H3 model, CLIP, video VAE, and audio VAE. The node automatically expands the tasks in sequence, so no manual segment-index loop is needed:
Read task and load media on demand → Encode prompts and reference conditioning → First pass
→ [Optional: upscale video latent → Second pass] → Decode audio/video → Trim context overlap
→ Save segment and context → Process next task → Output project name
Each segment uses its own prompts, references, and continuity mode. With Context enabled, the first pass uses the previous segment's context at the corresponding resolution; the second pass uses its final high-resolution result to preserve the join. Each segment is saved before the next one runs. You can choose a starting segment and count, or save a first-pass preview and resume the second pass later. See MultiTrack Project details for settings, upscaler dependencies, and the context implementation's source and adaptations.
Connect PROJECT_NAME to the combine node to preview saved segments, choose a version for each, and assemble the final video. Auto combine produces a complete video in one workflow run. Turn it off to review results first, then click Combine to run only assembly and downstream saving without re-running the model. If a segment has multiple generated versions, select two for synchronized comparison before choosing the final version. See MultiTrack Project Video Combine for details.
Tips: The advantage of the multi-track editor is its decoupling design — it is used solely for media editing and loading, and is not bound to any model. Users can freely choose any model node to process the media data output by the multi-track editor.
easy multiImagesLoader provides the same resolution choices for an ordered list of up to 25 images. Add images through the media selector or drop image files onto the grid; the node resizes each image and outputs an IMAGE list in grid order.
| Track Type | Description |
|---|---|
| Task Track | Supports multiple task type definitions such as t2v, i2v, r2v, v2v |
| Video Track | Import and manage video clips, supporting multi-segment video stitching and intelligent segmentation |
| Audio Track | Import and manage audio clips, supporting multi-segment audio stitching |
| Subtitle Track | Add subtitles recognized from audio/video |
- Task segments are the core of this node; workflows can be designed for automatic looping based on the number of task track segments
- When adding video clips to the video track, corresponding task segments will be automatically added with matching duration
- Selecting a task segment allows you to set image, task type, and user prompt / system prompt (defaults exist based on task type, or you can write your own)
- The MultiTrack Info Output node outputs video dimensions, total frame count, frame rate, and task count
- The MultiTrack Task Output node outputs user prompt & system prompt for corresponding segments; users can decide whether to connect LLM nodes for prompt expansion or use images in segments for reverse inference
The editor passes time ranges, prompts, media locations, and track settings through TRACKS_INFO. Downstream output nodes load the required media when processing each task, reducing the cost of decoding and assembling an entire long timeline in advance. Metadata probing, frontend thumbnails, and waveform previews are separate from loading complete media for workflow execution.
| Timeline Sources | Editor Outputs | How to Access Media |
|---|---|---|
| Files or URLs, without Slot references | TRACKS_INFO is available; media outputs are None |
Connect TRACKS_INFO to MultiTrack Project or MultiTrack Task Output for loading on demand |
| Slot references to upstream image, audio, or video inputs | Retains immediate loading and media outputs; requests only the input types actually referenced | Continue using the editor's media outputs and companion output nodes |
If an older workflow connects directly to the editor's media outputs, insert MultiTrack Task Output (easy multiTrackTaskOutput):
- Process one segment at a time: Connect
TRACKS_INFOand settask_index = 0, 1, 2…to retrieve media for the corresponding task range. - Retrieve all media at once: Set
task_index = -1to output the entire timeline's media, replacing the editor's former full-media output behavior. This restores the resource cost of loading the full timeline. - Read only dimensions, frame rate, total frames, and task count: Use
MultiTrack Info Output; complete media loading is unnecessary for this metadata. - Use the project pipeline: MultiTrack Project calls Task Output internally, so no extra Task Output node is needed between the editor and the project.
- Task track: Adds Shot / Context / Character Swap Context continuity modes. Select multiple task segments to change their task mode, continuity mode, and reference image size together.
- Lock audio: In MultiTrack Project, the input audio constrains visual generation, and the delivered video uses the original task audio instead of regenerating it. This differs from using audio only as a reference.
- Reuse audio: Share the same reference audio across tasks without copying it under every segment. Up to 15 seconds are used, without truncating it to a shorter task duration.
| Scenario | Description | Requirements |
|---|---|---|
| Video Generation | MiniMax H3/wan/bernini/ltx t2v, i2v, r2v | Task track segments only |
| Video Editing | bernini v2v, bernini vi2v, wan animate, ltx video replace, ltx iclora edit/inpaint/outpaint | Video track segments + task track segments |
| Video Reference | wan scail2, wan animate, ltx iclora guide | Video track segments + task track segments |
| Video Dubbing | wan infinititalk, longcat avatar, ltx ai2v | Task track segments + audio track segments |
| Video Subtitles | - | Task track segments + subtitle track segments |
- Only the most common open-source model generation types are listed; theoretically any video model pipeline can use the multi-track editor as a preprocessing tool
| Scenario | Description | Download | Local Path | Prerequisites |
|---|---|---|---|---|
| Video Subtitles (Whisper) | Audio/video recognition to generate subtitles | Whisper Large V3 | models/audio_encoders/ | pip install openai-whisper |
| Video Subtitles (Qwen3) | Audio/video recognition to generate subtitles | Qwen3-ASR Qwen3-ForcedAligner |
models/Qwen3-ASR/ | pip install qwen-asr torchaudio |
| Subtitle Narration | Convert subtitles to speech voiceover | VoxCPM2 | models/voxcpm/ | pip install voxcpm |
| Shot Detection | Intelligently segment video shots | OmniShotCut | models/checkpoints | - |
Note: Some models support automatic download via the built-in Easy-Media model download interface. Model files will be placed in the
ComfyUI/models/directory.
easy multitrackProject manages segment-by-segment generation for MiniMax H3 projects. The editor's dimensions are the first-pass dimensions; dual sampling scales up from this size for the second pass. Project files are stored in ComfyUI/output/easy_media/projects/<project_name>/, including segment media, project records, and context latents for continuation.
- Read and encode the task: Load its prompts and media, build conditioning for text-to-video, first/last frames, last-frame-only, or multimedia references according to the task mode, and create audio/video latents. When audio is locked, encode it and apply sampling constraints.
- First pass: Both
sampling_mode = singleanddualgenerate at the editor's configured dimensions. The first-pass size is not reduced based onupscale_by. - Upscale and second pass (
dualonly): Multiply each configured dimension byupscale_by, then align to the nearest multiple of 32 usinground(dimension * upscale_by / 32) * 32. Upscale the first-pass video latent to this size, recombine it with the audio latent, and sample again. Rebuild reference conditioning when dimensions change. Settingupscale_by = 1.000skips resizing but still allows a second pass. - Decode and save: Decode video and audio. For Context segments, trim the repeated guide prefix and extra trailing frames required by the temporal grid before saving the segment.
- Continue to the next segment: Save context and project records, then process subsequent tasks automatically. Shot starts a new shot; Context inherits the previous segment. Finally, output
PROJECT_NAMEfor the combine node.
| Setting | Description |
|---|---|
model_loader |
First-pass H3 model and shared CLIP, video VAE, and audio VAE; video projects also require the audio VAE |
model_loader_2nd |
Optional second-pass H3 model; defaults to the first-pass model. Even when connected, encoding and VAEs still come from the first-pass loader |
sampling_plan |
Built-in presets such as ultra_light, light, medium, and high select samplers and sigmas for Turbo / non-Turbo models; use custom for manual settings |
sampling_mode |
single, dual, or selflift; SelfLift expands transition_ratio, lowres_scale, and the optional highres_tiling switch. Pixel/VAE correction is disabled for ordinary segments and uses an internal conservative preset for context continuation. SelfLift also saves separate low- and high-resolution context lineages |
sampler / sigmas |
Connect both to override first-pass sampling. Second-pass overrides use sampler_2nd / sigmas_2nd, also as a pair. custom requires both inputs for every sampling pass that runs |
upscale_by |
Second-pass scale relative to the editor dimensions; default 1.250, three decimal places, step 0.001 |
disable_2nd_noise |
Disables added second-pass noise; it does not skip the second pass |
1st_pass_only |
In dual mode, runs and checkpoints only the first pass of the first selected segment. Disable it on the next run to resume that segment at the second pass |
Alignment uses Python round, matching the H3 latent upscaler; it is not always upward, and exact ties round to the even integer. For example, 1344 × 768 at 1.250 becomes 1680 × 960 before alignment and 1664 × 960 for the second pass. Both upscaling paths use this size, and project records and combined exports retain the resulting dimensions. Single-pass generation and first-pass-only previews keep the editor dimensions; audio-only projects (32 × 32) do not upscale. Existing workflows retain their saved multiplier rather than automatically adopting the new default.
To customize presets, copy h3_sample.json.example to presets/h3_sample.json and edit it. The current built-in dual-sampling presets use separate first- and second-pass sigma schedules. Do not assume that light always splits sigmas and leaves the first pass incomplete. A custom preset with split_step explicitly splits one sigma schedule across both passes. Existing custom files take precedence, so review your configuration after upgrading.
upscale_model |
Upscaling Path | Dependencies |
|---|---|---|
| An H3 latent upscaler model | Calls the bundled easy minimaxH3LatentUpscaler node to upscale the video latent directly to the scaled, aligned second-pass dimensions |
Place the H3 upscaler weights in ComfyUI/models/latent_upscale_models/; the external custom-node package is not required |
None |
Decode the first-pass video with the VAE → Resize images → Re-encode with the VAE → Second pass | Requires ImageResizeKJv2 from ComfyUI-KJNodes |
upscale_model selects H3 latent upscaler weights, which serve a different purpose from the H3 generation model in model_loader_2nd. Direct latent upscaling avoids the intermediate video VAE decode/encode round trip, but the second pass still runs at the target resolution; its VRAM requirements do not decrease proportionally. Easy Media includes the inference node and runtime; only the checkpoint is downloaded separately.
The standalone easy minimaxH3LatentUpscaler node exposes enable_temporal_chunking and force_unload. Both default to enabled; MultiTrack's internal latent-upscale path also enables both explicitly. Temporal chunking lowers peak memory for long latents, while forced unload releases the upscaler from VRAM after inference.
For example, download minimax_h3_latent_upscaler_3d_fp16.safetensors from the model repository above, place it in the specified directory, restart ComfyUI, and select it under upscale_model. Its documentation specifies a 1–4× scaling range. Follow the selected model's supported range even if the project parameter accepts larger values.
Context conditioning is based on NikoDemon80 / ComfyUI-H3-Motion-Context, adapted inside Easy Media for project loops and dual sampling. The current hard-continuity implementation requires native H3 audio/video keyframe support in ComfyUI (0.34.0+). Check ComfyUI compatibility when upgrading.
In addition to attaching the previous segment's context conditioning, the project pipeline applies these adaptations:
- First-pass hard continuity: Preserves native audio/video keyframes and existing multimedia references, while copying the previous segment's audio/video tail into the current starting latent. Separate video and audio masks lock and gradually release the copied region, allowing new content to emerge around the join.
- Separate low- and high-resolution context: In dual sampling, the first pass inherits the previous segment's first-pass context. After upscaling, the second pass copies the previous final high-resolution video tail into the current high-resolution latent, rather than merely enlarging the low-resolution join. Context second-pass sampling freezes the current audio to retain continuity established in the first pass.
- Context-specific second-pass sigmas: Built-in presets provide
sigmas_2nd_context = 0.50, 0.30, 0.14, 0.06, 0.0for second passes with previous-segment context. Explicit custom second-pass sampler or sigma inputs prevent this substitution. - Synchronized trimming and clean continuation sources: By default, the project takes the previous segment's last 22 frames as context and reserves 34 extra generation frames to satisfy H3's temporal grid. After decoding, it removes the repeated prefix and excess tail, retaining the task's required frame count. It then re-encodes context from the actual delivered audio/video range so subsequent segments do not inherit discarded tail frames.
- Memory-bounded high-resolution context: Before the next segment starts, the high-resolution continuation latent is reduced to the 22-frame video/audio tail, detached from the full sampling result, and moved to CPU. Project artifacts store the same context-only high-resolution payload after second-pass completion. The complete low-resolution first-pass latent remains saved so a
1st pass onlyrun can resume its second pass later; runtime context handoffs use a separate trimmed copy. - Automatic handoff and resuming across runs: Context passes automatically between segments in one run. When starting from a later segment, the project loads context from the previous segment's active saved version, without manually wiring Save / Load Latent nodes.
Note: Context continuation requires the previous segment's saved latents; an MP4 alone is insufficient. After changing editor dimensions, scaling factor, or the previous segment's version, check the downstream context chain. Regenerate first-pass checkpoints and context latents created with the old reduced-first-pass sizing before resuming with the new sizing behavior. Existing later segments are not automatically regenerated when an earlier segment changes.
| Setting | Behavior |
|---|---|
project_name |
Identifies the project directory and records; use the same name when resuming |
segment_start_number |
Starting segment, counting from 1, unlike the zero-based task_index in MultiTrack Task Output |
segment_count |
Maximum segments to generate in this run; -1 processes all remaining segments from the start |
project_save = new |
Preserves existing results in the same project and adds versions for regenerated segments, allowing comparison |
project_save = override |
Replaces the corresponding segment version. With segment_count = -1, clears saved segments from the start onward before regeneration; a first-pass checkpoint being resumed is preserved |
For example, to regenerate only segment 3, set segment_start_number = 3 and segment_count = 1. Choose project_save = new to retain the old result for comparison. If segment 3 uses Context, the project must contain compatible context from segment 2. After accepting the new segment 3, regenerate any later Context segments that depend on it.
easy multitrackProjectVideoCombine reads saved video segments from a project. Preview them sequentially on the timeline without first creating a merged file. Use the project selector to switch projects or refresh saved results.
- Auto combine (enabled by default): After project generation finishes, reads the current segments, concatenates them in timeline order, and outputs
VIDEOandFILENAME_PREFIX. Connect an output node such asSaveVideoto save the final video. - Manual combine: With Auto combine disabled, the project still generates and saves individual segments, while the combine node updates the preview without sending a merged result downstream. Review and select your clips, wait for the ComfyUI queue to empty, then click Combine. Only this node and its downstream nodes are queued; upstream encoding and sampling are not repeated.
- Saving requirements: Connect a downstream video save output before combining manually. The combine node provides a temporary merged video; the save node controls the final path and filename. Use
FILENAME_PREFIXas a naming prefix if desired.
After regenerating with project_save = new, click a timeline segment to open its video file list:
- Select one video: Use that version for project preview and assembly.
- Select two videos: Enter synchronized comparison to review results from different seeds, prompts, or sampling settings. Each segment allows at most two selected versions at a time.
- Keep one after comparing: Deselect the unwanted version, then combine manually. Selecting two is only for comparison: assembly still uses the primary version (the first in the current selection list), not the comparison layout or both versions together.
Segments display their Shot / Context mode to help review joins. Selecting an older version changes the assembled clip but does not repair already-generated downstream context. If versions end with different motion or audio, regenerate the affected later segments. Deleting a segment version or an entire project also deletes associated files and context; make sure they are no longer needed for resuming.
Scope: This node is for video projects. Audio-only projects with editor dimensions set to
32 × 32cannot use it to combine video.
Integrated and enhanced the video saving node from the SaveVideoRGBA node package. Supports video export with customizable output path, filename prefix, frame rate, and other parameters.
Load video files from a list of file paths (or URLs) and concatenate them into a single video output.
The trim_frame_count parameter defaults to -1, which keeps all frames of the merged video. When set to a value greater than 0, the node calculates the duration based on the merged video's frame rate and uses FFmpeg to trim the final video.
- Create a
config.yamlfile in the ComfyUI-Easy-Media directory and add the following content to enable frontend development mode:
WEB_VERSION: dev- Navigate to the frontend directory and compile the development code for debugging:
cd frontend && bun install && bun run dev- After modifying the code, compile for production:
bun run build:release| Category | Node ID | Description |
|---|---|---|
| 🎞️ Multi-Track Editor | easy multiTrackEditor | Edit multi-track timelines and pass track information; defer media loading to downstream nodes when there are no Slot references |
| easy multiTrackInfoOutput | Output multi-track dimensions, duration, frame rate, and task count | |
| easy multiTrackTaskOutput | Load and output task prompts and media on demand; task_index = -1 outputs media for the entire timeline | |
| easy multiTrackAddSubtitleToVideo | Add subtitle track to video track | |
| 🎬 MiniMax H3 | easy minimaxH3ToVideo | Build MiniMax H3 text-to-video, reference-to-video, or first/last-frame conditioning and latent inputs |
| easy MiniMaxH3ReferenceToVideoBridge | Bridge node for H3 reference conditioning without Autogrow expansion | |
| easy MiniMaxH3MotionContextHard | Apply H3 context conditioning with hard video/audio latent continuity | |
| easy MiniMaxH3HiResContinuity | Copy previous high-resolution video tail into current upscaled latent | |
| easy removeH3MotionContextLatent | Remove H3 Motion Context latent files after a loop finishes | |
| easy multitrackProject | Build and execute multi-track MiniMax H3 project with optional first/second-pass sampling | |
| easy multitrackProjectVideoCombine | Preview project segments, select versions, compare two videos, and combine automatically or manually | |
| 🎞️ LTX Video | LTXVAddGuidesFromBatchIndexes | Add guide images from batch images to specified frame indexes of latent variables |
| LTXVMakeRefVideo | Expand a reference image batch into an IC-LoRA reference video | |
| easy ltxMultiTrackEncode | Build Prompt Relay conditioning and LTX video/audio latents | |
| easy ltxI2VInplaceAndUpsample | Optionally upscale an LTX video latent and apply an image guide in place | |
| easy ltxSamplerSimple | Sample combined LTX audio/video latents and crop video guides | |
| 🎞️ Timeline Editor | easy timelineEditor | Load media timeline (prompt, image, audio tracks) and output structured data |
| easy timelineInfoOutput | Output timeline info including formatted prompt, dimensions, and image indexes | |
| easy timelineSegmentOutput | Output specific segment data from the timeline | |
| easy timelineSegmentCount | Output the total number of segments in the timeline | |
| 📋 Media List Operations | easy makeImageList | Combine multiple image inputs into an image list |
| easy makeAudioList | Combine multiple audio inputs into an audio list | |
| easy splitAudios | Split an audio list into multiple single-audio outputs | |
| easy audioMerge | Merge or concatenate up to six audio inputs | |
| easy makeVideoList | Combine multiple video inputs into a video list | |
| easy splitVideos | Split a video list into multiple single-video outputs | |
| easy imageIndexesToIntList | Convert comma-separated image index string to integer list | |
| easy splitImages | Split an image list or batch into multiple single-image outputs | |
| 🎬 Video Operations | easy saveVideo | Save images and optional audio as video file |
| easy getAudioFromVideo | Extract audio from a VIDEO input | |
| easy mergeVideos | Concatenate multiple compatible VIDEO segments | |
| easy mergeVideosFromPaths | Load and concatenate videos from file path list, optionally trimming the merged output by frame count | |
| easy compareVideos | Preview source and output VIDEO inputs side by side with an interactive comparison slider | |
| 📝 Subtitle | easy recognizeSubtitle | Recognize subtitles with Qwen3-ASR or Whisper Large V3; configure SRT/timestamp output, sentence length, and model unloading |
| easy addSubtitleToVideo | Normalize multiline SRT, timestamp, or bracket-formatted text and burn it into a video | |
| 🖼️ Reference & Image | easy makeRefsCompositeBySam3 | Detect subject in prompt using SAM3 and composite reference images onto canvas |
| 🔧 Utility | easy matchLine | Return zero-based index of the first line containing matching text |
| easy apiWorkflowGate | Determine if the workflow is an API call and pass through preceding input items | |
| 🗣️ Speech to Video (S2V) | easy berniniS2VConditioning | Unified Bernini + Wan S2V conditioning: original optional single-speaker audio, spatially masked single-speaker audio, or optional sequential two-speaker audio |






