Instructions to use Lightricks/LTX-2.5-22b-IC-LoRA-Refine-Details with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LTX-2
How to use Lightricks/LTX-2.5-22b-IC-LoRA-Refine-Details with LTX-2:
# Install the LTX-2 pipelines git clone https://github.com/Lightricks/LTX-2.git cd LTX-2 uv sync --extra natten
# Download the adapter weights from this repo # (base components come from Lightricks/LTX-2.5 β see Files and versions) hf download Lightricks/LTX-2.5-22b-IC-LoRA-Refine-Details --local-dir models/LTX-2.5-22b-IC-LoRA-Refine-Details
# Video-to-video with the IC-LoRA (runs on the distilled LTX-2.5 base) uv run python -m ltx_pipelines.ic_lora \ --transformer-path path/to/distilled-transformer.safetensors \ --text-encoder-path path/to/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \ --video-vae-path path/to/video-vae.safetensors \ --audio-vae-path path/to/audio-vae.safetensors \ --spatial-upsampler-path path/to/spatial-upsampler.safetensors \ --lora models/LTX-2.5-22b-IC-LoRA-Refine-Details/<weights>.safetensors 1.0 \ --video-conditioning reference.mp4 1.0 \ --prompt "your prompt here" \ --output-path output.mp4 - Notebooks
- Google Colab
- Kaggle
You need to agree to share your contact information to access this model
By clicking "Agree and Access" you acknowledge the Privacy Policy and consent to receive offers and updates. You can unsubscribe at any time.
Log in or Sign Up to review the conditions and access this model content.
LTX-2.5 22B IC-LoRA Refine Details
An In-Context LoRA that rebuilds fine detail in video: it takes a soft, compressed or upscaled clip as its reference and renders the same frames with the texture, edges and grain of a native high-resolution capture. Used tiled, it refines and upscales any clip to 4K.
It is based on the LTX-2.5 foundation model.
Example Outputs
- Prompt
- Locked-off wide shot of a clear mountain stream running over mossy granite boulders under a beech forest canopy, thousands of individual leaves, sunlight dappled evenly across the scene, water sparkling with fine ripples, ferns and lichen in the foreground, deep focus, tack sharp everywhere, nature documentary, photoreal, 8K, sharp photographic detail, crisp faces and clothing texture, natural grain, high resolution footage
- Prompt
- Slow dolly forward down a narrow neon-lit alley at night, glowing signs reading RAMEN, OPEN 24H, KARAOKE and BAR, wet pavement reflecting the colours, paper lanterns swaying gently, a few pedestrians with umbrellas walking away from camera, deep focus, photoreal, 8K, sharp photographic detail, crisp faces and clothing texture, natural grain, high resolution footage
Two generated clips, 1920x1088, refined twice with this LoRA through the tiled fusion sampler: to 3840x2176, then to 7680x4352. Every pixel is generated; the passes add detail the 1080p source never had. Each gallery clip is a 1:1 crop comparison (the source lanczos-upscaled on the left, the 8K result on the right, 800x800 result pixels); the full-frame wipes are next to them: forest stream and neon alley. The prompts shown are the ones the refine passes ran with: the generation prompt plus sharp photographic detail, crisp faces and clothing texture, natural grain, high resolution footage. Negatives: blurry, soft focus, shallow depth of field, bokeh, motion blur, low resolution, painting, watermark (plus fog, haze, text for the stream).
Model Files
ltx-2.5-22b-ic-lora-refine-details-1.0.safetensors
Final checkpoint, step 3000 (training v1). Runs on the LTX-2.5 distilled transformer.
Model Details
- Base Model: LTX-2.5-22B Video (trained on the 2.5 base with the Gemma-4 12B text encoder)
- Training Type: IC-LoRA (video-to-video)
- Control Type: reference video (the clip to refine, frame-aligned), plus an optional reference image at the -1 token (per tile)
- Reference Downscale Factor: 1
- Pipeline details: single pass at the trained tile size; per-step tiled latent fusion above it (see Pipeline Details)
- Audio: Not trained for audio generation
Intended Use & Out-of-Scope
Intended use: refining and upscaling delivery-grade video (1080p web, compressed, generated clips) to 4K; a second pass on generated video to replace VAE softness with real texture; the detail stage after the Restoration IC-LoRA on archive footage.
Out of scope: colour restoration or colourisation (the model sharpens what is there and slightly desaturates archive footage: run the Restoration IC-LoRA first); pristine cinema raw (it denoises sensor grain rather than adding detail); native 8K refinement in one pass (see Tips).
Control Signal Requirements
- Control signal type: the source video itself, resized to the output canvas (lanczos), fed as the IC-LoRA reference at downscale factor 1.
- Expected input: any resolution; the output canvas must be a multiple of 32 on both sides and the frame count 8n+1 (49, 97, 121, 241...). For a 1920x1080 source, refine at 3840x2176 or crop to 1920x1056 first.
- Preprocessing: none beyond the resize. Do not sharpen or denoise the input first.
- Alignment: 1:1, same frames, same frame rate. The output is the input's geometry with rebuilt detail.
- Mask support: none. Optional reference image at the -1 token: one to three native-resolution crops of the subject on a grey image, attached with a second
LTXAddVideoICLoRAGuideatframe_idx-1; it steers products, text, logos and faces and is not required (the per-tile rule is in Recommended Settings).
How It Works
The model was trained on pairs made from 4K footage: the target is a real downscale of each tile to 1024x576 (never an upscaled target), and the reference is the same tile degraded to look like a soft, upscaled source at a range of severities. It therefore learned exactly one thing: given a frame whose fine band is missing or smeared, rebuild that band consistently over time. It does not change framing, exposure or colour. Its working window is a 1024x576 tile; larger canvases are refined as overlapping tiles fused every denoising step, which keeps one coherent frame rather than a mosaic. At 4K the content inside a tile is at the scale the model trained on; at 8K a tile sees a 2x close-up and starts inventing texture, so 8K is reached progressively.
Usage
ComfyUI
- Copy the LoRA weights into
models/loras. - Open LTX-2.5_V2V_TiledFusion_Upscale.json from the ComfyUI-LTXVideo repository. It loads the LTX-2.5 distilled transformer, the Gemma-4 12B text encoder, the LTX-2.5 video VAE and this LoRA (
ltx-2.5-22b-ic-lora-refine-details-1.0.safetensorsat 1.0) and runsLTXVTiledFusionSamplerwith 1024x576 tiles. - Drop the clip on
LoadVideo, setoutput_size(FullHD, 4K or 8K), write a look-and-style prompt, run. The workflow's Preprocess group first upscales your clip to the output size with lanczos and feeds that as the guide; the model never sees the small original. Long clips are handled by 97-frame windows (use_streamingis on in the workflow). - Use the tiled workflow for every output size, including full HD. The LoRA was trained on 1024x576 tiles and gives its best result when each tile sees content at that scale; a 1920x1088 canvas is a 3x3 grid of those tiles. A single untiled pass over a full-HD frame puts the LoRA outside its training bucket. The single-tile LTX-2.5_V2V_ICLoRA_Single_Stage_Distilled.json is right only when the output itself is 1024x576 or smaller, i.e. when the whole frame is one tile.
- Building your own graph, in this order: lanczos-resize the clip to the output canvas (a multiple of 32) β
LTXAddVideoICLoRAGuide(latent_downscale_factor1) βLTXVTiledFusionSampler(tile 1024x576, overlap 0.5, blend 0.05,cfg1.0) on the 8-step distilled sigmas withLTXICLoRALoaderModelOnly(strength 1.0) on the distilled model βVAEDecodeTiled.
Pipeline Details
Refinement above the trained tile runs as per-step tiled latent fusion:
- One canvas, one noise field. The whole frame is a single latent; every denoising step crops overlapping 1024x576 windows (50% overlap), runs each window conditioned on its own crop of the reference, and merges the stepped windows back with a Gaussian weight (blend variance 0.05). Because the merge happens inside the step loop, the windows stay coupled and there are no seams and no per-tile flicker.
- Upscale ladder. Resize the source to the output canvas with lanczos and refine once. 1080p to 4K is one pass (about 22 minutes for 121 frames on an RTX PRO 6000). 8K is best reached as refine to 4K, lanczos to 8K, optional gentle second pass.
- Long clips. Up to about 240 frames can be one temporal extent. Beyond that, or for tighter memory, enable windowed guides (
use_streamingon the guide node withtile_frames97, and the same 97 on the sampler): 97-frame windows, each with its own encoded reference, fused the same way. - Decode tiled (1024 px, 128 overlap, temporal 256).
Recommended Settings
- LoRA strength / weight: 1.0
- Inference steps: 8 (distilled sigmas
1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875, 0.0, euler) - Guidance scale: 1.0
- Input preparation: upscale the clip to the output canvas with lanczos before it becomes the guide (the Upscale workflow does this); the guide should already be at the output size, not the small original.
- Resolution & frames: tile 1024x576 (or 576x1024 portrait), canvas any multiple of 32, 97-frame windows; 4K output is the sweet spot for a 1080p source
- Tiled at every output size: the closer each tile is to the training bucket (1024x576), the better the LoRA works, so run tiled even for a full-HD output rather than one untiled pass.
- Tiling: overlap 0.5, blend variance 0.05, grid cycle 1; never lower the overlap below 0.5 (phase breaks on fine periodic texture)
- Prompting: keep the prompt generic and about the rendering (sharpness, texture, grain, light), not the subjects. Every tile receives the whole prompt, so a subject named in it can be painted into tiles that do not contain it. Name a subject only when you prompt per tile: run the tile that holds it as its own single-tile pass with its own prompt. That is also how to use a generated clip's generation prompt (it helps heavily degraded objects, but only in the tile that holds them), never on the whole canvas.
- Reference image at the -1 token (products, text, logos, faces): when a subject is ambiguous after degradation, give the model the real thing: native-resolution crops of it on a grey image, attached with a second
LTXAddVideoICLoRAGuide(frame_idx-1, strength 1.0) chained after the clip guide, on the whole-clip path (use_streamingoff). The reference is per tile: the LoRA learned it against the 1024x576 tile that holds the subject, and the fusion sampler crops the reference image to each tile exactly as it crops the clip. So make the reference image the size of the output canvas and place each crop over its own subject; every tile that covers the subject then sees it. Crops placed anywhere else reach the wrong tiles: on a night street scene the placed reference restored a shop-sign word and brought a face back towards the truth, while the same crops in the wrong tiles garbled the word and lowered similarity to the truth. A reference also raises fine-detail energy across the whole frame, so use it for a subject, not by default.
Example positive prompt:
sharp photographic detail, crisp natural texture, fine surface detail, clean edges, natural film grain, high resolution footage
Example negative prompt:
blurry, soft, plastic, smeared detail, oversharpened halos, warped faces
References
- Code: GitHub Repository
- ComfyUI: ComfyUI-LTXVideo
Tips & Troubleshooting
- An upscale cannot exceed its source. Deep focus, even light and dense texture measured up to 3x the detail of a plain lanczos upscale; shallow depth of field measured 1.0x on the out-of-focus parts. Judge the in-focus region at 1:1.
- Archive or broadcast footage looks cleaner but emptier. The model was trained on clean downscales, so it reads codec damage in the mid band as noise and smooths it, and it desaturates old footage. Run the Restoration IC-LoRA first, then refine.
- Cinema raw (pristine sensor footage) comes out slightly softer than a lanczos upscale: the model removes grain it cannot rebuild. Not the tool for that source.
- 8K in one pass invents detail. A 1024x576 window covers 13% of an 8K frame, twice the trained scale; melons rearrange and sign text changes. Refine at 4K, then upscale.
- Stochastic texture (fireworks, sea foam, foliage in wind) can over-sharpen. Lowering strength costs more detail elsewhere than it saves there; prefer a descriptive prompt.
- Text, logos and signage can turn into glyph-shaped noise after heavy degradation: the model re-renders letters it cannot read. Give it the real sign as a reference at the -1 token, placed over the sign's own position (see Recommended Settings), or run that tile on its own with a prompt that names the text. Very small lettering is not recovered either way.
- The last one or two frames of any LTX generation lose about 10% detail. In windowed runs the overlap hides it; for a single extent, pad the clip by 8 frames and trim.
- The mp4 written by ComfyUI's
SaveVideo(about 8 Mbps at 4K) understates the result. Save the latents and re-decode, or encode from PNG frames.
Dataset
Paired clean and degraded 97-frame tiles built from 4K footage across 24 content families (people, animals, products, architecture, nature, textures, vehicles, crowds, generated footage). The target is a native downscale of each tile to 1024x576 (576x1024 for portrait); the reference is the same tile degraded to simulate a soft, upscaled source at a range of severities. Pairs also carry a reference image (the -1 token input): native-resolution crops of the subject on a grey image. 24 fps, one shared caption.
Training
- Technique: LTX-2 trainer flexible strategy with two conditions:
reference(the degraded clip, downscale factor 1, always present) andreference_frames(the reference image described under Recommended Settings, present at half the steps, so the model works both with and without a reference image at inference). - Hyperparameters: rank 128, alpha 128, targets attn1/attn2 q/k/v/out and ff, AdamW lr 1e-4, cosine schedule, batch 1, grad clip 1.0, bf16, no quantization, gradient checkpointing, seed 42
- Steps: 3000 (checkpoint every 250; step 3000 shipped)
- Resolution buckets:
1024x576x97; 576x1024x97 - Infrastructure: LTX-2 Community Trainer.
Held-out evaluation (14 clips, detail relative to ground truth): input 0.84, refined 0.97 with the clip guide alone, 0.98 with a reference image attached (the reference image helps a specific subject, not the average), SSIM 0.71. Tiled at 3840x2176: detail 0.79 to 0.95, SSIM 0.82, no visible tile boundaries.
License
See the LTX-2.x Community License Agreement for full terms.
Acknowledgments
- Base model by Lightricks
- Training infrastructure: LTX-2 Community Trainer
- Downloads last month
- 4,493
Model tree for Lightricks/LTX-2.5-22b-IC-LoRA-Refine-Details
Base model
Lightricks/LTX-2.5