You need to agree to share your contact information to access this model

By clicking "Agree and Access" you acknowledge the Privacy Policy and consent to receive offers and updates. You can unsubscribe at any time.

Log in or Sign Up to review the conditions and access this model content.

LTX-2.5 22B IC-LoRA Refine Details

An In-Context LoRA that rebuilds fine detail in video: it takes a soft, compressed or upscaled clip as its reference and renders the same frames with the texture, edges and grain of a native high-resolution capture. Used tiled, it refines and upscales any clip to 4K.

It is based on the LTX-2.5 foundation model.

Example Outputs

Prompt
Locked-off wide shot of a clear mountain stream running over mossy granite boulders under a beech forest canopy, thousands of individual leaves, sunlight dappled evenly across the scene, water sparkling with fine ripples, ferns and lichen in the foreground, deep focus, tack sharp everywhere, nature documentary, photoreal, 8K, sharp photographic detail, crisp faces and clothing texture, natural grain, high resolution footage
Prompt
Slow dolly forward down a narrow neon-lit alley at night, glowing signs reading RAMEN, OPEN 24H, KARAOKE and BAR, wet pavement reflecting the colours, paper lanterns swaying gently, a few pedestrians with umbrellas walking away from camera, deep focus, photoreal, 8K, sharp photographic detail, crisp faces and clothing texture, natural grain, high resolution footage

Two generated clips, 1920x1088, refined twice with this LoRA through the tiled fusion sampler: to 3840x2176, then to 7680x4352. Every pixel is generated; the passes add detail the 1080p source never had. Each gallery clip is a 1:1 crop comparison (the source lanczos-upscaled on the left, the 8K result on the right, 800x800 result pixels); the full-frame wipes are next to them: forest stream and neon alley. The prompts shown are the ones the refine passes ran with: the generation prompt plus sharp photographic detail, crisp faces and clothing texture, natural grain, high resolution footage. Negatives: blurry, soft focus, shallow depth of field, bokeh, motion blur, low resolution, painting, watermark (plus fog, haze, text for the stream).

Model Files

ltx-2.5-22b-ic-lora-refine-details-1.0.safetensors

Final checkpoint, step 3000 (training v1). Runs on the LTX-2.5 distilled transformer.

Model Details

  • Base Model: LTX-2.5-22B Video (trained on the 2.5 base with the Gemma-4 12B text encoder)
  • Training Type: IC-LoRA (video-to-video)
  • Control Type: reference video (the clip to refine, frame-aligned), plus an optional reference image at the -1 token (per tile)
  • Reference Downscale Factor: 1
  • Pipeline details: single pass at the trained tile size; per-step tiled latent fusion above it (see Pipeline Details)
  • Audio: Not trained for audio generation

Intended Use & Out-of-Scope

Intended use: refining and upscaling delivery-grade video (1080p web, compressed, generated clips) to 4K; a second pass on generated video to replace VAE softness with real texture; the detail stage after the Restoration IC-LoRA on archive footage.

Out of scope: colour restoration or colourisation (the model sharpens what is there and slightly desaturates archive footage: run the Restoration IC-LoRA first); pristine cinema raw (it denoises sensor grain rather than adding detail); native 8K refinement in one pass (see Tips).

Control Signal Requirements

  • Control signal type: the source video itself, resized to the output canvas (lanczos), fed as the IC-LoRA reference at downscale factor 1.
  • Expected input: any resolution; the output canvas must be a multiple of 32 on both sides and the frame count 8n+1 (49, 97, 121, 241...). For a 1920x1080 source, refine at 3840x2176 or crop to 1920x1056 first.
  • Preprocessing: none beyond the resize. Do not sharpen or denoise the input first.
  • Alignment: 1:1, same frames, same frame rate. The output is the input's geometry with rebuilt detail.
  • Mask support: none. Optional reference image at the -1 token: one to three native-resolution crops of the subject on a grey image, attached with a second LTXAddVideoICLoRAGuide at frame_idx -1; it steers products, text, logos and faces and is not required (the per-tile rule is in Recommended Settings).

How It Works

The model was trained on pairs made from 4K footage: the target is a real downscale of each tile to 1024x576 (never an upscaled target), and the reference is the same tile degraded to look like a soft, upscaled source at a range of severities. It therefore learned exactly one thing: given a frame whose fine band is missing or smeared, rebuild that band consistently over time. It does not change framing, exposure or colour. Its working window is a 1024x576 tile; larger canvases are refined as overlapping tiles fused every denoising step, which keeps one coherent frame rather than a mosaic. At 4K the content inside a tile is at the scale the model trained on; at 8K a tile sees a 2x close-up and starts inventing texture, so 8K is reached progressively.

Usage

ComfyUI

  1. Copy the LoRA weights into models/loras.
  2. Open LTX-2.5_V2V_TiledFusion_Upscale.json from the ComfyUI-LTXVideo repository. It loads the LTX-2.5 distilled transformer, the Gemma-4 12B text encoder, the LTX-2.5 video VAE and this LoRA (ltx-2.5-22b-ic-lora-refine-details-1.0.safetensors at 1.0) and runs LTXVTiledFusionSampler with 1024x576 tiles.
  3. Drop the clip on LoadVideo, set output_size (FullHD, 4K or 8K), write a look-and-style prompt, run. The workflow's Preprocess group first upscales your clip to the output size with lanczos and feeds that as the guide; the model never sees the small original. Long clips are handled by 97-frame windows (use_streaming is on in the workflow).
  4. Use the tiled workflow for every output size, including full HD. The LoRA was trained on 1024x576 tiles and gives its best result when each tile sees content at that scale; a 1920x1088 canvas is a 3x3 grid of those tiles. A single untiled pass over a full-HD frame puts the LoRA outside its training bucket. The single-tile LTX-2.5_V2V_ICLoRA_Single_Stage_Distilled.json is right only when the output itself is 1024x576 or smaller, i.e. when the whole frame is one tile.
  5. Building your own graph, in this order: lanczos-resize the clip to the output canvas (a multiple of 32) β†’ LTXAddVideoICLoRAGuide (latent_downscale_factor 1) β†’ LTXVTiledFusionSampler (tile 1024x576, overlap 0.5, blend 0.05, cfg 1.0) on the 8-step distilled sigmas with LTXICLoRALoaderModelOnly (strength 1.0) on the distilled model β†’ VAEDecodeTiled.

Pipeline Details

Refinement above the trained tile runs as per-step tiled latent fusion:

  • One canvas, one noise field. The whole frame is a single latent; every denoising step crops overlapping 1024x576 windows (50% overlap), runs each window conditioned on its own crop of the reference, and merges the stepped windows back with a Gaussian weight (blend variance 0.05). Because the merge happens inside the step loop, the windows stay coupled and there are no seams and no per-tile flicker.
  • Upscale ladder. Resize the source to the output canvas with lanczos and refine once. 1080p to 4K is one pass (about 22 minutes for 121 frames on an RTX PRO 6000). 8K is best reached as refine to 4K, lanczos to 8K, optional gentle second pass.
  • Long clips. Up to about 240 frames can be one temporal extent. Beyond that, or for tighter memory, enable windowed guides (use_streaming on the guide node with tile_frames 97, and the same 97 on the sampler): 97-frame windows, each with its own encoded reference, fused the same way.
  • Decode tiled (1024 px, 128 overlap, temporal 256).

Recommended Settings

  • LoRA strength / weight: 1.0
  • Inference steps: 8 (distilled sigmas 1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875, 0.0, euler)
  • Guidance scale: 1.0
  • Input preparation: upscale the clip to the output canvas with lanczos before it becomes the guide (the Upscale workflow does this); the guide should already be at the output size, not the small original.
  • Resolution & frames: tile 1024x576 (or 576x1024 portrait), canvas any multiple of 32, 97-frame windows; 4K output is the sweet spot for a 1080p source
  • Tiled at every output size: the closer each tile is to the training bucket (1024x576), the better the LoRA works, so run tiled even for a full-HD output rather than one untiled pass.
  • Tiling: overlap 0.5, blend variance 0.05, grid cycle 1; never lower the overlap below 0.5 (phase breaks on fine periodic texture)
  • Prompting: keep the prompt generic and about the rendering (sharpness, texture, grain, light), not the subjects. Every tile receives the whole prompt, so a subject named in it can be painted into tiles that do not contain it. Name a subject only when you prompt per tile: run the tile that holds it as its own single-tile pass with its own prompt. That is also how to use a generated clip's generation prompt (it helps heavily degraded objects, but only in the tile that holds them), never on the whole canvas.
  • Reference image at the -1 token (products, text, logos, faces): when a subject is ambiguous after degradation, give the model the real thing: native-resolution crops of it on a grey image, attached with a second LTXAddVideoICLoRAGuide (frame_idx -1, strength 1.0) chained after the clip guide, on the whole-clip path (use_streaming off). The reference is per tile: the LoRA learned it against the 1024x576 tile that holds the subject, and the fusion sampler crops the reference image to each tile exactly as it crops the clip. So make the reference image the size of the output canvas and place each crop over its own subject; every tile that covers the subject then sees it. Crops placed anywhere else reach the wrong tiles: on a night street scene the placed reference restored a shop-sign word and brought a face back towards the truth, while the same crops in the wrong tiles garbled the word and lowered similarity to the truth. A reference also raises fine-detail energy across the whole frame, so use it for a subject, not by default.

Example positive prompt:

sharp photographic detail, crisp natural texture, fine surface detail, clean edges, natural film grain, high resolution footage

Example negative prompt:

blurry, soft, plastic, smeared detail, oversharpened halos, warped faces

References

Tips & Troubleshooting

  • An upscale cannot exceed its source. Deep focus, even light and dense texture measured up to 3x the detail of a plain lanczos upscale; shallow depth of field measured 1.0x on the out-of-focus parts. Judge the in-focus region at 1:1.
  • Archive or broadcast footage looks cleaner but emptier. The model was trained on clean downscales, so it reads codec damage in the mid band as noise and smooths it, and it desaturates old footage. Run the Restoration IC-LoRA first, then refine.
  • Cinema raw (pristine sensor footage) comes out slightly softer than a lanczos upscale: the model removes grain it cannot rebuild. Not the tool for that source.
  • 8K in one pass invents detail. A 1024x576 window covers 13% of an 8K frame, twice the trained scale; melons rearrange and sign text changes. Refine at 4K, then upscale.
  • Stochastic texture (fireworks, sea foam, foliage in wind) can over-sharpen. Lowering strength costs more detail elsewhere than it saves there; prefer a descriptive prompt.
  • Text, logos and signage can turn into glyph-shaped noise after heavy degradation: the model re-renders letters it cannot read. Give it the real sign as a reference at the -1 token, placed over the sign's own position (see Recommended Settings), or run that tile on its own with a prompt that names the text. Very small lettering is not recovered either way.
  • The last one or two frames of any LTX generation lose about 10% detail. In windowed runs the overlap hides it; for a single extent, pad the clip by 8 frames and trim.
  • The mp4 written by ComfyUI's SaveVideo (about 8 Mbps at 4K) understates the result. Save the latents and re-decode, or encode from PNG frames.

Dataset

Paired clean and degraded 97-frame tiles built from 4K footage across 24 content families (people, animals, products, architecture, nature, textures, vehicles, crowds, generated footage). The target is a native downscale of each tile to 1024x576 (576x1024 for portrait); the reference is the same tile degraded to simulate a soft, upscaled source at a range of severities. Pairs also carry a reference image (the -1 token input): native-resolution crops of the subject on a grey image. 24 fps, one shared caption.

Training

  • Technique: LTX-2 trainer flexible strategy with two conditions: reference (the degraded clip, downscale factor 1, always present) and reference_frames (the reference image described under Recommended Settings, present at half the steps, so the model works both with and without a reference image at inference).
  • Hyperparameters: rank 128, alpha 128, targets attn1/attn2 q/k/v/out and ff, AdamW lr 1e-4, cosine schedule, batch 1, grad clip 1.0, bf16, no quantization, gradient checkpointing, seed 42
  • Steps: 3000 (checkpoint every 250; step 3000 shipped)
  • Resolution buckets: 1024x576x97; 576x1024x97
  • Infrastructure: LTX-2 Community Trainer.

Held-out evaluation (14 clips, detail relative to ground truth): input 0.84, refined 0.97 with the clip guide alone, 0.98 with a reference image attached (the reference image helps a specific subject, not the average), SSIM 0.71. Tiled at 3840x2176: detail 0.79 to 0.95, SSIM 0.82, no visible tile boundaries.

License

See the LTX-2.x Community License Agreement for full terms.

Acknowledgments

Downloads last month
4,493
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Lightricks/LTX-2.5-22b-IC-LoRA-Refine-Details

Adapter
(32)
this model

Space using Lightricks/LTX-2.5-22b-IC-LoRA-Refine-Details 1

Collection including Lightricks/LTX-2.5-22b-IC-LoRA-Refine-Details