Existing H3 latent
24 × 57 × 48 × 84
A convolutional adapter directly translates MiniMax H3 video latents into the latent space decoded by the frozen LTX-2.5 Conv Video VAE—without running the LTX diffusion transformer, latent upsampler, or refiner.
Deployment starts from an existing H3 latent—no H3 encoder is needed. Source and teacher videos below are evaluation references only.
24 × 57 × 48 × 84
384 × 25 × 24 × 42
Deterministic194.76M parameters
Learned128 × 25 × 24 × 42
192 × 768 × 1344
FrozenTwo preserved validation clips from the promoted checkpoint. The direct adapter decode is synchronized with both references.
Triptych assets are display encodes only. Metrics were computed from all 192 full-resolution decoded frames before web compression; SSIM follows the fixed deterministic frame-sampling protocol. “LTX-2.5” on this page always means the official Conv Video VAE, not the separate diffusion decoder.
Geometry is handled deterministically, leaving the network to learn the nonlinear H3-to-LTX channel basis and decoder-visible residual.
The anchor retains scale-4, native-8K, pooled-Jacobian, channel-Gram, and decoder-frequency gains. Doubling native pairs at matched exposure adds +0.01139 dB on formal exact-32. The deployed adapter remains the same 194.76M Conv3D network.
Fast-action blur is evaluated separately from average PSNR. Every new treatment must improve reconstruction without regressing the frozen highest-motion quartile.
Sixteen uniformly sampled teacher frames at 192×336 define four motion quartiles. The same mask and teacher Farneback flow are reused for parent, control, and treatment—never inferred from the candidate.
The prepared treatment replaces only the final three depthwise temporal convolutions with dense channel-time mixing: 194.759M → 199.842M parameters. Its diagonal initialization is bit-exact; a matched depthwise arm isolates the capacity effect.
A later matched A/B keeps the full 768×1344 VAE decode and adds teacher-motion-weighted moving-edge, velocity, and acceleration Charbonnier terms at 192×336. Their combined parameter-gradient budget is capped at 10% of 160×pixel-MSE.
Every promoted point uses decoded validation evidence. The latest points use the identical audited 768×1344 protocol.
Selection and reporting are deliberately separated from training proxies.
The same exact 32 videos are paired by sample ID for every formal comparison. Candidate and baseline geometries must match before deltas are computed.
Every clip contains 192 evaluated frames at 768×1344. PSNR uses all decoded pixels; temporal error compares first-frame differences.
Local and official LTX-2.5 Diffusers Conv VAE weights share SHA-256 425f0dfa…55c5 and byte size 1,452,233,194. No adapter rerun is needed to relabel the identical decoder.
Blind scale is closed: a 296.66M identity-expanded control produced no decoded gain. The remaining tests target unique-data exposure and selective temporal capacity.
One full exposure at the same 194.76M architecture adds +0.07897 dB with 32/32 PSNR wins.
Standard training adds +0.08061 dB over the prior leader with a positive exact-32 confidence interval and 32/32 PSNR wins.
The last controlled step adds +0.04270 dB over 2.00× (95% CI +0.03781 to +0.04749), with PSNR and SSIM gains on all 32 IDs.
The rebased native-data control adds +0.01266 dB (95% CI +0.00979 to +0.01584) and wins on 31/32 IDs.
Replacing the older sensitivity estimate adds +0.01035 dB (95% CI +0.00728 to +0.01358) and improves PSNR on 28/32 IDs.
The rebased direction adds +0.01249 dB on exact-32 (95% CI +0.01064 to +0.01438), with PSNR gains on all 32 IDs and SSIM gains on 31.
The sole selected point passes disjoint audit and formal exact-32, adding +0.02127 dB (95% CI +0.01610 to +0.02640). The scale grid is closed.
The stable 128×128 PSD metric adds +0.00486 dB on formal exact-32 (95% CI +0.00421 to +0.00556), with PSNR and SSIM gains on all 32 IDs and no inference-size change.
The fixed 78,400-step treatment passes fresh dev, disjoint audit, and formal gates. It adds +0.00133 dB (95% CI +0.00112 to +0.00153) with 31/32 PSNR wins and no inference-size change.
The fixed step-3,200 direction passes fresh dev, disjoint audit, and formal gates. It adds +0.01139 dB on formal exact-32 (95% CI +0.00851 to +0.01412), while SSIM adds +0.000461.
The first fixed block is complete. Internal validation changes by +0.0618 dB PSNR and +0.00146 SSIM, but this is diagnostic only. The low-frequency recovery monitor will resume the remaining audited blocks; only step 3,200 can enter formal gates.
The running endpoint covers 3,200/16,384 unique native pairs (19.53%). Only if consolidation passes all three gates but remains below 29 dB, the next hypothesis will preserve its optimizer and continue to exact step 16,384, then use new untouched development and audit splits before one formal decode.
Matched native-16K control/treatment at 1e-6 train only temporal parameters in blocks 19–21. Control has 9,024 trainable parameters; treatment has 5.092M inside a 199.842M model. An older parent showed +0.03139 dB on 32/32 IDs, but no result is claimed for the current parent.
Full-grid prediction and teacher decodes remain exact. Only the loss graph is reduced to 192×336; auxiliary weights are measured from the resolved parent's parameter gradients and the treatment is materialized only below a 76 GiB peak-allocation gate.
Standard paired PSNR/SSIM, frozen quartile-4 motion metrics, teacher-flow residuals, and eight-clip visual review are now one promotion contract. A standard-metric win alone cannot promote a future candidate.
The 12×21 crop passes the resource profile, but every bounded halo from 0 through 12 fails the unchanged worst-path full-decoder fidelity gate. Crop training is closed rather than relaxing fidelity.