Tech News
← Home  ·  All topics

Text To Image Models

1 GoKawiil brief on this topic

Linum's new JiT-DDT model cuts text-to-image training time 3.6x

Linum has released JiT-DDT, an experimental encoder-decoder architecture that trains text-to-image generation models using 3.6 times fewer GPU-hours than its previous Linum v2 system, while producing images with four times the pixel resolution. The approach merges compression and generation into a single pixel-space model rather than relying on separate VAE and diffusion transformer components, addressing a detail-loss problem seen in earlier pixel-space designs. Code and weights are being released under Apache 2.0 as a research checkpoint on the way to Linum v3.