Pyramid-JiT: Low-Res Drafts, High-Res Images A research team released Pyramid-JiT (P-JiT), a decoder-only pixel-space architecture that predicts target images at several resolutions along the DiT trunk, reaching Linum v2's FD-DINOv2 score with 11.3× fewer training samples and 4.3× fewer GPU-hours at 4× the pixels. The P-JiT code and model weights are available under the Apache 2.0 license, described as a research artifact rather than a full model release, en route to the Linum v3 video model. The team targets cutting video generation cost by an order of magnitude, noting a 720p, 5-second video costs 110K tokens for Linum v2 versus context sizes of 8K or less for 97% of LLM training regimes. In our last post, we developed a pixel-space encoder-decoder architecture JiT-DDT . By throwing away the VAE, we were able to compress tokens more aggressively. In turn, this slashed the cost of attention and directly led to faster training and inference. While effective, it's unclear how to scale encoder-decoder architectures under fixed parameter budgets. So today, we propose the Pyramid-JiT P-JiT , a novel decoder-only pixel space architecture that beats our prior models, by predicting the target image at several resolutions along the DiT trunk. It reaches Linum v2's FD-DINOv2 on 11.3× fewer training samples, and trains in 4.3× fewer GPU-hours at 4× the pixels. Research release P-JiT code and model weights are available under the Apache 2.0 license. We hope that by sharing our findings with the broader community, we can encourage others to also explore more efficient training methods. This should be treated as a research artifact, not a full model release. Stay tuned for more research checkpoints like this, en route to Linum v3. Where should the parameters go? Our goal is to build affordable animation tools, so that anyone can bring their story to screen. Right now the best video models in the world e.g. Seedance are expensive because they calculate attention on ridiculously large token sequences. For our own Linum v2 model https://linum.ai/field-notes/launch-linum-v2 , a short 720p, 5 second video cost 110K tokens. By comparison, a LLM trains on context sizes of 8K or less for 97% of their training regime https://arxiv.org/pdf/2512.13961 . To make our vision a reality, we need to cut the cost of video generation by an order of magnitude. We're bottlenecked by attention, which is an operation with respect to context-window token count. If we can cut down the number of tokens, we can slash the cost of attention and achieve massive savings across training and inference. En-route to our Linum v3 video model, we've been using text-to-image as a testbed to try out ideas that can help us condense context windows. Previously, we wrote https://linum.ai/field-notes/jit-ddt hitting-the-vae-compression-wall about how LDMs Latent Diffusion Models seem to have an empirical ceiling of 16x16 token compression unless you tweak the auto-encoder significantly. So, we pivoted towards pixel space models JiT https://arxiv.org/pdf/2511.13720 to achieve 32x32 token compression, unifying compression and generation into a single streamlined model. From there, we developed a novel pixel-space architecture JiT-DDT that produced better results than our prior LDM while training 3.6x faster https://linum.ai/field-notes/jit-ddt appendix . While effective, it's unclear how to scale encoder-decoder architectures. Do you increase the encoder's size so that it learns structure better? If so, it seems wasteful to spend so many parameters predicting a low resolution image. Or, do you increase the decoder's size so that it spends more time learning the fine-grained details that really matter to humans? In doing so, you might cap the model's ability to learn the framing and structure of a scene without which the details are meaningless . Language faced a similar dilemma ~10 years ago. If you dig through the archives, you'll see that the transformer was first proposed as the building block for an encoder-decoder trained on language translation. The field then split into encoder-only models BERT-style classifiers and decoder-only models GPT-style LLMs . The beauty of a decoder-only model is that it's really easy to scale, since there is no explicit architectural bottleneck. So the driving question is: Can we port over the advances from encoder-decoder back to the vanilla JiT and get similarly high quality results? Revisiting the JiT In our prior work, we found that we could dramatically improve the quality of our generations by widening the noise schedule during training and adding separate DiT blocks to refine the image and text streams independently, before merging them into a single shared stream. So, we copied and pasted these innovations over to the JiT. Additionally, we widened our linear patchification's bottleneck dimension from 256 to 1024 to reduce the effective compression rate within the network from 12x to 3x. And, we picked up the innovations from Z-Image https://linum.ai/field-notes/jit-ddt architecture-refinements . With these updates, we're able to recover a lot of performance, but we're still missing two big deltas from the JiT-DDT: multiresolution prediction and dual patchification. Replacing the encoder with readouts In the JiT-DDT, we had the encoder predict a 64x64 version of the 512x512 input image. Our hypothesis was this would help the model draft an initial low-res answer so that the decoder could allocate its capacity to filling in details. When we zoomed out https://linum.ai/field-notes/jit-ddt why-does-the-jit-ddt-work and looked at adjacent work like RAE v2 and Self-Flow, it seemed our encoder prediction served two more concrete purposes. First, it boosted the gradient signal in the early layers of the DiT that tend to adapt much more slowly. And second, by regressing a low resolution image it helped the early portions of the DiT explicitly learn structure something that hacks like REPA tap into to accelerate learning as well . Even if we throw away the encoder, we can still access these improvements. Since our model is trained with x-prediction, every intermediate DiT block is implicitly predicting a version of our target image. So we can slap K OutputHeads along the trunk of the DiT and have each regress a version of our target image either at full resolution or downsampled to a lower resolution, like we did in the JiT-DDT . We tried predicting full-resolution images along the network but that performed slightly worse than predicting images in an inverted pyramid of increasing resolutions from 128² to 256² to 512². Another hunch we had was that we might want to place less weight on the low-resolution readouts so the model could allocate more features for high-resolution detail. In practice that worked worse than equal-weighted readouts; so 3-Readout became our new baseline for our subsequent experiments. Multiview networks The JiT-DDT had two different patchifications, one for the encoder 64x64 and another for the decoder 32x32 . 3-Readout trains ~16% faster than JiT-DDT, even though it uses the more expensive 32x32 patchification. And aesthetically we seem to beat the JiT-DDT. So rather than bring the encoder's 64x64 patchification back in the mix, we were curious what if we could raise the ceiling on the model's generation quality by arming it with several 32x32 patchifications. Our hypothesis was that the model might be able to modulate between a set of patchifications according to timestep. Patchification A might be more active for details while Patchification B might be more active for structure. By enabling this type of soft specialization kind of like Mixture of Experts , we might be able to reduce the information lost to compression in a single linear bottleneck, improve the overall expressivity of the model, and enable both better structure learning and detail learning. Widening the residual stream Now for a slight detour we promise it'll all make sense soon . Over the past 12 months, LLM researchers out of China have recontextualized the skip connection as a memory buffer for the LLM. If we can increase the residual stream's expressivity e.g. widen it, introduce more complicated read/write rules , we can significantly improve the network's performance. DeepSeek mHC https://arxiv.org/pdf/2512.24880 , Kimi Attention Residuals https://arxiv.org/pdf/2603.15031 , and Qwen Gated Residuals https://github.com/QwenLM/Qwen3.8-Flash-Next/blob/main/tech report.pdf are all different articulations of this general idea. DeepSeek's mHC and Qwen's Gated Residuals are particularly exciting, because they don't increase the cost of attention. Moreover, the wider residual stream is a perfect place to bring our idea of multiple 32x32 patchifications to life. Qwen and DeepSeek initialized the wider residual stream by copying and pasting their representations along the channel dimension 4x at the start of their networks. Instead of copying it 4 times, we can just introduce 2 patchifications each duplicated twice or 4 distinct patchifications. Adapting gated residuals to flow matching In Qwen 3.8-Next's head-to-head ablations, the authors found that Gated Residuals worked better than mHC and matched Attention Residuals. It's the easiest to implement, so we ran with that instantiation of the expressive residual idea. The one roadblock to copy-and-pasting the Gated Residual into our flow matching model is timestep conditioning. In LLMs, you're predicting a series of auto-regressive tokens; there is no concept of timesteps or noised inputs. There's a paper from early 2025 https://arxiv.org/pdf/2502.13129 out of Kaiming He's lab that demonstrates that flow matching models should be able to implicitly estimate timestep from the noised sample . So for our first attempt at porting Gated Residuals, we gutted timestep modulation AdaLN scaling altogether. It produced notably worse results, so we went back to the whiteboard and brainstormed ways to reintroduce timestep conditioning. At a high level, timestep conditioning can be either a shift i.e., an added bias or a scale i.e., a multiplier on the existing features . A scale stretches or shrinks each feature across the image. Features with large absolute values get pitched up or down significantly, while features that hover around zero stay there regardless of the multiplicative factor. A shift slides every token by the same offset. That records the timestep in the hidden state's absolute values but leaves the relative differences between features untouched. In our 3-Readout model we rely entirely on scales, because we want to stretch and shrink the hidden states based on timestep e.g., turn up structure at high noise, turn up detail at low noise . So, we opted for scales here again. On the read, ① scales the hidden state within the sigmoid gate. This way timestep conditioning doesn't mess with the bounded read. We do this in the smaller 368-dimensional space to be parameter efficient. This costs us 94K parameters per gate vs. 3M parameters per gate in the full 11,776-dimensional gate-logit space i.e., if we did . ② is just a replication of our old AdaLN method. For writes, we can either condition the mixing function or scale the MLP/Attn outputs . ③ adds the timestep to the write gate's logit per-token-per-stream-per-channel.