{"slug": "neural-gpu-block-compression", "title": "Neural GPU block compression", "summary": "A developer has built a neural GPU texture block compression system that uses Evolution Strategies to train a small MLP decoder, achieving 5.0 bpp of amortized latent data across a texture stack. The 1,731-weight network takes five latent values plus local block UV coordinates as input and supports PBR materials across correlated texture arrays, with a variant adding in-loop deblocking that improves quality by roughly 0.9 dB in a 4x4 test. The developer notes that because ES is derivative-free, the GPU texture transcoder itself can be placed inside the optimization loop, allowing objectives such as neural decode followed by BC7 encode and decode without requiring differentiability.", "body_md": "Neural texture block decoding, trained with Evolution Strategies (ES), 3,000 iterations.\n\nThis example uses 2 source latent texture: A full-res 512x512 at 3-bpp, and 1/4 res (128x128) with 4 channels at 8 bpp=32 bpp. Latent is sampled using nearest sampling, i.e. it's just a trained block decoder. It uses 5.0 bpp of latent data amortized across the entire texture stack (in this example just 1 texture, but it supports up to 4).\n\nNeural net (MLP leaky ReLU, sigmoid on the output): 7→36→ 36→3, 1,731 weights\n\nMLP inputs: 5 latent values + local block UV texel coordinate as [-1,1]\n\nThe full-res latent texture can be 4-bpp, which looks noticeably better, or 2-bpp etc. Further lossy or lossless compression can be applied to the latent data.\n\nThis is a practical approach and it works on PBR materials (i.e. across multiple related textures in a correlated texture array). Prior art disclosure repo is next.\n\nI have a variant that adds in-loop deblocking using a 5 tap filter which ES factors in before computing loss, which boosts quality by ~0.9 dB in one 4x4 test.\n\nSide by side - right is compressed:\n\nTop-level (full resolution) 3-bit latent data, visualized as a 2D texture:\n\nThe top-level can have X channels, each a different number of bits, down to even 1-bit per channel. This roughly corresponds to dual or single plane modes in ASTC/BC7.\n\nSecond level (quarter resolution) \"control\" latent data, visualized as a 2D texture (4 channels, 8-bits per channel):\n\nThe \"control\" level can have a varying number of channels, not just 4, depending on the complexity of the texture or material.\n\nNote this per-block data can be stored into compressed block packets, not textures: much like a classic GPU texture.\n\nSame settings, but using a 4-bit top level latent texture:\n\nI have some ideas on how to fix the BC1-like block artifacts on chroma blocks.\n\nOne especially interesting consequence of using ES for training and inference on load (i.e. transcoding to a compressed texture like BC1-7/ASTC): the *GPU texture transcoder itself can be inside the optimization loop*. Since we're using derivative-free ES, the objective could literally be:\n\nneural bits → neural decode → BC7 encoder → BC7 decode → compare against source\n\nThe BC7 encoder can contain discrete mode decisions, bit quantization, endpoint selection, partition search, etc. None of that needs to be differentiable.", "url": "https://wpnews.pro/news/neural-gpu-block-compression", "canonical_source": "https://richg42.blogspot.com/2026/09/neural-gpu-block-compression.html", "published_at": "2026-09-05 05:38:59+00:00", "updated_at": "2026-09-14 17:25:17.413601+00:00", "lang": "en", "topics": ["machine-learning", "neural-networks", "computer-vision", "ai-research", "ai-infrastructure"], "entities": [], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/neural-gpu-block-compression", "markdown": "https://wpnews.pro/news/neural-gpu-block-compression.md", "text": "https://wpnews.pro/news/neural-gpu-block-compression.txt", "jsonld": "https://wpnews.pro/news/neural-gpu-block-compression.jsonld"}}