cd /news/machine-learning/neural-gpu-block-compression · home › topics › machine-learning › article
[ARTICLE · art-129394] src=richg42.blogspot.com ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

Neural GPU block compression

A developer has built a neural GPU texture block compression system that uses Evolution Strategies to train a small MLP decoder, achieving 5.0 bpp of amortized latent data across a texture stack. The 1,731-weight network takes five latent values plus local block UV coordinates as input and supports PBR materials across correlated texture arrays, with a variant adding in-loop deblocking that improves quality by roughly 0.9 dB in a 4x4 test. The developer notes that because ES is derivative-free, the GPU texture transcoder itself can be placed inside the optimization loop, allowing objectives such as neural decode followed by BC7 encode and decode without requiring differentiability.

by read2 min views24 publishedSep 5, 2026

Neural texture block decoding, trained with Evolution Strategies (ES), 3,000 iterations.

This example uses 2 source latent texture: A full-res 512x512 at 3-bpp, and 1/4 res (128x128) with 4 channels at 8 bpp=32 bpp. Latent is sampled using nearest sampling, i.e. it's just a trained block decoder. It uses 5.0 bpp of latent data amortized across the entire texture stack (in this example just 1 texture, but it supports up to 4).

Neural net (MLP leaky ReLU, sigmoid on the output): 7→36→ 36→3, 1,731 weights

MLP inputs: 5 latent values + local block UV texel coordinate as [-1,1] The full-res latent texture can be 4-bpp, which looks noticeably better, or 2-bpp etc. Further lossy or lossless compression can be applied to the latent data.

This is a practical approach and it works on PBR materials (i.e. across multiple related textures in a correlated texture array). Prior art disclosure repo is next.

I have a variant that adds in-loop deblocking using a 5 tap filter which ES factors in before computing loss, which boosts quality by ~0.9 dB in one 4x4 test.

Side by side - right is compressed:

Top-level (full resolution) 3-bit latent data, visualized as a 2D texture:

The top-level can have X channels, each a different number of bits, down to even 1-bit per channel. This roughly corresponds to dual or single plane modes in ASTC/BC7.

Second level (quarter resolution) "control" latent data, visualized as a 2D texture (4 channels, 8-bits per channel): The "control" level can have a varying number of channels, not just 4, depending on the complexity of the texture or material.

Note this per-block data can be stored into compressed block packets, not textures: much like a classic GPU texture.

Same settings, but using a 4-bit top level latent texture:

I have some ideas on how to fix the BC1-like block artifacts on chroma blocks.

One especially interesting consequence of using ES for training and inference on load (i.e. transcoding to a compressed texture like BC1-7/ASTC): the GPU texture transcoder itself can be inside the optimization loop. Since we're using derivative-free ES, the objective could literally be:

neural bits → neural decode → BC7 encoder → BC7 decode → compare against source

The BC7 encoder can contain discrete mode decisions, bit quantization, endpoint selection, partition search, etc. None of that needs to be differentiable.

── more in #machine-learning 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/neural-gpu-block-com…] indexed:0 read:2min 2026-09-05 · —