I rebuilt a retired photo-editing site so that every model runs client-side: object removal (MI-GAN, LaMa), background removal (RMBG-1.4) and 4× upscaling (Real-ESRGAN), all through ONNX Runtime Web with WebGPU and a WebAssembly fallback. No uploads, no server, models downloaded once after the user agrees and cached in Cache Storage.
It works, but three things broke in ways I did not expect. None of them threw an error that pointed at the real cause.
For background removal, BiRefNet-lite (MIT) beat RMBG-1.4 clearly in my offline comparison on ten images. In the browser on an Apple GPU, the first session.run() failed with:
Too many storage buffers in shader. Current: 11, Max is 10
WebGPU on Apple hardware reports maxStorageBuffersPerShaderStage = 10, and one of the fused kernels ONNX Runtime generates for this model needs 11. You cannot request more than the adapter offers. Dropping the graph optimisation level did not help, and the WebAssembly backend hit std::bad_alloc because the 1024×1024 transformer activations do not fit in a 4 GB wasm32 heap. BEN2 failed the same way.
Lesson: benchmark candidate models in the target browser on the weakest GPU you care about before comparing quality. A convolutional model (RMBG-1.4) runs in 0.25 s on WebGPU and ~6 s on WASM; the "better" model does not run at all.
LaMa ran on WebGPU without any error. The output tensor had the right shape and values in 0–255. But the inpainted hole was almost pure white:
webgpu hole mean=254.3 outside mean=127.0
wasm hole mean=107.3 outside mean=127.0
LaMa relies on Fourier convolutions (RFFT/IRFFT), and on the WebGPU execution provider those produced wrong values. My fallback chain (WebGPU → less optimised graph → WASM) only triggers on exceptions, so it never kicked in. LaMa now always runs on WASM, and the end-to-end tests assert on pixel colours of real outputs, not just "the run finished".
Lesson: test pictures, not the absence of errors.
Real-ESRGAN x4plus is shipped in fp16 with fp16 inputs and outputs. I encoded inputs into a Uint16Array and decoded outputs from raw half-float bits. On current Chrome the upscaled image came out solid black. Chrome now has a native Float16Array, and ONNX Runtime Web returns fp16 outputs as real numbers when it is available. Decoding 0.5 as if it were a bit pattern gives roughly zero.
The fix is to accept both:
const out = raw instanceof Uint16Array
? Float32Array.from(raw, halfBitsToNumber)
: Float32Array.from(raw as ArrayLike<number>);
The same model also failed on WebGPU with Shape mismatch attempting to re-use buffer until I pinned its symbolic dimensions with freeDimensionOverrides: {N: 1, H: 192, W: 192} and fed it fixed-size tiles.
env.wasm.wasmBinary. connect-src 'self' and no third-party scripts, embedded in the content site by iframe.
If you want to poke at it: https://inpainting.app/ — the full write-up with examples is on the site.
This post was written with AI assistance and reviewed by the author.