Wiring Android's Vulkan Backend to a Quantized On-Device Diffusion Model for Real-Time Texture Synthesis A developer has published a guide showing how to bypass Android's NNAPI dispatcher and drive a quantized INT8 diffusion model directly through Vulkan compute shaders on Adreno hardware, achieving sub-33ms inference per frame on Snapdragon 8 Gen 3. The writeup attributes NNAPI's runtime backend negotiation to mean latency of ~28ms with 99th-percentile spikes to ~52ms, and prescribes tightly scoped VkImageMemoryBarrier transitions, pre-allocated descriptor pools, and double-buffered tiles to hide blit latency behind inference. --- title: "Vulkan Compute for Diffusion Under 33ms on Snapdragon 8 Gen 3" published: true description: "Bypass NNAPI scheduling overhead and drive a quantized INT8 diffusion model through Vulkan compute shaders, staying under 33ms per frame on Snapdragon 8 Gen 3." tags: android, mobile, architecture, performance canonical url: https://mvpfactory.co/blog/vulkan-compute-diffusion-snapdragon-33ms --- By the end of this guide you will have a tile-based Vulkan compute pipeline that drives a quantized INT8 stable diffusion model directly on Adreno hardware — no NNAPI dispatcher in the hot path, sub-33ms inference per frame, and the memory barrier placement that keeps your 99th percentile latency from blowing your render budget. Let me show you a pattern I use in every project that needs deterministic GPU scheduling on Android. VkCommandBuffer submission and basic Vulkan synchronization NNAPI is easy to treat as a black box. Plug it in, let it handle hardware negotiation, then spend an afternoon debugging inconsistent latency under load with no obvious culprit. That inconsistency is baked into the design. NNAPI's internal dispatcher negotiates between DSP, NPU, and GPU backends at runtime. On Snapdragon 8 Gen 3, that negotiation introduces overhead that compounds badly in a render loop. Mean latency of ~28ms sounds workable until you see the 99th percentile spike to ~52ms. That is not a rounding error — that is a visible stutter. The fix is to bypass NNAPI entirely and drive your quantized model through Vulkan compute pipelines you own. Here is the minimal setup to get this working. The execution path for each frame looks like this: CPU: Tile Scheduler | v VkCommandBuffer submission Compute Shader: INT8 UNet Forward Pass | v VkImageMemoryBarrier COMPUTE SHADER - TRANSFER Transfer: Tile Blit to Texture Atlas | v VkImageMemoryBarrier TRANSFER - FRAGMENT SHADER Fragment Shader: Final Composite Barrier placement between compute and transfer is the single biggest lever you have on stall time. Here is the gotcha that will save you hours: a barrier using VK PIPELINE STAGE ALL COMMANDS BIT as its source stage serializes the entire GPU pipeline unnecessarily. For the inference-to-blit transition, scope it tightly: VkImageMemoryBarrier barrier{}; barrier.srcAccessMask = VK ACCESS SHADER WRITE BIT; barrier.dstAccessMask = VK ACCESS TRANSFER READ BIT; barrier.oldLayout = VK IMAGE LAYOUT GENERAL; barrier.newLayout = VK IMAGE LAYOUT TRANSFER SRC OPTIMAL; vkCmdPipelineBarrier cmd, VK PIPELINE STAGE COMPUTE SHADER BIT, // srcStageMask VK PIPELINE STAGE TRANSFER BIT, // dstStageMask 0, 0, nullptr, 0, nullptr, 1, &barrier ; The GPU's fixed-function hardware — texture units, rasterizer — keeps running while the barrier resolves only what needs to resolve. This single change can recover several milliseconds at the 99th percentile. Dynamic descriptor set allocation mid-frame is a silent killer. Every vkAllocateDescriptorSets call against an undersized pool can trigger a driver-side reallocation. The docs do not always make clear how badly this compounds over a tile grid. Pre-allocate at pipeline initialization time: VkDescriptorPoolCreateInfo poolInfo{}; poolInfo.maxSets = MAX TILES IN FLIGHT FRAMES IN FLIGHT; poolInfo.poolSizeCount = static cast