# Wiring Android's Vulkan Backend to a Quantized On-Device Diffusion Model for Real-Time Texture Synthesis

> Source: <https://dev.to/software_mvp-factory/wiring-androids-vulkan-backend-to-a-quantized-on-device-diffusion-model-for-real-time-texture-53a9>
> Published: 2026-10-02 13:52:41+00:00



```
---
title: "Vulkan Compute for Diffusion Under 33ms on Snapdragon 8 Gen 3"
published: true
description: "Bypass NNAPI scheduling overhead and drive a quantized INT8 diffusion model through Vulkan compute shaders, staying under 33ms per frame on Snapdragon 8 Gen 3."
tags: android, mobile, architecture, performance
canonical_url: https://mvpfactory.co/blog/vulkan-compute-diffusion-snapdragon-33ms
---
```

By the end of this guide you will have a tile-based Vulkan compute pipeline that drives a quantized INT8 stable diffusion model directly on Adreno hardware — no NNAPI dispatcher in the hot path, sub-33ms inference per frame, and the memory barrier placement that keeps your 99th percentile latency from blowing your render budget.

Let me show you a pattern I use in every project that needs deterministic GPU scheduling on Android.

`VkCommandBuffer` submission and basic Vulkan synchronization
NNAPI is easy to treat as a black box. Plug it in, let it handle hardware negotiation, then spend an afternoon debugging inconsistent latency under load with no obvious culprit.

That inconsistency is baked into the design. NNAPI's internal dispatcher negotiates between DSP, NPU, and GPU backends at runtime. On Snapdragon 8 Gen 3, that negotiation introduces overhead that compounds badly in a render loop. Mean latency of ~28ms sounds workable until you see the 99th percentile spike to ~52ms. That is not a rounding error — that is a visible stutter.

The fix is to bypass NNAPI entirely and drive your quantized model through Vulkan compute pipelines you own.

Here is the minimal setup to get this working. The execution path for each frame looks like this:

```
[ CPU: Tile Scheduler ]
        |
        v  VkCommandBuffer submission
[ Compute Shader: INT8 UNet Forward Pass ]
        |
        v  VkImageMemoryBarrier (COMPUTE_SHADER -> TRANSFER)
[ Transfer: Tile Blit to Texture Atlas ]
        |
        v  VkImageMemoryBarrier (TRANSFER -> FRAGMENT_SHADER)
[ Fragment Shader: Final Composite ]
```

Barrier placement between compute and transfer is the single biggest lever you have on stall time.

Here is the gotcha that will save you hours: a barrier using `VK_PIPELINE_STAGE_ALL_COMMANDS_BIT` as its source stage serializes the entire GPU pipeline unnecessarily.

For the inference-to-blit transition, scope it tightly:

```
VkImageMemoryBarrier barrier{};
barrier.srcAccessMask  = VK_ACCESS_SHADER_WRITE_BIT;
barrier.dstAccessMask  = VK_ACCESS_TRANSFER_READ_BIT;
barrier.oldLayout      = VK_IMAGE_LAYOUT_GENERAL;
barrier.newLayout      = VK_IMAGE_LAYOUT_TRANSFER_SRC_OPTIMAL;

vkCmdPipelineBarrier(
    cmd,
    VK_PIPELINE_STAGE_COMPUTE_SHADER_BIT,  // srcStageMask
    VK_PIPELINE_STAGE_TRANSFER_BIT,         // dstStageMask
    0, 0, nullptr, 0, nullptr, 1, &barrier
);
```

The GPU's fixed-function hardware — texture units, rasterizer — keeps running while the barrier resolves only what needs to resolve. This single change can recover several milliseconds at the 99th percentile.

Dynamic descriptor set allocation mid-frame is a silent killer. Every `vkAllocateDescriptorSets` call against an undersized pool can trigger a driver-side reallocation. The docs do not always make clear how badly this compounds over a tile grid.

Pre-allocate at pipeline initialization time:

```
VkDescriptorPoolCreateInfo poolInfo{};
poolInfo.maxSets       = MAX_TILES_IN_FLIGHT * FRAMES_IN_FLIGHT;
poolInfo.poolSizeCount = static_cast<uint32_t>(poolSizes.size());
poolInfo.pPoolSizes    = poolSizes.data();
```

For a fixed tile grid, skip `VK_DESCRIPTOR_POOL_CREATE_FREE_DESCRIPTOR_SET_BIT` — the simpler reset path is faster. Only add the flag if you genuinely need per-tile recycling.

With two tiles in flight and double-buffered descriptor sets, you can overlap the transfer of tile N with compute for tile N+1. This effectively hides blit latency behind inference latency.

On Adreno 750, a 128×128 INT8 tile inference pass through a quantized UNet fits inside 18–22ms in isolation. The remaining budget breaks down like this:

Total: comfortably under 33ms.

**Coarse barrier stages.** Replacing `VK_PIPELINE_STAGE_ALL_COMMANDS_BIT` with narrowest-applicable stage flags is not optional for a real-time pipeline. It is the first thing to audit.

**Descriptor pool sizing.** Size your pool to `tile_count × frames_in_flight` at startup. Any mid-frame reallocation invalidates your latency budget immediately.

**99th percentile, not mean.** NNAPI averages ~28ms but spikes to ~52ms at p99. Vulkan direct averages ~18ms and holds ~26ms at p99. A pipeline that stutters on every tenth frame is not acceptable regardless of its mean.

**Quantization path matters.** The barrier and descriptor set patterns above apply equally to TFLite GPU delegate and ONNX Runtime QNN backends — only the dispatch layer differs.

| Metric | NNAPI (GPU backend) | Vulkan Direct | Delta | 
|---|---|---|---|
| Mean tile inference | ~28ms | ~18ms | −10ms | 
| 99th percentile | ~52ms | ~26ms | −26ms | 
| Frame budget headroom | None | ~7ms | Available | 

If you own the target GPU and your model is quantized, Vulkan compute gives you deterministic scheduling that NNAPI cannot match. Scope your pipeline barriers to the narrowest applicable stages, pre-allocate descriptor sets at init time, and double-buffer your tile grid to hide transfer latency. Those three changes are what close the gap from 45ms to under 33ms on Snapdragon 8 Gen 3.
