{"slug": "vulkan-back-end-coming-to-sycl", "title": "Vulkan back end coming to SYCL", "summary": "AdaptiveCpp has added a Vulkan backend and Android cross-compilation support to SYCL, bringing portable GPU compute to Android devices that ship with Vulkan implementations. The project's GitHub CI now tests the Vulkan backend on Linux, Windows, and macOS on every commit, using Mesa llvmpipe on Linux and Windows and MoltenVK on macOS, with Ubuntu native builds requiring the LunarG Vulkan SDK 1.4.357.0, a Vulkan driver, the clspv compiler, and AdaptiveCpp itself. The backend closes a gap in SYCL's performance-portability story on mobile, where OpenCL-over-Vulkan projects such as clvk, pocl, and ANGLE had previously seen more success.", "body_md": "# \n[Bringing SYCL to Android: A Vulkan backend for portable GPU compute](https://adaptivecpp.github.io/hipsycl/adaptivecpp/sycl/vulkan/android/vulkan-android/)\n\n# Introduction\n\nSYCL’s promise is performance portability: write modern C++ once and execute across many different accelerators. But that promise only goes as far as the available backends. While desktop and HPC platforms capitalize on established OpenCL, CUDA, HIP, and Level Zero backends for accelerating SYCL applications on the GPU, mobile isn’t a domain commonly associated with SYCL. Yet almost every Android device ships with a capable Vulkan implementation, making mobile an unfulfilled chapter of the SYCL performance portability story.\n\nWhile projects such as [Sylkan](https://dl.acm.org/doi/10.1145/3456669.3456683) have demonstrated that SYCL over Vulkan is\nfeasible, the space has remained relatively unexplored in terms of feature completeness, Android support, and integration\nwith the wider SYCL ecosystem. OpenCL has had more success in this area, with projects such as\n[clvk](https://github.com/kpet/clvk), [pocl](https://github.com/pocl/pocl), and [ANGLE](https://github.com/google/angle)\nsuccessfully layering the OpenCL compute API over Vulkan.\n\nAdaptiveCpp closes this SYCL gap with its new Vulkan backend and Android cross-compilation support. Here is what we’ll cover in this blog post:\n\n- \n[Getting Started](#getting-started) : How to build AdaptiveCpp with the Vulkan backend on desktop so you can try it yourself.\n- \n[Benchmarks](#android-benchmarks) : Early performance results from bringing SYCL GPU acceleration to Android.\n- \n[Under the Hood](#under-the-hood) : A deep dive into how AdaptiveCpp layers over Vulkan, exploring the runtime and the compiler.\n\n# Getting Started\n\nWhen it comes to using the Vulkan AdaptiveCpp backend, thanks to the proliferation of Vulkan drivers,\nthere are many platforms on which the backend can be tested. We’re proud that AdaptiveCpp GitHub CI now has\nall of Linux, Windows, and macOS operating systems tested on the Vulkan backend on every commit,\nusing [Mesa llvmpipe](https://docs.mesa3d.org/drivers/llvmpipe.html) for Linux and Windows,\nand [MoltenVK](https://github.com/KhronosGroup/MoltenVK) for macOS.\n\nIn this article, we’ll only cover how to build and use the backend on Ubuntu using a native build flow.\nAndroid requires a more complex build process using the Android Native Development Kit (NDK) to\ncross-compile AdaptiveCpp and other dependencies. You can find the in-depth instructions for how to do that\n[here](https://github.com/AdaptiveCpp/AdaptiveCpp/blob/develop/doc/install-android.md).\n\n## Building AdaptiveCpp With The Vulkan Backend\n\nThere are four main pieces to assemble before you can compile and run your first SYCL program for Vulkan: the [LunarG SDK](#lunarg-sdk),\na [Vulkan driver](#vulkan-driver), the [clspv compiler](#clspv), and [AdaptiveCpp](#adaptivecpp) itself.\nThe first build dependency is the LunarG Vulkan SDK which contains SPIR-V Tools, Vulkan layers and the loader, and the VulkanHpp headers.\nThe second is a `clspv` executable, which can be built following the GitHub repo instructions. A Vulkan driver is also required to run the backend.\nFinally, we need to build AdaptiveCpp with the Vulkan backend enabled.\n\n### LunarG SDK\n\nDownload the Vulkan SDK tarball from the [LunarG website](https://vulkan.lunarg.com/sdk/home) and decompress it.\nThe LunarG SDK can then be made available in your system path after sourcing the `setup-env.sh` script that it\nships (we recommend automatically sourcing this as part of your `.bashrc`).\n\n``` bash\n$ wget https://sdk.lunarg.com/sdk/download/1.4.357.0/linux/vulkansdk-linux-x86_64-1.4.357.0.tar.xz\n$ tar -xvf vulkansdk-linux-x86_64-1.4.357.0.tar.xz\n$ source 1.4.357.0/setup-env.sh\n```\n\n### Vulkan Driver\n\nThe Vulkan SDK comes with a `vulkaninfo` tool for printing the Vulkan drivers on your system.\nAt least one driver is required to use as a SYCL backend device. If you don’t have any installed then the\neasiest way to reliably get a supported driver is to install the Mesa drivers with\n`apt install mesa-vulkan-drivers`. This will provide at least the llvmpipe CPU Vulkan driver\nwhich provides all the necessary capabilities for SYCL. For example:\n\n``` bash\n$ vulkaninfo --summary\nGPU0:\n\tapiVersion         = 1.4.318\n\tdriverVersion      = 25.2.8\n\tvendorID           = 0x10005\n\tdeviceID           = 0x0000\n\tdeviceType         = PHYSICAL_DEVICE_TYPE_CPU\n\tdeviceName         = llvmpipe (LLVM 20.1.8, 256 bits)\n\tdriverID           = DRIVER_ID_MESA_LLVMPIPE\n\tdriverName         = llvmpipe\n\tdriverInfo         = Mesa 25.2.8-0ubuntu0.25.10.2 (LLVM 20.1.8)\n\tconformanceVersion = 1.3.1.1\n```\n\n### clspv\n\nA `clspv` executable is required to be invoked at runtime as part of runtime kernel compilation,\nand can be built from [source](https://github.com/google/clspv). The exact commit of clspv should be checked in AdaptiveCpp CI\nor [doc/install-vulkan.md](https://github.com/AdaptiveCpp/AdaptiveCpp/blob/develop/doc/install-vulkan.md#requirements)\nfor the SHA hash.\n\n``` bash\n$ git clone https://github.com/google/clspv\n$ cd clspv\n$ git checkout <supported commit>\n$ python3 utils/fetch_sources.py\n$ mkdir build && cd build\n$ cmake .. -GNinja\n$ ninja\n$ export CLSPV_BIN_DIR=$PWD/bin\n```\n\nThe path to the directory with the `clspv` tool is exported as an environment variable so we\ncan reference it in a later build step.\n\n### AdaptiveCpp\n\nNow we can build AdaptiveCpp itself using these dependencies. If you haven’t got it already,\nyou first need to clone the [AdaptiveCpp GitHub repo](https://github.com/AdaptiveCpp/AdaptiveCpp).\nCombined with the `-DWITH_VULKAN_BACKEND=ON` option for enabling the Vulkan backend in the build,\nthe relevant parts of the CMake invocation are:\n\n``` bash\n$ git clone https://github.com/AdaptiveCpp/AdaptiveCpp.git\n$ cd AdaptiveCpp && mkdir build && cd build\n$ cmake -GNinja -DWITH_VULKAN_BACKEND=ON -DCMAKE_PROGRAM_PATH=$CLSPV_BIN_DIR -DCMAKE_INSTALL_PREFIX=$PWD/install\n$ ninja install\n$ export ACPP_BIN_DIR=$PWD/install/bin\n```\n\nYou can then check that a Vulkan device is indeed available. Note that other devices may be available too,\nbut the below is the minimum expected number of devices that `acpp-info` should output.\n\n```\n$ $ACPP_BIN_DIR/acpp-info -l\n=================Backend information===================\nLoaded backend 0: OpenMP\n  Found device: AdaptiveCpp OpenMP host device\nLoaded backend 1: Vulkan\n  Found device: llvmpipe (LLVM 20.1.8, 256 bits)\n```\n\n## Running Applications On The Vulkan Backend\n\nLet’s build and run a simple SYCL application with AdaptiveCpp to show the Vulkan backend in action. All this application does is use a 1D kernel to initialize a device USM allocation with the index value of each element, then verifies on the host that each element was initialized with its expected value.\n\n```\n// sycl_test.cpp\n#include <iostream>\n#include <vector>\n#include <sycl/sycl.hpp>\n\nint main() {\n    sycl::device d{sycl::default_selector{}};\n    sycl::queue q(d, sycl::property::queue::in_order());\n\n    std::string device = d.get_info<sycl::info::device::name>();\n    std::cout << \"Default-selected queue runs on device: \" << device << std::endl;\n\n    constexpr size_t N = 1024;\n    int *devicePtr = sycl::malloc_device<int>(N, q);\n\n    q.parallel_for(N, [=](sycl::id<1> idx) {\n      devicePtr[idx] = idx;\n    });\n\n    std::vector<int> dataHost(N);\n    q.copy(devicePtr, dataHost.data(), N).wait();\n\n    bool success = true;\n    for (int i = 0; i < N; i++) {\n      success = success && (dataHost[i] == i);\n    }\n\n    std::cout << (success ? \"SYCL application SUCCESS\" : \"SYCL application FAILED\")\n              << std::endl;\n\n    sycl::free(devicePtr, q);\n    return 0;\n}\n```\n\nOnce we have our compiled `sycl_test` binary we can run it using the `ACPP_VISIBILITY_MASK`\nenvironment variable to select the Vulkan backend as `ACPP_VISIBILITY_MASK=vk`. If there is more\nthan one compatible Vulkan driver on your system then you can ask for a specific device\nby name. For example, here we ask for llvmpipe with `ACPP_VISIBILITY_MASK=vk:llvmpipe`.\n\n``` bash\n$ $ACPP_BIN_DIR/acpp sycl_test.cpp -o sycl_test\n$ ACPP_VISIBILITY_MASK=vk:llvmpipe ./sycl_test\nDefault-selected queue runs on device: llvmpipe (LLVM 20.1.8, 256 bits)\nSYCL application SUCCESS\n```\n\n# Android Benchmarks\n\nWith the Ubuntu setup working, let’s look at what this enables on Android by measuring\nthe benefits of SYCL acceleration on one of the benchmarks from\n[HeCBench](https://github.com/ORNL/HeCBench). HeCBench provides multiple source code variants of\neach benchmark for different backends. We used the mandelbrot benchmark which has, among others, an OpenMP variant\n[mandelbrot-omp](https://github.com/ORNL/HeCBench/tree/master/src/mandelbrot-omp), and a SYCL variant\n[mandelbrot-sycl](https://github.com/ORNL/HeCBench/tree/master/src/mandelbrot-sycl).\n\nThe benchmarks were cross-compiled using release 27 of the Android NDK. The direct OpenMP benchmark was compiled as follows:\n\n``` bash\n$ cd mandelbrot-omp\n$ $NDK/toolchains/llvm/prebuilt/linux-x86_64/bin/clang++ *.cpp -O3 -o mandelbrot-omp-ndk --target=aarch64-linux-android34 -fopenmp=libomp --rtlib=compiler-rt -static-libstdc++\n```\n\nNote that the SYCL benchmarks in HeCBench require `-DUSE_GPU=1` to be set during compilation to enable a GPU\nSYCL queue selector, so for each benchmark we create two SYCL executables linked against an Android cross-compiled build of AdaptiveCpp\n`$ACPP_NDK_BUILD`. See the\n[install-android](https://github.com/AdaptiveCpp/AdaptiveCpp/blob/develop/doc/install-android.md)\ndoc for more details on how to achieve this.\n\n``` bash\n$ cd mandelbrot-sycl\n$ $ACPP_BIN_DIR/acpp *.cpp -O3 -o mandelbrot-gpu-ndk -DUSE_GPU=1 --target=aarch64-linux-android34 --sysroot=$NDK/toolchains/llvm/prebuilt/linux-x86_64/sysroot --rtlib=compiler-rt -static-libstdc++  -resource-dir=$NDK/toolchains/llvm/prebuilt/linux-x86_64/lib/clang/18/ -L $ACPP_NDK_BUILD/lib\n$ $ACPP_BIN_DIR/acpp *.cpp -O3 -o mandelbrot-cpu-ndk --target=aarch64-linux-android34 --sysroot=$NDK/toolchains/llvm/prebuilt/linux-x86_64/sysroot --rtlib=compiler-rt -static-libstdc++ -resource-dir=$NDK/toolchains/llvm/prebuilt/linux-x86_64/lib/clang/18/ -L $ACPP_NDK_BUILD/lib\n```\n\nThe SYCL acceleration results show the merit of GPU offload. On an Android 14 device with a Qualcomm Snapdragon 8 Gen 3 SoC containing an Adreno 750 GPU and Arm v8a CPU, we achieved a 46.5% speedup over raw OpenMP by using the GPU exposed by Vulkan. Using SYCL with the OpenMP backend showed a 7% overhead over the raw OpenMP equivalent, which matches expectations from the same comparison on a desktop machine.\n\nTaking the average parallel time over 1000 iterations, which is output by the benchmark, we observed the following:\n\n| Benchmark | Average parallel time (ms) | \n|---|---|\n| `./mandelbrot-omp-ndk 1000` | 101 | \n| `./mandelbrot-cpu-ndk 1000` | 108 | \n| `./mandelbrot-gpu-ndk 1000` | 54 | \n\n*Data taken from the median of 5 runs based on a build with\nAdaptiveCpp commit [8a75ba2410bc6fcbf2b41e4eaebaf2410c500f21](https://github.com/AdaptiveCpp/AdaptiveCpp/commit/8a75ba2410bc6fcbf2b41e4eaebaf2410c500f21)\n& HeCBench commit [8d66934f100d6d6c972ce476e7f67b3a0fdf0454](https://github.com/ORNL/HeCBench/commit/8d66934f100d6d6c972ce476e7f67b3a0fdf0454).*\n\n# Under the Hood\n\nNow that we’ve seen the backend in action, let’s find out how it works.\n\n## Runtime Implementation\n\nBefore any kernel code runs on a device, the runtime has to do a lot of heavy lifting to map SYCL concepts like queues,\nevents, and memory allocations onto Vulkan’s explicit, low-level API. To reduce implementation verbosity, the Khronos\n[VulkanHpp](https://github.com/KhronosGroup/Vulkan-Hpp) headers are used in the AdaptiveCpp source code,\nwhich provide useful concepts such as RAII-wrapped Vulkan object handles.\n\nEach SYCL device corresponds to a logical Vulkan device that meets the key capability criteria to implement SYCL,\nnamely [Timeline Semaphores](https://docs.vulkan.org/refpages/latest/refpages/source/VK_KHR_timeline_semaphore.html)\nand [Buffer Device Address](https://docs.vulkan.org/refpages/latest/refpages/source/VK_KHR_buffer_device_address.html).\nLet’s elaborate on why these features are necessary building blocks for our SYCL implementation.\n\n### Timeline Semaphores\n\nSynchronization in Vulkan is complicated, so we simplified everything down to a fundamental primitive that is powerful but also easy to reason about, the timeline semaphore. Rather than being in a binary signaled-or-not-signaled state, it uses a monotonically increasing 64-bit integer value to define synchronization order, eliminating the need to be reset before reuse. What’s more, a timeline semaphore can also be signaled from the host or device, which is useful for reasons we’ll discuss later.\n\nIn our implementation each SYCL queue that is created by a user has its own timeline semaphore, with a value that is\ninitialized to zero and incremented with each command submission. When a `sycl::queue::submit()` call is performed, a `vkCommandBuffer`\nsubmission is made to enqueue that work on the device. Crucially, rather than batching multiple commands into a single command buffer,\neach buffer contains just one command, allowing for complete synchronization using timeline semaphores.\nThis allows each command to be uniquely identified by a queue handle and timeline value pair.\n\n### Buffer Device Address\n\nMemory management is where SYCL’s USM abstraction and Vulkan’s explicitness collide. Exposing USM was a problem we needed to solve that had not been addressed by Sylkan or any of the OpenCL-on-Vulkan implementations to date.\n\nAdaptiveCpp builds SYCL buffers on top of USM. That makes USM support fundamental to the Vulkan backend. There are many types of USM in SYCL, but the minimum we need to support is device USM. Device USM gives the user a pointer to an allocation that can be dereferenced on the device, but not dereferenced on the host, unlike host/shared USM types, where a pointer can be dereferenced on both the host and device.\n\nVulkan asynchronous commands operate on `VkBuffer` objects, which are backed by `VkDeviceMemory` objects\nthat must be allocated and then bound to a `VkBuffer`. To implement USM we can’t give the user a device-dereferenceable\npointer to `VkDeviceMemory` because Vulkan only allows the memory to be mapped to a host pointer, not directly addressed\nfrom the device. Instead the solution is to implement device USM as a `VkBuffer` backed by `VkDeviceMemory`, and use the\nBuffer Device Address API\n[vkGetBufferDeviceAddress()](https://docs.vulkan.org/refpages/latest/refpages/source/vkGetBufferDeviceAddress.html) to\nget a 64-bit address that can be returned to the user from `sycl::malloc_device()`.\n\nInitially this seems straightforward, but there is a problem: the asynchronous SYCL `memcpy()` command takes pointer\nparameters that are either USM allocations **or host pointers**, but `vkCmdCopyBuffer` only takes `VkBuffer` objects.\nWe already have the `VkBuffer` associated with the USM allocations the user created, but we need a way to map a host\npointer operand to a `VkBuffer`. This also needs to be done asynchronously. If we create a `VkBuffer` internally and copy\nthe host data into it immediately, then we won’t respect previous asynchronous commands writing to the host data\nwhich have not yet completed. Likewise, we need a way to copy the data back from the internal `VkBuffer` to the\nhost pointer, if it was the destination `memcpy` operand, after the `vkCmdCopyBuffer` command has completed.\n\nThis is where timeline semaphores’ host signaling functionality becomes crucial! Through CPU multithreading, the runtime\ndoes asynchronous host-side work to copy data between a host pointer and a `VkBuffer` while respecting the SYCL command\ndependencies. Here a host worker thread is used to wait on and signal the timeline semaphore values of the queue.\n\nSee the diagram below for the sequence of operations in the case that both the\nsource and destination operations to a `sycl::memcpy` are host pointers.\n\n## Compiler Implementation\n\nLayering the runtime is only half the battle. The real challenge begins with compiling SYCL kernels into\nVulkan-consumable shader SPIR-V. While SYCL kernels can already be compiled to SPIR-V for OpenCL and Level Zero backends,\nthis is not the same dialect of SPIR-V that Vulkan drivers expect.\nFortunately, there is an established tool for generating Vulkan SPIR-V from compute kernels that we\ncan leverage, [clspv](https://github.com/google/clspv), which is used by all of the layered OpenCL-on-Vulkan implementations\nto compile OpenCL-C into Vulkan SPIR-V.\n\nThe clspv tool can accept LLVM IR input as well as OpenCL-C, making it suitable for integration into AdaptiveCpp’s\n[SSCP compilation flow](https://github.com/EwanC/AdaptiveCpp/blob/develop/doc/compilation.md).\nThis runtime JIT compiler allows backends to lower LLVM IR for the device using the most appropriate tooling for\nthat backend. For example, OpenCL/Level-Zero backends call into the\n[LLVM-SPIRV translator](https://github.com/khronosgroup/spirv-llvm-translator)\nto lower LLVM-IR to kernel capability SPIR-V, and the Vulkan backend calls into `clspv` in a similar way.\n\nclspv is typically used with OpenCL-C input rather than C++ single-source IR, so we had to build a\npipeline of LLVM passes for transforming the IR into a form that’s consumable by clspv.\nThe biggest challenge here is generic pointers, which are not part of the OpenCL-C 1.2 language that clspv was\ndesigned to handle. In OpenCL-C 1.2 pointers always have an [address space qualifier](https://registry.khronos.org/OpenCL/specs/unified/html/OpenCL_C.html#address-space-qualifiers)\nto tell the compiler if it’s a `__global`, `__local`, `__constant`, or `__private` address space pointer.\nIn SYCL, with the exception of the rarely used `sycl::multi_ptr`, there are no such qualifiers,\nmeaning the address space of all pointers must be inferred.\nThis inference is possible in most cases but breaks down when pointers themselves are loaded from memory, as there\nis no way to correctly infer the address space. So the SSCP Vulkan backend does a lot of work to try to\nrestore the LLVM address space in IR, but ultimately this is a fundamental mismatch between SYCL generic\npointers and SPIR-V.\n\nOnce we have generated the SPIR-V for our kernels, the SPIR-V code will advertise through\n[capabilities](https://registry.khronos.org/SPIR-V/specs/unified1/SPIRV.html#Capabilities)\nthe specific functionality that a Vulkan driver must support to run it. More advanced SYCL kernels require more capabilities,\nbut the SPIR-V capabilities required to run any kernel are as follows:\n\n- `physicalStorageBufferAddresses` - Use`PhysicalStorageBuffer64` memory\nmodel where physical pointers are allowed. This allows USM pointers to\nbe used in the kernel which are used in AdaptiveCpp to implement SYCL\nbuffers as well as USM.\n- `shaderInt64` - Enables 64-bit integers in SPIR-V. This allows the SYCL runtime\nto pass pointer parameters to a kernel as`i64` members of a SPIR-V\npush constant struct.\n- `variablePointers` - Allows a SPIR-V pointer to not statically know what\nobject it comes from. Enables use of`OpPtrAccessChain` SPIR-V instruction\nwhich is required to create a pointer to the physical pointer member inside\nthe push constant parameters struct.\n\n# Conclusion\n\nWith Vulkan as a foundation, SYCL can now reach every device class from HPC to Android phones.\nWhile this is a milestone worth celebrating, there’s still work left to do. Our roadmap of tasks is tracked as AdaptiveCpp GitHub issues under the\n[Vulkan label](https://github.com/AdaptiveCpp/AdaptiveCpp/issues?q=is%3Aissue%20state%3Aopen%20label%3Avulkan), including exciting features such as\n[backend interop](https://github.com/AdaptiveCpp/AdaptiveCpp/issues/2111)\nthat could bring SYCL and Vulkan graphics APIs together in the same application.\n\nIf this work interests you, please try it out and share your results on the [AdaptiveCpp GitHub repo](https://github.com/AdaptiveCpp/AdaptiveCpp).\nThere are many different Vulkan drivers and application workloads out there, and we’d love to know what works and what needs help.\n\n#### Share on\n\n[X](https://x.com/intent/tweet?text=Bringing+SYCL+to+Android%3A+A+Vulkan+backend+for+portable+GPU+compute%20https%3A%2F%2Fadaptivecpp.github.io%2Fhipsycl%2Fadaptivecpp%2Fsycl%2Fvulkan%2Fandroid%2Fvulkan-android%2F)\n\n[Bluesky](https://bsky.app/intent/compose?text=Bringing+SYCL+to+Android%3A+A+Vulkan+backend+for+portable+GPU+compute%20https%3A%2F%2Fadaptivecpp.github.io%2Fhipsycl%2Fadaptivecpp%2Fsycl%2Fvulkan%2Fandroid%2Fvulkan-android%2F)", "url": "https://wpnews.pro/news/vulkan-back-end-coming-to-sycl", "canonical_source": "https://adaptivecpp.github.io/hipsycl/adaptivecpp/sycl/vulkan/android/vulkan-android/", "published_at": "2026-10-07 13:23:27+00:00", "updated_at": "2026-10-07 13:50:24.087980+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-infrastructure", "developer-tools"], "entities": ["AdaptiveCpp", "SYCL", "Vulkan", "Android", "LunarG", "clspv", "Mesa llvmpipe", "MoltenVK"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/vulkan-back-end-coming-to-sycl", "markdown": "https://wpnews.pro/news/vulkan-back-end-coming-to-sycl.md", "text": "https://wpnews.pro/news/vulkan-back-end-coming-to-sycl.txt", "jsonld": "https://wpnews.pro/news/vulkan-back-end-coming-to-sycl.jsonld"}}