cd /news/artificial-intelligence/open-sourced-a-local-wan-2-1-mobile-… Β· home β€Ί topics β€Ί artificial-intelligence β€Ί article
[ARTICLE Β· art-83373] src=github.com β†— pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Open-sourced a local WAN 2.1 mobile pipeline for on-device video generation

Quartz, a from-scratch Rust inference engine developed by Saient, has been open-sourced, enabling on-device video generation with Wan2.1 T2V (1.3B parameters) on a Samsung Galaxy S24's Vulkan GPU, alongside LLM chat and SDXL image generation, with no cloud round-trip. The engine, previously named tinyq4, supports CPU, CUDA, and Vulkan backends and reads GGUF files directly.

read5 min views1 publishedAug 2, 2026
Open-sourced a local WAN 2.1 mobile pipeline for on-device video generation
Image: source

Quartz is the on-device inference engine behind Saient β€” a from-scratch Rust runtime (no llama.cpp, no ggml, no Candle, no PyTorch/libtorch) that runs LLM chat and image generation entirely on a phone or desktop, no cloud round-trip.

It's the engine that made this possible: Saient mobile generates full text-to-video clips β€” Wan2.1 T2V, 1.3B parameters β€” running end-to-end, locally, on a stock Samsung Galaxy S24's Vulkan GPU. Same phone also runs LLM chat and SDXL image generation locally through this engine. No server, no API key, nothing leaves the device.

(The Wan video model itself runs through a separate, MIT-licensed companion binary β€” see Relationship to the Wan video engine below for how the two fit together. This repo, Quartz proper, is the LLM chat + SDXL image half of the stack.)

This repo covers Quartz: a minimal, dependency-light Rust inference runtime for GGUF-format transformer language models, with a bundled Stable Diffusion (SD1.5/SDXL) image pipeline. CI-grade targets are a 4-core desktop CPU, a CUDA GPU, and an Android phone's Vulkan GPU, via a single ~15k-line, mostly-unsafe

-free codebase.

This project was previously developed under the name tinyq4. All crate, binary, and environment-variable names now use quartz

/ QUARTZ_*

; the git history and any external references to tinyq4

predate the rename.

Quartz ships as a single binary (quartz

) with two capabilities:

LLM chat inferenceβ€” GGUF model , tokenization (native GGUF vocab or aminijinja

-templated Jinja2 chat template, matching what the source model ships), forward pass, and an OpenAI-compatible HTTP server for streaming completions.Image generation (SD1.5 / SDXL)β€” a from-scratch CLIP text encoder, UNet, VAE, and scheduler implementation, driven by the same binary via--generate-sd15

/--generate-sdxl

.

Video generation (Wan2.1 T2V) is a separate binary and a separate codebase β€” see Relationship to the Wan video engine below. Don't conflate the two; they share the Quartz brand but not a source tree.

CPUβ€” portable Rust, with a hand-rolled AVX2/FMA dot-product path on x86_64 (auto-dispatched at runtime; falls back to scalar otherwise) and a hand-written NEON dot-product path for Q4_K on aarch64 (Android).CUDAβ€” optional (--features cuda

), custom GEMV kernels per quantization format (src/cuda_kernels.cu

,src/cuda.rs

), staged/paged weight , and a VRAM-aware KV cache that sizes itself to free VRAM rather than a hard token cap.Vulkanβ€” optional (--features vulkan

), used for both LLM and SDXL compute; this is the backend the Android build ships with, since phones have no CUDA.

Reads GGUF files directly (src/gguf.rs

) and dequantizes on the fly: Q4_K, Q6_K, Q8_0, Q5_0/Q5_1, Q4_0/Q4_1, and the IQ2/IQ3 k-quant formats (grid/sign tables in src/iq2_tables.rs

/ src/iq3_tables.rs

, generated from upstream ggml-quants.c

/ ggml-common.h

constant tables β€” the tables are data, not linked code).

quartz --server

exposes an OpenAI-compatible surface:

GET  /health
GET  /v1/health
GET  /v1/models
POST /v1/chat/completions   (SSE streaming)

This is what both the desktop and mobile Saient clients speak β€” same wire format either side.

cargo build --release                       # CPU only
cargo build --release --features cuda       # + CUDA GEMV kernels
cargo build --release --features vulkan     # + Vulkan (SD image gen, or Android)

Android (arm64-v8a), used by the Saient mobile app:

scripts/build-android-arm64.sh

Requires ANDROID_NDK_HOME

(or a discoverable NDK under $ANDROID_SDK_ROOT/ndk

). QUARTZ_ANDROID_API

(default 24) and QUARTZ_ANDROID_FEATURES

(default vulkan

) are overridable.

Variable Effect
QUARTZ_BIND
Server bind address. Defaults to 127.0.0.1 (loopback only); set to 0.0.0.0 to opt into LAN/network access.
QUARTZ_CTX
Upper-bounds the KV cache context length the host app requests.
QUARTZ_CUDA_ARCH
SASS target arch(es) for the CUDA build (build.rs ); CI/shipping builds set all-major .
QUARTZ_DENSE_FFN_CPU
Forces dense-FFN compute onto CPU (CUDA path), for debugging/comparison.
QUARTZ_DEBUG_PROMPT
Dumps the fully-rendered prompt before tokenization.
QUARTZ_NO_JINJA
Skips the Jinja2 chat-template path even when the model has no native template.
QUARTZ_NO_WARMUP
Set by the Android host app before every engine start (see note below).

Note on QUARTZ_NO_WARMUP: this flag is set by the mobile app (

plugins/quartz/QuartzEngine.kt

in the Saient mobile repo) to try to avoid a full-resident mmap warmup that can get the app OOM-killed on memory-constrained phones. As of this writing the Quartz binary does not actually read this variable anywhere β€” the mmap warmup always runs unconditionally (see the "quartz: mmap warmup complete"

log line in src/main.rs

). This was discovered while renaming TINYQ4_NO_WARMUP

β†’ QUARTZ_NO_WARMUP

; it is a pre-existing no-op on the engine side, not something this rename introduced or fixed. Worth following up on if low-memory devices are still seeing OOM kills at startup.The Saient mobile app also ships libquartz-wan.so

, a second, unrelated binary used only for Wan2.1 text-to-video generation. It is not built from this repository β€” it's a pinned fork of leejet/stable-diffusion.cpp (MIT-licensed, C++/ggml), patched with a small progress-reporting patch and cross-compiled for Android/Vulkan. Its build lock lives in the mobile repo at engine/wan/BACKEND_LOCK.json

and engine/wan/saient-progress.patch

. Both binaries carry the "Quartz" name in the mobile app's UI/branding (QuartzService

, QuartzEngine.kt

supervise both processes), but they are two separate codebases with two separate license situations β€” this repo (Quartz proper) is what's being open-sourced here; the Wan fork's own license terms (MIT) and upstream attribution travel with it separately. See the mobile repo's docs/VIDEO_PIPELINE.md

for how the two engines are orchestrated together on-device.

This repo has uncommitted local changes beyond the rename performed for open-sourcing (a chat-template fix, KV-cache sizing fix, NEON Q4_K dot product, and the SD1.5/SDXL modules sd15.rs

/sdxl*.rs

/tensor.rs

/ safetensors.rs

/vulkan.rs

are untracked as of this writing). Review and commit those separately before or as part of publishing β€” this README describes the working tree as it stands, not necessarily the last pushed commit.

MIT β€” see LICENSE. Matches the license of the bundled Wan video engine (see below), so the whole on-device stack is MIT.

── more in #artificial-intelligence 4 stories Β· sorted by recency
── more on @quartz 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/open-sourced-a-local…] indexed:0 read:5min 2026-08-02 Β· β€”