arXiv:2610.07026v1 Announce Type: new Abstract: The Offline AI Modules workstream enables practical, low-power, and community-accessible deployment of voice-first AI systems that operate fully offline. Designed for African language communities where speech is the dominant mode of interaction and internet connectivity is unreliable or absent, the workstream delivers three reinforcing components: a modular voice-first offline architecture, a low-cost hardware reference bill of materials, and a reproducible quantization and a reproducible quantization and benchmarking pipeline for instruction-tuned language models in the 2-5B parameter class. This paper presents the first end-to-end benchmark evaluation of the stack across two hardware tiers: an NVIDIA Jetson Orin NX (TierB) and a Raspberry Pi5 (TierA). Three instruction-tuned models are evaluated across four quantization formats, assessed for deployment metrics (decode throughput, chat latency, memory, power) and multilingual quality (topic classification accuracy on MasakhaNEWS across English, Hausa, Igbo, Nigerian Pidgin, and Yoruba; per-language perplexity drift). Speech recognition is evaluated using Ethio-ASR on Amharic and Oromo across both tiers. The principal finding is that Q4_K_M quantization represents the best size-to-quality trade-off for deployment on both tiers: gemma-4-E2B-it achieves 28.8t/s decode throughput and 89.2% topic classification accuracy at Q4_K_M on TierB, while all three models run within the 16GB memory budget on TierA.
Offline AI Modules: Voice-First Offline Architecture, Hardware Reference Stack, Quantization and Benchmarking
The Offline AI Modules workstream's first end-to-end benchmark evaluation found that Q4_K_M quantization offers the best size-to-quality trade-off for offline voice-first AI deployment, with gemma-4-E2B-it reaching 28.8 tokens/s decode throughput and 89.2% topic classification accuracy at Q4_K_M on an NVIDIA Jetson Orin NX (TierB), according to arXiv paper 2610.07026v1. The stack, aimed at African language communities with unreliable connectivity, pairs a modular offline architecture and low-cost hardware bill of materials with a reproducible quantization and benchmarking pipeline for 2-5B parameter instruction-tuned models. Three models were tested across four quantization formats on the Jetson Orin NX and a Raspberry Pi 5 (TierA), with all three running within the 16GB memory budget on TierA, and speech recognition evaluated via Ethio-ASR on Amharic and Oromo.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.