cd /news/large-language-models/juud-engine-an-experimental-strata-f… · home › topics › large-language-models › article
[ARTICLE · art-145821] src=discuss.huggingface.co ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Juud_engine: an experimental Strata fork for Qwen3.8-Flash-Next on an RTX 4090

A developer released Juud_engine, a free MIT-licensed experimental fork of Strata v0.1.39 for local Qwen3.8-Flash-Next inference, reporting median paired decode throughput 3.1–9.1% higher and paired full-response latency 0.5–5.6% better on a single RTX 4090 with an i7-13700. The author cautions that normal A/B output text matched in only 5/20 pairs, while a separate control with different cache and PCIe settings matched 12/12 pairs and showed 9.6–15.6% higher decode throughput, and that the two opt-in CPU-side decode changes were enabled together with individual contributions unmeasured. The repository provides the source diff, build instructions, model and binary hashes, request/output hashes, pair-level metrics and a verification script, and the author requests independent replications including negative results.

read1 min views1 publishedOct 6, 2026

I’m sharing Juud_engine, a free MIT-licensed experimental fork of Strata v0.1.39 for local Qwen3.8-Flash-Next inference. It preserves Strata’s local OpenAI-compatible API and adds two opt-in CPU-side decode changes:

I compared a source build with pinned Strata on one RTX 4090 (i7-13700), using the same Qwen3.8-Flash-Next IQ3_S pack. Each request generated 256 tokens. Four prompts covered code contexts of 4K, 32K and 128K tokens, plus a Korean 2K prompt. In five paired runs per prompt, median paired decode throughput was 3.1–9.1% higher; paired full-response latency improved 0.5–5.6%.

Important limits: Normal A/B output text matched in only 5/20 pairs. Different continuations and MTP acceptance can change timing, so those runs do not isolate the effect of the code. In a separate three-pair-per-prompt control, outputs matched in 12/12 pairs and decode throughput was 9.6–15.6% higher. That control used different cache and PCIe settings from the normal run, so I do not combine the results. Both changes were enabled together; their individual contributions are unmeasured. H100 and multi-user performance are also unmeasured.

The repository has the source diff, build instructions, model/binary hashes, request/output hashes, pair-level metrics and a verification script: GitHub - bnkaiteam/Juud_engine: Experimental Strata fork for Qwen3.8-Flash-Next with paired RTX 4090 benchmarks and opt-in CPU decode optimizations · GitHub . A source build and the RUN-JUUD wrapper are needed to enable the tested path; the inherited default setup can fetch a Strata prebuilt instead. Model weights are separate and have their own license.

I’d value independent replications (including negative results), suggestions for a per-change ablation, and feedback on the benchmark method.

Disclosure: I am sharing my project, and this post was drafted with AI assistance. The figures were checked against the published paired data.

── more in #large-language-models 4 stories · sorted by recency
── more on @juud_engine 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/juud-engine-an-exper…] indexed:0 read:1min 2026-10-06 · —