Juud_engine: an experimental Strata fork for Qwen3.8-Flash-Next on an RTX 4090 A developer released Juud_engine, a free MIT-licensed experimental fork of Strata v0.1.39 for local Qwen3.8-Flash-Next inference, reporting median paired decode throughput 3.1–9.1% higher and paired full-response latency 0.5–5.6% better on a single RTX 4090 with an i7-13700. The author cautions that normal A/B output text matched in only 5/20 pairs, while a separate control with different cache and PCIe settings matched 12/12 pairs and showed 9.6–15.6% higher decode throughput, and that the two opt-in CPU-side decode changes were enabled together with individual contributions unmeasured. The repository provides the source diff, build instructions, model and binary hashes, request/output hashes, pair-level metrics and a verification script, and the author requests independent replications including negative results. I’m sharing Juud engine, a free MIT-licensed experimental fork of Strata v0.1.39 for local Qwen3.8-Flash-Next inference. It preserves Strata’s local OpenAI-compatible API and adds two opt-in CPU-side decode changes: I compared a source build with pinned Strata on one RTX 4090 i7-13700 , using the same Qwen3.8-Flash-Next IQ3 S pack. Each request generated 256 tokens. Four prompts covered code contexts of 4K, 32K and 128K tokens, plus a Korean 2K prompt. In five paired runs per prompt, median paired decode throughput was 3.1–9.1% higher; paired full-response latency improved 0.5–5.6%. Important limits: Normal A/B output text matched in only 5/20 pairs. Different continuations and MTP acceptance can change timing, so those runs do not isolate the effect of the code. In a separate three-pair-per-prompt control, outputs matched in 12/12 pairs and decode throughput was 9.6–15.6% higher. That control used different cache and PCIe settings from the normal run, so I do not combine the results. Both changes were enabled together; their individual contributions are unmeasured. H100 and multi-user performance are also unmeasured. The repository has the source diff, build instructions, model/binary hashes, request/output hashes, pair-level metrics and a verification script: GitHub - bnkaiteam/Juud engine: Experimental Strata fork for Qwen3.8-Flash-Next with paired RTX 4090 benchmarks and opt-in CPU decode optimizations · GitHub https://github.com/bnkaiteam/Juud engine . A source build and the RUN-JUUD wrapper are needed to enable the tested path; the inherited default setup can fetch a Strata prebuilt instead. Model weights are separate and have their own license. I’d value independent replications including negative results , suggestions for a per-change ablation, and feedback on the benchmark method. Disclosure: I am sharing my project, and this post was drafted with AI assistance. The figures were checked against the published paired data.