cd /news/ai-tools/qwen-3-8-27b-at-2x-speed-on-a-5090 · home topics ai-tools article
[ARTICLE · art-98092] src=github.com ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Qwen 3.8 27B at 2x speed on a 5090

Adore LLC released Balto Speedrunner, a Windows app that turns Qwen 3.8 27B into a local coding agent for a single RTX 5090, achieving up to 300 tok/s on chat prompts and 150 tok/s during coding runs. The installer requires at least 90 GB free disk space and uses Docker Desktop, SGLang, and RadixArk models. Balto supports Tailscale Serve for private remote access and is proprietary software copyright 2026 Adore LLC.

read3 min views1 publishedAug 15, 2026
Qwen 3.8 27B at 2x speed on a 5090
Image: Michielbdejong (auto-discovered)

On a Mac? Get Balto for Mac

First launch downloads the inference engine and model. Keep at least 90 GB free.

How it works

Benchmarks

Support Balto Balto turns Qwen 3.8 27B into a fast local coding agent for one RTX 5090. Expect roughly 150 tok/s during real coding runs and up to 300 tok/s on clean chat prompts.

Run the installer, approve Windows if it asks to enable WSL, and Balto handles the rest. It resumes setup after a required restart, preserves partial downloads, and opens the coding workspace as soon as the model is ready.

The Windows app uses Tauri and the system WebView2 runtime. Model inference runs in Docker, and Balto keeps its Node.js workspace runtime in its own app data directory.

Workload Output speed Context Notes
Chat Up to 300 tok/s Short prompt Warm model
Code 150 tok/s Live agent session Real tool calls
  • Confirms that the PC has one RTX 5090, enough free disk space, and current NVIDIA support.
  • Installs Docker Desktop in its official per-user WSL 2 mode, or starts an existing installation.
  • Downloads the lmsysorg/sglang:qwen38-27b

image at the exact tested digest. - Creates a persistent Docker volume for model files, so interrupted downloads resume and updates do not erase weights.

  • Starts RadixArk/Qwen3.8-27B-NVFP4

withRadixArk/Qwen3.8-27B-DSpark

using the tested 80K configuration. - Installs the coding workspace into Balto's private app directory.

  • Reminds the owner to unload other local models before starting Balto.
  • Applies future performance configuration updates without deleting the persistent model cache.

First launch requires a large download and at least 90 GB of free disk space.

Balto can use Tailscale Serve after the owner signs in to Tailscale. It keeps the workspace bound to 127.0.0.1

and exposes private HTTPS endpoints only inside the user's tailnet. Balto does not enable Tailscale Funnel or open a public router port.

The onboarding screen shows the exact private URL and lets the owner turn remote access off without changing unrelated Tailscale routes.

The inference arguments live in runtime/balto.ps1. The important settings are:

model                  Qwen 3.8 27B NVFP4
context length         80000
attention backend      flashinfer
max running requests   1
speculation            DSpark, FP8 draft
sampling               temperature 0.6, top_p 0.95, top_k 20

Balto uses safe sampling defaults for coding. It does not force greedy temperature zero sampling, which can trap this model in repetitive reasoning loops.

Requirements:

  • Windows 11
  • Node.js 22 or newer
  • Rust stable with the MSVC target
  • WebView2
npm install
npm run check
npm run dev

Build the NSIS installer:

npm run build

Balto supports two different signatures:

  • Tauri updater signatures protect update artifacts and are required by the in-app updater.
  • Windows Authenticode identifies Adore LLC as the publisher and prevents the unsigned-app SmartScreen warning.

Every install shows its version in Settings. A green update arrow appears when GitHub publishes a newer signed release; one click verifies, installs, and relaunches it.

The release workflow uses Azure Artifact Signing when publisher credentials are configured. The in-app updater always verifies Tauri update signatures. Local development builds remain unsigned.

Balto Speedrunner is proprietary software, copyright 2026 Adore LLC. All rights reserved.

The coding agent interface integrates MIT-licensed software from DeepSeek AI. Inference is powered by SGLang. Qwen model weights remain under their own license. See THIRD_PARTY_NOTICES.md for the full notices.

Balto Speedrunner is not affiliated with or endorsed by DeepSeek, Qwen, Alibaba, SGLang, LMSYS, NVIDIA, Docker, Tailscale, Microsoft, or OpenAI.

If Balto saves you setup time, buy me a coffee.

── more in #ai-tools 4 stories · sorted by recency
── more on @adore llc 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/qwen-3-8-27b-at-2x-s…] indexed:0 read:3min 2026-08-15 ·