cd/entity/ninfer· home› entities› ninfer
grep -l @ninfer /news/*.json | wc -l → 4

ninfer

mentions 4 type Organization feed RSS

// recent coverage 4 mentions

19:18
2026-10-03
carteakey.dev
ai-infrastructure

The Rise of Overfit Inference Engines

A homelab test on an i5-12600K with 64 GB DDR5 and an RTX 4070 12GB found the narrow inference runtime Strata generated 512 tokens at 53.2 tok/s with 60,000 tokens of context on the 125B-parameter Qwe…

14:12
2026-08-26
forum.level1techs.com
large-language-models

Running qwen 3.6 / 3.8 on 3090+3080 over RPC?

A user running Qwen 3.8 27B on a 3090 reports that switching to ninfer, a Qwen-focused inference engine, doubled their token generation speed from 35-40 to 55-60 tokens per second. The user is conside…

13:32
2026-08-26
forum.level1techs.com
large-language-models

Make a custom qwen 3.8 27B abliterated for ninfer?

A user running ninfer's Qwen 3.8 27B model on an RTX 3090 is seeking guidance on creating an abliterated version of the model, noting that ninfer uses a custom file format that prevents simple configu…

07:58
2026-08-24
forum.level1techs.com
large-language-models

Sanity check my Qwen3.8 5090 Results; new to Local AI

A user running Qwen3.8-27B on an RTX 5090 with the ninfer runtime reported a 1.72× speedup over LM Studio/llama.cpp, achieving 124.08 tok/s decode and 8,143.82 tok/s prompt eval with mixed NVFP4/FP8 q…

// co-occurs with top 8 entities
// topics top 5 topics