Srovnání Qwen 3.8 27B, Nemotron Lightning 30B, Ornith 35B a Muse-Glimmer 30B v reálných testech. Qwen 3.8 ovládá coding, Ornith překvapuje rychlostí, Nemotron a Muse-Glimmer ztrácejí.
Uživatel provedl detailní srovnání čtyř populárních open-source modelů na coding úloze. Test běžel 25 minut a na Qwen 3.8 vyprodukoval 50 000 tokenů — jde o těžký thinking model, ale výsledek je fenomenální. Ornith se ukázal jako velmi schopný model, zejména vzhledem k dané rychlosti.
Výsledky detailně: llm-bench.io porovnání
Qwen 3.8 27B — jasný vítěz v codingu a architektuře. V testech computer use MCP byl bezchybný, na úrovni GPT-5.6 Luna. Jak uživatel popsal: Qwen3.8 was flawless, on par with gpt 5.6 luna. Ale je pomalý — na některých strojích sotva 15 TPS. Pro náročné úlohy, kde záleží na přesnosti, je to jasná volba. Podle dalšího uživatele: Glad to see i was justified in focusing mainly on 3.8 since its release.
Ornith 1.5 35B-A3B — překvapení soutěže. Ornith is seriously amazing especially given the fact it only has 3b active. Na Strix Halo dosahuje 50 TPS při 50k+ tokenech a zvládne plný kontext. Podle jednoho uživatele: After more testing Ornith is absolute sorcery, this is GPT oss20b levels of optimization. 50TPS on strix halo at 50k+ tokens and can load full context, compared to Qwen 3.8 that struggles to serve 15TPs, and I honestly can't notice the difference in front end work maybe 2% at best.
Nemotron 3.5 Lightning 30B-A3B — určen pro ne-technické agentic úlohy. Nemotron is for non technical agentic tasks. Solidní rychlost, ale v codingu zaostává.
Muse-Glimmer 30B — zaměřen na technické psaní. Muse Glimmer take care of 3 stages of technical writing. S dflash běží rychleji (150-200 TPS na 3090), ale v testech computer use dělal chyby: Muse glimmer was almost good, but would miss some step, misclick some button or forget to activate a window and that would doom it.
Zdroj: Reddit