cd/entity/Qwen· home entities Qwen
grep -l @qwen /news/*.json | wc -l → 687

Qwen

mentions 687 type Organization page 34/35 feed RSS

// recent coverage 687 mentions

00:31
2026-05-24
dev.to
large-language-models

Qwen 3.6 27B and 35B MTP vs Standard on 16GB GPU

The article summarizes tests of Multi-Token Prediction (MTP) on Qwen 3.6 27B and 35B models using a 16GB RTX 4080 GPU. For the 27B model, MTP at a draft depth of 2 provided a 67% speed increase (75 t/…

03:39
2026-05-23
dev.to
large-language-models

BeeLlama v0.2.0: 164 tok/s on a 27B model, one RTX 3090

BeeLlama v0.2.0 demonstrates that speculative decoding can achieve a 4.4x to 4.93x throughput multiplier on a single RTX 3090, running 27B and 31B parameter models at 37-36 tokens per second baseline …

13:06
2026-05-22
dev.to
artificial-intelligence

Run Powerful AI Coding Locally on a Normal Laptop

This article provides a step-by-step guide for developers to set up a private, offline AI coding assistant on a standard laptop (8GB or 16GB RAM) without a dedicated GPU. The setup uses Visual Studio …

00:00
2026-05-21
flox.dev
ai-tools

Run Frontier Models Anywhere with DeepSeek TUI

DeepSeek TUI, a terminal user interface for DeepSeek's models, has gained nearly 33,000 GitHub stars in two weeks. The tool can run against other providers including local models via Ollama or oMLX, a…

15:19
2026-05-20
dev.to
large-language-models

What did gemma see? - Thinking in comments...

The Gemma 4 26B model was the first local AI to achieve a perfect score on the HumanEval benchmark, including solving the notoriously difficult problem 145. This problem requires sorting integers by t…

19:54
2026-05-18
dev.to
large-language-models

First-call checklist before trying a new LLM gateway

A checklist the author uses when testing a new OpenAI-compatible LLM gateway, which helps catch integration failures before moving real workloads. For Chinese models like Qwen, DeepSeek, GLM, and Kimi…

00:00
2026-05-14
maltebuettner.eu
large-language-models

documentai bbox benchmark

Malte Buettner benchmarked bounding box accuracy for Document AI models using pages from the FlashAttention-3 paper, testing Qwen, Kimi, and Mistral via OpenRouter. The evaluation scored models on cov…

00:00
2026-05-10
jola.dev
large-language-models

Running local models on an M4 with 24GB memory

The article describes the author's successful setup for running local AI models on an M4 Mac with 24GB of memory, specifically highlighting Qwen 3.5-9B (Q4 quantized) as the best performing model at ~…

← prev page 34 / 35 next →
// co-occurs with top 8 entities
// topics top 6 topics