{"slug": "lessthink-qwen3-4b-the-same-model-with-far-less-thinking-p", "title": "LessThink-Qwen3-4B: the same model, with far less thinking [P]", "summary": "A developer posting as stey1r released LessThink-Qwen3-4B, a post-trained version of Qwen3-4B that uses 44% fewer tokens on reasoning while retaining the base model's knowledge and answer style, with the full pipeline run on a single GPU. The model is available at https://5ivatej.com/lessthink/.", "body_md": "I post-trained Qwen3-4B to spend 44% fewer tokens on reasoning, keeping its knowledge and answer style. The whole pipeline ran on one GPU. folks, you can check it out on : https://5ivatej.com/lessthink/ submitted by /u/stey1r", "url": "https://wpnews.pro/news/lessthink-qwen3-4b-the-same-model-with-far-less-thinking-p", "canonical_source": "https://aiflash.com/news/130262/", "published_at": "2026-10-03 09:00:45+00:00", "updated_at": "2026-10-03 09:06:07.097363+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "machine-learning", "artificial-intelligence"], "entities": ["LessThink-Qwen3-4B", "Qwen3-4B", "stey1r"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/lessthink-qwen3-4b-the-same-model-with-far-less-thinking-p", "markdown": "https://wpnews.pro/news/lessthink-qwen3-4b-the-same-model-with-far-less-thinking-p.md", "text": "https://wpnews.pro/news/lessthink-qwen3-4b-the-same-model-with-far-less-thinking-p.txt", "jsonld": "https://wpnews.pro/news/lessthink-qwen3-4b-the-same-model-with-far-less-thinking-p.jsonld"}}