{"slug": "an-exploratory-study-of-hallucination-in-qwen-3-8-27b-and-gpt-oss-20b", "title": "An exploratory study of hallucination in Qwen 3.8-27B and GPT-OSS-20B", "summary": "An exploratory, non-peer-reviewed study of 638 trials across 10 CSV files found two distinct hallucination failure modes in 2026 open-weight LLMs: Alibaba's Qwen 3.8-27B produced rigid confabulation with 100% consistency, while OpenAI's GPT-OSS-20B showed deliberation collapse with 50% empty responses. Qwen fabricated the same year (1991) for two different universities, and authority pressure on GPT-OSS increased reasoning tokens by up to 232% and triggered refusal under academic framing, while empathy pressure did not increase fabrication in either model.", "body_md": "An exploratory study of hallucination and abstention in 2026 open-weight LLMs.\n\n- Qwen 3.8-27B (Alibaba)\n- GPT-OSS-20B (OpenAI)\n\n1. \nTwo failure modes identified: \n  - Qwen: rigid confabulation (100% consistency)\n  - GPT-OSS: deliberation collapse (50% empty responses)\n2. \nTemplate-based confabulation in Qwen: \n  - Same fabricated year (1991) for two different universities\n3. \nAuthority pressure affects GPT-OSS: \n  - Up to 232% increase in reasoning tokens\n  - Academic framing triggers refusal\n4. \nEmpathy pressure is ineffective: \n  - Neither model increases fabrication under emotional pressure\n\n- README.md - this file\n- RESULTS.md - findings and tables\n- METHODOLOGY.md - protocol description\n- LIMITATIONS.md - what this study cannot claim\n- *.py - Python scripts\n- *.csv - raw data files\n\nTotal: 638 trials across 10 CSV files.\n\nexport GROQ_API_KEY=gsk_... python scientific_test.py\n\nExploratory pilot study. Not peer-reviewed. Small sample size.\n\nSee LIMITATIONS.md for full caveats.\n\nCC-BY-4.0", "url": "https://wpnews.pro/news/an-exploratory-study-of-hallucination-in-qwen-3-8-27b-and-gpt-oss-20b", "canonical_source": "https://github.com/alitenes2020-sys/Repository-name-llama-hallucination", "published_at": "2026-10-11 10:42:55+00:00", "updated_at": "2026-10-11 10:52:48.030075+00:00", "lang": "en", "topics": ["large-language-models", "ai-safety", "ai-research", "artificial-intelligence"], "entities": ["Qwen 3.8-27B", "GPT-OSS-20B", "Alibaba", "OpenAI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/an-exploratory-study-of-hallucination-in-qwen-3-8-27b-and-gpt-oss-20b", "markdown": "https://wpnews.pro/news/an-exploratory-study-of-hallucination-in-qwen-3-8-27b-and-gpt-oss-20b.md", "text": "https://wpnews.pro/news/an-exploratory-study-of-hallucination-in-qwen-3-8-27b-and-gpt-oss-20b.txt", "jsonld": "https://wpnews.pro/news/an-exploratory-study-of-hallucination-in-qwen-3-8-27b-and-gpt-oss-20b.jsonld"}}