An exploratory study of hallucination in Qwen 3.8-27B and GPT-OSS-20B An exploratory, non-peer-reviewed study of 638 trials across 10 CSV files found two distinct hallucination failure modes in 2026 open-weight LLMs: Alibaba's Qwen 3.8-27B produced rigid confabulation with 100% consistency, while OpenAI's GPT-OSS-20B showed deliberation collapse with 50% empty responses. Qwen fabricated the same year (1991) for two different universities, and authority pressure on GPT-OSS increased reasoning tokens by up to 232% and triggered refusal under academic framing, while empathy pressure did not increase fabrication in either model. An exploratory study of hallucination and abstention in 2026 open-weight LLMs. - Qwen 3.8-27B Alibaba - GPT-OSS-20B OpenAI 1. Two failure modes identified: - Qwen: rigid confabulation 100% consistency - GPT-OSS: deliberation collapse 50% empty responses 2. Template-based confabulation in Qwen: - Same fabricated year 1991 for two different universities 3. Authority pressure affects GPT-OSS: - Up to 232% increase in reasoning tokens - Academic framing triggers refusal 4. Empathy pressure is ineffective: - Neither model increases fabrication under emotional pressure - README.md - this file - RESULTS.md - findings and tables - METHODOLOGY.md - protocol description - LIMITATIONS.md - what this study cannot claim - .py - Python scripts - .csv - raw data files Total: 638 trials across 10 CSV files. export GROQ API KEY=gsk ... python scientific test.py Exploratory pilot study. Not peer-reviewed. Small sample size. See LIMITATIONS.md for full caveats. CC-BY-4.0