An exploratory study of hallucination and abstention in 2026 open-weight LLMs.
- Qwen 3.8-27B (Alibaba)
- GPT-OSS-20B (OpenAI)
Two failure modes identified:
- Qwen: rigid confabulation (100% consistency)
- GPT-OSS: deliberation collapse (50% empty responses)
Template-based confabulation in Qwen:
- Same fabricated year (1991) for two different universities
Authority pressure affects GPT-OSS:
- Up to 232% increase in reasoning tokens
- Academic framing triggers refusal
Empathy pressure is ineffective:
- Neither model increases fabrication under emotional pressure
- README.md - this file
- RESULTS.md - findings and tables
- METHODOLOGY.md - protocol description
- LIMITATIONS.md - what this study cannot claim
- *.py - Python scripts
- *.csv - raw data files
Total: 638 trials across 10 CSV files.
export GROQ_API_KEY=gsk_... python scientific_test.py
Exploratory pilot study. Not peer-reviewed. Small sample size.
See LIMITATIONS.md for full caveats.
CC-BY-4.0