cd /news/large-language-models/an-exploratory-study-of-hallucinatio… · home › topics › large-language-models › article
[ARTICLE · art-149120] src=github.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

An exploratory study of hallucination in Qwen 3.8-27B and GPT-OSS-20B

An exploratory, non-peer-reviewed study of 638 trials across 10 CSV files found two distinct hallucination failure modes in 2026 open-weight LLMs: Alibaba's Qwen 3.8-27B produced rigid confabulation with 100% consistency, while OpenAI's GPT-OSS-20B showed deliberation collapse with 50% empty responses. Qwen fabricated the same year (1991) for two different universities, and authority pressure on GPT-OSS increased reasoning tokens by up to 232% and triggered refusal under academic framing, while empathy pressure did not increase fabrication in either model.

read1 min views1 publishedOct 11, 2026
An exploratory study of hallucination in Qwen 3.8-27B and GPT-OSS-20B
Image: Michielbdejong (auto-discovered)

An exploratory study of hallucination and abstention in 2026 open-weight LLMs.

- Qwen 3.8-27B (Alibaba)
- GPT-OSS-20B (OpenAI)

Two failure modes identified:

  - Qwen: rigid confabulation (100% consistency)
  - GPT-OSS: deliberation collapse (50% empty responses)

Template-based confabulation in Qwen:

  • Same fabricated year (1991) for two different universities

Authority pressure affects GPT-OSS:

  • Up to 232% increase in reasoning tokens
  • Academic framing triggers refusal

Empathy pressure is ineffective:

  • Neither model increases fabrication under emotional pressure
- README.md - this file
- RESULTS.md - findings and tables
- METHODOLOGY.md - protocol description
  • LIMITATIONS.md - what this study cannot claim
- *.py - Python scripts
- *.csv - raw data files

Total: 638 trials across 10 CSV files.

export GROQ_API_KEY=gsk_... python scientific_test.py

Exploratory pilot study. Not peer-reviewed. Small sample size.

See LIMITATIONS.md for full caveats.

CC-BY-4.0

── more in #large-language-models 4 stories · sorted by recency
── more on @qwen 3.8-27b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/an-exploratory-study…] indexed:0 read:1min 2026-10-11 · —