cd /news/large-language-models/qwen-3-8-follows-gpt-5-5-pro-reasoni… · home topics large-language-models article
[ARTICLE · art-124898] src=gist.github.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Qwen 3.8 follows GPT-5.5 Pro reasoning prefills

A developer's v1.1 experiment reran reasoning-prefill tests using GPT-5.5 Pro as the teacher across 45 problems, finding that Qwen3.8 A95B's unigram source recall jumped from 33.92% to 54.50% (+20.58 percentage points) when prefilled with GPT-5.5 Pro's reasoning, with the largest gains on STEM problems (+27.55 pp). The results suggest Qwen may have learned from GPT-5.5 Pro or a closely related model, while Kimi K3 showed high overlap but smaller prefill gains (+4.31 pp).

read2 min views1 publishedSep 9, 2026

A follow-up to Reasoning prefills on a few open models and Stolen Thoughts This v1.1 reruns the reasoning-prefill experiment with GPT-5.5 Pro as the teacher.

For each problem, I generated two responses from each target model:

  1. an ordinary, unprefilled response; and
  2. a response starting with the first 1% of GPT-5.5 Pro's reasoning, inserted into the target model's reasoning channel.

The visible answer remained freely generated. I then measured how much of the teacher's visible answer appeared in the first 100 tokens of the target model's answer. The table below reports unigram source recall so the numbers are comparable to my previous post. Deltas are absolute percentage-point changes.

All problems #

The evaluation contains 45 problems: 15 STEM, 15 non-STEM, and 15 synthetic puzzles.

Model n Unprefilled GPT-5.5 Pro reasoning prefill Delta
DeepSeek V4 Flash 45 40.53% 40.89% +0.35 pp
Inkling 45 37.82% 38.67% +0.85 pp
Kimi K3 45 50.11% 54.42% +4.31 pp
Qwen3.8 A95B 45 33.92% 54.50% +20.58 pp

Qwen by category #

Category n Unprefilled GPT-5.5 Pro reasoning prefill Delta
STEM 15 36.21% 63.76% +27.55 pp
Non-STEM 15 38.26% 52.73% +14.46 pp
Puzzle 15 27.28% 47.00% +19.72 pp
All 45 33.92% 54.50% +20.58 pp

Discussion #

Qwen barely moved toward Opus 4.8 in the earlier experiment, but moved by +20.58 points toward GPT-5.5 Pro here, including a large effect on the private synthetic puzzles. The data suggest that Qwen may have learned from GPT-5.5 Pro, or from a closely related GPT model, rather than from Opus.

Kimi K3's overlap with GPT-5.5 Pro is also high both without and with the prefill (50.11% and 54.42%), although the prefill adds only +4.31 points.

── more in #large-language-models 4 stories · sorted by recency
── more on @gpt-5.5 pro 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/qwen-3-8-follows-gpt…] indexed:0 read:2min 2026-09-09 ·