cd /news/artificial-intelligence/birds-don-t-fly-like-planes-neither-… · home topics artificial-intelligence article
[ARTICLE · art-101687] src=tomtunguz.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Birds Don't Fly Like Planes. Neither Does AI.

Qwen3.8-27B, a 27-billion-parameter dense model from Alibaba, ranks #1 of 135 models on Artificial Analysis's Intelligence Index with a score of 52, one point above GLM-5.2, a 753-billion-parameter open-source model from Z.ai, demonstrating that a laptop-sized model can match a frontier cloud model roughly 28 times its size. In benchmarking against DeepSeek V4 and Qwen 3.6 35B on 25 venture-capital tasks, the Qwen3.8-27B matched quality but required more reasoning tokens, showing that smaller models achieve similar results through different inference paths.

read3 min views6 publishedAug 18, 2026
Birds Don't Fly Like Planes. Neither Does AI.
Image: Tomtunguz (auto-discovered)

Your laptop can now run a model as capable as nearly anything in the cloud. I swapped Qwen3.8-27B into my agent & it works brilliantly. This bird flies differently than a plane.

This little Qwen model ranks #1 of 135 models, scoring 52 on Artificial Analysis’s Intelligence Index, a point above GLM-5.2, the state-of-the-art open-source model from Z.ai, at 753b parameters. 1 A laptop model beats a recognizable, frontier-class cloud peer roughly 28 times its size.

How does a bumblebee achieve the same flight as an airliner? Bigger models can store more knowledge, so they can skip straight to an answer, like an expert in many different fields. Smaller models don’t have as much memorized, so they must reason more, almost from first principles, to close that gap.2

I saw this firsthand when benchmarking the DeepSeek V4 cloud model against two local models. I compared them on the same work, 25 venture-capital tasks (researching startups, summarizing articles, transcribing podcasts), scored by a judge model.3

Qwen3.8-27B is dense : it uses every chapter in the book on every question. Book skimmers DeepSeek & Qwen 3.6 35b (another local model I threw into the test), flips only to the relevant chapters for a question.4

model quality /9 tok/s avg tokens avg latency

qwen3.8-27b(bumblebee)qwen3.6-35b-a3b(hummingbird)These models provide identically good answers. But the speed varies. The local Qwen 35b shreds at top speed, but needs to think about 7.2x more than the cloud model, crossing the line 9 seconds after DeepSeek. The newest Qwen model is three seconds faster, & the cloud is 6 seconds faster yet.

The cloud model jumps to the right answer ; the local models contemplate & debate internally at different rates of speed & accuracy.

For example : on one triage task, the 35B spent 993 tokens to produce six words, “Classification: Scheduling / Action: Respond.” 1000 tokens of deliberation before the response is a hummingbird’s sprint to a honeysuckle. The bumblebee needed 369 thinking tokens, buzzing along at half the speed. Local models can achieve the same result as cloud models, but they’ll take a different flight path to get there.

Artificial Analysisranks the incumbent here, Qwen3.8-27B, #1 of 135 models on the Intelligence Index, scoring 52, a point above GLM-5.2’s 51, a 753b-parameter frontier model Z.ai shipped two months earlier. The same page ranks it #23 of 135 on output tokens per task, 160M weighted tokens against a class median of 43M. Intelligence rank & verbosity rank move independently, & that’s the trade this whole post is about.↩︎ - The imitation-gap explanation. Smaller models produce fluent chain-of-thought that’s more likely to drift logically inconsistent, because they have a sparser map of nearby correct examples to draw on once forced off the direct path to an answer. See

Chain of Thought in Large Language Models: Elicited Reasoning or Constrained Imitation?I covered the general shape of this tradeoff, trading inference-time compute for capability, inWhen Models Learn.↩︎ - Method. 25 venture-capital tasks (researching startups, summarizing articles, transcribing podcasts) drawn from my own agent queue. A separate judge model, deepseek-v4-pro, scored outputs blind on completeness, accuracy & conciseness, 3 points each for 9 total. max_tokens was 4096 for every run. Both local models were served through Ollama on the same MLX runtime, so the comparison isn’t confounded by runtime differences. I established the judge’s noise floor by re-scoring identical outputs, which returned a mean absolute difference of 0.16.

↩︎ - Qwen3.6-35B-A3B is a 35b parameter model with 3b active parameters per token, a sparse mixture-of-experts architecture, 256 total experts with 8 routed & 1 shared active per token. Per

Qwen’s model card on Hugging Face&vLLM’s model recipe.↩︎

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @qwen3.8-27b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/birds-don-t-fly-like…] indexed:0 read:3min 2026-08-18 ·