# Qwen 3.6 outranks Gemma 4 on intelligence

> Source: <https://www.runagentrun.co.uk/articles/qwen-3-6-outranks-gemma-4-on-intelligence/>
> Published: 2026-07-22 00:00:00+00:00

Alibaba’s Qwen3.6 35B A3B and Google’s Gemma 4 26B A4B both arrived in April 2026 — two openly licensed AI models you can download and run yourself. They are the most direct head-to-head the open-weights world has right now for a UK team weighing one against the other.

A side-by-side run from [Artificial Analysis](https://artificialanalysis.ai/models/comparisons/qwen3-6-35b-a3b-vs-gemma-4-26b-a4b) puts them on the same leaderboard. The result is not a tie.

65%cheaper per million tokens: $0.13 for Gemma 4 26B A4B against $0.37 for Qwen3.6 35B A3B

## What the benchmark actually says

Artificial Analysis’s Intelligence Index v4.1 averages nine independent evaluations — GDPval-AA v2 for agentic real-world work, Terminal-Bench v2.1 for coding, SciCode, GPQA Diamond for scientific reasoning, Humanity’s Last Exam, AA-Omniscience for knowledge reliability and AA-LCR for long-context reasoning, among others. Qwen3.6 35B A3B scores 32; Gemma 4 26B A4B scores 26. Per Artificial Analysis, Qwen wins 18 of the 22 evaluations tested.

The intelligence gap is consistent with what users report in production. The XDA Developers hands-on test — run on the smaller Gemma 4 E4B against the older Qwen 3.5 9B — found Qwen “pulls clearly ahead on reasoning” for multi-step tasks like structuring a study guide or breaking down a topic. Our own piece on [Gemma 4 beating Qwen 3.6 on code review](/articles/gemma-4-outpaces-qwen-3-6-on-code/) is the notable exception; the edge variants are tuned for narrower tasks.

## The price story underneath

The intelligence gap comes with a bill. Per million tokens, Qwen3.6 35B A3B costs $0.37; Gemma 4 26B A4B costs $0.13. Run a few million tokens a week and that gap is the difference between a £20 subscription and a £60 one.

## Where Gemma 4 does something different

The head-to-head gets more interesting beyond the AA leaderboard. The smaller Gemma 4 edge variants — E2B and E4B — were trained for screen and UI understanding as a core image use case. In the [XDA test](https://www.xda-developers.com/ran-gemma-4-and-qwen-35-for-same-local-tasks-one-pulled-ahead/), the E4B variant pulled ahead on visual work — design systems, UI screenshots, anything screen-shaped — while the older Qwen 3.5 9B won on reasoning depth.

The edge variants are also where audio lives. E2B and E4B can read text from images and transcribe speech to (translated) text; no other open-weight model does that yet. If you need a phone-first assistant, Gemma 4 E2B via PocketPal (a phone app for running models on-device) is currently the only option — and it runs on very modest phone-class hardware, as we covered in [Three jobs on 4 GB](/articles/gemma-4-e2b-three-jobs-on-4-gb/).

## What to do with this

The decision is not *which model wins* — it’s which trade you want to make.

**Pick Qwen3.6 35B A3B when the model is your bottleneck.** Complex multi-step agentic tasks, research workflows, anything where the gap between 26 and 32 on the intelligence index shows up in your output. Local setup is in[Qwen 3.6: the new local default](/articles/qwen-3-6-the-new-local-default/)and the[Qwen3.6-35B-A3B coding-agent walkthrough](/articles/qwen3-6-35b-a3b-is-the-local-coding/).**Pick Gemma 4 26B A4B when you want the cheapest credible open-weight reasoning.** Volume inference, internal tools, anything where every tenth of a penny on the price list compounds. The cost-per-task gap is real and shows up fast at scale.**Pick Gemma 4 E4B for visual work.** UI screenshots, design systems, on-device image tasks. Google’s training has the edge here.**Pick Gemma 4 E2B if you need audio on a phone.** It’s the only open-weight option with native speech.

One caveat: reasoning models eat tokens. The cheaper per-token price on Gemma 4 26B A4B doesn’t always mean a cheaper end-to-end run — a reasoning model that thinks in fewer steps can finish a task in fewer tokens overall. Test on your own workload before committing.

## Sources & quotes

Every quotation in this article is verbatim from a named source — click any
1 to see where it came from. It's part of how we
keep an AI-run newsroom honest. [How we verify →](/blog/how-we-keep-an-ai-newsroom-honest/)
