cd /news/artificial-intelligence/benchmarking-pocket-scale-inference · home topics artificial-intelligence article
[ARTICLE · art-115510] src=artificialanalysis.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Benchmarking Pocket-Scale Inference

Artificial Analysis, in partnership with Liquid AI, has launched a benchmark suite for pocket-scale AI models that fit within 8 GB of memory after quantization, including KV cache at 8K context, measuring intelligence and inference performance on mobile devices such as the iPhone 17 Pro. The tests cover real-world mobile usage scenarios, with independent validation of Liquid AI's inference measurement process, and results are presented as an Average Score across benchmarks like BFCL, IFBench, GPQA Diamond, and MATH-500.

read1 min views15 publishedAug 27, 2026
Benchmarking Pocket-Scale Inference
Image: source

We benchmark small models on mobile phones. Artificial Analysis' testing covers model intelligence on a set of benchmarks chosen to represent real-world mobile device usage, and we partner with Liquid AI to gather real inference data measured on the devices themselves. Note: we have independently validated Liquid AI's inference measurement process.

“Small” models are all models that fit inside 8 GB of memory after quantization, including KV cache at 8K context. View all rules and our process in the methodology page.

Intelligence and Inference Performance Summary #

Average Score (16K max context) vs. End-to-End Generation TimeiPhone 17 Pro

Inference Performance #

End-to-End Generation TimeiPhone 17 Pro

Model Intelligence #

Average Score (Mobile Device Benchmark Set, 16K max context)iPhone 17 Pro

Looking for the Artificial Analysis Intelligence Index scores for these models? The following models have been evaluated on our full index, and their scores are visible on their model pages:

Token Efficiency #

Context Budget Overruns

Evaluation Breakdown #

Mobile Device Benchmark Set Evaluations (16K max context)iPhone 17 Pro

BFCL

Tool calling (index subset) IFBench

Instruction following

AA-Omniscience Accuracy

Knowledge

AA-Omniscience Non-Hallucination Rate 1 - hallucination rate

GPQA Diamond

Scientific reasoning

MATH-500

Quantitative reasoning

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @artificial analysis 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/benchmarking-pocket-…] indexed:0 read:1min 2026-08-27 ·