cd /news/artificial-intelligence/a-10m-token-context-efficient-agenti… · home topics artificial-intelligence article
[ARTICLE · art-92305] src=console.pokee.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

A 10M-token context efficient agentic model

Pokee AI's Isaac 28B v0 model is the only one in a six-model comparison to return usable scores at every context length up to 10M tokens, achieving 93.3 on the RULER benchmark at 10M, while competitors like GPT-5.6 Luna and Gemini 3.5 Flash Lite fail at 1M or beyond. Isaac also leads in agentic benchmarks (BFCL v4: 70.94, τ³-bench: 0.662) and security (DTAP ASR: 35.6), with API pricing at $0.15 per 1M input tokens and $1.00 per 1M output tokens.

read9 min views2 publishedAug 11, 2026

Isaac reasons, plans, and uses tools across context windows up to 10M tokens, under a compute budget small enough to serve inside a VPC, on customer premises, or on a device. Every figure below was measured by Pokee AI in a single controlled environment, for Isaac and for every baseline alike, unless marked otherwise.

Serving api.pokee.ai · Results from Pokee-Isaac 28B v0: A 10M-Token Context Efficient Agentic Model, Pokee AI, August 3, 2026.

The smallest model in the comparison panel, by a wide margin.

Usable end to end, not merely addressable.

Below every baseline that can be bought at this length.

Benchmark overview

The full comparison at a glance. Each section below takes one row of this table and shows how the result was reached — where the panel diverges, and where a baseline finishes ahead.

API pricing: USD per 1M input / output tokens

Benchmark Pokee-Isaac28B (v0)In/Out$0.15 / $1.00 GPT-5.6 LunaAzureIn/Out$0.40 / $1.80>272K context Gemini 3.5 Flash LiteVertex AIIn/Out$0.30 / $2.50 Claude Haiku 4.5BedrockIn/Out$1.00 / $5.00 Nemotron 3 SuperAmazon Bedrock · USIn/Out$0.15 / $0.65 Qwen 3.5 122BOpenRouterIn/Out$0.26 / $2.08
Long context
RULER256K / 512K / 1M higher is better 96.9 / 96.7 / 95.0 (best in row) 95.0 / 91.4 / 0.0 94.5 / 94.6 / 29.4 0.0 / 0.0 / 0.0 96.3 / 95.7 / 91.8 0.0 / 0.0 / 0.0
RULER2M / 4M / 10M higher is better 95.8 / 96.7 / 93.3 (best in row) 0.0 / 0.0 / 0.0 0.0 / 0.0 / 0.0 0.0 / 0.0 / 0.0 0.0 / 0.0 / 0.0 0.0 / 0.0 / 0.0
MRCR v2256K / 512K / 1M, 8 needles higher is better 0.607 / 0.743 / 0.500 (best in row) 0.208 / 0.173 / 0.050 0.474 / 0.473 / 0.205 0.000 / 0.000 / 0.000 0.145 / 0.161 / 0.067 0.000 / 0.000 / 0.000
Agentic capabilities
BFCL v4overall higher is better 70.94 (best in row) 70.61 64.85 67.52 33.13 64.88
τ³-bench4-domain average higher is better 0.662 (best in row) 0.527 0.631 0.408 0.426 0.611
Terminal-Bench 2.1text-only subset higher is better 65.1% 69.8% (best in row) 46.5% 34.9% 24.4% 46.5%
MCP-Atlasclaim coverage higher is better 74.59% 77.90% (best in row) 76.67% 56.45% 48.95% 70.24%
Security
DTAPattack success rate (ASR) lower is better 35.6 (best in row) 50.1 66.3 37.9 60.4 54.0
DTAPbenign success rate (BSR) higher is better 82.5 85.1 (best in row) 83.3 71.3 63.3 79.4

🥇 marks the best value in each row. ASR is attack success rate and is lower-is-safer; BSR is benign task success rate; higher is better for every other benchmark. A 0.0 / 0.000 means the model returned nothing usable at that length.

context-overflow error at 1M. vendor self-reported, not measured by Pokee — excluded from the row comparison. Every other figure was produced on one installation, for Isaac and each baseline alike.Pricing: Pokee and Luna rates are as supplied/official; Gemini and Haiku use standard public API rates; Nemotron uses Amazon Bedrock on-demand pricing for US East / US West; Qwen uses the OpenRouter headline rate, which varies by provider. Haiku, Nemotron, and Qwen cannot be purchased at the context lengths above 262K that this comparison covers.

Long context

RULER holds task difficulty fixed and scales only the context length, so the curve shows how far a model's usable context tracks its nominal one. Isaac is the only model in the panel that returns a score at every length.

Model 256K 512K 1M 2M 4M 10M
Pokee-Isaac 28B 96.9 96.7 95.0 95.8 96.7 93.3
GPT-5.6 Luna(Azure) 95.0 91.4 0.0, context-overflow error 0.0, no usable score 0.0, no usable score 0.0, no usable score
Gemini 3.5 Flash Lite(Vertex AI) 94.5 94.6 29.4 (context-overflow error) 0.0, no usable score 0.0, no usable score 0.0, no usable score
Claude Haiku 4.5(Bedrock) 0.0, no usable score 0.0, no usable score 0.0, no usable score 0.0, no usable score 0.0, no usable score 0.0, no usable score
Nemotron 3 Super 120B 96.3 (self-reported by vendor) 95.67 (self-reported by vendor) 91.75 (self-reported by vendor) 0.0, no usable score 0.0, no usable score 0.0, no usable score
Qwen 3.5 122B 0.0, no usable score 0.0, no usable score 0.0, no usable score 0.0, no usable score 0.0, no usable score 0.0, no usable score

Pokee-IsaacBaselineVendor self-reported No usable score (0.0)

RULER score (%, averaged across task configurations), ten samples per configuration. The 256K and 512K columns average all 13 configurations; common-words extraction is unavailable from 1M onward, so those columns average the remaining 12.

marks a context-overflow error at that length. Nemotron's 256K–1M figures are self-reported by NVIDIA, not measured by Pokee, and are excluded from the comparison.RULER measures how deep a model can reach; MRCR measures whether it can tell several buried targets apart. GPT-5.6 Luna scores 95.0 on RULER at 256K but 0.050 here at 1M — multi-needle disambiguation splits the panel far more sharply than single-target recall does.

Agentic capability

Each probes something the others cannot mask: deterministic function calling, sustained multi-turn coherence, execution in a real shell, and discovery across live tool servers. Isaac leads two, places second on one, and third on one.

Security

DTAP places an agent in simulated environments and measures whether injected attacks succeed. Against the five baselines run on the same installation with the judge held fixed, Isaac is the safest of six on both axes while placing third of six on capability.

Model Direct ASRlower is safer Indirect ASRlower is safer Combined ASRlower is safer BSRhigher is better
Pokee-Isaac 28B 36.0 (best in column) 35.2 (best in column) 35.6 (best in column) 82.5
GPT-5.6 Luna 54.4 46.1 50.1 85.1 (best in column)
Gemini 3.5 Flash Lite 84.1 49.5 66.3 83.3
Claude Haiku 4.5 38.2 37.1 37.9 71.3
Nemotron 3 Super 120B 80.3 42.0 60.4 63.3
Qwen 3.5 122B 60.8 47.8 54.0 79.4

DTAP (DecodingTrust-Agent Platform) places an agent in simulated environments and measures whether injected attacks succeed. Direct ASR is the attack success rate when the harmful request is in the user prompt; Indirect ASR when it arrives through tool output or the environment. BSR is benign task success — utility. Macro-average over 12 Linux domains and 6,195 judged tasks, same judge for every row.

Isaac is the safest of the six on both attack axes, and its Direct and Indirect rates differ by 0.8 points — the tightest balance in the set. Two limits are stated plainly in the report: the indirect figures are a guards-inactive measurement, because the harness matches native tool names while DTAP's attacks arrive over MCP; and explicit refusals fired on only 1.5% of malicious tasks, so most of the current defense is incidental rather than declined. Against the broader published leaderboard of sixteen systems, Isaac places #5 on Direct ASR, #6 on Indirect, and #9 on capability — a middle-of-the-field result, and the report states it as one.

Efficiency

Measured under the same RULER workload whose accuracy is reported above, on a single B200-class GPU — so the serving profile is directly comparable to the capability curve rather than benchmarked on a friendlier task.

Top of the measured sweep on a single NVIDIA B200; 42,400 at 1M.

For a full 10M-token prompt on that B200.

Unchanged from 1M to 10M context.

Context Concurrency TTFT Prefill (tok/s) Decode (tok/s)
1M 1 23.6 s 42,400 335
1M 4 49.3 s 81,200 322
10M 1 72.9 s 137,200 337

Pricing

Published rates per million tokens. Three models in the panel carry a rate but cannot be bought at the context lengths this evaluation covers — no commercially available endpoint serves them beyond 262K.

Model Max context Input $/M Output $/M
Pokee-Isaac 28B 10M $0.15 $1.00
GPT-5.6 LunaAzure 1.05M $0.40 $1.80>272K context
Gemini 3.5 Flash LiteVertex AI 1M $0.30 $2.50
Claude Haiku 4.5Bedrock 200Knot sold above 262K $1.00 $5.00
Nemotron 3 Super 120BAmazon Bedrock · US 262Knot sold above 262K $0.15 $0.65
Qwen 3.5 122BOpenRouter 262Knot sold above 262K $0.26 $2.08

Retrieved from each provider's public pricing page on 3 August 2026; long-context rates are quoted where a provider meters them separately. Nemotron uses Amazon Bedrock on-demand pricing for US East / US West; Qwen uses the OpenRouter headline rate, which varies by provider. Pokee-Isaac rates are provisional and subject to confirmation at launch. For in-boundary deployments the per-token comparison understates the difference — cost becomes a fixed function of the hardware provisioned rather than a variable function of tokens consumed.

Deployment

Portability is a first-class property: Isaac is adapted to run on heterogeneous accelerator hardware, served natively through the Pokee SDK. A 28B footprint is what makes the small end of this range possible.

137,200 tok/s prefill at 10M context, decode flat near 335 tok/s.

A single consumer GPU is enough to run Isaac privately — no datacenter part required.

3.6–5× the prefill and 2.3× the decode of stock llama.cpp on the same card.

On-device at extended context on NPU-class mobile silicon; AMD in progress.

Point any chat-completions client at the Pokee base URL — no new SDK to learn.

A purpose-built agentic architecture for function calling and long-horizon execution.

Stream over SSE, or run generation in background mode that survives a disconnect.

Licensed to run inside a VPC, on-premises, or on-device — no request leaves your perimeter.

Get started

Swap in your key and the base URL below. Everything else is a standard chat-completions request.

Request bodies over 16 MiB require SSE. Set stream: true

and send Accept: text/event-stream

. See long-context requirements.

curl https://api.pokee.ai/v1/chat/completions \
  -H "Authorization: Bearer pk-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "pokee-isaac",
    "messages": [{"role": "user", "content": "hello"}]
  }'

Pokee-Isaac is a text model for now. Image, audio, and video inputs are not supported.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @pokee ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-10m-token-context-…] indexed:0 read:9min 2026-08-11 ·