Isaac reasons, plans, and uses tools across context windows up to 10M tokens, under a compute budget small enough to serve inside a VPC, on customer premises, or on a device. Every figure below was measured by Pokee AI in a single controlled environment, for Isaac and for every baseline alike, unless marked otherwise.
Serving api.pokee.ai · Results from Pokee-Isaac 28B v0: A 10M-Token Context Efficient Agentic Model, Pokee AI, August 3, 2026.
The smallest model in the comparison panel, by a wide margin.
Usable end to end, not merely addressable.
Below every baseline that can be bought at this length.
Benchmark overview
The full comparison at a glance. Each section below takes one row of this table and shows how the result was reached — where the panel diverges, and where a baseline finishes ahead.
API pricing: USD per 1M input / output tokens
| Benchmark | Pokee-Isaac28B (v0)In/Out$0.15 / $1.00 | GPT-5.6 LunaAzureIn/Out$0.40 / $1.80>272K context | Gemini 3.5 Flash LiteVertex AIIn/Out$0.30 / $2.50 | Claude Haiku 4.5BedrockIn/Out$1.00 / $5.00 | Nemotron 3 SuperAmazon Bedrock · USIn/Out$0.15 / $0.65 | Qwen 3.5 122BOpenRouterIn/Out$0.26 / $2.08 |
|---|---|---|---|---|---|---|
| Long context | ||||||
| RULER256K / 512K / 1M higher is better | 96.9 / 96.7 / 95.0 (best in row) | 95.0 / 91.4 / 0.0 | 94.5 / 94.6 / 29.4 | 0.0 / 0.0 / 0.0 | 96.3 / 95.7 / 91.8 | 0.0 / 0.0 / 0.0 |
| RULER2M / 4M / 10M higher is better | 95.8 / 96.7 / 93.3 (best in row) | 0.0 / 0.0 / 0.0 | 0.0 / 0.0 / 0.0 | 0.0 / 0.0 / 0.0 | 0.0 / 0.0 / 0.0 | 0.0 / 0.0 / 0.0 |
| MRCR v2256K / 512K / 1M, 8 needles higher is better | 0.607 / 0.743 / 0.500 (best in row) | 0.208 / 0.173 / 0.050 | 0.474 / 0.473 / 0.205 | 0.000 / 0.000 / 0.000 | 0.145 / 0.161 / 0.067 | 0.000 / 0.000 / 0.000 |
| Agentic capabilities | ||||||
| BFCL v4overall higher is better | 70.94 (best in row) | 70.61 | 64.85 | 67.52 | 33.13 | 64.88 |
| τ³-bench4-domain average higher is better | 0.662 (best in row) | 0.527 | 0.631 | 0.408 | 0.426 | 0.611 |
| Terminal-Bench 2.1text-only subset higher is better | 65.1% | 69.8% (best in row) | 46.5% | 34.9% | 24.4% | 46.5% |
| MCP-Atlasclaim coverage higher is better | 74.59% | 77.90% (best in row) | 76.67% | 56.45% | 48.95% | 70.24% |
| Security | ||||||
| DTAPattack success rate (ASR) lower is better | 35.6 (best in row) | 50.1 | 66.3 | 37.9 | 60.4 | 54.0 |
| DTAPbenign success rate (BSR) higher is better | 82.5 | 85.1 (best in row) | 83.3 | 71.3 | 63.3 | 79.4 |
🥇 marks the best value in each row. ASR is attack success rate and is lower-is-safer; BSR is benign task success rate; higher is better for every other benchmark. A 0.0 / 0.000 means the model returned nothing usable at that length.
context-overflow error at 1M. vendor self-reported, not measured by Pokee — excluded from the row comparison. Every other figure was produced on one installation, for Isaac and each baseline alike.Pricing: Pokee and Luna rates are as supplied/official; Gemini and Haiku use standard public API rates; Nemotron uses Amazon Bedrock on-demand pricing for US East / US West; Qwen uses the OpenRouter headline rate, which varies by provider. Haiku, Nemotron, and Qwen cannot be purchased at the context lengths above 262K that this comparison covers.
Long context
RULER holds task difficulty fixed and scales only the context length, so the curve shows how far a model's usable context tracks its nominal one. Isaac is the only model in the panel that returns a score at every length.
| Model | 256K | 512K | 1M | 2M | 4M | 10M |
|---|---|---|---|---|---|---|
| Pokee-Isaac 28B | 96.9 | 96.7 | 95.0 | 95.8 | 96.7 | 93.3 |
| GPT-5.6 Luna(Azure) | 95.0 | 91.4 | 0.0, context-overflow error | 0.0, no usable score | 0.0, no usable score | 0.0, no usable score |
| Gemini 3.5 Flash Lite(Vertex AI) | 94.5 | 94.6 | 29.4 (context-overflow error) | 0.0, no usable score | 0.0, no usable score | 0.0, no usable score |
| Claude Haiku 4.5(Bedrock) | 0.0, no usable score | 0.0, no usable score | 0.0, no usable score | 0.0, no usable score | 0.0, no usable score | 0.0, no usable score |
| Nemotron 3 Super 120B | 96.3 (self-reported by vendor) | 95.67 (self-reported by vendor) | 91.75 (self-reported by vendor) | 0.0, no usable score | 0.0, no usable score | 0.0, no usable score |
| Qwen 3.5 122B | 0.0, no usable score | 0.0, no usable score | 0.0, no usable score | 0.0, no usable score | 0.0, no usable score | 0.0, no usable score |
Pokee-IsaacBaselineVendor self-reported No usable score (0.0)
RULER score (%, averaged across task configurations), ten samples per configuration. The 256K and 512K columns average all 13 configurations; common-words extraction is unavailable from 1M onward, so those columns average the remaining 12.
marks a context-overflow error at that length. Nemotron's 256K–1M figures are self-reported by NVIDIA, not measured by Pokee, and are excluded from the comparison.RULER measures how deep a model can reach; MRCR measures whether it can tell several buried targets apart. GPT-5.6 Luna scores 95.0 on RULER at 256K but 0.050 here at 1M — multi-needle disambiguation splits the panel far more sharply than single-target recall does.
Agentic capability
Each probes something the others cannot mask: deterministic function calling, sustained multi-turn coherence, execution in a real shell, and discovery across live tool servers. Isaac leads two, places second on one, and third on one.
Security
DTAP places an agent in simulated environments and measures whether injected attacks succeed. Against the five baselines run on the same installation with the judge held fixed, Isaac is the safest of six on both axes while placing third of six on capability.
| Model | Direct ASRlower is safer | Indirect ASRlower is safer | Combined ASRlower is safer | BSRhigher is better |
|---|---|---|---|---|
| Pokee-Isaac 28B | 36.0 (best in column) | 35.2 (best in column) | 35.6 (best in column) | 82.5 |
| GPT-5.6 Luna | 54.4 | 46.1 | 50.1 | 85.1 (best in column) |
| Gemini 3.5 Flash Lite | 84.1 | 49.5 | 66.3 | 83.3 |
| Claude Haiku 4.5 | 38.2 | 37.1 | 37.9 | 71.3 |
| Nemotron 3 Super 120B | 80.3 | 42.0 | 60.4 | 63.3 |
| Qwen 3.5 122B | 60.8 | 47.8 | 54.0 | 79.4 |
DTAP (DecodingTrust-Agent Platform) places an agent in simulated environments and measures whether injected attacks succeed. Direct ASR is the attack success rate when the harmful request is in the user prompt; Indirect ASR when it arrives through tool output or the environment. BSR is benign task success — utility. Macro-average over 12 Linux domains and 6,195 judged tasks, same judge for every row.
Isaac is the safest of the six on both attack axes, and its Direct and Indirect rates differ by 0.8 points — the tightest balance in the set. Two limits are stated plainly in the report: the indirect figures are a guards-inactive measurement, because the harness matches native tool names while DTAP's attacks arrive over MCP; and explicit refusals fired on only 1.5% of malicious tasks, so most of the current defense is incidental rather than declined. Against the broader published leaderboard of sixteen systems, Isaac places #5 on Direct ASR, #6 on Indirect, and #9 on capability — a middle-of-the-field result, and the report states it as one.
Efficiency
Measured under the same RULER workload whose accuracy is reported above, on a single B200-class GPU — so the serving profile is directly comparable to the capability curve rather than benchmarked on a friendlier task.
Top of the measured sweep on a single NVIDIA B200; 42,400 at 1M.
For a full 10M-token prompt on that B200.
Unchanged from 1M to 10M context.
| Context | Concurrency | TTFT | Prefill (tok/s) | Decode (tok/s) |
|---|---|---|---|---|
| 1M | 1 | 23.6 s | 42,400 | 335 |
| 1M | 4 | 49.3 s | 81,200 | 322 |
| 10M | 1 | 72.9 s | 137,200 | 337 |
Pricing
Published rates per million tokens. Three models in the panel carry a rate but cannot be bought at the context lengths this evaluation covers — no commercially available endpoint serves them beyond 262K.
| Model | Max context | Input $/M | Output $/M |
|---|---|---|---|
| Pokee-Isaac 28B | 10M | $0.15 | $1.00 |
| GPT-5.6 LunaAzure | 1.05M | $0.40 | $1.80>272K context |
| Gemini 3.5 Flash LiteVertex AI | 1M | $0.30 | $2.50 |
| Claude Haiku 4.5Bedrock | 200Knot sold above 262K | $1.00 | $5.00 |
| Nemotron 3 Super 120BAmazon Bedrock · US | 262Knot sold above 262K | $0.15 | $0.65 |
| Qwen 3.5 122BOpenRouter | 262Knot sold above 262K | $0.26 | $2.08 |
Retrieved from each provider's public pricing page on 3 August 2026; long-context rates are quoted where a provider meters them separately. Nemotron uses Amazon Bedrock on-demand pricing for US East / US West; Qwen uses the OpenRouter headline rate, which varies by provider. Pokee-Isaac rates are provisional and subject to confirmation at launch. For in-boundary deployments the per-token comparison understates the difference — cost becomes a fixed function of the hardware provisioned rather than a variable function of tokens consumed.
Deployment
Portability is a first-class property: Isaac is adapted to run on heterogeneous accelerator hardware, served natively through the Pokee SDK. A 28B footprint is what makes the small end of this range possible.
137,200 tok/s prefill at 10M context, decode flat near 335 tok/s.
A single consumer GPU is enough to run Isaac privately — no datacenter part required.
3.6–5× the prefill and 2.3× the decode of stock llama.cpp on the same card.
On-device at extended context on NPU-class mobile silicon; AMD in progress.
Point any chat-completions client at the Pokee base URL — no new SDK to learn.
A purpose-built agentic architecture for function calling and long-horizon execution.
Stream over SSE, or run generation in background mode that survives a disconnect.
Licensed to run inside a VPC, on-premises, or on-device — no request leaves your perimeter.
Get started
Swap in your key and the base URL below. Everything else is a standard chat-completions request.
Request bodies over 16 MiB require SSE. Set stream: true
and send Accept: text/event-stream
. See long-context requirements.
curl https://api.pokee.ai/v1/chat/completions \
-H "Authorization: Bearer pk-..." \
-H "Content-Type: application/json" \
-d '{
"model": "pokee-isaac",
"messages": [{"role": "user", "content": "hello"}]
}'
Pokee-Isaac is a text model for now. Image, audio, and video inputs are not supported.