cd /news/artificial-intelligence/how-image-and-video-input-actually-b… · home topics artificial-intelligence article
[ARTICLE · art-115259] src=digitalapplied.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

How Image and Video Input Actually Bills Across Vendors

A new reference from Artificial Analysis, based on vendor documentation fetched August 26, 2026, finds that image and video input billing across AI APIs is highly inconsistent, with the same 1024×1024 image costing anywhere from 384 to 1,601 tokens depending on the vendor's conversion rule—a 4.2× spread. Dollar costs for that image range from roughly $0.0008 on Gemini Flash to roughly $0.0068 on Claude Opus 5's high-resolution tier, and three vendors (Mistral, xAI, Z.ai) publish per-token prices without any conversion formula, making costs impossible to budget upfront.

read20 min views38 publishedAug 26, 2026
How Image and Video Input Actually Bills Across Vendors
Image: Digitalapplied (auto-discovered)

Image and video input billing is the least standardized corner of AI API pricing: every vendor charges per token, but a token of image is not a defined quantity. At least five different pixel-to-token conversion rules coexist across the major APIs, and three vendors publish a per-token price while publishing no conversion at all — which means their image costs cannot be budgeted before the invoice arrives.

The stakes compound quietly. The same 1024×1024 screenshot bills anywhere from 384 to 1,601 tokens depending purely on which published rule processes it — a roughly 4.2× spread before any per-token price difference is applied. Where both a published formula and a published rate exist, the dollar cost of that one image runs from roughly $0.0008 on Gemini Flash’s standard rate to roughly $0.0068 at Claude Opus 5’s high-resolution tier — nearly a nine-fold gap for identical pixels.

One boundary sentence before anything else: this reference prices what it costs to feed images and video in — our cost-per-finished-second reference is the output-side companion that prices generating video, and our image generation pricing comparison owns the output side for images. Likewise, our tokenizer cost variance reference covers how text becomes tokens; this page covers how pixels do. Every number below was fetched from vendor documentation on August 26, 2026, and every derived cell is recomputed from the vendor’s own stated formula.

  • 01Text has one unit. Pixels have at least five rulebooks.OpenAI runs two regimes (512px tiles and 32×32px patches), Anthropic counts 28×28px patches, Google counts flat 258-token tiles, Qwen publishes a pixel-area formula, and DeepSeek publishes a flat ceiling instead of a formula. None of these agree.
  • 02The same 1024×1024 image spans 384 to 1,601 tokens.Recomputed from each vendor’s own published rule: DeepSeek’s ceiling bills 384 tokens, Together.ai’s tile grid bills 1,601, and OpenAI, Anthropic, Google, and Qwen land between — a roughly 4.2× spread before per-token prices even enter.
  • 03Three vendors publish a price but no conversion.Mistral’s own docs pose the token-per-image question in an FAQ whose answer did not render on two independent automated fetches; xAI’s models page lists only size limits; Z.ai’s GLM-5.3-Flash guide documents the request format and nothing about token counts.
  • 04Gateways bill their own accounting, not the model author’s.Together.ai publishes its own 560px-tile formula for the open-weight vision models it hosts, and Fireworks publishes an empirical per-resolution table for Qwen2.5-VL — the same model can cost different tokens depending on who serves it.
  • 05Video input has exactly one published per-second meter.Google publishes per-frame and per-second-audio token rates for Gemini; Qwen’s video cost is derivable from its image formula via frame sampling. No other vendor in this set publishes a separate video input rate we could locate.

01 — The ProblemOne per-token price, five conversion rulebooks. #

When a pricing page says “$X per million input tokens,” that number is only half the bill for multimodal input. The other half is the conversion: how many tokens does your image become? For text, the tokenizer answers that (imperfectly — see the variance reference above). For pixels, the answer depends entirely on vendor-specific preprocessing: tiling, patching, resizing, caps, and detail parameters that can move the token count by multiples without any change to the headline rate.

Across the vendors and gateways in this reference, the answers sort into three honesty classes — and the third class is the reason this page exists.

A published formula

You can compute the exact token count of any image before sending it. Every formula-derived number in this reference comes from this class, recomputed from the vendor’s own stated rule.

A published ceiling

No formula, but a hard upper bound: past a resize threshold, every image bills the same. You can budget the worst case exactly, even though you cannot derive intermediate counts.

Nothing located

A per-token price is published; the pixel-to-token conversion is not. Image cost on these models cannot be computed before the invoice — the absence is a documented finding of this pass, not a gap we papered over.

dropped, not estimated. Where a vendor bills images through an unpublished conversion, the cell reads “not published,” and the model is named in prose. Aggregator-sourced conversion figures that will not verify against a vendor primary are excluded entirely — identical wording across third-party pricing blogs is copy-propagation, not corroboration.

02 — OpenAITwo regimes: tiles and patches. #

OpenAI is the only vendor in this set running two accounting regimes at once, split by model generation. Older vision models bill a base token count plus a per-tile charge; the current generations count 32×32-pixel patches with a multiplier. Both are fully documented in the images and vision guide, retrieved August 26, 2026.

Base + 512px tiles

At detail: low, only the base tokens apply regardless of image size. At high or auto, the image is scaled to fit 2048×2048, then — if the shortest side still exceeds 768px — scaled so that side is 768px; each 512px square needed to cover the result adds the per-tile tokens. Outlier: gpt-4o-mini bills 2,833 base + 5,667 per tile.

32×32px patches × 1.2

The gpt-5.2 through gpt-5.6 families count patches, multiply by 1.2 on every patch-based model listed, and round up. Detail levels cap the pixel dimensions and patch budget — gpt-5.4 at detail: high allows up to 2,500 patches and a 2,048px maximum dimension.

Run the patch rule on our test image, on gpt-5.4 at detail: high

: ceil(1024/32) × ceil(1024/32) = 1,024 patches, inside that detail level’s 2,048px dimension cap so no resize applies, and ceil(1024 × 1.2) = 1,229 billable tokens — our arithmetic from OpenAI’s published rule.

Under the tile regime, the same image resizes to 768×768 (shortest side over 768px), which four 512px squares cover: gpt-4o bills 85 + 4×170 = 765 tokens, and gpt-5.1 bills 70 + 4×140 = 630 tokens — our arithmetic from the published rule. Note the direction: OpenAI’s newer patch regime bills roughly twice the tokens of its older tile regime for this image. A model upgrade inside one vendor can move your image bill more than switching vendors would.

03 — Anthropic28×28 patches and two resolution tiers. #

Anthropic’s vision documentation publishes the cleanest single formula in this set.

“Claude views images in patches instead of pixels. Each patch is a 28×28-pixel block of the image, referred to as a visual token. An image, therefore, costs ⌈width / 28⌉ × ⌈height / 28⌉ visual tokens.”— Anthropic vision documentation, retrieved August 26, 2026

The formula is constant; what changed in 2026 is the resolution ceiling it operates under. The standard tier — all models before Claude 4.7 — downsizes any image whose long edge exceeds 1,568px and caps out at 1,568 visual tokens. The high-resolution tier — Claude 4.7 and later — raises the long-edge limit to 2,576px and the cap to 4,784 visual tokens. Per the same docs, “High-resolution support is automatic on the listed models and requires no beta header or client-side opt-in,” and a high-resolution image can use “up to roughly three times more visual tokens than the same image on a standard-tier model.” Read that as a billing event: upgrade models with no code change, and large-image workloads can cost up to roughly 3× more input tokens.

Claude visual tokens by resolution · standard vs high-res tier

Source: Anthropic vision docs, resolution-to-token table, retrieved 2026-08-26Anthropic also publishes its own dollar arithmetic, which we cite rather than restate as ours: at Claude Haiku 4.5’s $1 per million input tokens on the standard tier, a 1000×1000 image costs about $1.30 per thousand images; at Claude Opus 5’s $5 per million on the high-resolution tier, the same image costs about $6.48 per thousand, and a 4K image about $23.92 per thousand. For our 1024×1024 test image, the formula yields ⌈1024/28⌉² = 37×37 = 1,369 visual tokens on both tiers — about $0.0068 per image at Opus 5’s rate, our arithmetic from Anthropic’s published formula and published price.

04 — Google GeminiFlat 258-token tiles — and the only published video meter. #

Google’s image understanding docs publish a tile rule with a twist: images with both dimensions at or under 384px cost a flat 258 tokens, and larger images are cut into 768×768px tiles that each also cost 258 tokens. The tiling uses a crop unit of floor(min(width, height) / 1.5)

— the docs’ own worked example runs a 960×540 image to a 360px crop unit and 3×2 = 6 tiles. A media_resolution

parameter “determines the maximum number of tokens allocated per input image or video frame” — detail traded for tokens, vendor-documented.

Video is where Gemini stands alone in this set: the video understanding docs publish an actual per-second meter. Default sampling is 1 frame per second; at default media resolution each frame costs 258 tokens and audio adds 32 tokens per second (≈290 per second of video by component sum — the docs round to ≈300); at low media resolution each frame costs 66 tokens plus the same 32 for audio.

On price, Gemini’s structure is worth naming precisely: the conversion differs per modality (the formulas above), but the resulting tokens bill at one uniform rate. As of August 26, 2026, Gemini 3.7/3.6 Flash standard input is $0.75 per million tokens through December 31, 2026, with a published rise to $1.50 per million on January 1, 2027. Our 1024×1024 test image works out to a 682px crop unit, 2×2 = 4 tiles, 1,032 tokens — about $0.00077 at the standard rate, our arithmetic from Google’s published rule and published price. That published price step-up also means every image and video figure on this page priced in Gemini dollars is scheduled to double in about four months, with no change to the token math.

05 — Qwen & DeepSeekQwen’s formula, DeepSeek’s ceiling. #

Alibaba’s Model Studio vision docs publish Qwen’s conversion outright: Image Tokens = h_bar × w_bar / token_pixels + 2

, where the dimensions are the image’s size after Qwen’s internal scale-down and token_pixels

is the pixel area per visual token — 32×32 pixels on the qwen3.x families and Qwen3-VL, 28×28 on Qwen2.5-VL and QVQ. A 16,384-token-per-image ceiling applies, with a vl_high_resolution_images

switch and per-family max_pixels defaults governing how large an image stays before preprocessing shrinks it. Our test image needs no downscale: 1024×1024/1,024 + 2 = 1,026 tokens.

What makes Qwen singular in this set is video. The docs describe an fps

parameter — frames extracted every 1/fps seconds — and run those sampled frames through the same published image formula. That makes Qwen the one vendor here whose video input accounting is fully derivable from its own image math, rather than requiring a separately published video rate. (We deliberately publish no dollar figure for Qwen here: the per-million pricing we found is corroborated only by third-party trackers, not by a direct Alibaba pricing-page quote in this pass, so per the rules of this reference it stays out.) For where the hosted Qwen API line sits against today’s open-weights release, see our Qwen3.8-Flash open-versus-hosted split post.

DeepSeek takes the opposite documentation strategy: no formula, a hard ceiling. Its vision guide for deepseek-v4-flash-vision-exp states: “there is an upper bound of 384 tokens per image: for example, a 2000×2000 image and a 5000×5000 image consume the same number of tokens after resizing.” Small images scale up toward roughly 384×384px, large ones scale down toward roughly 800×800px before the cap applies — so our 1024×1024 test image, above that threshold, bills the flat 384. This is the one case in this reference where more pixels provably buy zero extra billed tokens once past a threshold. The model, its peak/off-peak pricing clock, and the launch context are covered in our DeepSeek V4 Flash Vision launch reference — the single-vendor instance of exactly the problem this page treats class-wide.

Pixel area per visual token

On qwen3.x families and Qwen3-VL; Qwen2.5-VL and QVQ count 28×28-pixel tokens instead. The pixel area, not the model’s price, is what the formula keys on.

Token ceiling per image

The vl_high_resolution_images switch and per-family max_pixels defaults decide how much resolution survives preprocessing before the formula runs.

Tokens per image, flat upper bound

No formula published — a 2000×2000 and a 5000×5000 image bill identically after resizing. Worst case is exactly budgetable; intermediate counts are not derivable.

06 — GatewaysGateways bill their own accounting. #

A detail most cost models miss: when an open-weight vision model is served through a gateway, the gateway’s image accounting applies — not the model author’s. Two gateways document theirs.

Together.ai publishes its own formula for the vision models it serves: “Each image breaks into a tile grid (capped at 2×2 of 560-pixel tiles)” at 1,601 tokens per tile — image_tokens = min(2, max(width//560, 1)) × min(2, max(height//560, 1)) × 1601

, capping at 6,404 tokens per image. Run the arithmetic and the cap binds early: any image measuring 1,120px or more on both sides — a 4K screenshot included — bills the identical 6,404-token maximum. Our 1024×1024 test image lands in a single tile: 1,601 tokens, the most expensive published conversion in this reference for that image.

Fireworks.ai maintains a dedicated FAQ page titled “How many tokens per image?” and answers it with an empirical table — but only for Qwen2.5-VL. For other vision families it hosts, Fireworks states the count “varies by model and resolution” and recommends running the model’s own tokenizer. That is a notably honest posture: publish what you measured, name what you did not.

Fireworks’ published image tokens · Qwen2.5-VL by resolution

Source: Fireworks AI billing FAQ, empirical table for Qwen2.5-VL, retrieved 2026-08-26shared patch size, not shared billing: both architectures happen to use 28-pixel patching internally, but Anthropic’s number is a formula on its own models and Fireworks’ is an empirical measurement of a different vendor’s model. A reader skimming both tables could mistake architecture for accounting — do not.

07 — The GapsThe vendors that publish nothing. #

Three vendors in this set sell multimodal input at a per-token price while publishing no way to compute how many tokens an image becomes. Naming them precisely — including exactly what was and was not checkable — is the most useful thing this reference does.

Mistral is the strangest case: a locatability failure rather than a documented absence. Mistral’s own vision docs carry an FAQ entry with the literal heading “How many tokens correspond to an image and/or what is the maximum resolution?” — and on two independent automated fetch attempts on August 26, 2026 (a plain fetch and a scrape with a JavaScript wait), the question heading rendered while the answer body did not. The same page does publish other limits precisely: up to eight images per request and PNG/JPG/JPEG/GIF/WEBP support. A “500 to 2,000 tokens depending on resolution” range circulates across third-party pricing sites for Mistral’s vision models, but we could not trace it to any quotable sentence on Mistral’s own docs — it is excluded from every table here and should be treated as unverified. To be clear about what this is not: we are not claiming Mistral hides its conversion deliberately, only that the answer was not retrievable by automated fetch on this date.

xAI publishes no conversion at all. The models reference states only size constraints — “Maximum image size: 20MiB,” “Maximum number of images: No limit” — with no tile size, patch size, or per-resolution token table anywhere on the fetched page. Secondary coverage describes Grok image input as billed at the model’s standard per-token input rate with no separate per-image fee, but we did not confirm that against a vendor pricing page on this pass, so no dollar figure for Grok appears anywhere on this page. The verified finding is the absence: there is no published way to know what an image costs on Grok before sending it.

Z.ai’s GLM-5.3-Flash — revealed today as the model behind the ox-alpha stealth listing; our reveal-day post has the full story — documents the request format and nothing more. Its API guide shows how to pass an image (“Add a content block with type: image_url

…”) with no formula, tile size, or per-image token count anywhere in the guide. On the OpenRouter listing as of August 26, vision is folded into one unified per-token price regardless of modality: $0.075 per million input tokens promotional through September 9, 2026, and $0.15 per million list price after — a price you can quote, on a conversion you cannot.

08 — The Reference TableThe same screenshot, ten rows. #

One 1024×1024 screenshot, run through every accounting regime that publishes one. Every derived count is recomputed from the vendor’s own stated rule — the basis column says whose arithmetic each number is. Where no conversion is published, the cell says so; nothing in this table is estimated.

Vendor · accounting regime Published conversion rule Tokens · 1024×1024 image Basis
Vendor-published formulas
OpenAI · tile-based (gpt-4o, gpt-4.1) 85 base + 170 per 512px tile after resize to 768px shortest side 765 Our arithmetic from OpenAI’s published rule
OpenAI · tile-based (gpt-5.1, gpt-5) 70 base + 140 per 512px tile 630 Our arithmetic from OpenAI’s published rule
OpenAI · patch-based (gpt-5.2 → gpt-5.6 families) ceil(w/32) × ceil(h/32) patches × 1.2, rounded up 1,229 Our arithmetic from OpenAI’s published rule
Anthropic Claude (both tiers, no resize at this size) ⌈w/28⌉ × ⌈h/28⌉ visual tokens 1,369 Our arithmetic from Anthropic’s published formula
Google Gemini 258 per 768px tile · crop unit floor(min(w,h)/1.5) 1,032 Our arithmetic from Google’s published rule
Qwen (Alibaba Model Studio) h̄ × w̄ / token_pixels + 2 · 32×32px tokens on qwen3.x 1,026 Our arithmetic from Alibaba’s published formula
Gateway-published accounting
Together.ai (hosted open-weight vision models) 1,601 per 560px tile · 2×2 grid cap (6,404 max) 1,601 Our arithmetic from Together’s published formula
Fireworks.ai (Qwen2.5-VL) Empirical per-resolution table, this model only 1,369 Published directly in Fireworks’ table
Published ceiling, not a formula
DeepSeek (deepseek-v4-flash-vision-exp) Flat 384-token upper bound after resize 384 DeepSeek’s published ceiling · image exceeds the ~800×800px threshold
No conversion located
Mistral · xAI Grok · Z.ai GLM-5.3-Flash Per-token price only not published No vendor-side conversion located, 2026-08-26

The trend this table interprets is a divergence, not a convergence. Text token billing slowly standardized around comparable BPE-family tokenizers; image billing is moving the other way — OpenAI just migrated its own line from tiles to patches, Anthropic added a second resolution tier this year, and the newest multimodal launches (GLM-5.3-Flash today among them) ship with no conversion documentation at all. The unit of purchase is getting less comparable across vendors over time, not more — which makes per-vendor unit references like this one, and the discipline of reading a pricing page correctly, more load-bearing every quarter.

09 — Video InputWhat a minute of video costs to feed in. #

Video input is images at scale plus audio, and the documentation picture is starker than for images: one vendor publishes a per-second meter, one vendor’s video cost is derivable from its image formula, and everyone else in this set publishes no separate video input rate we could locate. Again — this table is about feeding video in for understanding; generating video out is priced in the output-side companion reference.

Vendor · mode Sampling & inputs Tokens / second ≈ Tokens / minute Basis
Published per-second video meter
Gemini · default media resolution 1 fps · 258 per frame + 32 per second audio 290 ≈17,400 Our arithmetic from Google’s published per-second breakdown
Gemini · low media resolution 1 fps · 66 per frame + 32 per second audio 98 (docs round to ≈100) ≈5,880 Our arithmetic from Google’s published per-second breakdown
Derivable from the vendor’s own image formula
Qwen · fps-sampled frames Example: 1 fps of 1024×1024 frames → 1,026 tokens per frame ≈1,026 ≈61,560 Our arithmetic from Alibaba’s published image formula; frame size and fps are our example assumptions
No published video meter located
xAI Grok · Mistral · Z.ai GLM-5.3-Flash · DeepSeek not published not published No separate video input rate located, 2026-08-26

Two things jump out of the normalization. First, at Gemini’s published $0.75-per-million standard rate (through December 31, 2026), a minute of default-resolution video costs about $0.013 of input tokens, and a minute at low resolution about $0.0044 — our arithmetic from Google’s published components. Second, the Qwen example row shows why frame-sampled billing punishes resolution: a minute of 1fps 1024×1024 frames runs roughly 3.5× Gemini’s default-resolution token count for the same wall-clock minute, because each Qwen frame bills like a full image while Gemini’s per-frame rate is a flat 258. The lesson generalizes: for video, the sampling and per-frame rule dominate the bill far more than the per-token price does.

Use the low-detail levers

OpenAI’s detail: low bills base tokens only regardless of size; Gemini flat-rates anything at or under 384px at 258 tokens. If the task is UI navigation rather than OCR, the cheap modes are vendor-sanctioned — use them.

Price the tier before the upgrade

Claude 4.7+’s high-resolution tier is automatic and can bill up to roughly 3× the visual tokens of the standard tier on large images, per Anthropic’s own comparison. A model upgrade with no code change is also a billing change — model it first.

Route to published meters

Gemini is the only published per-second video meter in this set; Qwen is derivable from its image formula. Everything else is unmeterable pre-flight — route video workloads to vendors whose meter you can read, or accept unbudgeted invoices.

Treat as post-hoc only

On Mistral, Grok, and GLM-5.3-Flash, image cost is knowable only from usage dashboards after the fact. Run a fixed image set through the API, measure billed tokens empirically, and re-measure on every model update — the vendor gives you nothing to pin a forecast on.

One capability-side note to keep the frame honest: none of this says which model reads images best — token cost and vision quality are independent axes, and the capability side has its own reference in our multimodal benchmarks roundup. For teams running image- or video-heavy AI workloads in production, this is exactly the class of unit-economics audit our AI transformation engagements start with: measure the real billed tokens per workload unit on your own traffic, then route by meter, not by headline rate.

10 — ConclusionThe pixel has no standard exchange rate. #

Budget the conversion, not the rate — and name what cannot be converted.

The per-token price is the visible half of multimodal input billing; the pixel-to-token conversion is the half that actually moves the bill, and it varies by multiples across vendors for identical pixels. Where the conversion is published — OpenAI, Anthropic, Google, Qwen, and the two gateways — every image cost on this page is computable before you send the request. Where it is not, no amount of pricing-page reading will produce a budget.

The honest gaps are the part of this reference we expect to age best. Three vendors currently sell image input at a price with no published unit behind it, and today’s launches suggest the gap is widening rather than closing: the newest multimodal models shipped with request-format docs and nothing about token accounting. Our working projection is that conversion disclosure becomes a competitive surface within a few quarters — the vendors that publish formulas are already the only ones whose multimodal costs enterprises can model, and procurement teams notice that asymmetry faster than benchmarks move.

Until then, the practical rule is the one this page is built on: trust only the vendor’s own stated rule, recompute every derived number from it, route meterable workloads to meterable vendors, and treat any model whose conversion is unpublished as a measure-after-the-fact line item — never a forecastable one.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-image-and-video-…] indexed:0 read:20min 2026-08-26 ·