DeepSeek V4 Flash Vision Exp Launches: Multimodal AI at $0.66 per Million Tokens DeepSeek launched DeepSeek-V4-Flash-Vision-Exp on August 21, 2026, adding native image understanding to its budget flagship at $0.66 per million output tokens off-peak, with a 1M-token context window. The model scores within 5 points of Claude Opus 4.8 on most agent benchmarks while costing roughly 38x less per output token, according to vendor-run tests. DeepSeek claims the model brings multimodal agent performance close to Opus-4.8, though independent evaluations are pending. Industry News DeepSeek V4 Flash Vision Exp Launches: Multimodal AI at $0.66 per Million Tokens DeepSeek adds native vision to its budget flagship with V4-Flash-Vision-Exp. Off-peak output at $0.66/1M tokens, 1M context, and benchmark scores close to Claude Opus 4.8. Some links are partner links: if you subscribe through them, we may earn a commission, at no extra cost to you. The crowd verdicts stay independent. DeepSeek-V4-Flash-Vision-Exp launched on August 21, 2026, adding native image understanding to the cheapest frontier model on the market. The experimental release extends V4-Flash's text capabilities into multimodal territory at the same $0.66 per million output tokens off-peak , making vision-capable AI accessible to developers who previously had to choose between cost and capability. The short answer V4-Flash-Vision-Exp processes images and text in a single model at $0.66/1M output tokens off-peak $1.32 peak . It scores within 5 points of Claude Opus 4.8 on most agent benchmarks while costing roughly 38x less per output token. Why this release matters for multimodal agents Until now, DeepSeek-V4 was text-only, a significant gap when competing against Claude Opus 4.8 57% crowd recommend in the GLAD-AI-TOR arena and GPT-5.5, both of which have had vision for months. The new model fills that hole with first-class image input: up to 384 tokens per image, billed at the standard input rate with no surcharge. The context window stays at 1,048,576 tokens with a 384,000-token output limit. Thinking mode is on by default and the API is available through both OpenAI-compatible and Anthropic-format endpoints. | Model | Output price per 1M | Context | Vision | Crowd score | |---|---|---|---|---| | DeepSeek-V4-Flash-Vision-Exp | $0.66 off-peak | 1M | Yes | 57% V4 base | | DeepSeek-V4 Pro | $0.87 | 1M | No | 57% | | Claude Opus 4.8 | $25.00 | 1M | Yes | 57% | | GPT-5.5 | $30.00 | 1M | Yes | 57% | | Gemini 3 Pro | $12.00 | 1M | Yes | 57% | The price gap is stark: V4-Flash-Vision-Exp costs 38x less than Claude Opus 4.8 and 45x less than GPT-5.5 per output token. Cache hits drop the input rate to $0.007/1M off-peak, roughly 99% below the standard input price. Benchmark reality check: close to Opus, not equal DeepSeek's announcement claims the model "brings multimodal agent performance close to Opus-4.8." The self-reported numbers mostly support this, with caveats. On Terminal Bench 2.1, V4-Flash-Vision-Exp scores 83.9 versus Claude Opus 4.8's 85.0. The gap widens on NL2Repo 57.7 vs 69.7, a 12-point deficit and DSBench-Hard 63.6 vs 71.7 . It leads Opus on three benchmarks: DeepSWE +1.3 points , Agents' Last Exam +1.6 , and ZeroBench +1.0 . These are vendor-run benchmarks, not independent evaluations. Treat "close to Opus" as accurate and "matches Opus" as marketing. The text-only V4 in the GLAD-AI-TOR arena shows a 57% crowd recommend rate from 3 votes. For context, Claude Opus 4.8 also sits at 57% from 3 votes, though Opus brings a richer feature set mid-conversation system messages, parallel subagent spawning in Claude Code, 69.2% on SWE-Bench Pro versus V4's agentic coding scores in the high 50s to low 60s . The peak pricing catch DeepSeek introduced time-of-day pricing with the V4 GA release in mid-July 2026, and V4-Flash-Vision-Exp inherits it. During peak hours 01:00 to 04:00 UTC and 06:00 to 10:00 UTC , all rates double: output jumps to $1.32/1M, cache-miss input to $0.44/1M. For teams running production agents, this means budgeting for worst-case peak rates or scheduling batch jobs around Beijing business hours. The off-peak window covers most US and European working hours, but global teams will hit the surcharge regularly. DeepSeek-V4 Open-weights 1.6T MoE with 1M-token context and near-frontier agentic coding at $0.87 per 1M output tokens Partner link. The crowd verdicts stay independent. What V4-Flash-Vision-Exp does well The model is engineered for multimodal agent workflows. According to DeepSeek's documentation, it excels at: UI automation : navigating complex software interfaces by identifying interactive elements visually Chart and data analysis : interpreting visual data with precision that rivals top-tier models Visual reasoning : executing tasks requiring both textual logic and visual context This is not a general-purpose image captioner paired with a text model. It is native multimodal processing, similar to how Claude and GPT-5.5 handle vision internally. The 384-token-per-image billing keeps costs predictable for document-heavy or chart-heavy workflows. Known limitations from V4 carry over The base V4 model has documented issues that likely persist in the vision variant: Malformed tool calls : function calls sometimes appear as plain text in the content field instead of the tool calls structure GitHub issue deepseek-ai 1244 Thinking mode breaks long agent chains : 400 errors in multi-turn tool-call sequences, with the fix still incomplete OpenClaw issue 72044 Hallucinated APIs : developers report fabricated endpoints in custom codebases and acting on imagined user input in agent loops Verbosity : V4 produces roughly 180M eval output tokens versus a 95M median, which erodes cost savings on latency-sensitive workloads The MIT-licensed open weights for V4-Pro and V4-Flash remain text-only. No announcement yet on when or whether vision weights will be released. Claude Opus 4.8 Anthropic's flagship Opus-tier model for long-horizon agentic coding; 1M context at $5/$25 per 1M tokens. Partner link. The crowd verdicts stay independent. The verdict - V4-Flash-Vision-Exp is the cheapest way to run a vision-capable frontier model: $0.66/1M output off-peak versus $25 Claude Opus 4.8 , $30 GPT-5.5 , or $12 Gemini 3 Pro - Benchmark scores trail Opus 4.8 by 5 to 12 points on most agent tasks, but lead on a few - Peak pricing doubles all rates during 7 hours of the day, so budget accordingly or schedule batch jobs off-peak - Existing V4 issues malformed tool calls, hallucinated APIs, verbosity likely carry over - Best fit: high-volume visual agent workloads where cost matters more than the last few benchmark points For a head-to-head breakdown, see the DeepSeek V4 vs Claude Opus 4.8 comparison /vs/deepseek-v4-vs-claude-opus-4-8 or explore the full LLM leaderboard /hall-of-fame/llm-models . Keep exploring Every claim above is backed by the arena's live data: crowd votes, verified pricing, honest pros & cons. More from the arena journal Industry News DeepSeek V4 vs Claude Opus 4.8: Can the $0.87 Model Compete in Agentic Coding? DeepSeek V4 costs 28.7x less than Claude Opus 4.8 per output token. Both score 57% crowd approval, but the real tradeoff lies in tool-call reliability vs raw price. Here is what the data shows. Jul 27, 2026 · 4 min read Industry News /blog/industry-news/suno-bmg-deal-ai-music-licensing Industry News Suno Signs Landmark BMG Deal: What It Means for AI Music Licensing Suno's global licensing partnership with BMG marks the second major label deal for the AI music generator. Here's what changes for creators, from download caps to Studio 2.0. Aug 24, 2026 · 3 min read Industry News /blog/industry-news/claude-sonnet-5-vs-opus-when-cheaper-wins Industry News Claude Sonnet 5 vs Opus 4.8: When the $2 Model Beats the $25 One Sonnet 5 costs 60% less than Opus 4.8 but matches it on knowledge work. Per-task cost analysis reveals when each model wins. Aug 19, 2026 · 4 min read