DeepSeek v4 Pro Price
DeepSeek announced new peak/off-peak pricing for its DeepSeek-V4-Flash and DeepSeek-V4-Pro models, effective August 17, 2026, with off-peak input (cache miss) prices at 1.5 yuan and 4.5 yuan per milli…
DeepSeek announced new peak/off-peak pricing for its DeepSeek-V4-Flash and DeepSeek-V4-Pro models, effective August 17, 2026, with off-peak input (cache miss) prices at 1.5 yuan and 4.5 yuan per milli…
DeepSeek is raising prices for its flagship V4 models by more than four times, with new peak-hour pricing taking effect Aug. 16. The Hangzhou-based AI lab will charge $1.32 per 1 million output tokens…
DeepSeek announced the launch of DeepSeek-V4-Pro and V4-Flash, introducing flexible reasoning effort levels and native OpenAI Responses API support. The company also updated its API pricing with peak …
DeepSeek released DeepSeek V4 Pro 0813 on OpenRouter, following the open-weights releases of DeepSeek-V4-Pro in April and DeepSeek-V4-Flash-0731 in July, though open weights for the new version are no…
A new benchmark, Manager Coercion Bench (MCB), from Compassion in Machine Learning (CaML) finds that Anthropic's Claude models neither escalate to threats nor fabricate success, while all non-Anthropi…
Lasso Security reported on August 3 that changing only the agent harness altered outcomes across a 1,000-attack autonomous red-team test. The two harnesses averaged similar objective success rates, ab…
DeepSeek-V4-Pro defeated Phi-4-reasoning 12 tasks to 0 with an aggregate score of 105.5 to 33.0 and 100% confidence in a head-to-head benchmark of 12 fresh text tasks scored by gpt-5.4. The evaluation…
DeepSeek released DeepSeek-V4-Flash in public beta on 2026-07-31, showing significantly enhanced agent capabilities with benchmark results far exceeding V4-Pro-Preview, including Terminal Bench 2.1 at…
DeepSeek-V4-Pro defeated Llama-4-Maverick-17B-128E-Instruct-FP8 with an aggregate score of 101.3 to 86.2, winning 9 of 12 tasks with 98% confidence in a head-to-head benchmark conducted by an unnamed …
A new framework called DARWIN uses a genetic algorithm to evolve jailbreak prompts for large language models, achieving nearly 100% success rates on DeepSeek-V4-Pro and over 90% on GPT-5.5. The DARWIN…
Google's Gemini 3.6 Flash defeated DeepSeek-V4-Pro in a head-to-head text task evaluation, winning 5 tasks to 2 with 86% confidence and an overall score of 111.7 to 98.8, according to tests run by an …
DeepSeek-V4-Pro defeated Sarvam M 113.0 to 44.4 in a 12-task benchmark, achieving a 12-0 sweep with 100% confidence. The test, judged twice by gpt-5.4 to cancel position bias, found DeepSeek-V4-Pro re…
Gpt-oss-120b defeated DeepSeek-V4-Pro 106.3 to 93.7 across 12 text tasks, winning 10 of 12 matchups with 99% confidence, according to a head-to-head benchmark scored by gpt-5.4. The open-source model …
An independent test of DeepSeek-V4-Pro found sharply different behavior depending on whether prompts contained cyber or biology content, raising new questions about whether some requests were being an…
EvoClawBench, a new benchmark testing whether AI agents can learn reusable skills from experience, shows mixed results across 100 tasks and 502 sub-problems. Nanobot's GPT-5.4 model consistently excee…
Muse Spark 1.1 defeated DeepSeek-V4-Pro 11-1 in a head-to-head benchmark of 12 text tasks, scoring 110.8 to 84.3 with 100% confidence. The model outperformed in localization, proofreading, JSON extrac…
DeepReinforce released Ornith-1.0, an open-source family of coding models that learn to build their own task scaffolds, moving intelligence from the surrounding harness into the model itself. The mode…
Inferize announced DeepSeek-V4-Pro, claiming it can serve the model in 20 seconds with highly optimized, elastic AI inference. The company is building fast, efficient LLM serving that scales with dema…
DeepReinforce open-sourced Ornith-1.0, a family of self-improving coding models ranging from 9B to 397B parameters, which learn to generate their own task-specific scaffolds during reinforcement learn…
AMD shipped ATOM + ATOMesh, a ROCm-native LLM serving stack for Instinct GPUs that implements prefill/decode disaggregation, splitting the two inference phases onto separate GPU pools to optimize for …