Run DeepSeek V4 Flash locally or in the cloud (US-hosted) LM Studio has released DeepSeek V4 Flash 0731, a 284B-parameter Mixture-of-Experts model, available for local download or US-hosted cloud inference with Zero Data Retention enabled by default. According to DeepSeek, V4 Flash outperforms GLM 5.2 on all reported benchmarks while using about 62% fewer parameters, making it a cost-efficient option for agentic coding and tool-use tasks. The model is released under the MIT License and requires at least 156 GB of system memory for local use. Run DeepSeek V4 Flash locally or in the cloud US-hosted DeepSeek V4 Flash is now available in LM Studio Bionic. Run it in the cloud /settings/cloud for agentic work, or download the model https://lmstudio.ai/models/deepseek/deepseek-v4-flash to run locally in LM Studio. DeepSeek V4 Flash cloud inference in LM Studio is hosted on US-based servers, with Zero Data Retention ZDR enabled by default. Leap in capabilities for the size class, and for the cost DeepSeek reports that V4 Flash outperforms GLM 5.2 on every benchmark below where both models have a score. It does so with 284B total parameters versus GLM 5.2's 753B—about 62% fewer parameters—bringing capable coding and tool-use performance at a remarkably low inference cost. Built for agentic work DeepSeek V4 Flash 0731 https://lmstudio.ai/models/deepseek/deepseek-v4-flash is the official release of DeepSeek V4 Flash, replacing the preview with substantially stronger agentic capabilities. This makes it a strong fit for Bionic workflows that require sustained tool use: navigating a repository, implementing changes across files, debugging, or working through complex research and document tasks. Benchmark results The following results are reported by DeepSeek in the official model card: | Benchmark | DeepSeek V4 Flash 0731 | V4 Flash Preview | V4 Pro Preview | GLM-5.2 | Opus-4.8 | |---|---|---|---|---|---| | Terminal Bench 2.1 | 82.7 | 61.8 | 72.1 | 81.0 | 85.0 | | NL2Repo | 54.2 | 39.4 | 38.5 | 48.9 | 69.7 | | Cybergym | 76.7 | 38.7 | 52.7 | — | 83.1 | | DeepSWE | 54.4 | 7.3 | 12.8 | 46.2 | 58.0 | | Toolathlon-Verified | 70.3 | 49.7 | 55.9 | 59.9 | 76.2 | | Agents' Last Exam | 25.2 | 15.8 | 16.5 | 23.8 | 25.7 | | AutomationBench Public | 25.1 | 10.8 | 12.8 | 12.9 | 27.2 | | DSBench-FullStack † | 68.7 | 37.0 | 41.8 | 61.8 | 71.6 | | DSBench-Hard † | 59.6 | 25.8 | 31.1 | 54.5 | 71.7 | For public code-agent tasks, DeepSeek evaluated V4 Flash 0731 using the minimal mode of DeepSeek Harness, max reasoning effort, temperature = 1.0 , and top p = 0.95 . † DSBench-FullStack and DSBench-Hard are internal DeepSeek test sets. See the official model card https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 for methodology and details. Run it locally or in the cloud DeepSeek V4 Flash is a 284B-parameter Mixture-of-Experts model. You can download it https://lmstudio.ai/models/deepseek/deepseek-v4-flash and run it on your own hardware in LM Studio, but plan for at least 156 GB of system memory. The model and its weights are released under the MIT License. If you would rather not manage that hardware, select the cloud model in Bionic. It delivers highly capable coding and agentic performance while remaining incredibly cost-efficient. LM Studio hosts it on US-based infrastructure, with Zero Data Retention enabled by default. Get started Download LM Studio Bionic https://lmstudio.ai , create an LM Studio account, and select DeepSeek V4 Flash from the cloud model picker. Or open the DeepSeek V4 Flash model page https://lmstudio.ai/models/deepseek/deepseek-v4-flash to download it for local use.