cd/entity/DeepSeek V4 Flash· home› entities› DeepSeek V4 Flash
grep -l @deepseek v4 flash /news/*.json | wc -l → 201

DeepSeek V4 Flash

mentions 201 type Person page 2/11 feed RSS

// recent coverage 201 mentions

08:56
2026-09-10
geopolitechs.org
large-language-models

DeepSeek V4.1 Flash: Stronger, Faster, More Accessible

DeepSeek released DeepSeek V4.1 Flash, a 552B-parameter MoE model built on a new Causal-Encoder-Decoder architecture with 8B activated parameters on the input side and 16B on the output side, availabl…

00:00
2026-09-10
mindstudio.ai
artificial-intelligence

DeepSeek V4.1 Flash Specs: KV Cache Compression Explained

DeepSeek AI's model card for DeepSeek V4.1 Flash reports a global KV cache footprint of roughly 890 bytes per token, about a quarter of the 4x-larger footprint of predecessor DeepSeek V4 Flash and a 4…

17:24
2026-09-09
gist.github.com
large-language-models

Qwen 3.8 follows GPT-5.5 Pro reasoning prefills

A developer's v1.1 experiment reran reasoning-prefill tests using GPT-5.5 Pro as the teacher across 45 problems, finding that Qwen3.8 A95B's unigram source recall jumped from 33.92% to 54.50% (+20.58 …

14:23
2026-09-08
blog.sshh.io
artificial-intelligence

I Asked 100 Agents to Hack Me

In a five-hour experiment, ~100 self-hosted abliterated AI agents compromised 3 accounts via software vulnerabilities, 2 via password brute forcing, made 16 social engineering attempts, and found sens…

00:00
2026-09-08
mindstudio.ai
large-language-models

Tencent Hunyuan Hi-4 Preview: Specs, Benchmarks, and Pricing

Tencent released Hunyuan Hi-4 Preview, a 770B-parameter mixture-of-experts model with 49B active parameters and a 1M-token context window, under the Apache 2.0 license. The model scores 92.3 on GPQA D…

20:46
2026-09-07
inference.academy
large-language-models

DeepSeek V4 Flash across 14 providers: cost, speed and caching

DeepSeek V4 Flash serving measurements across 14 providers show that a repeated prompt can reduce 100k-input, 100-output-budget request costs by 14.9× compared to cold requests, with Telnyx leading me…

00:00
2026-09-03
wagtail.org
artificial-intelligence

Open models only: back to school challenge

Z.ai's GLM 5.3 Flash open-weight model is the focus of a September challenge urging developers to abandon AI coding subscriptions and use only open-weight models, citing benchmark scores at the top of…

13:00
2026-09-02
vettedconsumer.com
large-language-models

How Much RAM Do You Need to Run a Local LLM in 2026?

Running large Mixture-of-Experts local LLMs in 2026 requires at least 64GB of system RAM, ideally 128GB or more, because sparse MoE models keep rarely-used expert weights in system memory rather than …

11:13
2026-08-30
1endpoint.dev
ai-infrastructure

Show HN: 1endpoint – Cheaper access to AI models

1endpoint, a new API gateway, offers access to multiple AI models through a single endpoint with usage-based pricing, starting at $0.0420 per 1M input tokens for GLM 5.2. The service includes prompt c…

← prev page 2 / 11 next →
// co-occurs with top 8 entities
// topics top 6 topics