ls /news/large-language-models · home › news›large-language-models
grep -r --recent /news/large-language-models | head -20

Large Language Model News

Large language model (LLM) news — GPT-4, Claude, Gemini, Llama, Mistral and the latest research on training, fine-tuning, RLHF, and deployment of LLMs.

31195 articles page 244 of 1560 0 sources 30 min sync cycle updated 2026-09-12

// latest articles 31195 indexed

05:56
2026-09-12
latent.space
artificial-intelligence · ↑ pos

[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale

DeepSeek released DeepSeek v4.1-Flash, a 763B-parameter model with a novel causal encoder-decoder architecture that splits 8B parameters to prefill and 16B to decode, per the model's tech report on Hugging Face. The rele…

← prev page 244 / 1560 next →
LIVE [news/large-language] indexed:31195 page:244/1560 en · ua 2026-05-20 · —