ls /news/large-language-models · home newslarge-language-models
grep -r --recent /news/large-language-models | head -20

Large Language Model News

Large language model (LLM) news — GPT-4, Claude, Gemini, Llama, Mistral and the latest research on training, fine-tuning, RLHF, and deployment of LLMs.

20677 articles page 726 of 1034 0 sources 30 min sync cycle updated 2026-06-16

// latest articles 20677 indexed

04:00
2026-06-16
arxiv.org
large-language-models · 1m read · neu

Equity with Efficiency: An Empirical Study of Tokenizers for Multilingual Large Language Models

A new empirical study systematically compares tokenizers for multilingual large language models across 11 Southeast Asian languages, finding that Parity-aware BPE achieves the best balance of compression efficiency and c…

04:00
2026-06-16
arxiv.org
machine-learning · 1m read ↑ pos

FastMix: Fast Data Mixture Optimization via Gradient Descent

Researchers introduced FASTMIX, a framework that automates data mixture optimization for training large models by reformulating mixture selection as a bilevel optimization problem and using gradient descent to jointly op…

04:00
2026-06-16
arxiv.org
machine-learning · 1m read ↑ pos

Transformers Learn the Mestre-Nagao Heuristic

Researchers trained a two-layer transformer encoder to classify rational elliptic curves by rank, achieving over 99% accuracy. The model learned the Mestre-Nagao heuristic from data alone, with input weights matching the…

04:00
2026-06-16
arxiv.org
machine-learning · 1m read ↑ pos

Temporal Difference Learning for Diffusion Models

Researchers introduced a temporal difference (TD) objective for diffusion models that enforces cross-time consistency along the denoising trajectory, improving sample quality especially with few sampling steps. The metho…

← prev page 726 / 1034 next →
LIVE [news/large-language] indexed:20677 page:726/1034 en · ua 2026-05-20 ·