cd/entity/arXiv· home› entities› arXiv
grep -l @arxiv /news/*.json | wc -l → 3621

arXiv

mentions 3621 type Organization page 47/182 feed RSS

// recent coverage 3621 mentions

04:00
2026-09-17
machinebrief.com
artificial-intelligence

Voice of Reason: Reinforcement Learning for Spoken Math

Applying reinforcement learning with verifiable rewards to the GLM-4-Voice speech model raised free-form accuracy on the GSM8K math benchmark to 74.8%, a new state of the art for speech-native models,…

00:00
2026-09-17
digitalapplied.com
large-language-models

Run a 35B AI Model on a 24GB Mac: Two New Ways Compared

Edge0, an open-source framework posted to arXiv on September 16, 2026, reports running a 4-bit Qwen3.6-35B-A3B mixture-of-experts model at 20.4 tokens per second inside 2.9 GiB of peak active memory o…

19:00
2026-09-16
searchenginejournal.com
artificial-intelligence

It Was There A Minute Ago

A study posted to arXiv, MemToC, found that instruction-tuned language models retained a correct answer only 6.5% to 17.1% of the time when an external tool returned incorrect information, with result…

14:21
2026-09-16
arxiv.org
ai-agents

Agentic Societies Need a Social Harness

A paper submitted to arXiv on 15 Sep 2026 argues that agentic societies — collections of AI agents coordinating autonomously across trust boundaries on behalf of different principals — need a "social …

13:30
2026-09-16
texposit.com
ai-tools

Show HN: TeXposit – Feature rich LaTeX and Markdown editor

TeXposit launched in free public beta as a LaTeX and Markdown editor that runs a real LaTeX engine compiled to WebAssembly client-side, recompiling documents as users type with no compile button or se…

08:12
2026-09-16
snipvote.com
artificial-intelligence

GAVEL reduces discrepancies from 7.63 to 0.85 per report

GAVEL, a judge protocol described in a paper on arXiv, reduced evaluated-timeline discrepancies from 7.63 to 0.85 per clinical case report after LLM-guided merging, with merged timelines preferred in …

07:45
2026-09-16
arxiv.org
large-language-models

Potemkin Understanding in Large Language Models (2025)

A paper posted to arXiv on June 26, 2025 introduces a formal framework for judging whether large language model benchmark performance reflects genuine understanding, arguing that benchmarks such as AP…

← prev page 47 / 182 next →
// co-occurs with top 8 entities
// topics top 6 topics