cd/entity/arXiv· home› entities› arXiv
grep -l @arxiv /news/*.json | wc -l → 3733

arXiv

mentions 3733 type Organization page 117/187 feed RSS

// recent coverage 3733 mentions

04:00
2026-08-17
arxiv.org
artificial-intelligence

AI Evaluation Should Work With Humans

A position paper submitted to arXiv on 6 Jul 2026 argues that the dominant paradigm of AI evaluation, which focuses on superhuman autonomous performance, is guiding AI development in the wrong directi…

20:30
2026-08-16
discuss.huggingface.co
robotics

Independent researchers

Independent researcher Chance, founder of CharacterOS LLC, is seeking an arXiv endorser in cs.RO or cs.AI for a paper on social robot character safety certification under the RoboSafe standard, which …

20:01
2026-08-16
cst.cam.ac.uk
artificial-intelligence

Red queen hypothesis – a new way forward for self-improving AI

Researchers from the University of Cambridge, NVIDIA, and Flower Labs have developed the Red Queen Gödel Machine, a method for recursive self-improving AI agents that co-evolves the agent and its eval…

18:16
2026-08-16
math.columbia.edu
artificial-intelligence

HEP-TH and AI

The rate of hep-th submissions to arXiv is now about one-third higher this year than in previous years, according to data gathered by Peter Woit, a mathematician at Columbia University. Woit notes tha…

07:04
2026-08-16
arxiv.org
artificial-intelligence

Is this the end of human code review?

A case study reports that an AI coding agent successfully dismantled a core architectural invariant across 189 files in a 717,725-line TypeScript codebase with no human code review and no test oracle,…

14:43
2026-08-15
arxiv.org
artificial-intelligence

Paper on Architecture for AI-Assisted Software Development

Researchers have proposed the Spec Growth Engine, a framework for AI-assisted software development that addresses context explosion and silent spec-code drift through a machine-readable spec graph, a …

13:51
2026-08-15
arxiv.org
artificial-intelligence

QuoteBench: Matched Scores Can Hide Command-Path Failures

A new benchmark, QuoteBench, reveals that matched execution scores for LLM coding agents can hide command-path failures, with replaying the same reply through an added parser lowering success by 55.4 …

12:31
2026-08-15
promptcube3.com
artificial-intelligence

BDH-CQ hits 29.5% on ARC-AGI-1 with only 150M parameters

BDH-CQ, a 150M-parameter model, achieved a 29.5% pass@2 score on the ARC-AGI-1 benchmark at a cost of $0.00070 per task, demonstrating that recurrent latent state reasoning can outperform larger model…

11:15
2026-08-15
arxiv.org
artificial-intelligence

Agentic Reasoning for Large Language Models

A new survey on arXiv organizes agentic reasoning for large language models into three layers: foundational, self-evolving, and collective multi-agent reasoning, distinguishing in-context from post-tr…

06:07
2026-08-15
arxiv.org
artificial-intelligence

The Embedder's Dilemma: LLMs Are Better, but at What Cost?

A new study comparing ten LLMs and 26 embedding models across 37 tasks finds the best LLM (Gemini 3.1 Pro, 77.6) and the best embedding model (77.2) are effectively tied, but LLMs cost up to 1,431x mo…

← prev page 117 / 187 next →
// co-occurs with top 8 entities
// topics top 6 topics