AI Benchmarks: When Is Enough Truly Enough?
A study examining partial evaluations on AI agent benchmarks including SWE-bench, AppWorld, and tau-bench finds that partial budgets are only valid when they replicate the full benchmark's final decis…
A study examining partial evaluations on AI agent benchmarks including SWE-bench, AppWorld, and tau-bench finds that partial budgets are only valid when they replicate the full benchmark's final decis…
A new approach using Quadratic Unconstrained Binary Optimization (QUBO) for evidence selection in retrieval-augmented generation (RAG) systems achieves competitive exact-match and token-F1 performance…
Researchers have developed a Bayesian model that explains how speakers understand unfamiliar languages through intercomprehension, outperforming larger zero-shot language models in aligning with human…
A study of 44 language models found that 41% of the time, AI models choose the word 'serendipity' when asked to pick any word, revealing a strong tendency toward conformity. In seven out of 31 categor…
Researchers have introduced a new pipeline that automates the evaluation of financial reports by generating rubrics without human experts, producing 14,450 candidate rubrics from 104 user queries. The…
The Roundtable Context Window Test (RCWT) reveals that large language models (LLMs) suffer a performance cliff when coordination tokens displace task instructions within a fixed 4096-token context win…
Researchers have developed a multi-feature fusion framework that combines static lexical representations (Word2Vec) with dynamic contextual representations (GPT) via a non-linear cross-attention mecha…
Large language models (LLMs) handle irrelevant context well in aggregate but remain fragile on specific examples, with random gibberish sometimes boosting performance and other times dragging it down,…
Researchers have developed WikiSTAR (Scientific Tracking of Article Revisions), an AI tool that uses an LLM classifier to analyze how scientific articles evolve on Wikipedia by tagging changes such as…
Researchers have proposed a new AI-based health misinformation detector designed for culturally and linguistically diverse (CALD) communities, using Bangla-translated health misinformation data to tes…
South Korea's government has issued a tender for private companies to develop a universal AI chatbot and a government services agent, backed by up to 256 Nvidia B200 GPUs for successful bidders who ma…
ASML, a heavyweight in semiconductor manufacturing, raised its forecast for the second time in 2023, betting on surging demand for AI chips as tech giants expand production capacities. The company's a…
A study testing large language models on 72 real breast cancer cases found that the top performer, Claude Opus 4.8 with a D&C+SA pipeline, achieved a global score of only 0.594 ± 0.025, revealing pers…
A new study comparing human semantic memory retrieval with that of large language models GPT-4o, Gemini 2.5-Pro, and Claude-Sonnet-4.5 found that humans exhibit higher entropy, larger semantic steps, …
AgentLens, a new open-source benchmark for interactive code agents, evaluates not just task completion but also instruction-following, verification, and error recovery by blending formal verification …
Researchers have introduced LOD-MSNO (LOD-Multiscale Neural Operator), a hybrid model that combines the Localized Orthogonal Decomposition method with neural operators to improve accuracy on multiscal…
A new method called Class-Contrastive Influence (C2I) quantifies synthetic sample usefulness by gradient-based influence on classifiers, outperforming traditional realism-focused approaches in few-sho…
A new evaluation technique called reversal-drop reveals that temporal video models often rely on positional encoding rather than understanding visual sequences, with Molmo2 failing to distinguish even…
Fin-Analyst, an AI trading agent, achieved a 13.51% return on Tesla (TSLA), outperforming the Buy-and-Hold strategy by 28.33 percentage points, according to its developers. The agent uses an eight-spe…
A new study improving the Huth et al. fMRI language decoding pipeline achieved a mean METEOR score of 0.149 and BLEU-1 score of 0.200, an 11% relative METEOR gain over baseline, but the fMRIFlamingo m…