Real-Time AI Models Face a Hidden Collapse Risk
Real-time AI interaction models like Moshi and Qwen-Omni face a hidden collapse risk due to swelling KV cache, causing sudden performance drops under sustained load. A simple fix—bounding each session…
Real-time AI interaction models like Moshi and Qwen-Omni face a hidden collapse risk due to swelling KV cache, causing sudden performance drops under sustained load. A simple fix—bounding each session…
Researchers propose a novel method for selecting diverse and informative subsets of synthetic images, reducing redundancy and improving AI model performance without retraining generators. The approach…
OpenClaw's GLM-5 inference optimization study found that adjusting parameters like chunked prefill size and request concurrency improved throughput and reduced latency, cutting serving costs by 10.4% …
Researchers introduced Agentic SABRE, a neuro-symbolic multi-agent framework for ransomware detection that combines semantic and behavioral analysis. Tested on RDset and RanSMAP, it achieved perfect d…
Researchers introduced SALT, a deterministic benchmark for evaluating AI uncertainty estimation in long-form text, eliminating reliance on fallible labels. Analysis of over 50 LLMs revealed that curre…
Oyster-II, a new reinforcement learning-based framework for large language model safety, outperforms its predecessor and competitors like Qwen3-14B while balancing safety and utility. The model addres…
Researchers introduced the Agent Step Value (ASV) framework, which evaluates AI agents by scoring each action's impact on task outcomes, offering granular insights into decision-making. In tests on 10…
VideoAgent, a new AI framework for long-form video editing, automates shot creation and uses multi-agent orchestration to generate coherent narratives, achieving an 87-95% orchestration success rate a…
Researchers introduced ELBO-T2IAlign, a plug-and-play method that improves text-image alignment in diffusion models without retraining. The approach uses zero-shot referring image segmentation to enha…
Researchers have developed ASCEND, a constraint-based causal discovery framework that leverages hierarchical structures to infer gene regulatory networks from high-dimensional genomic data. The method…
Researchers introduced Agent Step Value (ASV), a framework that evaluates each action in an AI agent's sequence by scoring its impact on outcome distributions, tested on 100 open-QA tasks with PubMed …
Researchers developed CIPHER, a new AI framework that reduces bias in medical diagnoses by systematically addressing four pathways through which sensitive attributes influence image content. Tested on…
VideoAgent, a new AI-driven video editing framework, achieves orchestration success rates of 87-95% and reduces API costs by 60%, outperforming existing multimodal LLMs on the VideoEdit benchmark. The…
Researchers have discovered a unique signature in AI-generated text called the 'Vestigial Heuristic,' which stems from large language models' aversion to repeating words. They developed a new metric, …
A new benchmarking framework for Graph Explainable AI (G-XAI) evaluates Graph Neural Networks (GNNs) using tabular explainability metrics that separate graph topology and node features, identifying no…
Researchers introduced CIPHER, a framework that reduces bias in medical AI by addressing multiple causal pathways linking sensitive attributes to image features. In tests on chest X-ray and dermoscopy…
Researchers have identified a 'Vestigial Heuristic' in large language models, a bias against token repetition that leaves a unique signature in AI-generated text. A new metric, Telescope Perplexity, l…
Researchers introduced KARMA, a new AI model that uses Slot-Parallel Alignment to focus on detailed slots rather than broad templates, outperforming existing models in biomedical, computer science, an…
New research finds that large language models (LLMs) tend to defect in social dilemmas like the prisoner's dilemma, even with advanced reasoning. Game-theoretic mechanisms such as conditional contract…
A new benchmark called Incompressible Knowledge Probes (IKPs) evaluates AI models' factual recall accuracy across 201 models from 27 vendors, revealing a strong correlation between parameter count and…