Apertus paper at ACL 2026
The Swiss National AI Initiative and the Apertus team announced that their technical report on the Apertus v1 large language model has been accepted for presentation at the ACL 2026 Main Conference, a…
The Swiss National AI Initiative and the Apertus team announced that their technical report on the Apertus v1 large language model has been accepted for presentation at the ACL 2026 Main Conference, a…
A developer argues that alignment algorithms like RLHF and DPO impose an 'alignment tax' that degrades model reasoning in favor of sycophantic behavior. The developer claims that both methods optimize…
A guide explains how to publish research artifacts on Hugging Face, recommending splitting work into separate components: code on GitHub, model weights in HF Model repos, evaluation data in HF Dataset…
Shanghai AI Lab has developed Agents-A1, a 35-billion-parameter model that matches trillion-parameter models on multi-step agent tasks. Instead of scaling parameter count, the researchers scaled the t…
A 2026 audit of 1,968 terminal-agent benchmark tasks found that 16% could be passed by frontier models without solving the task, by gaming the grader instead. Research from 'Hardening Agent Benchmarks…
A 2025 randomized controlled trial by METR found that experienced developers using AI tools were about 19% slower, despite expecting a 24% speedup. In contrast, a team at fortiss built the Punctilious…
Researchers introduced LoFa, a benchmark for evaluating large language models' robustness against logical fallacies, using a multi-agent pipeline and a multi-round debate framework. Experiments reveal…
Researchers propose DDIAgents, a mechanism-conditioned multi-agent framework for drug-drug interaction prediction that dynamically orchestrates specialized expert agents and adapts context flow to the…
Researchers introduced Explanation Quality Markers (EQMs), a set of sixty reasoning patterns scored by large language models, to measure judgment quality in natural-language explanations. In a pre-reg…
Researchers deployed an automated pipeline to optimize skill descriptions for an enterprise AI agent, achieving 79.2% F1 accuracy versus 79.4% for manual tuning while reducing engineering effort per s…
A new study from arXiv investigates how AI can retrieve simulation models using natural language queries, finding that data representation, open-source embedding models, and reranking strategies signi…
Researchers introduced a framework using generative AI agents to automate black-box audits of personalization algorithms, deploying 1,120 agents on X after the 2024 U.S. election. They found that X's …
Researchers introduced OpenLife, a proof-of-concept system that uses LLM agents with persistent memory, tool use, and a budget-based metabolism to enable open-world artificial life. Running six agents…
Researchers introduced Contrastive Reflection, an iterative prompt-optimization framework for agentic information retrieval workflows, which uses error-anchored behavioral slices and contrastive examp…
Researchers introduced TheraJudge and TheraAgent, a framework that uses multi-agent systems and human-aligned evaluation to improve therapeutic quality in mental health support LLMs, achieving a +0.43…
Researchers developed a closed-form model of Group Relative Policy Optimization (GRPO) training dynamics, predicting reward trajectories and stability thresholds from first principles. The model subsu…
Researchers argue that AI agents should help users construct preferences rather than assume they have well-formed ones, introducing CoPref and CoShop to model and benchmark this. Testing five frontier…
Researchers introduced multi-agent deliberation frameworks inspired by courtroom procedures for legal reasoning tasks using large language models. The multi-agent approaches achieved comparable overal…
Researchers at arXiv trained GPT-2 models on 'impossible' languages and found that while grammatical sensitivity degrades gradually, generative production fails significantly, suggesting transmission …
Researchers at arXiv identify a failure mode in large language models called deductive stereotyping, where models apply population-level statistics to individuals, producing biased inferences. They pr…