Accelerating Business
In-house legal teams are increasingly adopting AI tools to streamline operations and reduce costs, according to a Financial Times Technology series. The report highlights how legal departments are usi…
In-house legal teams are increasingly adopting AI tools to streamline operations and reduce costs, according to a Financial Times Technology series. The report highlights how legal departments are usi…
Bond markets are increasingly dominated by a bet on the same artificial intelligence thesis as equities, according to the Financial Times, raising concerns about concentration risk across asset classe…
CuspAI co-founder Max Welling said the UK-based science start-up is using AI to design molecules that remove 'forever chemicals' from water, aiming to address complex environmental challenges.…
ByteDance, the Chinese company behind TikTok, is pouring resources into artificial intelligence in a bid to dominate the technology, a move some analysts consider a big gamble, according to the Financ…
A new study from arXiv introduces OptimismBench, a benchmark that detects directional bias in large language models' probability judgments by using inverted pairs to measure asymmetry between P(succes…
A study using the GPT-4-Turbo model and four other LLMs found that AI-generated stories about people with intellectual disabilities (ID) contain implicit biases, depicting them as younger, more depend…
A controlled study from arXiv (2607.26541v1) finds that varying speech delivery presets while holding transcript content fixed can jailbreak audio LLMs, with the Q=1 Panic preset achieving 38/95 succe…
A systematic analysis of language sensitivity in codec-based self-supervised learning (SSL) models shows that downstream performance is insensitive to the neural audio codec (NAC) training language bu…
A new study from arXiv argues that large language models (LLMs) converge with human cognition on five key principles: inferential organization, computational architecture, representational structure, …
A new training-free inference method called V-Steer restores instruction hierarchy compliance in large language models by editing cached value vectors, raising primary constraint accuracy from under 1…
Researchers introduced AgentGUI, a locally hosted graphical interface for observing and steering long-running AI agents across multiple concurrent sessions, achieving a 38% faster identification of ke…
Mercor, in partnership with Ramp, introduced APEX-Accounting, a benchmark to assess whether frontier models can perform real accounting tasks such as reconciling accounts and posting transactions. Acr…
A study using BERT, a transformer-based AI model, achieved 97.05% accuracy in multilabel classification of 14,590 Mpox research articles into topics such as outbreaks, vaccination, and epidemiology, a…
A new benchmark study from arXiv preprint 2607.26348 finds that large language models (LLMs) fail to replicate real human survey responses across two independent domains—U.S. general social attitudes …
A new study from arXiv finds that fine-tuned compact encoders match or exceed large language models in classifying fine-grained inconsistencies in financial disclosure text. Using a 5,940-instance sna…
Researchers introduced CMT-RAG, a complementary memory framework for multi-turn multi-hop retrieval-augmented generation that aligns conversational memory with retrieval by representing dialogue conte…
Researchers propose ForgetBench, a benchmark to systematically characterize forgetting behavior in large language models (LLMs) under continual knowledge editing. The benchmark introduces concept-base…
A new study from arXiv reveals that filesystem-based memory for LLM agents—where agents store long-term memory as a directory tree of markdown files—can roughly halve retrieval costs when organized, b…
A new analysis of lossy verification in speculative decoding reveals that truncation-based methods can degrade generation quality due to distributional distortion, while collaborative verification req…
A new diagnostic benchmark, EC-Reason-Bench, reveals that general large language models (LLMs) score near zero on complete enzyme classification (EC) number prediction, with accuracy dropping sharply …