The Agent Stack
Vercel launched The Agent Stack, a suite of tools including AI SDK, AI Gateway, Workflow SDK, and Vercel Sandbox, to help developers build production-grade AI agents without vendor lock-in. The stack provides unified mod…
Large language model (LLM) news — GPT-4, Claude, Gemini, Llama, Mistral and the latest research on training, fine-tuning, RLHF, and deployment of LLMs.
Vercel launched The Agent Stack, a suite of tools including AI SDK, AI Gateway, Workflow SDK, and Vercel Sandbox, to help developers build production-grade AI agents without vendor lock-in. The stack provides unified mod…
Emily Dalton Smith, Meta's head of product for its internal AI for Work effort, is leaving the company after 11 years, according to an internal announcement seen by Reuters. Her departure comes two months after she was n…
A developer built an autonomous incident investigation agent called FRIDAY that reduced mean time to resolution by 65% on a platform serving 30+ million end users. The agent, triggered by PagerDuty alerts, first checks G…
Forbes launched an AI-powered daily audio briefing, The Daily Brief, on May 29, 2025, which uses its internal AI tool Bertie to summarize and voice the top three stories of the day. The product keeps a human in the loop …
EPFL researchers released MeditronFO, the first fully open framework for building medical large language models, making every stage of development publicly available to ensure transparency and accountability in AI health…
On April 16, 2026, NAVI-Orbital became the first system to demonstrate a zero-shot vision-language model performing autonomous multi-modal inference entirely onboard a Low Earth Orbit spacecraft. The system uses Gemma 3 …
Researchers introduced CaVe-VLM-CoT, a modular reflection-based agentic-RAG framework that enforces evidence-grounded reasoning in vision-language models through a five-stage closed-loop pipeline. The framework achieves …
Researchers introduced CEO-Bench, a benchmark that evaluates language model agents on long-horizon tasks by simulating operating a startup for 500 days. The strongest agents, including Claude Opus 4.8 and GPT-5.5, strugg…
Researchers introduced DeFAb, a benchmark for defeasible abduction in foundation models, converting knowledge bases into 372,648+ logically verifiable instances. Frontier language models achieved at most 65% accuracy, dr…
Researchers from DiDi introduced ProfiLLM, an agentic LLM data pipeline that generates utility-aligned user profiles for industrial ride-hailing dispatch. Deployed on DiDi's production system, ProfiLLM achieved up to +6.…
Researchers introduced SciRisk-Bench, a benchmark evaluating AI safety in scientific contexts across 7 disciplines and 10 risk dimensions. Tests on mainstream and science-oriented LLMs revealed where models remain unsafe…
Researchers introduced Decoupled Search Grounding (DSG), a vendor-agnostic architecture that separates search from reasoning in LLM agents, enabling independent control over retrieval policy, provider routing, and cachin…
Researchers propose ThinkDeception, a novel framework that uses multimodal large language models and chain-of-thought reasoning to detect deception with interpretability. It introduces a progressive training strategy and…
Researchers introduced ARIADNE, a training-free routing framework for dynamic adapter selection in parameter-efficient fine-tuning ecosystems. It selects task-specialized adapters at inference time by measuring input pro…
Researchers propose a fully local AI cascade for de-identifying educational dialogue that achieves 0.958 macro F1 on math tutoring transcripts, outperforming a commercial API (0.706) and LLM-only baselines (0.767), while…
Researchers introduced SproutRAG, a hierarchical retrieval-augmented generation framework that uses learned inter-sentence attention to organize sentence-level chunks into a binary tree for multi-granularity retrieval wi…
Researchers propose activation steering as an alternative to few-shot prompting for generating synthetic data in low-resource languages. The method improves data diversity and downstream model performance by steering lan…
Researchers from Hao AI Lab introduced JetFlow, a speculative decoding framework that breaks the scaling ceiling of autoregressive LLMs by combining one-forward drafting efficiency with branch-wise causal conditioning. J…
A new study from arXiv reveals that large language models use partially shared parameters for mathematical reasoning across languages, with the strongest overlap in intermediate layers. English yields the largest set of …
Researchers at an undisclosed institution have found that large language models (LLMs) poorly preserve diagnostic uncertainty in clinical text, altering phrases like "possible pneumonia" more than half the time, which co…