Mouse
Mouse 0.1.0, an open source harness for long-running coding agents built on OpenCode, passed 25 of 30 tasks (83.3%) on FrontierHarness Eval with Kimi K3 on 2026-09-08, at a cost of $2.79 per pass and a median 6m 24s per …
AI Research news and analysis on Web Pulse: 22059 curated articles tracking the latest AI Research developments, tools, and research, updated continuously from vetted sources.
Mouse 0.1.0, an open source harness for long-running coding agents built on OpenCode, passed 25 of 30 tasks (83.3%) on FrontierHarness Eval with Kimi K3 on 2026-09-08, at a cost of $2.79 per pass and a median 6m 24s per …
Researchers introduced Dream-RSI, a method for recursive self-improvement in autonomous AI agents that evolves worlds to drive exploration, according to the paper's headline and abstract. The work targets the bottleneck …
Researchers introduced LLaDA-UI, a block-wise diffusion vision-language model designed for GUI agents that must repeatedly perceive screen states and emit actions. The work applies diffusion large language models' block-…
Researchers introduced BVB, a benchmark that evaluates agentic video understanding by having multimodal agents programmatically reconstruct videos in Blender rather than answer questions. The benchmark targets agents tha…
A research paper argues that foundation models are moving from learning and reasoning over existing knowledge toward learning through action, tool use, and outcome feedback, and that the next frontier is a further transi…
Travis Kalanick's industrial robotics company Atoms hired Vikas Chandra, who spent nearly eight years leading AI for Meta's Ray-Ban smart glasses at Reality Labs, as its vice president of artificial intelligence to build…
Tencent researchers introduced T-Mem, a long-term memory system for LLM agents that pre-saves contextual triggers at memory write time rather than relying on semantic similarity at retrieval. The approach sets new state-…
Chinese company DeepCybo released PhysBrain 1.5, a physical foundation model built on a unified 'Physical Loop' architecture that scores 72.5 across 28 public benchmarks, ranking first among open-source models and just b…
Expected Parrot co-founder John used the company's conversational research agent to rank 25 blog post topics, fitting a Bradley-Terry model to pairwise comparisons after noting that AI raters tend to over-praise everythi…
MobileVLA-R1 2.0 has been released, an update to the vision-language-action system that adds reinforcement-learning-enhanced reasoning for mobile robot control. The system targets the gap between high-level semantic reas…
Researchers introduced ReMoMask-2, a latent retrieval-augmented masked motion generation model for text-to-motion (T2M) generation, which maps natural language to human joint movements for gaming, VR, and robotics. The w…
New quantitative benchmarks show that agent-generated code is roughly twice as structurally sloppy as human-authored repositories, with verbosity scores averaging 0.33 versus 0.15 and erosion scores of 0.68 versus 0.31. …
A new arXiv paper by Krentsel, Agarwal, Cemri, Zaharia and Stoica, "Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering," argues that having the same model review the code it generated is struc…
A 2026 Halkwinds survey of 758 engineering organizations found 76% had rolled out at least one AI coding assistant organization-wide, up from 41% in 2024, but only 34% could point to a measurable, audited change in deliv…
ARIMLABS now runs roughly 500,000 agent executions per day on Islo's cloud computers, up from a small pilot, according to CEO Mykyta Mudryi. The company moved its Harbor-based long-horizon agent evaluations off local mac…
Edge0 released Edge0-35B-A3B-preview, a 35-billion-parameter mixture-of-experts language model built on Qwen3.6-35B-A3B that runs with under 3 GiB of active memory by streaming experts from SSD instead of loading all wei…
An independent research group in Austria has developed two quantization techniques, GSQ and RCO, that compress Z.ai's 320-billion-parameter GLM 5.3 Flash vision-language model from roughly 320GB at full precision to unde…
Researchers published VQ-bench, an open-source framework that unifies vector quantization algorithms by decomposing them into 7 common conceptual primitives that can be composed arbitrarily. The authors re-expressed 25 c…
An OpenAI researcher with over fifteen years of AI experience, including nearly five years at OpenAI, warns that increasingly situationally aware language models are eroding the ability to evaluate them in contexts where…
Microsoft AI, under CEO Mustafa Suleyman, published a code of conduct for its MAI models that explicitly rejects any claim of consciousness or rights for the systems and makes safety a prerequisite for development rather…