Agents love prefill
DeepSeek's V4.1 Flash technical report introduces a "Causal Encoder–Decoder" architecture that runs only the first 20 of the model's 40 layers during prefill, cutting prefill compute in half for its 552B-parameter MoE mo…
AI Research news and analysis on Web Pulse: 22644 curated articles tracking the latest AI Research developments, tools, and research, updated continuously from vetted sources.
DeepSeek's V4.1 Flash technical report introduces a "Causal Encoder–Decoder" architecture that runs only the first 20 of the model's 40 layers during prefill, cutting prefill compute in half for its 552B-parameter MoE mo…
A 33-organization Korean consortium led by Naver Cloud is developing K-MYTHOS, a security-domain AI foundation model that fuses threat intelligence from malware, dark web, network, and cloud sources into a dual-model att…
A developer measured the "exchange rate" between training data and document length for a small transformer versus a zero-parameter count-table cache, finding that 16x more training data (from 500K to 8M tokens) moved the…
Bruce Schneier argues that AI models remain far less capable than experienced academic mathematicians, despite recent results such as OpenAI's frontier model disproving the 80-year-old unit distance conjecture in mid-May…
A from-scratch autoregressive transformer trained in 1.5 hours on a single Nvidia RTX 5090 GPU reached 44% on the ARC-1 benchmark, matching specialized architectures TRM and HRM, according to a writeup covered by Mango D…
A research note argues that the AI release slowdown debate conflates release cadence with training cadence, contending that coding, math and cyber are being remade where RL data and feedback loops exist while most other …
A report titled "The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation" was submitted to arXiv on 20 Feb 2018 and last revised 1 Dec 2024, surveying potential security threats from malicio…
A developer has reverse-engineered Anthropic's Claude tokenizer, finding it is not a standard byte-level BPE but instead uses a minimum-piece tokenization scheme similar to MinGram or PathPiece. The reconstruction reveal…
An opportunistic AI system detected colorectal cancer from routine, noncontrast CT scans, according to a report published by Radiology Business. The tool works on imaging already collected for other purposes, meaning no …
A machine-learning practitioner built a horse-racing ranking model trained on 1.18 million runners and evaluated it with walk-forward validation, finding that the betting market remains a very strong baseline. The projec…
IEEE Spectrum published a multi-part explainer examining when humanoid robots will reach homes, tracing the field from Honda's Asimo and iRobot's Roomba through modern machines such as Tesla's Optimus, Boston Dynamics' A…
A developer outlined four common AI agent architecture patterns — ReAct, SOP, Reflection, and Multi-Agent — in a deep dive on modern agent design. The writeup explains how each pattern handles reasoning, reliability, sel…
A developer outlined best practices for evaluating AI models, emphasizing a multi-layered framework combining standardized benchmarks like MMLU and HumanEval, adversarial red teaming for prompt injection and jailbreaks, …
Specific Labs released Real-SWE, an enterprise-code SWE benchmark whose leaderboard shows that the same CLI harness can produce more than a 2x score gap depending on the underlying model. Results such as Codex CLI scorin…
Real-SWE benchmarked frontier coding agents against licensed, private enterprise codebases in billing, tax, and multi-service work, finding that extending rollout duration barely improves resolution rates. Rollouts finis…
A developer argues that current AI systems, including those from OpenAI, are not AGI because their binary, stateless HTTP-based architecture forces deterministic input-output processing rather than genuine thinking, with…
A Ukraine-led paper titled "DF26: We Cannot Tell Fake From Real Anymore" found that automated deepfake video detectors dropped from 94% to 48% accuracy when tested against 2026-generation generative AI video models, a de…
GPT-6 Astra scored 63% on ARC-AGI-3 with the standard agent harness and 99% with a custom one, versus 8% for GPT-5.6 Sol and 30% for Claude Opus 5 on default harnesses, according to the ARC Prize blog. In a follow-up Bab…
AMD's Helios GPU, the Instinct MI455X, packs 432 GB of HBM4, 23 TB/s of HBM bandwidth per GPU, and 40 PFLOPs of FP4 compute, according to AMD's product brochure. A new educational blog post builds a ladder of BF16 genera…
Anthropic CEO Dario Amodei published a roughly 9,000-word essay, "We Must Pace the Frontier," arguing AI labs should slow the rate at which they advance capability without halting development, citing recursive self-impro…