AI/ML Research Digest — Aug 02, 2026
New research highlights advances in AI reliability and efficiency. Σ-Mem and LedgerMind introduce explicit reliability modeling and structured evidence ledgers to reduce hallucinations and enable prov…
New research highlights advances in AI reliability and efficiency. Σ-Mem and LedgerMind introduce explicit reliability modeling and structured evidence ledgers to reduce hallucinations and enable prov…
Moonshot AI's PerceptionBench benchmark shows that no frontier multimodal AI model reaches 60 percent accuracy in visual perception tasks, with GPT-5.6 Sol leading by a narrow margin. The benchmark se…
Researchers introduce PerceptionBench, a benchmark designed to evaluate atomic visual perception in multimodal large language models (MLLMs), finding that no model reaches 60% accuracy across ten atom…
Kimi Team released PerceptionBench, a benchmark that isolates atomic visual perception in multimodal large language models by attributing failures across 40+ benchmarks to 10 perceptual capabilities a…