04:00
2026-08-25
machinebrief.com
artificial-intelligence
K-Bench: measuring model performance on real scientific agent requests
K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web, ran 1,602 completed agent runs by nine frontier models in identical sandboxes, with three blindeβ¦