Frontier Models Suck Sometimes (A Benchmark)
A developer building a realtime note-taking app found that Apple Intelligence's on-device language model failed to process a 27-minute conversation, forcing a hierarchical map-reduce approach to summarize long audio reco…