llms, the jagged frontier, and health
OpenAI's GPT-6 Astra shows only incremental gains over GPT-5.6 Sol on health benchmarks, with HealthBench Professional rising from 60.5 to 63.4 and HealthBench Hard from 33.1 to 36.3, while Consensus …
OpenAI's GPT-6 Astra shows only incremental gains over GPT-5.6 Sol on health benchmarks, with HealthBench Professional rising from 60.5 to 63.4 and HealthBench Hard from 33.1 to 36.3, while Consensus …
PhenoBench, a new benchmark built on the Human Phenotype Project cohort of more than 13,000 participants, introduces 90 tasks across 15 clinical domains to evaluate how AI models use longitudinal heal…
A researcher and developer, Jamon, reports that working with AI agents late in the day leads to lapses in judgment and reduced quality, prompting a personal rule of 'no agents after 8pm.' He notes tha…
Frontier AI labs have become attractive buyers of biological data, but selling to them can distract bio companies from their core mission and may not guarantee long-term value, according to a Substack…
A Nature Medicine paper claiming general-purpose LLMs outperform specialized clinical tools on medical benchmarks is criticized for flawed methodology. The benchmark, Real Clinical Queries, evaluated …
A developer reports sustained use of PEEK, a system that maintains a compact context map in an agent's prompt to orient it within recurring external contexts. The map, a few hundred tokens appended to…
A developer describes two modes of working with AI coding assistants: 'synced mode' for building shared understanding and 'delegate mode' for offloading tasks once boundaries are solid. The author arg…
The medARC group released Medmarks v1.0, the largest fully open medical LLM evaluation suite, featuring 30 benchmarks across verifiable and open-ended subsets, covering 61 models on 71 configurations.…