How funny are the frontier AI models?
A researcher at the lab Lossfunk ran a blind human-rating study testing whether frontier LLMs can produce original funny jokes, generating jokes from six closed OpenAI and Anthropic models (Haiku 4.5,…
A researcher at the lab Lossfunk ran a blind human-rating study testing whether frontier LLMs can produce original funny jokes, generating jokes from six closed OpenAI and Anthropic models (Haiku 4.5,…
GitHub has switched Copilot from request-based to token-based billing, metering usage through AI Credits, while Claude Code retains a flat-rate model with throttling. A developer's analysis of publish…
OpenAI's $200-per-month ChatGPT Pro plan has been silently routing some Pro model sessions to the mini model for several months, according to a user who analyzed chat metadata and contacted support fo…
A study using the ASAP 2.0 dataset found that GPT-5 mini achieved the highest agreement with human ratings for automated essay scoring, while GPT-5 produced the strongest summarization quality, reveal…
A developer built a personal algorithm using a browser extension that saves articles, posts, videos, and tools, then generates a profile.md file via Claude to filter content based on the user's saved …
GitHub switched Copilot to usage-based AI credit billing on June 1, causing some Pro+ subscribers to deplete their monthly allocation within two days. The model price spread is 24x, with GPT-5.5 costi…
Meta Platforms released developer access to its Muse Spark AI model and an upgraded version, Muse Spark 1.1, on Thursday, positioning it against Anthropic and OpenAI. The model, priced at $1.25 per mi…
Researchers introduced EnterpriseMem-Bench, a multi-turn Text-to-SQL benchmark of 300 sessions and 1,400 queries built from three enterprise domains, to evaluate how five frontier AI models handle mem…