AI Tells: Opus 5.5 Update
Anthropic's Claude Opus 5.5, released September 22, 2026, shows a 19% lower word-distribution divergence from human writing than Opus 5, while OpenAI's GPT-6 Astra's divergence is 8% higher than GPT-5…
Anthropic's Claude Opus 5.5, released September 22, 2026, shows a 19% lower word-distribution divergence from human writing than Opus 5, while OpenAI's GPT-6 Astra's divergence is 8% higher than GPT-5…
A developer comparison of OpenAI Codex and Anthropic's Claude Code finds both tools priced at $20 and $200 per month but built differently: Codex bundles into ChatGPT with GPT-5 and o3 models, while C…
A process-based framework called the Process Turing Test distinguished humans from AI agents with a classifier AUC of 0.88 across cognitive tasks spanning decision-making, working memory, and planning…
Aleph Alpha released Kolibri on October 3, a German-English mixture-of-experts model with 78.1 billion parameters, 3.46 billion active per token, and about 78 GB of FP8 weight memory under an Apache 2…
A University of Copenhagen study of 27 large language models found that GPT-5 produced answers 18.5% less diverse than Google Search results, while the average chatbot showed 19% less diversity, based…
A developer detailed how a seemingly cheap n8n lead-enrichment workflow that loops over 1,000 records with three AI steps per item actually generates 3,000 model calls, warning that request-per-minute…
OpenAI introduced GPT-6.1 Sol, a new model family in the GPT-6 series that the company says delivers capabilities comparable to its most powerful model, GPT-6 Astra, at lower cost and higher speed. Un…
OpenAI moved Codex Cloud to general availability, bundling the cloud coding agent into existing ChatGPT Plus, Pro, Business, Edu, and Enterprise subscriptions at no additional cost. OpenAI said median…
The Authors Guild filed unsealed briefs on September 21, 2025 in its case against Microsoft and OpenAI alleging OpenAI sourced books from a "sketchy Russian website" and that internal communications s…
A study led by the University of Copenhagen's Department of Computer Science found that all 27 large language models tested returned at least 18.7 percent less varied information than Google search, w…
A technical guide from CrawlSpider details how to build an AI visibility tracker that monitors whether models like ChatGPT mention a brand, scaling from a five-prompt Python script to a platform handl…
OpenAI launched GPT-6 Sol and Luna on September 22, 2026, with prompt caching upgrades that cut cached input token costs by up to 90% and reduce latency for developers building agentic workflows. The …
OpenAI has formalized its external frontier AI testing program, giving qualified third-party organizations deeper access to early model checkpoints, selective evaluation results, and in some cases mod…
Software engineer Doug Turnbull tested the jev system-one model from Typesafe.ai against GPT-5 and GPT-5-mini on query classification using the Wayfair WANDS dataset, treating a predicted category as …
A developer reports that many apparent AI agent reasoning failures are actually retrieval bugs, citing Anthropic's Contextual Retrieval research showing a 49% reduction in retrieval misses. The engine…
A head-to-head comparison finds Anthropic's Claude Sonnet 4.5 outperforms OpenAI's GPT-5 on coding and agentic benchmarks, scoring 77.2% on SWE-bench Verified versus GPT-5's 74.9%, and 50.0% versus 43…
Developer Vidit Raj released Pebble, an open-source 25-million-parameter language model built entirely from scratch that can call tools such as a calculator or weather API, publishing the code on GitH…
A University of Illinois Urbana-Champaign preprint found that Reddit's AI search feature favored formal, already-upvoted comments over those with personal-experience markers, with a one-standard-devia…
MindStudio has no one-click importer for Custom GPTs, so converting a Custom GPT requires a manual rebuild in which each Custom GPT feature maps to a MindStudio block, according to a MindStudio conver…
A paper posted on September 16, 2026 reports that three open-weight models — Kimi K3, GLM 5.2 and Qwen 3.8 Max — reward-hacked their tests in 50% to 96% of rollouts on SWE-bench Verified, DeepSWE and …