When chat is the wrong UI
GitHub technologist Burke Holland argues in a GitHub Blog post that chat is the wrong UI for most AI interactions, proposing "canvases" — full-stack applications that run inside the GitHub Copilot app…
GitHub technologist Burke Holland argues in a GitHub Blog post that chat is the wrong UI for most AI interactions, proposing "canvases" — full-stack applications that run inside the GitHub Copilot app…
Ringg, a voice and chat agent platform, announced on September 23, 2026 that its AI agents resolve up to 65% of routine customer inquiries without human intervention across more than 7 million connect…
GitHub will retire six Copilot models on October 19, 2026, including GPT-5.5, GPT-5.4, GPT-5.4 mini, GPT-5 mini, Grok 4.5, and Gemini 3.7 Flash, according to a deprecation list GitHub published on Sep…
OpenAI introduced GPT-6 Sol and GPT-6 Luna on September 22, 2026, two models built on methods behind GPT-6 Astra that cut API prices by 50% versus their GPT-5.6 counterparts. GPT-6 Sol is priced at $2…
An Epoch AI report by Emberson and Roodman found that the cost of a given level of AI performance has fallen an average of about 47% per quarter over the past three years, a 13-fold drop every year. T…
An independent evaluation of TypeSafe's jev-1.13.0 judge model on 495 claims from the LLM-AggreFact benchmark found it matched a human oracle on all 500 repeated decisions in LangChain's earlier test,…
TypeSafe's new System One model, Jev, was adopted by the open-source Matrix agent platform MindRoom within two days of release, where it now powers three decisions including adaptive agent participati…
OpenAI announced price and performance updates to its GPT-5.6 model lineup, cutting the cost of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, while adding a Fast mode to GPT-5.6 Sol. The company frame…
Artificial Analysis's first full evaluation of OpenAI's GPT-6 Sol and GPT-6 Luna found both models cost roughly half their GPT-5.6 predecessors but deliver only modest intelligence gains, with GPT-6 S…
A retrieval-augmented assistant built on GPT-5.6 Luna answered a Quadient Exstream PDF/A-3 configuration question three out of four times without any supporting document in its knowledge base, accordi…
Inception Labs' Mercury 2.5 is the fastest LLM available via API as of September 2026, posting 1,107 tokens per second on vendor benchmarks and 440 tok/s P50 at 1.17s latency in OpenRouter telemetry, …
OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026, 19 days after GPT-6 Astra, pricing Sol at $2 per million input tokens and $10 per million output tokens and Luna at $0.10 and $0.50 — ha…
The Rails Foundation reported on September 21st that GPT-6 Astra kept first place in its feature-development benchmark and raised its solve rate from 35% at medium reasoning effort to 53.3% at maximum…
Notch, an AI startup that makes video ads, cut its median agent harness cost from $4.44 to $0.50 per video-producing session — nearly 90% — by switching its harness from a Sonnet-powered model to GPT-…
OpenAI's Codex CLI is open source under Apache-2.0 and its Free plan is listed at $0 a month, but OpenAI publishes no usage allowance for that tier — its published estimates begin with Plus at $20 a m…
Five of 16 AI models tested with a web search tool attached still named King Harald V as Norway's current monarch after his death on 28 August, according to a test published by stale.jock.pl creator j…
Independent evaluator Vals found that Google's Gemini 3.8 Flash scored 71.7% on BioMysteryBench's human-solvable tasks and 21.6% on its hard tasks in production runs, versus the 88.8% and 56.5% Google…
A Microsoft developer's evaluation of GPT-5.6 Luna's knowledge of Dev Proxy versions was invalidated because the coding agent located a local Dev Proxy installation and source checkout on the host mac…
Alibaba's Qwen team released Qwen 3.8 27B, a 27-billion-parameter model distributed as a 17GB GGUF file that scores 52 on the Artificial Analysis Intelligence Index, matching OpenAI's cloud-hosted GPT…
An analysis of the Entelligence benchmark comparing GPT-5.6 Luna and GPT-6 Astra on code review found that labeling methodology, not model choice, drives reported precision. Run on 2026-09-14 against …