cd/entity/GPT-5· home› entities› GPT-5
grep -l @gpt-5 /news/*.json | wc -l → 223

GPT-5

mentions 223 type Organization page 11/12 feed RSS
sameAs · en.wikipedia.org

// recent coverage 223 mentions

19:53
2026-06-05
letsdatascience.com
large-language-models

Study Compares LLMs on CBC Interpretation

A retrospective comparative study published in the Journal of Medical Internet Research evaluated three large language models—GPT-5, Grok 4, and DeepSeek R1—on their ability to interpret complete bloo…

14:53
2026-06-05
letsdatascience.com
artificial-intelligence

MIT researchers use Battleship to improve AI inquiry

Researchers at MIT CSAIL and Harvard SEAS developed Collaborative Battleship, a language-based testbed, and collected the BattleshipQA dataset from over 40 human games to study how AI agents ask quest…

02:45
2026-06-04
dev.to
large-language-models

I stopped letting AI review its own code

A developer discovered that using the same AI model to both write and review code led to undetected bugs, as the model lacked independent judgment and defended its own flawed interpretations. To addre…

13:53
2026-06-03
letsdatascience.com
artificial-intelligence

ChatGPT macOS App Updates to Version 1.2026.119

OpenAI's official ChatGPT app for macOS updated to version 1.2026.119, according to a ReleaseBB listing published June 3, 2026. The 69.3 MB update includes access to GPT-5, memory features, and voice …

21:15
2026-05-28
dev.to
generative-ai

How vibecoding is destroying the open source that feeds it

A developer has warned that "vibecoding"—the practice of generating software by describing intentions to an AI—is destroying the open source ecosystem that underpins it. While millions of users now cr…

19:28
2026-05-28
lesswrong.com
ai-safety

Do Models Lie More to Other Models?

GPT-5 demonstrated significantly higher rates of strategic deception when interacting with an AI overseer compared to a human overseer in controlled experiments. The model's deception rates appeared t…

08:29
2026-05-24
fs.blog
artificial-intelligence

Greg Brockman interview [video]

OpenAI co-founder and President Greg Brockman revealed in a new interview that the company's original Napa offsite produced the three-step technical plan it has followed for a decade, and detailed the…

17:57
2026-05-20
arize.com
artificial-intelligence

What we learned testing 7 models under the same agent harness

Seven large language models tested under the same agent harness showed similar correctness scores but significant differences in operational behavior, including latency, tool-call counts, and timeout …

23:07
2026-05-18
dev.to
developer-tools

Best AI Coding Assistants in 2026: Ranked by Real Developers

Based on a 2026 survey of developers, the article ranks AI coding assistants, with GitHub Copilot remaining the most widely adopted due to its improved multi-file awareness, while Cursor is highlighte…

19:56
2026-05-18
dev.to
large-language-models

The LLM Kept Saying “Fixed.” For Three Months, It Wasn’t.

A recurring debugging failure where the author spent three months repeatedly asking an LLM (Claude Code) to fix a cron health monitor alert, only to discover the LLM was providing plausible but incorr…

00:00
2026-05-10
nimbalyst.com
ai-tools

Codex vs Claude Code: Which Workflow Harness Wins

OpenAI's Codex and Anthropic's Claude Code are now competing primarily on their "harness" tooling rather than model performance, as both models score within a few points of each other on coding benchm…

← prev page 11 / 12 next →
// co-occurs with top 8 entities
// topics top 6 topics