cd/entity/o3· home entities o3
grep -l @o3 /news/*.json | wc -l → 18

o3

mentions 18 type Organization feed RSS

// recent coverage 18 mentions

03:43
2026-07-22
pub.towardsai.net
artificial-intelligence

The Next AI Breakthrough May Not Be a Bigger Model

OpenAI's o3 model scored 87.5% on the ARC-AGI benchmark in December 2024, up from GPT-4o's 5%, by using 5.5 billion tokens of inference-time compute instead of scaling model parameters. The benchmark,…

15:22
2026-07-21
lesswrong.com
ai-safety

Measuring Reward-Seeking via Contrastive Belief Updates

Researchers at Redwood Research and Anthropic have developed a method called Contrastive Synthetic Document Finetuning to measure reward-seeking behavior in AI models, finding that intermediate checkp…

18:59
2026-07-11
dev.to
large-language-models

Model Kombat: The LLM Fighting Game!

A developer built Model Kombat, a 2D fighting game that visualizes Large Language Model architectures, parameter scales, and hardware constraints as playable mechanics. The game features fighters repr…

00:00
2026-07-07
seangoedecke.com
artificial-intelligence

Blog about things you don't understand yet

Blogger Sean Goedecke argues that writing controversial blog posts forces deeper learning and clearer thinking, contrary to common advice favoring unstructured self-expression. He claims that structur…

19:15
2026-06-24
letsdatascience.com
large-language-models

OpenAI updates GPT-5.5 Instant behaviour in ChatGPT

OpenAI updated ChatGPT's default model to GPT-5.5 Instant in May 2026, replacing GPT-5.3 Instant for all users. The update delivers 52.5% fewer hallucinated claims and 37.3% fewer inaccurate answers, …

00:00
2026-06-10
mindstudio.ai
artificial-intelligence

ChatGPT vs Claude in 2026: Which AI Should You Actually Use?

ChatGPT leads in image generation and voice interaction, while Claude excels in long-form writing, document analysis, and agentic tasks, according to a 2026 comparison of the two leading AI models. Us…

17:11
2026-05-27
twitter.com
large-language-models

GPT 5.5 aces 20x20 multiplication that o3 couldn't handle

OpenAI's unreleased GPT 5.5 model has successfully performed 20x20 multiplication, a task that the company's previous o3 model failed to complete. The achievement marks a significant advancement in la…

00:00
2026-05-21
seangoedecke.com
artificial-intelligence

The famous o3 "GeoGuessr" prompt did not work

OpenAI's o3 model performed worse at geolocation when using a widely-circulated "GeoGuessr" prompt than with a basic default prompt, according to a benchmark test of 200 images. The median distance fr…

19:05
2026-03-28
muratbuffalo.blogspot.com
artificial-intelligence

Measuring AI Ability to Complete Long Software Tasks

Based solely on the provided article, researchers at METR introduced a new metric called the "50%-task-completion time horizon" to track AI progress, finding that this horizon—the length of a software…

// co-occurs with top 8 entities
// topics top 6 topics