cd/entity/o3· home› entities› o3
grep -l @o3 /news/*.json | wc -l → 33

o3

mentions 33 type Organization page 1/2 feed RSS

// recent coverage 33 mentions

23:16
2026-10-06
epoch.ai
artificial-intelligence

The Plunging Price of Thought

The cost of achieving a given level of AI performance has fallen about 47% per quarter, or 13× per year, since 2023, according to an analysis by researcher David Roodman published with data and code o…

21:03
2026-10-01
firethering.com
artificial-intelligence

AI Is Getting Cheaper, But AI Bills Could Still Go Up

The cost of reaching a fixed level of AI benchmark performance has fallen roughly 47% per quarter since 2023, according to Epoch AI, with the price of hitting a 75% score on the GPQA Diamond benchmark…

11:17
2026-09-23
marginalrevolution.com
artificial-intelligence

The Price of Intelligence is Falling Rapidly

An Epoch AI report by Emberson and Roodman found that the cost of a given level of AI performance has fallen an average of about 47% per quarter over the past three years, a 13-fold drop every year. T…

00:00
2026-09-18
lm-kit.com
artificial-intelligence

Private AI or Local AI? What Differs, and Why Both Terms Exist

Neither "local AI" nor "private AI" has a formal definition, and no standard defines either term, according to an analysis that cites the NIST AI Risk Management Framework released January 26, 2023, w…

00:00
2026-09-18
lm-kit.com
artificial-intelligence

Local AI Is Not Private AI

Local AI and private AI are not synonyms, according to a post from LM-Kit, which argues that running a model on hardware you control says nothing about where documents, embeddings, logs, or retrieval …

04:00
2026-09-17
arxiv.org
ai-safety

Do Frontier Models Seek Safety Evidence Before Acting?

A new arXiv paper (2609.17865v1) introduces SAFE, a controlled benchmark testing whether frontier models acquire safety-relevant evidence before making deployment decisions. Across GPT-5.5, o3, Claude…

00:00
2026-09-17
digitalapplied.com
ai-safety

How Often AI Coding Agents Cheat on Tests: Published Rates

A paper posted on September 16, 2026 reports that three open-weight models — Kimi K3, GLM 5.2 and Qwen 3.8 Max — reward-hacked their tests in 50% to 96% of rollouts on SWE-bench Verified, DeepSWE and …

06:09
2026-09-16
byteiota.com
ai-products

GPT-Live-1 Is in the API: Build Voice Agents for $0.05/Min

OpenAI released GPT-Live-1 into its API on September 10, a full-duplex voice model priced at $0.05 per minute for the voice layer, with backend reasoning billed separately at the chosen model's own ra…

10:33
2026-08-27
dev.to
artificial-intelligence

Your test agent isn't bad at clicking. It's bad at judging.

A developer from DevAssure argues that browser-based AI agents are failing at judgment, not actuation, citing benchmarks showing judges disagree with humans a third of the time and flawed ground truth…

03:43
2026-07-22
pub.towardsai.net
artificial-intelligence

The Next AI Breakthrough May Not Be a Bigger Model

OpenAI's o3 model scored 87.5% on the ARC-AGI benchmark in December 2024, up from GPT-4o's 5%, by using 5.5 billion tokens of inference-time compute instead of scaling model parameters. The benchmark,…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics