cd /news/artificial-intelligence/openai-research-scientist-says-ai-ag… · home topics artificial-intelligence article
[ARTICLE · art-132645] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

OpenAI research scientist says AI agents still can’t figure out what’s worth researching

OpenAI research scientist Noam Brown said current AI models are "poor" at "research taste," the ability to identify which research directions are promising, and expects improvement within one to two model generations. OpenAI confirmed it hit its "automated research intern" milestone by September 2026, with agents performing 3.1 times the human labor days on multi-day research tasks, though over 50% of those tasks still require human intervention, and the median OpenAI researcher spends over $600 per day on agent inference while the 90th percentile exceeds $7,000 per day. OpenAI has set a target of a fully autonomous AI researcher by March 2028 and has acknowledged it does not currently know how to safely achieve the alignment level such a system would require.

read3 min views3 publishedSep 17, 2026
OpenAI research scientist says AI agents still can’t figure out what’s worth researching
Image: Cryptobriefing (auto-discovered)

OpenAI official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment

Noam Brown admits current models lack 'research taste' while OpenAI spends thousands per day on AI inference costs that still require heavy human oversight

Your AI assistant can write code, summarize papers, and synthesize data across dozens of sources. What it can’t do is look at a messy, open-ended research landscape and decide which problem actually matters. That gap has a name: research taste. And according to one of OpenAI’s own scientists, today’s models are pretty bad at it.

Noam Brown, a research scientist at OpenAI, recently stated that current AI models are “poor” at research taste, the intuitive ability to identify which research directions are promising and which are dead ends. He expects improvements within one to two model generations, but for now, the limitation is stark enough that Brown is highlighting it publicly.

The automated intern that still needs a boss #

OpenAI has made real progress on the execution side of AI research. The company confirmed it achieved its milestone for building an “automated research intern” by September 2026, with agents performing 3.1 times the human labor days on multi-day research tasks.

That sounds impressive until you hear the rest. Over 50% of those tasks still require human intervention.

The cost structure tells a revealing story too. The median researcher at OpenAI spends over $600 per day on agent inference. Researchers at the 90th percentile blow through more than $7,000 per day.

AI, tech, and the markets they move—in one daily briefing.

Daily. Free. Join 34,000+ readers across crypto, finance, and policy.

The agents are effective at engineering-heavy tasks: writing code, running experiments, processing large datasets. Where they fall apart is in the softer, harder-to-define skills. Ideation, knowing when to abandon a failing approach, deciding which experiment to run next.

Research taste is the bottleneck #

Independent studies have found that AI models perform nearly randomly when trying to predict citation velocity, a rough proxy for whether a piece of research will actually matter to the field. They also struggle with judging whether work is publishable, assessing originality, and executing the kind of complex pivots that define breakthrough research.

The road to March 2028 #

OpenAI has set an ambitious target: a fully autonomous AI researcher by March 2028. That would mean an agent capable of not just executing research tasks but identifying which problems to work on, designing experimental approaches, and knowing when to change direction.

The company has been candid about the difficulty. OpenAI has acknowledged it does not currently know how to safely achieve the level of alignment required for such a system.

Brown’s expectation of improvement within one to two model generations is optimistic but not unreasonable. The question is whether research taste, which relies on something closer to intuition built from deep domain experience, is the kind of capability that scales with model size and training data, or whether it requires fundamentally different approaches.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-research-scie…] indexed:0 read:3min 2026-09-17 ·