{"slug": "openai-research-scientist-says-ai-agents-still-cant-figure-out-whats-worth", "title": "OpenAI research scientist says AI agents still can’t figure out what’s worth researching", "summary": "OpenAI research scientist Noam Brown said current AI models are \"poor\" at \"research taste,\" the ability to identify which research directions are promising, and expects improvement within one to two model generations. OpenAI confirmed it hit its \"automated research intern\" milestone by September 2026, with agents performing 3.1 times the human labor days on multi-day research tasks, though over 50% of those tasks still require human intervention, and the median OpenAI researcher spends over $600 per day on agent inference while the 90th percentile exceeds $7,000 per day. OpenAI has set a target of a fully autonomous AI researcher by March 2028 and has acknowledged it does not currently know how to safely achieve the alignment level such a system would require.", "body_md": "OpenAI official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment\n\n# OpenAI research scientist says AI agents still can’t figure out what’s worth researching\n\nNoam Brown admits current models lack 'research taste' while OpenAI spends thousands per day on AI inference costs that still require heavy human oversight\n\nYour AI assistant can write code, summarize papers, and synthesize data across dozens of sources. What it can’t do is look at a messy, open-ended research landscape and decide which problem actually matters. That gap has a name: research taste. And according to one of OpenAI’s own scientists, today’s models are pretty bad at it.\n\nNoam Brown, a research scientist at OpenAI, recently stated that current AI models are “poor” at research taste, the intuitive ability to identify which research directions are promising and which are dead ends. He expects improvements within one to two model generations, but for now, the limitation is stark enough that Brown is highlighting it publicly.\n\n## The automated intern that still needs a boss\n\nOpenAI has made real progress on the execution side of AI research. The company confirmed it achieved its milestone for building an “automated research intern” by September 2026, with agents performing 3.1 times the human labor days on multi-day research tasks.\n\nThat sounds impressive until you hear the rest. Over 50% of those tasks still require human intervention.\n\nThe cost structure tells a revealing story too. The median researcher at OpenAI spends over $600 per day on agent inference. Researchers at the 90th percentile blow through more than $7,000 per day.\n\n### AI, tech, and the markets they move—in one daily briefing.\n\nDaily. Free. Join 34,000+ readers across crypto, finance, and policy.\n\nThe agents are effective at engineering-heavy tasks: writing code, running experiments, processing large datasets. Where they fall apart is in the softer, harder-to-define skills. Ideation, knowing when to abandon a failing approach, deciding which experiment to run next.\n\n## Research taste is the bottleneck\n\nIndependent studies have found that AI models perform nearly randomly when trying to predict citation velocity, a rough proxy for whether a piece of research will actually matter to the field. They also struggle with judging whether work is publishable, assessing originality, and executing the kind of complex pivots that define breakthrough research.\n\n## The road to March 2028\n\nOpenAI has set an ambitious target: a fully autonomous AI researcher by March 2028. That would mean an agent capable of not just executing research tasks but identifying which problems to work on, designing experimental approaches, and knowing when to change direction.\n\nThe company has been candid about the difficulty. OpenAI has acknowledged it does not currently know how to safely achieve the level of alignment required for such a system.\n\nBrown’s expectation of improvement within one to two model generations is optimistic but not unreasonable. The question is whether research taste, which relies on something closer to intuition built from deep domain experience, is the kind of capability that scales with model size and training data, or whether it requires fundamentally different approaches.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/openai-research-scientist-says-ai-agents-still-cant-figure-out-whats-worth", "canonical_source": "https://cryptobriefing.com/openai-ai-agents-research-taste-shortcomings/", "published_at": "2026-09-17 14:12:31+00:00", "updated_at": "2026-09-17 14:24:50.270584+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-research", "ai-safety"], "entities": ["OpenAI", "Noam Brown"], "alternates": {"html": "https://wpnews.pro/news/openai-research-scientist-says-ai-agents-still-cant-figure-out-whats-worth", "markdown": "https://wpnews.pro/news/openai-research-scientist-says-ai-agents-still-cant-figure-out-whats-worth.md", "text": "https://wpnews.pro/news/openai-research-scientist-says-ai-agents-still-cant-figure-out-whats-worth.txt", "jsonld": "https://wpnews.pro/news/openai-research-scientist-says-ai-agents-still-cant-figure-out-whats-worth.jsonld"}}