Your AI assistant can write code, summarize papers, and synthesize data across dozens of sources. What it can’t do is look at a messy, open-ended research landscape and decide which problem actually matters. That gap has a name: research taste. And according to one of OpenAI’s own scientists, today’s models are pretty bad at it.
Noam Brown, a research scientist at OpenAI, recently stated that current AI models are “poor” at research taste, the intuitive ability to identify which research directions are promising and which are dead ends. He expects improvements within one to two model generations, but for now, the limitation is stark enough that Brown is highlighting it publicly.
The automated intern that still needs a boss
OpenAI has made real progress on the execution side of AI research. The company confirmed it achieved its milestone for building an “automated research intern” by September 2026, with agents performing 3.1 times the human labor days on multi-day research tasks.
That sounds impressive until you hear the rest. Over 50% of those tasks still require human intervention.
The cost structure tells a revealing story too. The median researcher at OpenAI spends over $600 per day on agent inference. Researchers at the 90th percentile blow through more than $7,000 per day.
The agents are effective at engineering-heavy tasks: writing code, running experiments, processing large datasets. Where they fall apart is in the softer, harder-to-define skills. Ideation, knowing when to abandon a failing approach, deciding which experiment to run next.
Research taste is the bottleneck
Independent studies have found that AI models perform nearly randomly when trying to predict citation velocity, a rough proxy for whether a piece of research will actually matter to the field. They also struggle with judging whether work is publishable, assessing originality, and executing the kind of complex pivots that define breakthrough research.
The road to March 2028
OpenAI has set an ambitious target: a fully autonomous AI researcher by March 2028. That would mean an agent capable of not just executing research tasks but identifying which problems to work on, designing experimental approaches, and knowing when to change direction.
The company has been candid about the difficulty. OpenAI has acknowledged it does not currently know how to safely achieve the level of alignment required for such a system.
Brown’s expectation of improvement within one to two model generations is optimistic but not unreasonable. The question is whether research taste, which relies on something closer to intuition built from deep domain experience, is the kind of capability that scales with model size and training data, or whether it requires fundamentally different approaches.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
28








English (US) ·