OpenAI’s “deep research” is the latest artificial intelligence (AI) feature drawing attention, with claims it can complete in minutes the sort of work that would take a human specialist hours.
Packaged inside ChatGPT Pro and promoted as a research assistant capable of matching a trained analyst, it independently searches the web, gathers sources and produces structured reports. In testing, it achieved 26.6 percent on Humanity’s Last Exam (HLE), a demanding AI benchmark, beating many competing models.
Even so, deep research does not fully deliver on the publicity. Although it can generate professional-looking write-ups, it comes with notable weaknesses. Journalists who have trialled it report that it can overlook important details, falter on very recent information and, at times, fabricate claims.
OpenAI acknowledges this in its own description of the tool’s constraints. The company also says it "can sometimes hallucinate facts in responses or make incorrect inferences, though at a notably lower rate than existing ChatGPT models, according to internal evaluations".
That unreliable material can appear is unsurprising: AI models do not “know” things in the way people do.
The concept of an AI “research analyst” also prompts a host of questions. Can a machine - however capable - genuinely stand in for a trained expert? What does this mean for knowledge work? And does AI help us reason more effectively, or does it simply make it easier to switch our brains off?
Why OpenAI’s deep research is attracting attention
On the surface, the pitch is compelling for knowledge workers: a tool that can do the legwork of searching, compiling and writing up a report rapidly, with sources attached.
But a closer inspection shows there are real limits to what it can do well.
What is ‘deep research’ and who is it for?
Aimed at professionals working in finance, science, policy, law and engineering - as well as academics, journalists and business strategists - deep research is the newest “agentic experience” OpenAI has introduced in ChatGPT. It is presented as a way to compress substantial research work into minutes.
At present, deep research is available only to ChatGPT Pro users in the United States, priced at US$200 a month. OpenAI says it will extend access to Plus, Team and Enterprise users over the next few months, and that a cheaper version is planned later on.
Rather than behaving like a typical chatbot that replies quickly, deep research runs through several stages to produce a structured report:
- The user submits a request. This might range from a market analysis to a summary of a legal case.
- The AI clarifies the task. It can ask follow-up questions to narrow and define the scope.
- The agent searches the web. It autonomously browses hundreds of sources, including news coverage, research papers and online databases.
- It synthesises its findings. The AI pulls out the main points, arranges them into a structured report and provides citations.
- The final report is delivered. Within five to 30 minutes, the user receives a multi-page document - potentially even a PhD-level thesis - summarising the results.
Initial impressions may suggest a perfect solution for research-heavy roles. Yet early trials have highlighted major shortcomings:
- It lacks context. AI can condense material, but it does not fully grasp what matters most.
- It ignores new developments. It has failed to capture significant legal decisions and scientific updates.
- It makes things up. As with other AI models, it can generate false information with confidence.
- It can’t tell fact from fiction. It does not reliably separate authoritative sources from untrustworthy ones.
Although OpenAI argues the tool can rival human analysts, AI cannot replicate the judgement, scrutiny and subject expertise that underpin high-quality research.
What AI can’t replace
ChatGPT is not alone in offering web-scraping, report-writing capabilities triggered by a handful of prompts. In particular, just 24 hours after OpenAI’s launch, Hugging Face released a free, open-source alternative that comes close to matching its performance.
The central danger with deep research - and similar AI products sold as “human-level” research - is that they can create the impression that human thinking is no longer needed. AI can summarise content, but it cannot interrogate its own assumptions, draw attention to gaps in knowledge, think creatively or properly appreciate different viewpoints.
Nor do AI-produced summaries reach the depth achieved by an experienced human researcher.
Any AI agent, regardless of speed, remains a tool rather than a substitute for human intelligence. For people whose jobs depend on knowledge, it is increasingly vital to develop capabilities AI cannot reproduce: critical thinking, careful fact-checking, deep expertise and creativity.
Using AI research tools without losing rigour
If you decide to use AI research tools, it is possible to do so in a responsible way. Used thoughtfully, AI can speed up parts of research without undermining accuracy or depth. For instance, you might rely on AI to improve efficiency - such as summarising documents - while keeping human judgement for decisions.
Check sources every time, because AI-generated citations can be deceptive. Avoid accepting outputs at face value; apply critical thinking and verify claims against reputable sources. For high-stakes areas - including health, justice and democracy - pair any AI findings with expert input.
Despite marketing that suggests otherwise, generative AI still faces many limits. People who can creatively combine information, challenge assumptions and think critically will continue to be needed - AI cannot replace them yet.
Raffaele F Ciriello, Senior Lecturer in Business Information Systems, University of Sydney
This article is republished from The Conversation under a Creative Commons licence. Read the original article.
Comments
No comments yet. Be the first to comment!
Leave a Comment