The Deep Research Moment
For most people in September 2026, Gemini Deep Research is the best all-round pick: the strongest published benchmarks and the cheapest paid entry at $4.99. Perplexity Pro at $20 gives the best volume for the money, and OpenAI and Claude produce the deepest reports on ambiguous work. Developers should start with Google's Deep Research API, the one turnkey option that isn't on a retirement clock, because Perplexity retires sonar-deep-research on September 27, 2026 and OpenAI shut down its dedicated research models in July.
The rest of this guide shows how those verdicts hold up across pricing, benchmarks, APIs, and real jobs.
The category itself is not yet two years old. On February 2, 2025, OpenAI announced Deep Research. It was the first agent most people had used that could take a one-sentence prompt, plan a 30-minute investigation, browse dozens of sources on its own, and return with a cited report.
The industry reaction was telling. Within six weeks, Perplexity shipped its own Deep Research (February 14) and opened the Sonar Deep Research API to developers weeks later. Google, which had launched Gemini Deep Research quietly in December 2024, accelerated its rollout and kept upgrading the backbone. Anthropic made Claude's web search generally available on May 7, 2025, packaging the Research feature in the same spring window.
Four labs, one product category, two quarters. That doesn't happen by accident. 2024 was the year context windows crossed 200K tokens, tool use became reliable, and agentic loops stopped silently failing halfway through. Deep research was the first consumer-facing app that made all three feel worth paying for. It's closely tied to the broader shift toward agent protocols we cover in The Agentic Web: Inside the MCP Protocol Wars.
If you write, study, analyze markets, or evaluate products, you're already at a disadvantage if you don't use one. The question is which one, and when.
What "Deep Research" Actually Does
It's easy to confuse deep research with chat search. You type a question, you get an answer with links. The mechanics are different.
A chat search (like regular ChatGPT with browsing) runs one or two web queries and synthesizes the top results in seconds. A deep research agent does something closer to what a junior analyst does over an afternoon. It breaks your question into sub-questions, runs dozens or hundreds of searches, reads full pages, follows citations, updates its plan as it learns, and produces a structured report with footnotes.
Ask chat search "what are the main critiques of the Phillips curve?" and you'll get a three-paragraph summary. Ask a deep research agent the same thing and you'll get a 15-page report covering Friedman's natural rate hypothesis, the 1970s stagflation breakdown, rational expectations revisions, post-2008 flattening debates, and recent papers from 2023-2025, each with a source you can click.
The trade-off is time. Runs take minutes to an hour, and Google is the only vendor that publishes a budget: most tasks inside 20 minutes, hard-capped at 60. That's the point. You queue one up, work on something else, and come back to a report that would have taken you half a day to assemble manually. For more on restructuring research habits around AI agents, see How to Build an AI-Powered Research Workflow in 2026.
Head-to-head: The 4 Consumer Tools Compared
Here's the matrix, with model backing and pricing read off the vendors' own pages in September 2026. One column has changed more than any other since this guide first ran: the "backing model" column, which is now a moving target at three of the four.
| Tool | Backing model (Sept 2026) | Price and access | Published run time |
|---|---|---|---|
| OpenAI Deep Research | Not disclosed; "the latest model by default" | Free and Go ($8/mo): Limited; Plus ($20/mo): Expanded; Pro (from $100/mo): Maximum | 5-30 min (OpenAI, Feb 2025) |
| Perplexity Deep Research | Claude Opus, routed inside Perplexity Computer | Pro $20/mo; Max $200/mo | 2-4 min (Perplexity, Feb 2025) |
| Gemini Deep Research | Gemini 3.1 Pro | AI Plus $4.99/mo; AI Pro $19.99/mo; AI Ultra from $99.99/mo | 20 min typical, 60 min cap (API) |
| Claude Research | Not disclosed | All paid plans: Pro $20/mo, Max $100 or $200/mo, Team, Enterprise | Not published |
That table has more "not published" in it than it used to, and that's the honest state of things. Only Google publishes a hard time budget for a current product. OpenAI's widely quoted "5 to 30 minutes" sits on its February 2025 launch page and has never been restated. Perplexity published 2 to 4 minutes at that same launch and hasn't updated it either, so the "under 3 minutes" this guide used to print was a February 2025 number describing a product that has since been rebuilt twice. Anthropic says only that Research delivers answers "in minutes."
The one-paragraph profiles:
OpenAI Deep Research is still the heavyweight for ambiguous, open-ended questions. Reports run long and the reasoning is visibly deeper when the question has no clean answer. Two things changed in 2026. The pricing page stopped quoting run counts and now says Limited, Expanded, and Maximum, so the "5 free, 25 on Plus, 250 on Pro" figures still circulating come from an April 2025 blog update rather than from anything OpenAI currently publishes per plan. And on September 10, 2026, OpenAI paused new signups to its $200 Pro tier. There are two Pro tiers: $100 unlocks 5x Plus usage and $200 unlocks 20x, and the pause leaves $100 as the top plan a new customer can buy.
Perplexity Deep Research is the one whose engine swapped out underneath it. It launched on Perplexity's own Sonar models. Since February 2026 the Advanced Deep Research mode has run on Claude Opus, with Max subscribers getting a newer Opus revision than Pro subscribers, and on June 11, 2026 Perplexity moved Deep Research inside Perplexity Computer, which routes subtasks across more than 20 frontier models. Perplexity attributes a jump from 40.7% to 83.8% on BrowseComp to that architecture move rather than to the model swap. If you read an article calling Perplexity Deep Research a "Sonar" product, it was written before February 2026. Sonar is now the name of the API line, not the consumer agent.
Gemini Deep Research has the strongest published results of the big four. Google upgraded the agent to Gemini 3.1 Pro in April 2026, split it into a standard and a Max tier, and exposed both through a real API. It still pulls from your Gmail, Drive, and Docs alongside the web, and it shows a visible research plan you can edit before the agent runs. Google publishes allowances as multipliers rather than counts, which at least tells you what an upgrade buys.
Claude Research is the patient one, good at holding contradictory sources side by side instead of resolving them too early. Two corrections worth making: Research is available on every paid Claude plan (Pro, Max, Team, and Enterprise), not Pro alone, and Anthropic's lineup has moved past Sonnet 5 and Opus 5 to include Claude Fable 5.1, released September 1, 2026. Anthropic doesn't say which model backs Research in the apps. What it does document is that Claude 4.6 and later carry the full 1M-token context at standard pricing, so large source sets don't get truncated or surcharged.
ChatGPT Deep Research Alternatives: Who Else Ships One
"ChatGPT deep research alternative" is one of the most common ways people arrive at this topic, and the four-way comparison above doesn't answer it. Here's who else actually ships a research agent in September 2026, and who only looks like they do.
| Challenger | Product and API | Entry price | The distinguishing fact |
|---|---|---|---|
| xAI Grok | Multi-agent mode; grok-4.20-multi-agent (beta) | SuperGrok $30/mo | Publishes no HLE or BrowseComp score |
| Moonshot Kimi | Deep Research (Kimi-Researcher); no endpoint for it | Plus $19/mo | Publishes per-run telemetry nobody else does |
| Manus | Autonomous agent, Wide Research; task-shaped REST | $20/mo (4,000 credits) | Publishes no named public benchmark |
| You.com | Consumer product mostly gone; Research API | $12 per 1,000 (lite) | Five effort tiers spanning a 100x price range |
| Mistral | Deep Research skill in Vibe; no standalone endpoint | Pro $14.99/mo | Cheapest paid tier outside the big four |
| DeepSeek | Web search in the app; nothing agentic in the API | Free | The API ignores every tool you pass it |
xAI Grok quietly renamed the thing everyone still searches for. "DeepSearch" and "DeeperSearch" no longer appear anywhere in xAI's developer documentation. What exists now is a multi-agent model, grok-4.20-multi-agent, still labeled beta, where reasoning.effort selects 4 agents at low or medium and 16 at high or xhigh. The docs tell you plainly to use 16 for deep research. The catch for anyone comparing on quality: xAI published no Humanity's Last Exam, BrowseComp, or GAIA score for the underlying Grok 4.6, so there's nothing to put in the usual benchmark table. It does report 81.6% on DeepSearchQA, where it loses to both GPT-5.5 and Claude Opus.
Moonshot's Kimi is the best-documented challenger and the one most people outside Asia have never tried. Its Deep Research mode runs on an in-house Kimi-Researcher model, and Moonshot publishes per-task telemetry that the American labs don't: an average of 23 reasoning steps, 74 planned search keywords, and 206 URLs found per task, of which only the top 3.2% survive into the report. Runs take 10 to 25 minutes and paid tiers start at $19/month. Moonshot ships search and fetch endpoints you can call, but nothing that exposes the Deep Research agent itself, so the research product is consumer-only.
Manus is the interesting business story and the thinnest technical one. Meta acquired it in December 2025, China's NDRC blocked the deal on April 27, 2026, and Manus announced it had resumed independent operations on September 1, 2026. Its Wide Research mode runs 20 subtasks concurrently, though you can't trigger it manually. Worth knowing before you cite it: Manus publishes no named public benchmark anywhere, only charts it labels internal.
You.com has effectively left the consumer market. ARI, its research agent, no longer appears on the homepage, and there's no consumer subscription on the pricing page. What's left is a genuinely good Research API with five effort levels from lite to frontier, priced from $12 to $1,200 per thousand requests. A 100x spread on a single endpoint is the clearest illustration of something this whole category has converged on, which the API section below picks up.
Mistral moved its research agent and renamed the app around it. Le Chat became Vibe in May 2026, and Deep Research was sunset in the Chat tab; starting one now redirects you to the Work tab with the Deep Research skill preselected. At $14.99/month, Mistral Pro is the cheapest paid tier outside the big four, though Google's $4.99 AI Plus undercuts it, and for European teams with data-residency requirements Mistral is often the only option that clears procurement.
DeepSeek is the useful negative result. Despite constant inclusion in "best deep research tools" listicles, DeepSeek ships no deep research product. Its app has web search and a think mode, but its API documentation marks web_search and every other tool as ignored, so there's nothing agentic to build on. It's an excellent cheap reasoning model you can point at text you supply, and it isn't a research agent. The same goes for Qoob and Arc Search, both of which show up in comparison queries and both of which are fast-answer search apps, not agents that plan and browse for half an hour.
The open-source options, and a warning
If you want to self-host, the honest 2026 answer is that maintenance and benchmark performance now live in different repositories.
Actively maintained and worth starting with: GPT Researcher (Apache-2.0, roughly 29,500 stars, shipping releases through August 2026) is the one you can clone and run today without first replacing a retired model name in its config. ByteDance's DeerFlow is the largest by far at around 82,000 stars and the cheapest to try, since it runs with no search API keys at all using DuckDuckGo and a keyless reader, though its 2.0 rewrite repositioned it as a general agent harness with deep research as one skill among many. Alibaba's Tongyi DeepResearch is the credible open-weights option under Apache-2.0, but its repository has been silent since February 2026.
Now the warning, because two projects that topped every 2025 recommendation list are dead and don't look it. LangChain's open_deep_research was archived on August 21, 2026. GitHub shows the archive banner, but the README itself carries no notice, so the project reads as live to anyone who arrives from a tutorial. Stanford's STORM had 146 pull requests opened in 2026 and merged exactly zero of them; its last commit was September 2025, and its documented quickstart depends on the Bing Search API, which was retired in August 2025. The STORM papers are still worth reading. The software should not be your starting point.
Among the best-known projects, the pattern is that the open-weight agents which posted competitive benchmark scores have gone quiet, while the ones still shipping code publish almost no benchmarks. Pick for maintenance, then measure yourself.
If your research is bounded to documents you already have rather than the open web, that's a different tool category with different winners, which we cover in NotebookLM in 2026.
Deep Research APIs: What You Can Actually Call
This is the question the comparison articles keep skipping, and the one developers keep searching for: which deep research agents can you call from your own code, how much do they cost, and how good are the citations? Most of this changed through 2026, and one of the biggest changes lands on September 27, 2026.
| Provider | Callable deep research? | What you call | Price |
|---|---|---|---|
| Google Gemini | Yes, turnkey | deep-research-preview-04-2026 or deep-research-max-preview-04-2026 on the Interactions API | ~$1 to $3 per task; Max ~$3 to $7 per task (preview rates) |
| Perplexity | Yes, but migrating | Agent API POST /v1/agent with effort presets; sonar-deep-research ends support Sept 27, 2026 | Legacy Sonar: $2/1M in, $8/1M out, plus $2/1M citation and $3/1M reasoning tokens, plus $5 per 1,000 searches |
| OpenAI | Not as a dedicated model | Flagship model on the Responses API with web_search and high reasoning effort, or the new Agents API | Model token rates plus web_search at $10 per 1,000 calls |
| Anthropic Claude | No research-branded API | Messages API web_search tool, or Claude Managed Agents (beta) | $10 per 1,000 searches plus tokens; Managed Agents add $0.08 per session-hour |
| You.com | Yes | Research API, five effort levels | $12 (lite) to $1,200 (frontier) per 1,000 |
| Exa | Build-your-own | Search, Deep Search, and Agent endpoints | Search $7 per 1,000; Deep Search $12 to $15 per 1,000; Agent $0.012 to $1.00 per request at fixed effort |
| Tavily | Build-your-own | Search and Research endpoints | Free 1,000 credits/mo; PAYG $0.008/credit; Research runs 4 to 250 credits |
Four things stand out.
Google is now the frontier lab's turnkey option, and that's a reversal. For most of this category's life, Gemini Deep Research was a consumer feature you couldn't call programmatically. It's now the cleanest of the big four: you call a deep research model on the Interactions API, with background=true required rather than optional, and Google publishes both an estimated per-task price band and a real time budget. Enterprise access runs through the Gemini Enterprise Agent Platform, which replaced Vertex AI as the product name for this in 2026, so ignore older articles pointing you at Vertex.
Perplexity's turnkey API is being retired right now. sonar-deep-research was for two years the answer to "which deep research API can I just call," and Perplexity announced in July 2026 that Sonar Chat Completions endpoints retire on September 27, 2026. The replacement is the Agent API, where you POST /v1/agent and choose an effort preset. Perplexity maps the old deep research model to high and points at xhigh for its best results. Every Sonar price in the table above holds only until September 27, so if you're costing out a build, cost the Agent API. Note also that the lifecycle changed shape: the legacy async endpoint used uppercase statuses like IN_PROGRESS while the Agent API uses a different lowercase set, which means a polling loop written against one fails quietly against the other.
OpenAI moved deep research into the model line, then into a managed agent. The dedicated o3-deep-research and o4-mini-deep-research models were deprecated in April 2026 and shut down on July 23, 2026. Their replacement path is a current flagship model on the Responses API with the web_search tool in background mode. As of September 10, 2026 there's also a managed Agents API in public beta, where OpenAI runs the session orchestration, context compaction, and recovery for you. One trap worth flagging: OpenAI's own deep research guide page and its o3-deep-research model page are both stale and still present those models as available, contradicting the deprecations page. Trust the deprecations page.
Anthropic no longer makes you build it all yourself. The old framing, which this guide used to print, was that Claude has no research API and you assemble your own loop. That's out of date. Claude Managed Agents has been in public beta since April 2026: long-running sessions, server-sent events, steer and interrupt, cron-scheduled deployments, and built-in web search, web fetch, and MCP, billed at token cost plus $0.08 per session-hour. Anthropic's own worked example puts a one-hour Opus session at well under a dollar, and web fetch, unlike web search, costs nothing. There's still no model called "deep research," but the harness exists.
For the "best API for research reports with citations" question specifically, three approaches work. Google and Perplexity give you citations out of the box. You.com, Exa, and Tavily give you source-attributed results that you feed into your own model, which is the pattern most production research features are actually built on, because you control cost and citation formatting. And any model with a web search tool can cite as it goes if you prompt it to. Teams that need audit trails tend to prefer the build-your-own route, because every claim traces to a URL you fetched rather than to a model's memory. For why grounding and retrieval quality matter more than raw context size, see Context Rot: Why RAG Still Beats a Giant Context Window.
Two ownership notes for anyone picking a vendor to depend on. Tavily was acquired by the AI cloud provider Nebius in February 2026, so the "independent search API" framing needs an asterisk. And Exa removed its /research endpoint on April 1, 2026, folding it into /search with a deep-reasoning type, which means any tutorial telling you to call exa-research is broken.
How to Automate Deep Research With an API
Calling one of these APIs is not like calling a chat completion. Runs take minutes to an hour, so every provider makes you deal with asynchrony, and they all do it differently. If you're wiring a research agent into a product or a pipeline, this is the part that costs you a day.
Assume background execution, not a blocking call. Google requires it outright: the deep research models only work with background=true, and you poll the interaction or resume a stream. OpenAI supports background mode on the Responses API with polling, resumable streaming, and webhooks on the Standard Webhooks spec. Perplexity documents polling and no webhooks. xAI is the outlier. background appears in its schema but is marked unsupported, so the real async path there is the batch API, which returns within 24 hours rather than minutes.
Budget by task, not by token, and check what the search costs. Search surcharges are where research API bills come from. They're priced per call, not per token. OpenAI charges $10 per 1,000 web search calls, Anthropic $10 per 1,000 searches, Perplexity $5 per 1,000 on the retiring Sonar deep research line, and xAI $5 per 1,000 tool calls. Google gives paid-tier projects 5,000 free grounding searches a month across its Gemini 3.x models and then charges $14 per 1,000, and Deep Research tasks draw on that same pool at roughly 80 to 160 searches each, so about 60 tasks exhausts it. Perplexity's own sample response is the most useful worked example published: a single deep research call costing $0.816, of which 71% was reasoning tokens and only 11% was output. The report you read is the cheapest part of the invoice.
Expect effort tiers, because they replaced model IDs. This is the structural shift hiding behind every individual change above. OpenAI retired its research models in favor of a flagship plus reasoning effort. Perplexity is retiring Sonar in favor of Agent API presets. Google ships standard and Max variants. xAI makes you pick 4 or 16 agents through reasoning.effort. Exa, You.com, and Tavily all price the same shape of ladder, from a cheap floor to a per-request ceiling. If you're designing around these APIs, design for a dial you can turn per request, not a model name you hardcode.
The clarification step is yours to build. OpenAI says this explicitly: deep research through the API includes no clarification or prompt-rewriting step, and the model expects a fully formed prompt up front. The consumer apps ask you follow-up questions before they start; the APIs don't. The standard fix, which OpenAI's own cookbook demonstrates, is a small cheap model that triages the request and rewrites it into a research instruction before the expensive agent ever runs.
For fact-checking specifically, verification is a pipeline stage rather than a prompt. A pattern has settled across both OpenAI's and Anthropic's published cookbooks: decompose the output into standalone claims, verify each claim in its own clean context, then have a separate skeptic pass challenge the verdicts. Anthropic's version makes the case for why bluntly, noting that "double-check your findings" is also just an instruction, and under context pressure it gets skipped. If you're building the real-time fact checker that people keep searching for, that decomposition is the architecture, and the search API underneath it matters less than the structure around it.
Deep Research Benchmarks: What the Leaderboards Show
Three numbers get quoted most: Humanity's Last Exam, BrowseComp, and SimpleQA. They're useful, they're widely misquoted, and they carry two structural problems that almost nobody mentions.
Start with the first. No independent leaderboard ranks the shipped research agents against each other with tools. Scale's own evaluation runs no tools variant, the vendor boards are self-reported, and the one serious third-party harness benchmarks model-plus-search configurations you can't buy rather than products you can. When you see a vendor's agent scoring in the fifties or sixties "with tools," you're reading that vendor's own harness, search budget, and judge. That doesn't make the numbers fake. It makes them incomparable.
The second problem is tier labeling. This guide previously printed 54.6% HLE and 85.9% BrowseComp as "Gemini Deep Research." Both figures belong to Deep Research Max, Google's higher-effort tier, against 50.4% and 61.9% for the standard tier. That's a 24-point gap on BrowseComp between two products sold under one name, and comparison articles collapse them constantly. Worth knowing that Google published these in a chart in its April 2026 announcement rather than as text you can quote, which is part of why they get garbled downstream.
| Agent or configuration | HLE (with tools) | Reported by | When |
|---|---|---|---|
| Perplexity Sonar Deep Research | 21.1% | Vendor | Feb 2025 |
| OpenAI Deep Research (o3) | 26.6% | Vendor | Feb 2025 |
| Gemini Deep Research (3 Pro) | 46.4% | Vendor | Dec 2025 |
| Gemini Deep Research (3.1 Pro) | 50.4% | Vendor | Apr 2026 |
| Gemini Deep Research Max (3.1 Pro) | 54.6% | Vendor | Apr 2026 |
| Kimi K3 (the model, not the Deep Research agent) | 56.0% | Vendor | Jul 2026 |
| Claude Opus 5 + Perplexity search, 25 turns | 77.4% | Third party | Aug 2026 |
That last row is the one worth sitting with. It comes from OpenRouter's own benchmark harness, which runs production endpoints rather than collecting vendor claims, and publishes cost and latency alongside accuracy: $0.46 per question, 83 seconds, on an 84-question sample. It isn't a product you can buy. It's a model plus a search provider plus a turn budget, assembled by a third party, and it beats every shipped agent's self-reported number. One caveat on who's doing the measuring: OpenRouter resells most of the models and search engines it benchmarks.
The single most useful finding in the whole category comes from that same harness. On BrowseComp, one configuration scored 35.8% with a 1-turn search limit and 89.0% with a 25-turn limit, in the same run series. Same model, same benchmark. The scaffolding around it, specifically how many searches it's allowed, moved the result by 53 points. This is why a plain browsing model scores under 2% on BrowseComp while an agent built on it scores in the eighties, and it's why "which model is smartest" is close to the wrong question. Buy search budget.
The cost curve is flatter than the marketing implies, too. In the same runs, the top BrowseComp configuration cost $0.99 per question, while a cheap model paired with the same search provider hit 77.0% at $0.076, roughly 13 times cheaper for 12 points less accuracy. For most real work, that trade is obviously correct.
The benchmarks built for research agents
HLE and BrowseComp measure whether an agent can find a hard fact. They don't measure whether the report is any good. Three newer benchmarks do, and they're far less flattering.
ResearchQA, an independent academic effort spanning 21,000 queries and 160,000 rubric items across 75 fields, validated by 31 PhD annotators in eight of them, found that not one of the plain language models or retrieval systems it tested cleared 70% rubric coverage. Only the purpose-built deep research agents did better, and the best of them stopped at 75.3%. More pointedly, that top system fully addressed under 11% of citation rubric items and roughly half of the limitation and comparison items. Agents find sources. They're much worse at stating what the sources don't establish.
ResearchRubrics, built on 2,800-plus hours of expert labor, put both Gemini Deep Research and OpenAI Deep Research under 68% average rubric compliance. DeepResearch Bench II, with 9,430 binary rubrics distilled from expert investigative articles, reported in January 2026 that even the strongest models satisfied fewer than half of them. Its live leaderboard has since climbed past that mark, which is the usual trajectory.
Two more cautions before you use any of these numbers to pick a tool. Which model grades the rubric moves scores by more than ten points, and it can reorder the top two, so any table mixing judges is telling you about judges rather than agents. And agents browsing the live web can retrieve a benchmark's own answers: one 2026 study measured up to 4% inflation from this. Google's published evaluation methodology now blocks sites like Hugging Face during testing, which tells you the problem is real enough to engineer around.
SimpleQA deserves a footnote rather than a headline now. Perplexity's 93.9% is genuine, but it dates to February 2025 and describes a Sonar-based product that, as covered above, no longer exists in that form.
The most durable independent finding is also the oldest. The Tow Center at Columbia Journalism Review tested eight AI search engines across 1,600 queries and found they failed to retrieve correct information more than 60% of the time, with error rates ranging from 37% to 94% depending on the engine. That study is from March 2025 and the Tow Center has not published a repeat at that scale, which is its own commentary on how thin independent evaluation remains.
The honest read: benchmarks rank tools reliably at the extreme high end of difficulty and badly below it. Run the same real prompt through two tools on their cheapest paid tier and compare. Benchmarks are suggestive. Your own test is decisive. For a longer argument about why benchmark obsession misleads, see The AI Thinking Trap.
Free Tier Reality Check
The marketing pages all highlight free access. Here's what "free" actually means in September 2026, and the honest answer starts with an admission: the numbers got worse to find, not better.
OpenAI is the only one with a free allowance, and it stopped saying how big. The "5 free per month, 25 on Plus, 250 on Pro" figures still circulating come from an April 2025 blog update. The current pricing page says Limited on Free and on the $8/month Go tier, Expanded on Plus, and Maximum on Pro, and the help center confirms an in-product counter that resets every 30 days without saying what it starts at. The only current published number sits on the business rate card, where one deep research task costs 50 credits. If you need a predictable quota, this is the least predictable vendor.
Perplexity moved Deep Research behind the paywall. Free accounts get search. Deep Research is a Pro feature, and since it now runs inside Perplexity Computer it's Pro and Max only. This guide used to print "5 per day" for the free tier, which was a 2025 fact about a product that has been rebuilt since.
Gemini's cheap tier is the real story, not a free one. Google publishes plan allowances as multipliers rather than counts, with limits refreshing every five hours, and its own plan pages put Deep Research on the paid tiers. What makes Google the value pick isn't a free allowance, it's that AI Plus costs $4.99/month, well under everyone else's entry price.
Claude has no free Research, but it isn't Pro-only either. Research is unavailable on the free plan and available on every paid plan: Pro at $20, Max at $100 or $200, plus Team and Enterprise. Earlier versions of this guide described it as Pro-only, which understated who gets it.
The practical read: there's no longer a deep research tool you can lean on for real work without paying. If you only pay for one, Gemini's AI Plus at $4.99 is the cheapest route to a benchmark-leading agent, and Perplexity Pro at $20 gives the best volume for the money. If you want the deepest output on ambiguous questions, ChatGPT Plus at $20 buys OpenAI Deep Research plus everything else in the bundle. Claude Pro makes most sense if you already use Claude for reading and writing and want one subscription.
Which Tool for Which Job
After running hundreds of queries across all of these, clear patterns emerge. Here's how I'd route work now.
Academic literature review. Claude Research. The 1M-token context matters when the agent needs to hold 20-plus papers at once, and Claude is noticeably better at distinguishing superficially similar claims. Runs take longer, but literature reviews aren't time-sensitive.
Market sizing and competitive intelligence. OpenAI Deep Research. The depth of reasoning on ambiguous strategic questions comes through clearly here. It's the one I trust most for "help me understand this industry" prompts.
Hardest factual retrieval. Gemini Deep Research Max. If the benchmark leaders tell you anything, it's that Google's high-effort tier is the strongest shipped agent at finding obscure, multi-hop facts. Use the Max tier deliberately, since the standard tier is a meaningfully different product.
Quick briefings before a meeting. Perplexity. Reports are shorter and more encyclopedic, which is what you want when you need orientation rather than an essay.
Research involving your own documents. Gemini Deep Research. The Workspace integration is the moat. If half your source material is in Drive, Gmail, and old meeting notes, nothing else compares.
Developer integrations and bulk runs. Google's Deep Research API is now the turnkey pick, since Perplexity's Sonar line retires September 27, 2026 and OpenAI has no dedicated research model. If you want control over citation formatting and cost, build the loop yourself on Exa, Tavily, or You.com's Research API.
Cheap, high-volume research inside a product. A mid-tier model paired with a good search provider and a generous turn budget. The benchmark evidence above says this gets you most of the accuracy at a fraction of the cost, and it's the configuration most production features actually ship.
Synthesizing contradictory evidence. Claude. When sources disagree, Claude is the most willing to surface the disagreement rather than pick a side prematurely.
One pattern that might surprise people: no single tool dominates. I run the same prompt through two agents for high-stakes work. The cost is about $40/month for two subscriptions, and the output is noticeably better than either alone. If you're deciding how AI fits your study or work routine more broadly, AI study modes compared covers the adjacent question.
The Missing Piece: Turning Research Reports Into Usable Knowledge
Here's what almost no comparison article mentions. The report the agent produces is raw material, and the research isn't finished until you've read it properly.
A 20-page Claude Research output or a 15-page OpenAI Deep Research report is the start of the work, not the end. Read it once, skim the conclusion, close the tab, and you've paid an agent to summarize something you didn't actually learn. A 2025 MIT Media Lab study (tracked in our analysis of AI's impact on learning) measured EEG activity while people wrote essays with and without ChatGPT, and found markedly lower cognitive engagement in the assisted group. It was a small preprint about writing rather than reading, but the direction it points is the one every heavy AI user recognizes.
The fix is what researchers have done for centuries: annotate. Highlight the claims that matter. Flag the sources you want to verify. Link insights across reports.
This is where Glasp's web highlighter fits into the workflow. Run your research on OpenAI, Perplexity, Gemini, or Claude. Paste the report into a readable page. Highlight directly in the browser as you read. Your highlights sync to your Glasp library, searchable and organized, alongside everything else you've read that month.
A few specific workflows that work:
Highlight, then re-query. Read the report, highlight the 10-15 claims that matter most. Paste those highlights back into the same agent with "dig deeper on these specific points." Iterative rather than one-shot.
Stack reports by topic. When you research the same topic across two tools (say, OpenAI and Claude), highlighting both reports in Glasp lets you see where they converge and diverge. Disagreements are often the most interesting parts.
Use YouTube alongside text. When the best sources are podcasts or talks, YouTube Summary gives you transcript-level summaries with timestamps. Pairing a text deep research report with 3-4 annotated YouTube talks covers a topic more thoroughly than either alone.
Chat with your highlights. Glasp's AI chat can answer questions using your annotations as the source. It's the difference between "what did the agent say about X?" and "what have I actually concluded about X?"
Publish what you learned. The community on Glasp is full of other people researching similar topics. Sharing highlighted reports is a forcing function to finish the research, not just queue more of it. For a step-by-step guide, see How to Annotate Articles the Right Way.
A report you read once is a receipt, not knowledge. The highlight-and-annotate step is what converts agent output into something you actually know.
Frequently Asked Questions
Which deep research tool is the most accurate?
Among the big four, Google's Gemini Deep Research Max leads at 54.6% on Humanity's Last Exam (Google, April 2026). Be careful with that name: the standard Deep Research tier scored 50.4% on the same test, and most comparison articles quote the Max number under the plain product name. Moonshot's Kimi K3 claims a higher figure still, at 56.0%. All of these are vendor self-reports, and in third-party testing a model-plus-search configuration nobody sells outscored every shipped agent. In practical use, accuracy differences between the top tools are smaller than leaderboards suggest.
What is the best ChatGPT deep research alternative?
For most people, Gemini Deep Research, which has the best published benchmark results and the cheapest paid entry at $4.99/month, or Perplexity, which is broader and faster. If you want something outside the big four: Moonshot's Kimi has a well-documented Deep Research mode, Mistral's is the cheapest of the challengers at $14.99/month, and xAI's Grok offers a multi-agent mode. If you want to self-host, GPT Researcher and ByteDance's DeerFlow are the actively maintained open-source options. Avoid LangChain's open_deep_research, archived in August 2026, and Stanford's STORM, which hasn't merged a pull request in over a year.
What is the best deep research API?
Google's Deep Research models on the Interactions API, for most teams. It's the only turnkey research API from a frontier lab that isn't being retired, it publishes a per-task price estimate of roughly $1 to $3 (or $3 to $7 for Max), and it documents a real time budget. Perplexity's Agent API is the natural landing spot if you already use Sonar, since sonar-deep-research ends support on September 27, 2026. If you want control over citation formatting and cost, build the loop yourself on Exa, Tavily, or You.com rather than buying a finished report.
What is the fastest deep research API?
Google is the fastest to document: most Deep Research tasks finish within 20 minutes, capped at 60. Nobody else publishes a current latency figure, which is itself the answer. OpenAI says only "tens of minutes," and Perplexity's 2 to 4 minutes dates to its February 2025 launch. If you need speed, the lever is effort, not vendor: every one of these APIs now exposes a tier dial, and the low tiers return in seconds to a couple of minutes. For raw sourced results in seconds to feed your own model, Exa, Tavily, and You.com's lower effort levels are the fast path.
How long do deep research runs take?
Using only figures the vendors publish: Google says most tasks finish within 20 minutes and caps runs at 60. Moonshot says Kimi runs 10 to 25 minutes. OpenAI's launch page still says 5 to 30 minutes, a sentence it hasn't revised since February 2025, and Perplexity's 2 to 4 minutes comes from that same month. Anthropic says only that Research answers arrive "in minutes." Treat the older numbers as historical, because all three products have been rebuilt since.
Can I use deep research tools via API?
Yes, and the options changed substantially in 2026. Google's Deep Research models on the Interactions API are the turnkey option, priced at roughly $1 to $3 per task and $3 to $7 for the Max tier. Perplexity's sonar-deep-research was the developer favorite but ends support on September 27, 2026, replaced by Agent API effort presets. OpenAI retired its dedicated o3-deep-research and o4-mini-deep-research models on July 23, 2026, so you now run a flagship model on the Responses API with the web search tool, or use its Agents API, in public beta since September 2026. Anthropic has no research-branded API but does offer Claude Managed Agents in beta, billed at tokens plus $0.08 per session-hour. Exa, Tavily, and You.com are the popular build-your-own options.
Which deep research API has the best citations?
Google and Perplexity return fully cited reports out of the box. For production systems that need every claim traced to a URL, teams usually build on Exa, Tavily, or You.com, which return source-attributed results you feed into your own model, so you control formatting and can show an audit trail. Any model with a web search tool can cite inline if you prompt it to. Bear in mind the independent research on this: the best evaluated system fully addressed under 11% of citation rubric items, so citations being present is not the same as citations being sufficient.
Is there a free deep research tool?
Barely, and it got worse in 2026. OpenAI is the only one of the four that still gives free accounts a deep research allowance, described as "Limited" with no published count. Perplexity moved Deep Research behind Pro, Claude has never offered it free, and Google puts it on paid plans. For regular work the realistic entry points are Gemini AI Plus at $4.99 or a $20 plan from any of the big three. If you want something genuinely free and don't mind running it yourself, the open-source agents covered above are the honest answer.
How do I stop hallucinations in deep research reports?
Three practical tactics. First, click at least the top three to five cited sources and confirm the claim is actually in the source, since mis-citing a real source is far more common than inventing a fake one. Second, run the same prompt through a second tool; disagreements are usually where one of them went wrong. Third, if you're building this into software, make verification a separate pipeline stage rather than a line in your prompt: extract each claim, verify it in a clean context, then have a skeptic pass challenge the verdicts. That's the pattern both OpenAI's and Anthropic's published cookbooks converged on.
Can deep research tools read my private documents?
Gemini Deep Research has the deepest integration, with native access to your Gmail, Drive, and Docs with permission. Claude supports Google Workspace connectors. OpenAI Deep Research can read files you upload during a session but doesn't integrate directly with cloud storage. Perplexity primarily works against the web. If your source material is largely in Google Workspace, Gemini is the obvious pick.
What's the best way to save and reuse deep research reports?
Export the report as PDF or Markdown, open it in a readable view, and highlight it like you would any long article. Glasp is built for exactly this workflow: highlights sync to a library you can search, link to other highlights, and revisit. Without a highlighting step, most deep research reports get read once and forgotten. This connects to what educators call the generation effect: information you process actively is retained far better than information you passively receive.
Conclusion: The Research Stack, Not the Research Tool
Nineteen months after OpenAI's launch, the category has clarified and then started reorganizing itself. Deep research agents aren't a winner-take-all market. They're a mix where the right answer depends on what you're researching, how much time you have, where your source material lives, and whether you're clicking a button or calling an API.
If I had to pick one for most knowledge workers right now, it's Gemini: the best published benchmark results, the cheapest paid entry at $4.99, and the only frontier-lab research API that isn't on a retirement clock. For heavier or more ambiguous work, pair it with OpenAI Deep Research or Claude Research. Developers migrating off Sonar should look at Perplexity's Agent API before September 27.
The deeper change is structural. The dedicated deep research model is on its way out at every lab: OpenAI retired its research models, Perplexity is retiring Sonar, Google ships an agent rather than a model, Anthropic ships a managed harness, and xAI makes you choose an agent count. What replaced them is an effort dial. That's good news for anyone building, because effort is a parameter you can tune per request, and it's why the most useful benchmark finding of 2026 was that raising a search budget moved accuracy by 53 points while the model stayed the same.
But the tool choice matters less than what you do with the output. The biggest mistake I see people make is treating a deep research report as finished work. It isn't. It's raw material. The actual knowledge gets built when you highlight the claims that matter, link them to other things you've read, and return to them later when the topic comes up again.
That's the workflow Glasp is designed for. Highlight any report, any article, any YouTube transcript. Build a searchable library of what you actually thought was important. Chat with your highlights later when you need to recall something specific. Share your work with others doing the same research.
The deep research agents will keep getting better, and the APIs will keep multiplying. The ones that don't also get a highlighting layer on top will keep producing reports that get read once and forgotten. Don't build your 2026 research workflow around a single tool. Build it around a stack, and make sure the last link in that stack is the one where your own understanding gets recorded.
Start by running one real research question through two of the four tools this week. Highlight both reports. Compare what you learned. That's the workflow. Everything else is a feature list.