Research

AI Search Benchmarks May Be Measuring Memory More Than Search

LiveBrowseComp finds that static browsing benchmarks can reward models for verifying what they already know. Its live, recent questions expose a gap between memory-backed verification and evidence-led search.

Research feature image

A search agent can look busy without learning very much from search. It can issue queries, open pages and cite sources while still relying on a hypothesis it already formed from knowledge inside the model. A new benchmark study argues that this distinction matters because current browsing leaderboards can reward both behaviors as if they were the same capability.

LiveBrowseComp, from researchers at Harbin Institute of Technology and Xiaohongshu, asks a deceptively simple question: when an AI agent scores well on a web-search benchmark, did it actually discover the answer through external evidence, or did it already know enough to guess the answer and use the web mainly for confirmation?

The paper's answer is not that search agents are fake or that browsing tools do not help. It is narrower and more important for evaluation: static benchmarks can mix two capabilities that should be measured separately — intrinsic knowledge coverage and evidence-led search.

A search benchmark can be partly solved before search begins

The researchers begin by removing search tools entirely. Across BrowseComp, BrowseComp-ZH, Humanity's Last Exam and GAIA, closed-book pass@4 averages 38.9 across 24 model-benchmark pairs. On BrowseComp specifically, MiniMax M2.5 reaches 44.5 without retrieval. On BrowseComp-ZH, Kimi K2.6 reaches 62.0.

Closed-book success does not prove that benchmark questions were copied into training data. The authors explicitly distinguish their concern from ordinary benchmark contamination. A model can know enough surrounding facts from broad pretraining to infer an answer even if it never saw the exact question.

That creates what the paper calls Intrinsic Knowledge Dependence, or IKD. A high browsing score may reflect a combination of knowing a plausible answer in advance and then using search to verify it. If so, the final number says less about the agent's ability to discover information outside its existing knowledge boundary.

Removing the right evidence makes search actively harmful

The strongest diagnostic is not the closed-book test. It is what happens when agents are allowed to search but the retrieval environment is stripped of answer-supporting documents.

Using BrowseComp-Plus, the researchers build a controlled dense-retrieval index and remove its evidence and gold documents, leaving only irrelevant material and hard negatives. Every tested model then performs worse than it did with no search tools at all. Average pass@4 falls from 26.1 in the closed-book condition to 6.2 with evidence-blocked search. MiniMax M2.5 drops from 44.5 to 8.0; Kimi K2.6 falls from 25.5 to 2.3.

The result suggests that current agents do not always treat retrieval as an independent evidence-discovery process. When the web cannot confirm the model's prior hypothesis, misleading but superficially relevant results can pull the search trajectory away from an answer the model could previously produce on its own.

The search loop is often model-led rather than evidence-led

Trajectory analysis explains why. The researchers trace where the key information behind each query first appeared: in the model's own reasoning or in retrieved material. For every model they analyze, more than half of queries are model-originated, and that share rises above 60% in later browsing rounds.

Even retrieving the right evidence does not guarantee that the agent will use it. In the analyzed trajectories, the evidence-use rate stays below one-third: 32.2% for DeepSeek v3.2, 24.7% for GLM-5.1, 30.8% for MiniMax M2.5 and 31.5% for Kimi K2.5.

That matters for products marketed as deep research. Tool access alone is not the capability. A reliable search agent has to let retrieved evidence change its working hypothesis, abandon dead ends and distinguish a failed search from evidence that its initial guess was wrong.

LiveBrowseComp tries to remove the memory shortcut

The benchmark is designed to place questions outside the model's likely intrinsic knowledge coverage. It contains 335 human-authored questions built from recent facts across continuously updated sources including GDELT, TMDB, RAWG, CVE/NVD, TheSportsDB and USGS.

At least one clue in each question must depend on information produced within the previous 90 days. The construction pipeline also filters out globally salient events, removes facts whose answers are still changing, and requires multi-step web research rather than a direct lookup.

The human validation is important because recent questions could simply be harder. The study reports nearly identical human solve rates on BrowseComp and LiveBrowseComp — 30% versus 31% — with similar completion-time distributions. That gives the authors evidence that the large model drop is not merely a consequence of making the questions harder for everyone.

Once memory is neutralized, the leaderboard changes

Without tools, every evaluated model scores below 2% on LiveBrowseComp. With search enabled, avg@4 ranges from 28.0 for MiniMax M2.5 to 43.2 for GPT-5.4. The same models score between 51.4 and 77.3 on BrowseComp.

More revealing is the ranking shift. GLM 5.1 scores 68.0 on BrowseComp but 33.9 on LiveBrowseComp. DeepSeek v3.2 scores only 51.4 on BrowseComp yet reaches 37.6 on LiveBrowseComp, overtaking several models that ranked above it on the static benchmark. The Pearson correlation between BrowseComp and LiveBrowseComp falls to 0.53, compared with 0.79 between BrowseComp and BrowseComp-ZH.

The implication is not that BrowseComp is useless. It is that a static search benchmark can gradually become a different test as frontier models absorb more of the facts needed to solve it. The benchmark remains fixed while the amount of relevant information already encoded inside models keeps increasing.

A better search agent needs to learn from retrieval, not just cite it

LiveBrowseComp points to a broader evaluation problem for agentic AI. As models become more knowledgeable, simply attaching search tools and measuring final-answer accuracy makes it harder to tell whether the tool changed the answer or merely supplied supporting links for a conclusion the model already preferred.

For real research tasks, the harder cases are precisely the ones where the model does not already know what it is looking for: a new vulnerability, an obscure event, an unfamiliar company or a fact that changed after training. In those cases, the useful capability is not memory-backed verification. It is the ability to follow external evidence into a hypothesis the model could not have produced at the start.

The benchmark still has limits

The authors caution that the 90-day boundary is only a heuristic. Some information may leak or be announced earlier, and model training cutoffs differ. The evaluation also uses a single search backend, serper.dev, so part of measured performance can reflect the search index rather than the agent alone.

The benchmark is also expensive to maintain because its questions require professional annotation and multiple stages of human verification. And the evaluation uses a shared research scaffold rather than every vendor's full proprietary deep-research product; the authors note that their setup omits context-management techniques used by some production agents, which may lower absolute scores.

Those caveats do not remove the central measurement problem. A search system should be rewarded for finding what the underlying model did not already know. LiveBrowseComp's strongest contribution is making that boundary visible.

As AI products increasingly promise deep research rather than simple question answering, the important metric may no longer be whether an agent used search. It may be whether search caused the agent to know something it could not have known without it.

Sources and further reading

  1. LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know? - arXiv
  2. LiveBrowseComp paper (HTML) - arXiv
  3. LiveBrowseComp dataset - Hugging Face
  4. BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents - arXiv
  5. BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent - arXiv