Pith. sign in

Paper Citation Record · LEDGER

BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2504.12516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12516 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 100 of 221 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:47:14.793292Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cb9c2891-a70c-4fec-90e0-fc94261749eb · inbound

BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese cites this paper.

BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T22:04:49.966529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T22:04:49.915916Z digest=sha256:ea00dda395cbff543b8df0a935b9f313c531da87e83306c40d017cafb93953ff

Observation 8ae84d09-d6c1-4c32-90a0-2b398ade77d9 · inbound

MedBrowseComp: Benchmarking Medical Deep Research and Computer Use cites this paper.

MedBrowseComp: Benchmarking Medical Deep Research and Computer Use BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:50.706006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:30:50.706006Z digest=sha256:7c0a1549797cdf77c3c86d38bd62ccb6c515f3c5a8113bd915617817196b3f85

Observation 40e101cd-6d9b-4b0c-b069-2f46f3b24295 · inbound

InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation cites this paper.

InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:19.113302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:19.113302Z digest=sha256:1fc363ac34d823bee3dd7687273953555575287231fe62209938dfd8e2dcd399

Observation b480e41b-d137-4ead-aad7-5d2f0264e4f4 · inbound

Agent-Environment Alignment via Automated Interface Generation cites this paper.

Agent-Environment Alignment via Automated Interface Generation BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:33.912877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:33.912877Z digest=sha256:4592426b87715b0b1c1b9eb633cce67eebe1997b2cd04bd2d2160cd89c0a8f1e

Observation 7ceb12c9-7bcc-44c2-aaa6-b3a848a8f9c2 · inbound

EvolveSearch: An Iterative Self-Evolving Search Agent cites this paper.

EvolveSearch: An Iterative Self-Evolving Search Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.907430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.907430Z digest=sha256:04070860fdd51033b10ecc3b4462ebfb0221185feb4d616ea7a08f7fe66fc0cc

Observation b559645e-74b5-4ef2-b83d-2e0665baebae · inbound

WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks cites this paper.

WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:47.805483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:34:47.805483Z digest=sha256:2204c403482eaa2168819b6a65e859602cb232515f8da1c8c4f2885c8cb70135

Observation 0a077d01-0f0b-4170-aff1-153ca9cf2068 · inbound

Real-Time Execution of Action Chunking Flow Policies cites this paper.

Real-Time Execution of Action Chunking Flow Policies BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:18:51.715959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T14:18:51.613045Z digest=sha256:74271aaeb8b21021f0fc70da4797b1ac70b705d4961286f32db375b84d439ed8

Observation 11bc6327-6b0a-4785-9486-a84afdc122fc · inbound

TaskCraft: Automated Generation of Agentic Tasks cites this paper.

TaskCraft: Automated Generation of Agentic Tasks BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:19.960609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:19.960609Z digest=sha256:64e437b2488d5b8890ab4dfa7d88f00d7ce18bc5a12c60c369b31a121c28e251

Observation 7a5c3690-980a-4587-bf2d-dd0bebaa3480 · inbound

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents cites this paper.

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:07:39.517594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T08:07:39.384613Z digest=sha256:5bab29aab8c1b48304e7ac996c3cca4444838786eacd47e298bd2dba53c8ee0f

Observation 58efdf6a-6899-4dee-a80b-573e578ee62a · inbound

ScholarSearch: Benchmarking Scholar Searching Ability of LLMs cites this paper.

ScholarSearch: Benchmarking Scholar Searching Ability of LLMs BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:22.723274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:22.723274Z digest=sha256:33ea2a42985be4a42fbd608f665729b541148e695209d54f0ff683a643cbcc57

Observation fbeb1393-1d01-4ce8-b7ee-eda0c4b5eda1 · inbound

OAgents: An Empirical Study of Building Effective Agents cites this paper.

OAgents: An Empirical Study of Building Effective Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:50.210552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:50.210552Z digest=sha256:c280558170243486d98bfd60930073599ba4dac6e95d47e19ba77c2ae7fd67d7

Observation 825b5c51-08b5-4314-aa48-d0d0adfcc73c · inbound

From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents cites this paper.

From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T18:47:14.793292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:47:14.793292Z digest=sha256:cc0aac1e2548aa1481fa5d4d86b84f70ec9293361616a85f3bbbddf5d8b263a2

Observation 0d38bb0d-ff7b-4f2a-af0f-8d3c0b968969 · inbound

Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge cites this paper.

Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.935985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.935985Z digest=sha256:f93124fdcaeab5f8982928ca8906253b80dd23a8467216e5db3c8bdd3f4c7662

Observation f7735579-3806-44ba-b6a1-c829227839cd · inbound

WebSailor: Navigating Super-human Reasoning for Web Agent cites this paper.

WebSailor: Navigating Super-human Reasoning for Web Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:37:09.732186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T15:37:09.572241Z digest=sha256:c9c03e1ebd72474067438cd40ff39921d57b29f4019ea654cab9a435193e3f11

Observation 2107abac-dd79-4ba3-86ea-9410807c6e93 · inbound

WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization cites this paper.

WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:06.463820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:06.463820Z digest=sha256:7f702fcc9f051492d3e16887492d123e59417f7062e7ad85f544cbb3aad2168c

Observation 7c7ef63f-33e8-4f12-84d9-a5c67d0c1b52 · inbound

ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry cites this paper.

ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:53.479271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:17:53.479271Z digest=sha256:c3b0300550a847c24d6abcd5dc6cccdc00746dae1c4f4ead0abb04e1f5be5db1

Observation a34f4c58-cc1c-459a-90d5-13e7d529220e · inbound

RAVine: Reality-Aligned Evaluation for Agentic Search cites this paper.

RAVine: Reality-Aligned Evaluation for Agentic Search BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:08:33.146164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:08:33.146164Z digest=sha256:6f6c9be50e2e31dec0b2b37b1797e1ca4e806c7480cf32fe3925728754c87f41

Observation d034cd4b-c2be-4c45-a573-1555235ad47d · inbound

Beyond Context Limits: Subconscious Threads for Long-Horizon Reasoning cites this paper.

Beyond Context Limits: Subconscious Threads for Long-Horizon Reasoning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:18.588483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:18.588483Z digest=sha256:c8b6ebf82112c0f15892df5b458bdbe2a6346eb5f06902543470ef78bb0bd5fe

Observation 3f35406d-e002-4523-83b1-ffb6137c12fa · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 75

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T22:23:15.670194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:d1d2c49148a3d11fac4c356ba8d9eb62763fed9a635269637b9e194c9d25d9c5

Observation e87c215b-c00d-4e70-8414-029689052a61 · inbound

How Far Are AI Scientists from Changing the World? cites this paper.

How Far Are AI Scientists from Changing the World? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 176

Resolution
unresolved
no resolver link, observed 2026-08-06T10:55:15.091176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:55:15.091176Z digest=sha256:22cf9ba459cf18ab16ea6b23d866c7aac6e94b28fc4929e17df3a59be8f3167c

Observation a53f149a-7e84-4e6d-a52a-cb9a4b22006b · inbound

TextQuests: How Good are LLMs at Text-Based Video Games? cites this paper.

TextQuests: How Good are LLMs at Text-Based Video Games? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.679463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.679463Z digest=sha256:df4cf086056a27259b837667fc6f9ab1b5fef9ba96efadc17e16473d6cb45d79

Observation 50778f16-28cb-4ecb-b74d-e786c5d15fd7 · inbound

MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning cites this paper.

MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T10:17:13.681710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:17:13.681710Z digest=sha256:bd8bca40cffd577af858d241d1fbebfd45e57139c52734708c04a0b125b312e4

Observation 6318c7e8-e969-43a5-8ba5-90c57f2b68d7 · inbound

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts cites this paper.

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:44.134958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:44.134958Z digest=sha256:255614f2e2fd0e1f6d1cf0bdd310bd8305daf4c44e7709dcd0a580262478a650

Observation 1f1b34e4-deb5-4000-b553-9b6971b139b1 · inbound

Efficient Agents: Building Effective Agents While Reducing Cost cites this paper.

Efficient Agents: Building Effective Agents While Reducing Cost BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:58.258010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:58.258010Z digest=sha256:dc0b24003294a75fc5e58a2dc3c094ef46432d892400c998bf70df77b91b2c1d

Observation bb0a689d-f746-4cfe-ab35-ab456294fade · inbound

Characterizing Deep Research: A Benchmark and Formal Definition cites this paper.

Characterizing Deep Research: A Benchmark and Formal Definition BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T00:56:06.248077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:56:06.248077Z digest=sha256:03a9bd7700aa69850da320394300d89a27d11e329e83d6fd4633c3b6103147fb

Observation 90c1f86a-9076-421d-9b75-43c0b280e8b1 · inbound

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent cites this paper.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T18:56:24.021133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:7bc39da21277c7bf6fe840645ef1a689f45570bacc984a73cfa73957ce031765

Observation 402d821e-d025-4f7e-9f7c-15a470d63a1d · inbound

Observation of momentum dependent charge density wave gap in EuTe4 cites this paper.

Observation of momentum dependent charge density wave gap in EuTe4 BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T22:45:16.591776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:45:16.591776Z digest=sha256:585992930218707b785011b93e902220d142e99dbb15033e0a3ada4e880e40fb

Observation 019096d1-88d5-4ada-85c1-f298425b77cc · inbound

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models cites this paper.

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T17:50:08.399160Z digest=sha256:f0d207480e5dbf9d51bc835894bba6552cc74611c4d9aa2bb2f823182c0afa58

Observation d0c2b477-bdd0-45ab-8e41-79f46b29fcd4 · inbound

BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent cites this paper.

BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T22:46:12.460911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:46:12.460911Z digest=sha256:bf561178b031d433ea7ee89ff19254bdcb297ec7e018aa7b5231eb4935d5cd7e

Observation 3921d393-b249-445a-9c10-f426ec02bdeb · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 102

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T23:21:42.179576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:c1f3b25778152bc7c06c0dcd0975e98f614091f0239c3c24b02b1a37d6809a03

Observation 35d4c42a-e2a4-4727-8889-b87742c9e5e4 · inbound

Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty cites this paper.

Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:36:49.150437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:36:49.150437Z digest=sha256:1873c52813caf6a0440e3242970512e9fc33c47e73400d0fc57a5ec5eaaf81b0

Observation 65abb7e1-0218-4aa4-80d4-013c5f0659fc · inbound

SSRL: Self-Search Reinforcement Learning cites this paper.

SSRL: Self-Search Reinforcement Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.008368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.008368Z digest=sha256:e893bbc8f8316e877f639d435a9e35659ff9b50bcc0d9df3c5624ff12245d007

Observation e36dc413-0095-4501-8c10-8d271cbf4040 · inbound

FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction cites this paper.

FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:31:13.199520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:31:13.199520Z digest=sha256:dd78a1f246d9bba0f105fb15d8ea89e8d3d7fca1ada95ae14059980c53a6eb02

Observation e8ed3d6f-8058-46c5-8684-edc8330fa21e · inbound

Exploring Spatial-Temporal Dynamics in Event-based Facial Micro-Expression Analysis cites this paper.

Exploring Spatial-Temporal Dynamics in Event-based Facial Micro-Expression Analysis BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:33:37.253946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:33:37.253946Z digest=sha256:2993b3b4e8a5139d0b3fe629b5372975f2663e47eeae31e069ff5e493f006f75

Observation aa2d15bd-0624-4483-af3c-7557f484ea49 · inbound

WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents cites this paper.

WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T22:31:52.850992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T22:31:19.752228Z digest=sha256:e20e776de1e18aa932ba1fab8ba3c9cb4d6e4ff010e014626b92b9f8cc03f73b

Observation c4d45d17-24ac-4d7b-9ab3-b79be03868be · inbound

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL cites this paper.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.713967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.713967Z digest=sha256:cae013f68b1b191fc9cd10881a92e4db8703025bbf9ab72e78945029166a3f7c

Observation f6c93b97-821d-4ced-b1ff-6ea8bb2bd3cd · inbound

Search-Time Data Contamination cites this paper.

Search-Time Data Contamination BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T21:08:06.123994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:08:06.123994Z digest=sha256:55c6061b0bc098d203f9193a9b787af291f98d43ea17704352b22167e3d64e6e

Observation 9d8fe418-16a1-4b2b-930b-8e26a4138302 · inbound

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents cites this paper.

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:18:58.349502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:18:58.349502Z digest=sha256:762d58bed5ab678e554c7c75eed43b0715a132a84fedb87bb6d2041d6f84cb99

Observation 18a23df0-5f8f-4c0a-9cbf-4f3bcb23d3b6 · inbound

UQ: Assessing Language Models on Unsolved Questions cites this paper.

UQ: Assessing Language Models on Unsolved Questions BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.844366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.844366Z digest=sha256:3559144ad4f677445ffab04b918e41fcccc7904d318bb5a7b83c8f6e672f7ad1

Observation 5c94c0c2-c0ea-407f-a219-8f4a9a2dea81 · inbound

Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning cites this paper.

Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:11.436613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:00:11.436613Z digest=sha256:dd6bbfea050fb19f6fed530943f0fa6059b60f1d1ada2e9a3f21f370377b7e6b

Observation 488771f2-8483-4d15-b829-1b80f323ac68 · inbound

Open Data Synthesis For Deep Research cites this paper.

Open Data Synthesis For Deep Research BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T13:45:13.108221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:45:13.108221Z digest=sha256:e15abd14b2bd8a6e36ab51b767e285903479ea571fb557b8acb8f138a0756b06

Observation e03d7d87-8c84-4017-a82a-6caa432876bc · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:58.926566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:2e69635863d5dfa31bff09e007589039ebbbcad33b92c1e49742a3b14951bfe4

Observation 3163fa74-307e-4fb7-a646-a816d38ddff8 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 300

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.193835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:66b74a2db6c09385f2c42626d07cce5cf4f1a854ff27e73790bd2c3d2665dd63

Observation 1179dc46-41ec-458a-9c62-c58a6f670ef7 · inbound

From Long to Short: LLMs Excel at Trimming Own Reasoning Chains cites this paper.

From Long to Short: LLMs Excel at Trimming Own Reasoning Chains BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T00:04:02.104563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:04:02.104563Z digest=sha256:b81acb826ea6914e9a4e2243c7932fa128a69cb7ce24560e8df19c48a2dbc243

Observation 95356520-f7dc-4c93-9e37-f41cd17eef1b · inbound

AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning cites this paper.

AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T16:08:41.200957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:08:41.200957Z digest=sha256:adbda90d169f2ebf7744ebd9f2decccae29715f8dfe1a8e0e1c623e716bd873b

Observation 3d5d4c74-0af9-4ac7-8d46-475da075b5b0 · inbound

ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents cites this paper.

ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:26:22.484803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T12:22:47.695255Z digest=sha256:01a10f5cc641a2cd7a7e034c329f2929c9753683174c7ba243858dad6641246b

Observation 23e07b29-ba27-4f23-9870-ea8a98d96346 · inbound

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents cites this paper.

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:54.497956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:54.497956Z digest=sha256:8f7d3d33682057a90848d9d594174d80ab6698bdb97d47c697c1920f57d22aa4

Observation 30a28683-7f48-45d9-bf20-9d390b19f67f · inbound

When Should Users Check? Modeling Confirmation Frequency inMulti-Step Agentic AI Tasks cites this paper.

When Should Users Check? Modeling Confirmation Frequency inMulti-Step Agentic AI Tasks BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 91

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T09:01:08.962311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T08:59:35.944554Z digest=sha256:8e8500d2d8cb06cbd369151de51cdc2424285d9fd898158ac2b2f6a56ea00281

Observation f859f58a-86bb-439c-b203-cd87289ef507 · inbound

Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search cites this paper.

Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T08:48:43.544718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:48:43.544718Z digest=sha256:e84e52f6e1ff7c94f4a4a6314dd5a8c7147d66d8273c7c31d1b877c4fdce88b2

Observation f8002ee8-bd8a-40c5-abb2-53d049c61ad8 · inbound

InteractComp: Evaluating Search Agents With Ambiguous Queries cites this paper.

InteractComp: Evaluating Search Agents With Ambiguous Queries BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:45:00.286343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:45:00.286343Z digest=sha256:816d8b64362d5c827a7a0a97b9e52ec17246fbe6d59372d20d38f0a0ee265f07

Observation 0ee09f8d-558b-4186-9312-29f992a97f85 · inbound

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling cites this paper.

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:45:17.813468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T21:44:18.744201Z digest=sha256:42a8da03ee09d579c4f540dad1b60b548b7299209782851c85877278bf68b535

Observation 94e6e948-fa2b-4a94-aa93-b03638d9fb59 · inbound

ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents cites this paper.

ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 981

Resolution
malformed identifier
no resolver link, observed 2026-08-03T19:23:16.581499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:23:16.581499Z digest=sha256:da4866ced72262ccbb90a58d97444014ce317d12c1cec05c58e33f9f77ad51c1

Observation 82d6bac5-19ae-41da-9fef-8ff7897f842f · inbound

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services cites this paper.

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T18:00:19.681078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:00:19.681078Z digest=sha256:430590e3b49d2b80c6fc50d8b74d073b5da5b3b495be69bbee571693df2edddd

Observation e462099b-cbe3-4c46-97a3-5ee33c1f7f9d · inbound

MemEvolve: Meta-Evolution of Agent Memory Systems cites this paper.

MemEvolve: Meta-Evolution of Agent Memory Systems BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:18:15.241614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T22:18:15.149975Z digest=sha256:b201d3fe2e3a6fe5b2ba65b26f8011ce6a57b335ba9eddc056565aa76f7a51bb

Observation fb2454a0-d5b8-4349-9316-bf5834683b13 · inbound

MiMo-V2-Flash Technical Report cites this paper.

MiMo-V2-Flash Technical Report BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:33:32.777828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T11:33:32.568261Z digest=sha256:6298c4844c0f85b67212185530ca7f19b2d74f7d7a2d700b3fa17e9a3f6fc993

Observation 3662bfe7-2c0b-4ca5-901a-9079d5d45016 · inbound

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning cites this paper.

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T16:40:22.782416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T16:35:24.557809Z digest=sha256:bf1f511765676b33e95a24bc06d5088e7d8378b48224a55a4471e6185037a5df

Observation bde73516-2ff7-4988-a70d-eb795dfe8342 · inbound

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts cites this paper.

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:12:58.699728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:11:14.123806Z digest=sha256:daaceec98376dfcb02260c1aab079626d9cac44a361abe88930124be67b39aed

Observation 01348109-343f-4fb5-be57-24d66fbdb1e2 · inbound

Toward Efficient Agents: Memory, Tool learning, and Planning cites this paper.

Toward Efficient Agents: Memory, Tool learning, and Planning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:44.896240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:44.896240Z digest=sha256:72b1b8cbd834f22b1f32a178a96b12b9c352be3d699e5c6c5ec2927b67f6f3b2

Observation b885cc4b-c826-4328-8b02-ac9eea9dfe58 · inbound

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification cites this paper.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:30:53.895063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:fa347fb740b544cfbf6c3f0080140f24497c61467070ba8894f1be418549ea27

Observation 996dfc2f-9ca1-4e16-9913-59d4e1f133b5 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:327b05c92225d00d891ed5110000db339f4105159edf45347ea76404b2f3fc22

Observation 7f7c4e4c-2dc0-4641-9dc0-38283dc68d21 · inbound

"LLM Agent Performance" Is Not a Single Evaluation Target cites this paper.

"LLM Agent Performance" Is Not a Single Evaluation Target BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T05:04:15.262351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:04:15.262351Z digest=sha256:03c5221599e0f82243bd92fe72c2dc6bea2050f999c350faa4b58b431d95313f

Observation 7dc572a8-0df3-451e-b9ac-a00441bee7b8 · inbound

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions cites this paper.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.409116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.409116Z digest=sha256:09100baf2875c0e16592a5731b2d1db78c9d966479217f492d4843ff43626f60

Observation 2fcf0696-a33a-4fad-a8b2-b9fe0dcfbb66 · inbound

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies cites this paper.

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:57:24.609605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T05:53:29.860037Z digest=sha256:2f455612aebe73743655bac428fa7bc6cfe8bdf0aa7a67a6848460d408ffbd1d

Observation 44ba6d0b-b908-43ff-8e89-918a4eaf850e · inbound

Hunt Globally: Wide Search AI Agents for Drug Asset Scouting in Investing, Business Development, and Competitive Intelligence cites this paper.

Hunt Globally: Wide Search AI Agents for Drug Asset Scouting in Investing, Business Development, and Competitive Intelligence BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:30:19.629414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T21:30:04.131940Z digest=sha256:cfa1fc3e963b78f138d8041ce3d95e63eb66e4153c11de132fca0d86427d8d06

Observation 0f007349-f4da-4ee8-9f97-003c3910cfed · inbound

GLM-5: from Vibe Coding to Agentic Engineering cites this paper.

GLM-5: from Vibe Coding to Agentic Engineering BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T05:46:40.836161Z digest=sha256:be0656dfeadf0a06fe5f7ca4c44f533d3df75381af86c4b0cc9d07fa9c523ce9

Observation 4a1fc836-d265-4d0a-b0fe-fb1744016c9c · inbound

Revisiting Text Ranking in Deep Research cites this paper.

Revisiting Text Ranking in Deep Research BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T21:06:41.680094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:06:41.680094Z digest=sha256:59b371f7ef1d9b0df0c73ec061c7d28f3990d5204f37bc84324e0329d5a1e593

Observation 226996e0-590f-4bf1-9419-74cc107b499c · inbound

DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent cites this paper.

DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T19:46:43.536223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:46:43.536223Z digest=sha256:62c7c382c2579b3f8ca19f629f540598837bd02a703fb3a7450d786a615b880b

Observation 09d8f00c-27fa-40d1-8385-fe3147b5ebe5 · inbound

Evaluating the Search Agent in a Parallel World cites this paper.

Evaluating the Search Agent in a Parallel World BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:06:19.309885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T17:02:10.963839Z digest=sha256:5b9adc7052e2fe30d6c29a59b8066a64619b70c94d5468c8b1e2a71ef9175984

Observation 10f4c6f1-c8fd-4c43-9aeb-2f31cef235d8 · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:45:14.455877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:0fe72aaf8df346e22fcf762e5181f7310c741c16c870f0547abf0d3916f7efb6

Observation 566c80f3-b2f1-46d7-b515-c800fe6880db · inbound

LightThinker++: From Reasoning Compression to Memory Management cites this paper.

LightThinker++: From Reasoning Compression to Memory Management BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:28:02.493066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T17:25:28.432170Z digest=sha256:42291a65fa31f391e09c74a3ac699607642ec228b43e95db9ef7738ae9e22140

Observation 9dae44f4-491c-4868-847a-c443cdbba20b · inbound

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces cites this paper.

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:33:02.339705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T17:31:08.575993Z digest=sha256:36d0cbf50c95e5511bf9e46f89fb90555ef9f0ca982f41179657195cb12599f2

Observation f8459255-830d-4580-870d-d5b7e83cf5df · inbound

Towards Trustworthy Report Generation: A Deep Research Agent with Progressive Confidence Estimation and Calibration cites this paper.

Towards Trustworthy Report Generation: A Deep Research Agent with Progressive Confidence Estimation and Calibration BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T19:41:00.530274Z digest=sha256:22341ba0e67ab23754a142eba19d809b0d284ef4b58f70213424e4e7983a6a2b

Observation d971aec6-1c05-49f8-a08a-2f194c032995 · inbound

TEC: A Collection of Human Trial-and-error Trajectories for Problem Solving cites this paper.

TEC: A Collection of Human Trial-and-error Trajectories for Problem Solving BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:52:09.041721Z digest=sha256:8179b546a30264c0251fcb722f55e92aa57d55f6042f7fc1a469d21559ff2e16

Observation aa63eaaf-7143-41ff-893f-466a4841df4f · inbound

Towards Knowledgeable Deep Research: Framework and Benchmark cites this paper.

Towards Knowledgeable Deep Research: Framework and Benchmark BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:17:35.879705Z digest=sha256:930c0bfa4967b534322208fe94d08ab9ac928fed1aab595678f13d5e844cd7eb

Observation 21141e13-c002-42a3-8167-90e561d037e7 · inbound

DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? cites this paper.

DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:46:24.780270Z digest=sha256:3a803da89c8bcc9c1f6066b62d374883305d6bb0bc465b7df69724cdf17262c4

Observation fc8c3796-4ae5-44a3-a691-59794622832c · inbound

LABBench2: An Improved Benchmark for AI Systems Performing Biology Research cites this paper.

LABBench2: An Improved Benchmark for AI Systems Performing Biology Research BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:17:30.255532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T07:16:57.796927Z digest=sha256:10801371085b0c126bd09e50fd649e5071100c7aca6b01a92fdf37a607beeadb

Observation 813d494a-e6e8-4ce6-bc8f-18e8949d23f8 · inbound

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation cites this paper.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:6780a5291f2759b0fc789616f3e55591bf932e980265480a660916324de5f93c

Observation 5cb17d06-2a09-4d8e-af46-f902a3c81eec · inbound

WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark cites this paper.

WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:00:20.385808Z digest=sha256:b281d2a3c2bc88efe84dfd52f42a8063901753d9cc7b0df29dac775cbf310083

Observation 4ee44383-207c-46ed-8d0a-5c6610146b1f · inbound

AlphaEval: Evaluating Agents in Production cites this paper.

AlphaEval: Evaluating Agents in Production BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:aa1c1a3cd49f5f9621520ae4d1a0a70d10b00a6fa4366d7518996137c07d8429

Observation 763bf919-fd8a-4897-abec-9b8320a28b5b · inbound

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization cites this paper.

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:17:32.290531Z digest=sha256:68d9c37f785db5b034ee4fe5fd94b8659a1245dcdfc0d268905ea9417dfe2a98

Observation 5344862e-9ba0-43c8-b8c5-98d18d80176e · inbound

Towards Long-horizon Agentic Multimodal Search cites this paper.

Towards Long-horizon Agentic Multimodal Search BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T15:40:32.137708Z digest=sha256:c01e8a41087628fde82756b0b59b53cafcd7654998518e45a092e98348b1823e

Observation a7d32d62-3a0f-4c4f-9dcd-fd05c9f3aa63 · inbound

MARCA: A Checklist-Based Benchmark for Multilingual Web Search cites this paper.

MARCA: A Checklist-Based Benchmark for Multilingual Web Search BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T12:54:27.110593Z digest=sha256:3d326d4b0fab9e35d5cd475e0a3cd4c4117108044a38599030a0d63944596814

Observation 9ed03c0b-280e-4077-98fb-efe429502964 · inbound

Mind DeepResearch Technical Report cites this paper.

Mind DeepResearch Technical Report BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T11:46:49.178896Z digest=sha256:86ec99f7083262dd3f7adfe93c9173886e4096f2bb062ad945c6eb5e82399eec

Observation 333a2ca7-a976-481e-a7c7-eb5bff37573c · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-05T17:51:14.807535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-05T17:45:55.631459Z digest=sha256:179ff3f1fa2e3e4563dc234d61abaeeb15d59f675919e4fbeac007ed7a549912

Observation 718233f9-b37f-4c57-afdb-45e4bf6f6bdb · inbound

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent cites this paper.

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-05T15:11:10.890514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-05T15:03:50.420072Z digest=sha256:ee196d4a3170120494377a8fdf3f1bfe7b551c31b42810d3ebc559aa5a91d1f1

Observation 69c9040f-4975-4619-9ab5-f1419096bdd0 · inbound

DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data cites this paper.

DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T02:32:29.401023Z digest=sha256:8071beb3e94959e5cf9499df06889ba03a45a1b6e690ce5738e04efdf35e1ed8

Observation a3f43076-ebdf-42c3-81fe-bebdb9e05e20 · inbound

GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation cites this paper.

GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T03:23:25.390218Z digest=sha256:a6674195cfe3340f9c82e9885af353490e38516cd0e08603d02e1529dd1d7322

Observation d577f46d-6349-46cd-b75c-6b7b39cab7d2 · inbound

AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery cites this paper.

AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T16:38:40.938722Z digest=sha256:2aff3b364a8d4db5b38ddb3cfb043462bbeb849e4090cfa9e664b2ce849458e1

Observation 2fe4ca2d-c598-48a7-b30f-e06da792442f · inbound

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning cites this paper.

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T14:18:14.048230Z digest=sha256:b972774e9e7427ec3ebc2bc7fe813428e194a609aa323297ee7a31b37665b1cf

Observation 0184b442-1ac2-499f-ac7f-32fdc7dbebe4 · inbound

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces cites this paper.

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T18:44:27.685266Z digest=sha256:c6414e09c2bca04380232a40289d865acc6f83a943aae06cb6242d9584d5b7c9

Observation 5196cd4d-717c-4335-bac8-48a3d5da3345 · inbound

PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization cites this paper.

PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T17:49:39.533090Z digest=sha256:2be2415007b9ba36395a90926f82efd6a666275bda2b70b37cc974b1c4119fe8

Observation 1c6a0ad6-d504-4fa3-a2a8-a845992dc0bf · inbound

Inference-Time Budget Control for LLM Search Agents cites this paper.

Inference-Time Budget Control for LLM Search Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T11:51:02.872129Z digest=sha256:040d8d747a1f1b3922e557955bf353a9125b99b48b5a14d1f5037a4f22bf8b96

Observation 970abbb4-88e3-4ec4-a4a4-edb8b8d6a29e · inbound

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents cites this paper.

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T09:53:00.677605Z digest=sha256:d3a33e52566911df572f28ff5635c6cad49ac154adaaed787886fe87d0b53071

Observation 2592e845-1789-458d-91f1-92977cdbb2ff · inbound

TeamBench: Evaluating Agent Coordination under Enforced Role Separation cites this paper.

TeamBench: Evaluating Agent Coordination under Enforced Role Separation BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T00:55:51.358828Z digest=sha256:57598e20beff260f24c4f1422436f0c422b7a266fe3f28502095f8ec17d0aca2

Observation 5c59efcf-0482-46ef-b593-fbdfb7d1b8bc · inbound

Slipstream: Trajectory-Grounded Compaction Validation for Long-Horizon Agents cites this paper.

Slipstream: Trajectory-Grounded Compaction Validation for Long-Horizon Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:31:24.292654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T01:04:55.618417Z digest=sha256:63be0af42671d73aaf45267670bad49aa4656435aba07a037612cd987fb0a532

Observation 17c4b39f-4402-4d36-8ff9-88a17f06ff36 · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 108

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:25.291682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T01:13:35.990078Z digest=sha256:e5a345f09e4566207ee6d490065390d58d0a15545558b5ce14368c0b78604f30

Observation aa7b7f4c-be8e-417f-a491-8e52acb8d490 · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:25:45.983058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T23:12:57.154537Z digest=sha256:bbb98e778bc8e70a295f0f711712b5e1b95870380101b38251913ddbdef02353

Observation 68d452ba-3d55-4582-8745-62298dac744c · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 109

Resolution
unresolved
no resolver link, observed 2026-07-12T17:14:49.310598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:14:49.310598Z digest=sha256:6334ce44a0024c6f45ddf2e1c469d98e3157d658d9fb5dd9576c9354056e1679

Observation 0dbf2785-7ae6-4f9f-ab1d-67557ff474f0 · inbound

EvoMAS: Learning Execution-Time Workflows for Multi-Agent Systems cites this paper.

EvoMAS: Learning Execution-Time Workflows for Multi-Agent Systems BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T02:29:55.683565Z digest=sha256:a15c769b087e2ca7185c7ca071a62df7084bbd5a4fdc503d135a28bd52e0383b

Observation 7f7cc863-5c73-4bae-858e-ba5ae65aaa24 · inbound

TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems cites this paper.

TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T04:53:54.754878Z digest=sha256:529e82a15b3e516651cef886332abe9310b3402f0b44a82c694de6a4daff4011