Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2504.12516.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:47:14.793292Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation cb9c2891-a70c-4fec-90e0-fc94261749eb · inbound
BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8ae84d09-d6c1-4c32-90a0-2b398ade77d9 · inbound
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e101cd-6d9b-4b0c-b069-2f46f3b24295 · inbound
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b480e41b-d137-4ead-aad7-5d2f0264e4f4 · inbound
Agent-Environment Alignment via Automated Interface Generation BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ceb12c9-7bcc-44c2-aaa6-b3a848a8f9c2 · inbound
EvolveSearch: An Iterative Self-Evolving Search Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b559645e-74b5-4ef2-b83d-2e0665baebae · inbound
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a077d01-0f0b-4170-aff1-153ca9cf2068 · inbound
Real-Time Execution of Action Chunking Flow Policies BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 11bc6327-6b0a-4785-9486-a84afdc122fc · inbound
TaskCraft: Automated Generation of Agentic Tasks BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a5c3690-980a-4587-bf2d-dd0bebaa3480 · inbound
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 58efdf6a-6899-4dee-a80b-573e578ee62a · inbound
ScholarSearch: Benchmarking Scholar Searching Ability of LLMs BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbeb1393-1d01-4ce8-b7ee-eda0c4b5eda1 · inbound
OAgents: An Empirical Study of Building Effective Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 825b5c51-08b5-4314-aa48-d0d0adfcc73c · inbound
From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d38bb0d-ff7b-4f2a-af0f-8d3c0b968969 · inbound
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7735579-3806-44ba-b6a1-c829227839cd · inbound
WebSailor: Navigating Super-human Reasoning for Web Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2107abac-dd79-4ba3-86ea-9410807c6e93 · inbound
WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c7ef63f-33e8-4f12-84d9-a5c67d0c1b52 · inbound
ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a34f4c58-cc1c-459a-90d5-13e7d529220e · inbound
RAVine: Reality-Aligned Evaluation for Agentic Search BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d034cd4b-c2be-4c45-a573-1555235ad47d · inbound
Beyond Context Limits: Subconscious Threads for Long-Horizon Reasoning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f35406d-e002-4523-83b1-ffb6137c12fa · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e87c215b-c00d-4e70-8414-029689052a61 · inbound
How Far Are AI Scientists from Changing the World? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 176
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a53f149a-7e84-4e6d-a52a-cb9a4b22006b · inbound
TextQuests: How Good are LLMs at Text-Based Video Games? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50778f16-28cb-4ecb-b74d-e786c5d15fd7 · inbound
MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6318c7e8-e969-43a5-8ba5-90c57f2b68d7 · inbound
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f1b34e4-deb5-4000-b553-9b6971b139b1 · inbound
Efficient Agents: Building Effective Agents While Reducing Cost BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb0a689d-f746-4cfe-ab35-ab456294fade · inbound
Characterizing Deep Research: A Benchmark and Formal Definition BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90c1f86a-9076-421d-9b75-43c0b280e8b1 · inbound
WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 402d821e-d025-4f7e-9f7c-15a470d63a1d · inbound
Observation of momentum dependent charge density wave gap in EuTe4 BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 019096d1-88d5-4ada-85c1-f298425b77cc · inbound
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d0c2b477-bdd0-45ab-8e41-79f46b29fcd4 · inbound
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3921d393-b249-445a-9c10-f426ec02bdeb · inbound
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 35d4c42a-e2a4-4727-8889-b87742c9e5e4 · inbound
Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65abb7e1-0218-4aa4-80d4-013c5f0659fc · inbound
SSRL: Self-Search Reinforcement Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e36dc413-0095-4501-8c10-8d271cbf4040 · inbound
FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ed3d6f-8058-46c5-8684-edc8330fa21e · inbound
Exploring Spatial-Temporal Dynamics in Event-based Facial Micro-Expression Analysis BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2d15bd-0624-4483-af3c-7557f484ea49 · inbound
WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c4d45d17-24ac-4d7b-9ab3-b79be03868be · inbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6c93b97-821d-4ced-b1ff-6ea8bb2bd3cd · inbound
Search-Time Data Contamination BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d8fe418-16a1-4b2b-930b-8e26a4138302 · inbound
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a23df0-5f8f-4c0a-9cbf-4f3bcb23d3b6 · inbound
UQ: Assessing Language Models on Unsolved Questions BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c94c0c2-c0ea-407f-a219-8f4a9a2dea81 · inbound
Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 488771f2-8483-4d15-b829-1b80f323ac68 · inbound
Open Data Synthesis For Deep Research BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e03d7d87-8c84-4017-a82a-6caa432876bc · inbound
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3163fa74-307e-4fb7-a646-a816d38ddff8 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 300
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1179dc46-41ec-458a-9c62-c58a6f670ef7 · inbound
From Long to Short: LLMs Excel at Trimming Own Reasoning Chains BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95356520-f7dc-4c93-9e37-f41cd17eef1b · inbound
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d5d4c74-0af9-4ac7-8d46-475da075b5b0 · inbound
ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 23e07b29-ba27-4f23-9870-ea8a98d96346 · inbound
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30a28683-7f48-45d9-bf20-9d390b19f67f · inbound
When Should Users Check? Modeling Confirmation Frequency inMulti-Step Agentic AI Tasks BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f859f58a-86bb-439c-b203-cd87289ef507 · inbound
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8002ee8-bd8a-40c5-abb2-53d049c61ad8 · inbound
InteractComp: Evaluating Search Agents With Ambiguous Queries BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ee09f8d-558b-4186-9312-29f992a97f85 · inbound
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 94e6e948-fa2b-4a94-aa93-b03638d9fb59 · inbound
ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 981
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82d6bac5-19ae-41da-9fef-8ff7897f842f · inbound
LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e462099b-cbe3-4c46-97a3-5ee33c1f7f9d · inbound
MemEvolve: Meta-Evolution of Agent Memory Systems BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fb2454a0-d5b8-4349-9316-bf5834683b13 · inbound
MiMo-V2-Flash Technical Report BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3662bfe7-2c0b-4ca5-901a-9079d5d45016 · inbound
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bde73516-2ff7-4988-a70d-eb795dfe8342 · inbound
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 01348109-343f-4fb5-be57-24d66fbdb1e2 · inbound
Toward Efficient Agents: Memory, Tool learning, and Planning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 140
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b885cc4b-c826-4328-8b02-ac9eea9dfe58 · inbound
Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 996dfc2f-9ca1-4e16-9913-59d4e1f133b5 · inbound
Kimi K2.5: Visual Agentic Intelligence BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7f7c4e4c-2dc0-4641-9dc0-38283dc68d21 · inbound
"LLM Agent Performance" Is Not a Single Evaluation Target BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc572a8-0df3-451e-b9ac-a00441bee7b8 · inbound
OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fcf0696-a33a-4fad-a8b2-b9fe0dcfbb66 · inbound
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 44ba6d0b-b908-43ff-8e89-918a4eaf850e · inbound
Hunt Globally: Wide Search AI Agents for Drug Asset Scouting in Investing, Business Development, and Competitive Intelligence BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0f007349-f4da-4ee8-9f97-003c3910cfed · inbound
GLM-5: from Vibe Coding to Agentic Engineering BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4a1fc836-d265-4d0a-b0fe-fb1744016c9c · inbound
Revisiting Text Ranking in Deep Research BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 226996e0-590f-4bf1-9419-74cc107b499c · inbound
DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d8f00c-27fa-40d1-8385-fe3147b5ebe5 · inbound
Evaluating the Search Agent in a Parallel World BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 10f4c6f1-c8fd-4c43-9aeb-2f31cef235d8 · inbound
Seed1.8 Model Card: Towards Generalized Real-World Agency BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 566c80f3-b2f1-46d7-b515-c800fe6880db · inbound
LightThinker++: From Reasoning Compression to Memory Management BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9dae44f4-491c-4868-847a-c443cdbba20b · inbound
GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f8459255-830d-4580-870d-d5b7e83cf5df · inbound
Towards Trustworthy Report Generation: A Deep Research Agent with Progressive Confidence Estimation and Calibration BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d971aec6-1c05-49f8-a08a-2f194c032995 · inbound
TEC: A Collection of Human Trial-and-error Trajectories for Problem Solving BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation aa63eaaf-7143-41ff-893f-466a4841df4f · inbound
Towards Knowledgeable Deep Research: Framework and Benchmark BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 21141e13-c002-42a3-8167-90e561d037e7 · inbound
DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fc8c3796-4ae5-44a3-a691-59794622832c · inbound
LABBench2: An Improved Benchmark for AI Systems Performing Biology Research BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 813d494a-e6e8-4ce6-bc8f-18e8949d23f8 · inbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5cb17d06-2a09-4d8e-af46-f902a3c81eec · inbound
WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4ee44383-207c-46ed-8d0a-5c6610146b1f · inbound
AlphaEval: Evaluating Agents in Production BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 763bf919-fd8a-4897-abec-9b8320a28b5b · inbound
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5344862e-9ba0-43c8-b8c5-98d18d80176e · inbound
Towards Long-horizon Agentic Multimodal Search BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a7d32d62-3a0f-4c4f-9dcd-fd05c9f3aa63 · inbound
MARCA: A Checklist-Based Benchmark for Multilingual Web Search BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9ed03c0b-280e-4077-98fb-efe429502964 · inbound
Mind DeepResearch Technical Report BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 333a2ca7-a976-481e-a7c7-eb5bff37573c · inbound
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 718233f9-b37f-4c57-afdb-45e4bf6f6bdb · inbound
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 69c9040f-4975-4619-9ab5-f1419096bdd0 · inbound
DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a3f43076-ebdf-42c3-81fe-bebdb9e05e20 · inbound
GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d577f46d-6349-46cd-b75c-6b7b39cab7d2 · inbound
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2fe4ca2d-c598-48a7-b30f-e06da792442f · inbound
SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0184b442-1ac2-499f-ac7f-32fdc7dbebe4 · inbound
Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5196cd4d-717c-4335-bac8-48a3d5da3345 · inbound
PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1c6a0ad6-d504-4fa3-a2a8-a845992dc0bf · inbound
Inference-Time Budget Control for LLM Search Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 970abbb4-88e3-4ec4-a4a4-edb8b8d6a29e · inbound
Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2592e845-1789-458d-91f1-92977cdbb2ff · inbound
TeamBench: Evaluating Agent Coordination under Enforced Role Separation BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5c59efcf-0482-46ef-b593-fbdfb7d1b8bc · inbound
Slipstream: Trajectory-Grounded Compaction Validation for Long-Horizon Agents BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 17c4b39f-4402-4d36-8ff9-88a17f06ff36 · inbound
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation aa7b7f4c-be8e-417f-a491-8e52acb8d490 · inbound
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 68d452ba-3d55-4582-8745-62298dac744c · inbound
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dbf2785-7ae6-4f9f-ab1d-67557ff474f0 · inbound
EvoMAS: Learning Execution-Time Workflows for Multi-Agent Systems BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7f7cc863-5c73-4bae-858e-ba5ae65aaa24 · inbound
TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.