Pith. sign in

Paper Citation Record · LEDGER

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

As of 17 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2607.06118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06118 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T16:11:49.725631Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact8
  • verified fuzzy30
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch16

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 06787661-ff35-492f-80f1-7df22e8c88fa · outbound

This paper cites Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.173245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:17f0ad683c690727568404b428950d4f4dad424f899aa27665948a9a8e318ef0

Observation 03029bd2-3506-45e5-a82a-d1085db6bbc4 · outbound

This paper cites Official Product Announcement (2025), https://www.anthropic.com/news/claude-sonnet-4-5.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Official Product Announcement (2025), https://www.anthropic.com/news/claude-sonnet-4-5

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.917404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:b874f20ff4b4ba6992e5b18da92e75b6a40b42fe1320a74723992d99e14c7e38

Observation e2bfcd8a-8a34-4fd2-a13a-84538db6e30c · outbound

This paper cites Qwen2.5-VL Technical Report.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:15:06.142033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:a41303da8465a1838ba8b1e98b6b110e752febb55b07bd4a597cdb0efd3fb7cf

Observation 81e72636-117a-4819-abc0-cdc2aa7a0ff2 · outbound

This paper cites Advances in Neural Information Processing Systems36, 78142–78167 (2023).

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Advances in Neural Information Processing Systems36, 78142–78167 (2023)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.915716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:5a7cdb33ed3279ef0a40fe983464c621553a15c0691953ee11351c2b15d912d6

Observation 6c694b16-c7fb-4195-8cfc-72e337039286 · outbound

This paper cites Advances in Neural Information Processing Systems37, 5996–6051 (2024).

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Advances in Neural Information Processing Systems37, 5996–6051 (2024)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.919095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:7d261fd8888c81d9ceaab2de733b01cb44ce9c8bea9f68665138e89518e241b9

Observation d741c6f2-e174-4bd6-80bc-d446962e2e5a · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation In: Proceedings of the AAAI Conference on Artificial Intelligence

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.920815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:59ee9e557591cb0059c2e5b1f897e12fa763af89becf8609748e90418a85b7aa

Observation acd6c39e-27c5-4cbd-a9f1-c7ebdf899815 · outbound

This paper cites ChatShop: Interactive Information Seeking with Language Agents.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation ChatShop: Interactive Information Seeking with Language Agents

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:15:06.167723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:d19d39ecdaabe7f65b252ee66092afa2008ea2304cfa0d8b03262f94b30f40a6

Observation b97de35a-31ec-42ef-82b8-00e54a4e5271 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.191481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:d9eaa41b2ff1b014b946592a76f8d8ebd62e29e9feae344e8ed5f49f0a8f6dd8

Observation f83d7c26-54af-453b-9335-665cc9a33324 · outbound

This paper cites Advances in Neural Information Processing Systems36, 28091–28114 (2023).

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Advances in Neural Information Processing Systems36, 28091–28114 (2023)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.951561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:cc6bca038f2707856909c9e84f07624088479834e524d37670275a771eaa4efa

Observation a7c9e5f4-7a98-470e-be9d-f41799d08407 · outbound

This paper cites In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.910544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:61c65e1b5bfc449bce3c263a322e75a8df03c16d7d4faa7b9e4b72b4acd55149

Observation b969623a-d1b6-4606-a722-9fbd93d51e5c · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:15:06.163619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:2e898bfc17b3902cea89b40c79e55a7d79980ab92f317a53c89f7607d8ac4ff2

Observation ea6a8cf4-8e1d-4d4c-9e47-442a199a34f7 · outbound

This paper cites In: European Conference on Computer Vision.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation In: European Conference on Computer Vision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.912202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:300d457cf2cdf79750ea5f4bc80113e7270c503a9db9013e7d8187d0ba99a2c2

Observation 7ce5f196-44fe-4ecf-8b36-9062be85e6db · outbound

This paper cites The Devil is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation The Devil is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.175905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:8eb5d670823eec2df60545ff4318794fbdb4efab34a9fd667b5e84caf23e9e7a

Observation b2db4765-23b1-4a68-9599-1ea8aec701e4 · outbound

This paper cites Mano technical report.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Mano technical report

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T16:15:06.152082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:1c8841c2a369085b061e26f8d08b95165345761990411a62dbb63ab0724124b9

Observation 667e2d04-628f-458e-9eed-4398d0ce090c · outbound

This paper cites In: NeurIPS 2023 Foundation Models for Decision Making Workshop (2023).

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation In: NeurIPS 2023 Foundation Models for Decision Making Workshop (2023)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.908800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:e7d14ee58f5c9540111e59047a5ed9ea7327345ea1b2e01e9cda9636fea2ce1c

Observation 6dc4e38a-c8f0-4bd9-a5f7-1d1a654824fe · outbound

This paper cites REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:15:06.149393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:7b82b59b62d751f42735ba5104e1912002ea6f71dfc31ed21e21a596b76e355b

Observation 6c815a94-bf2d-4715-85f8-5ee9de747242 · outbound

This paper cites Seed1.5-VL Technical Report.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Seed1.5-VL Technical Report

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:15:06.199956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:01ea03203ee10a5641ce5309c747efb9831e3934144d661ba8b7627720b8a94a

Observation 45f44b0b-c39f-4463-88ca-a83e56f48f70 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.178735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:5bb84320beed4be7d4d10c33fb7fd36a3e9942f659cd9b09b1f46038d70dcf6f

Observation afbee48b-a87f-4657-91e7-41ea4ef6db7d · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.906911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:149c6ed577440e5e04f25c3e7865454fa94fc5090ef4ee94174525ca60e02e4e

Observation 8fa6923c-366e-448c-8a72-23f1bf12a43c · outbound

This paper cites GPT-4o System Card.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation GPT-4o System Card

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.181170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:648620d539e72bf832fd07946f51712c468130bb3b31b110c5362614c2c787c7

Observation d3bc22cb-88cd-41e1-9e57-34ab7b0d3b40 · outbound

This paper cites In: European Conference on Computer Vision.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation In: European Conference on Computer Vision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.903479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:3b54d34c05e507ce8e4404a0dd93c4be92d21dce9a9d6b4e4bbabe8d06bcdff7

Observation bc0802a3-232e-4968-891b-747ded284f0a · outbound

This paper cites In: ICLR 2025 Workshop on Foundation Models in the Wild (2025).

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation In: ICLR 2025 Workshop on Foundation Models in the Wild (2025)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.901836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:dfe44dbec2d349a6b665439d6058fbdf3c160677942714e79e430ed9541d1089

Observation 218ca6f1-97d4-4419-9fba-32b69fd714ca · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.139548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:f2295ad98c4a9d528a1c3bccb72df761c903db27fc874c53503fa5fd31f62698

Observation d6e16495-0b25-44fb-9f5e-424020b0dbd2 · outbound

This paper cites In: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation In: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.898757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:2150e410bdb4ef657373976a22d31495e82e9479d6afcb9b87233bc3de3f729a

Observation 94304a56-1081-4421-8e21-e07f0d33824b · outbound

This paper cites ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.158251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:3f773b4d58955a634e8be210c4308ea358ec7bbd5d211be839f5ef3a50674213

Observation 663d5d8c-ad67-4e79-8c0b-17d002ee5598 · outbound

This paper cites an unresolved cited work.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-07-08T16:15:06.897003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:fde181a5cc636765594b5d7d463c9e7d72c421674873dc6023136a803558635b

Observation 21336d68-4879-4e8c-a23b-0f4a87e51b00 · outbound

This paper cites Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.183705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:fb710bc9bb88f150f7757b29b0dce36d0192dc5d6785d44c333086d030adb190

Observation 8e37b665-4417-4cd4-acbf-bbb52ec8eb24 · outbound

This paper cites VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:15:06.155483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:a4e1f249f783b734e03974129a8fbaef6f83403cff73cdfebdc94257facb3d56

Observation 0df533bd-d95b-4737-b9f7-946f921a1e05 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation AgentBench: Evaluating LLMs as Agents

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.137100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:bc11d2c21d0f05c8465e4115f3a9b453b9ec239e3d52d8f4d454c6ce1194b21c

Observation cf9507c0-f886-4110-9269-0b6b154c13d8 · outbound

This paper cites VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.197306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:8eb6536bf83ba0dad265ff7d98c3bbcf1df578cb3dced4a51ed419cc035e14d8

Observation 60cad104-1700-4f90-8350-af856bf446ed · outbound

This paper cites an unresolved cited work.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-07-08T16:15:06.900247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:64e8f21cf1d38fc8b25a11475c0e4f470056cf725a9c2c650da1f2866b4ba504

Observation 4187d94f-fe8e-4a58-9650-083fca57cdd2 · outbound

This paper cites Official Product An- nouncement (2025),https://openai.com/index/introducing-o3-and-o4-mini/.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Official Product An- nouncement (2025),https://openai.com/index/introducing-o3-and-o4-mini/

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.905175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:80ae82ad46cf36b0702a376763ffb3df425d11b8fff9365d32ff48394b2c0834

Observation 95811a26-e939-4d91-85a2-3406407db7f3 · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Autonomous Evaluation and Refinement of Digital Agents

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.186003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:daebf3f7cd5ee3405a72bbcf11fab1639273dece744cb360c468ddd3f785cf4e

Observation 58662504-66b9-4f0e-a865-87b08c2261b5 · outbound

This paper cites WebCanvas: Benchmarking Web Agents in Online Environments.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation WebCanvas: Benchmarking Web Agents in Online Environments

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.188595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:123a816a95411dd78e8f8467ca887a5371c502abad0b5c8806d700793af54ceb

Observation db0d6835-8efb-4445-b0c2-56b1ca0abb44 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.170568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:3e5f32c097123458023dc47f8a6d3beecf704f27b9fc0da2cd9e5887378b909a

Observation 3bfc5e02-81cd-4efc-bc02-3ebd6594fbd3 · outbound

This paper cites In: International Conference on Machine Learning.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation In: International Conference on Machine Learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.888355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:c1e25a7cb640370ca343ab26ae0e0294bdc8f2eea0fe94cfb4a954d542dd493f

Observation d125ef64-429e-4da5-bc64-b6189d4cf838 · outbound

This paper cites BEARCUBS: A benchmark for computer-using web agents.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation BEARCUBS: A benchmark for computer-using web agents

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.194530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:5eed3bf69f9221ae495ea9991c33a84afabdb4fd174e20fd8e6ccf20885088fc

Observation 782a0652-941a-42a7-94c1-a4bff4b9b72e · outbound

This paper cites In: Findings of the Association for Computational Linguis- tics: ACL 2025.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation In: Findings of the Association for Computational Linguis- tics: ACL 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.890235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:ed8e34186d7301334a711a639c42266ce048f19fe7b5da719e4ef0894d0ea739

Observation 1e54e159-d65a-4440-9237-ee6b61101bf2 · outbound

This paper cites Advances in Neural Information Processing Systems37, 115963–116021 (2024).

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Advances in Neural Information Processing Systems37, 115963–116021 (2024)

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.891960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:fec419d708a9b2e69c1c4e35982efd9ef0e085151f9f4e011971fa59e5314f53

Observation fa945c79-1228-4942-babe-d85312edaf71 · outbound

This paper cites AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:15:06.161036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:f4ee4e10e9910e38c6ad68f8b7cbbbff154ff78c38516f992d776f8154437ba4

Observation 1af37a66-1fd2-4ab9-9037-b4127c28a134 · outbound

This paper cites In: Second Conference on Language Modeling (2025),https://openreview.net/forum?id=6jZi4HSs6o.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation In: Second Conference on Language Modeling (2025),https://openreview.net/forum?id=6jZi4HSs6o

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.893683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:f2ca2265772d808878a3c5a92089e45b5f03e18ed9438095c839a8064ba6327b

Observation e11ff1da-4135-49d2-b472-6fd3c58faa93 · outbound

This paper cites In: ArXiv (preprint).

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation In: ArXiv (preprint)

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.895362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:1ac2af45a4d18b9425cfa4e181bffde5155838544f857e82a5e453473d49cfd0

Observation 622653ac-1134-467a-a48a-ad690fafde7a · outbound

This paper cites AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-08T16:15:06.146949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:ebfb0e0645e521d7774923c964f96f0c365196a0b712c90be66e7a3df804ba2b

Observation ffeec677-9f78-4e36-84b2-2fc5813ac791 · outbound

This paper cites GPT-4V(ision) is a Generalist Web Agent, if Grounded.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation GPT-4V(ision) is a Generalist Web Agent, if Grounded

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T16:15:06.144645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:f057a17e2e0ab10e1c183344def4f4f7c38e265d90de4e0d65f3713c8a0cf607

Observation 3e0f2e3a-b877-4246-940a-94a767323bdb · outbound

This paper cites Advances in neural information processing systems36, 46595–46623 (2023).

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Advances in neural information processing systems36, 46595–46623 (2023)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.913921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:293048b16c2ef112879cdced8d4dbd045ad65b7c506e0bd30263bcdae75d0178

Observation 8c649a3e-3a86-434e-9684-06070b64d526 · outbound

This paper cites ICLR (2024).

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation ICLR (2024)

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.945888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:f665e882f666bde7aa00f56b9353879a4c16cb977d4e8e3485882b715f3c9161

Observation 5981e5f8-5be1-4743-a4d1-1f1d460925b4 · outbound

This paper cites Support” l Click “Predicting with your Models.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Support” l Click “Predicting with your Models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.947714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:af814caee509541f6fd5daa087c5028e14cb54a59d3580e9cead6c7a6ff485bd

Observation 18f31368-c06c-48e3-8a3f-610ae5a182c6 · outbound

This paper cites Every requirement, keyword, constraint, or objective in the task is treated as a criterion.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Every requirement, keyword, constraint, or objective in the task is treated as a criterion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.944193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:6470f78a27f10a41c1199e4f9cbe759143d30209ae653d411380d9a369f4a2e9

Observation 865eb141-290c-44c8-9137-d93f14e7b531 · outbound

This paper cites an unresolved cited work.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-07-08T16:15:06.937622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:118d74e563b62d9f91f0a2bfd4df0d9702cc04b03d1c80ea302522df43d8597a

Observation b9a1771d-a6f9-46c4-b1ac-ddfbcc34adbb · outbound

This paper cites an unresolved cited work.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-07-08T16:15:06.939160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:b2a40e3287af4fdcbcd6ba2d8d58e3abb387ece87335c15d428bdb17e0d2522e

Observation aa1eab71-1776-4e32-99f3-f3769a6f05e8 · outbound

This paper cites submit”, “search.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation submit”, “search

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.942401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:bda651657be3edff63b252f256fed220b40c795e7ebc672680545f1a15702cea

Observation 083a542b-cd0a-4961-9f6d-97715acd8557 · outbound

This paper cites No quotes, no bullets, no numbering.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation No quotes, no bullets, no numbering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.940670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:0acb66f1380149e97e4231d0bdb3dfa878bb60f1541d69b5d5208c3049490285

Observation 09d6f12a-2fef-4671-9399-547ec14337d3 · outbound

This paper cites Do NOT mechanically list actions; instead, describe actions as steps that move toward completing the goal.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Do NOT mechanically list actions; instead, describe actions as steps that move toward completing the goal

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.949657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:ab40c9852b572633289a3a901bb4798732ba86cc1c7ae29efb1c7e69007ab75b

Observation 6c8f433f-f831-4047-a4a3-d7ae0f4b4183 · outbound

This paper cites an unresolved cited work.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-07-08T16:15:06.953115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:0e6161bd94c81dafb077ff74933e2b187da9f985eb351cef42ad0b235976926b

Observation 11bd375c-0bdd-479c-a0ae-20b849e8d731 · outbound

This paper cites an unresolved cited work.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-07-08T16:15:06.932390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:883788fbe9bfc433c62207d23f4bb2e98ed9965486d7aeea48f631210baa4703

Observation cb65d2f7-e168-45f6-87b9-e157a6ceaf5e · outbound

This paper cites ** Click/selection wording constraints:.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation ** Click/selection wording constraints:

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.936056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:e293004d940aaf32d6c8152a80524a937e12e115d85ab16ad1893296ac06c6f2

Observation ab9283f2-dff7-4d71-bb78-5df406407828 · outbound

This paper cites an unresolved cited work.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-07-08T16:15:06.929298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:0ee80d6ab99e4c842df8b8c917659b0c7d7f3c0a0bfdd0ceb0c8f8a2d05253e2

Observation 1abdfeaf-e13b-4350-bdc2-ef94be832507 · outbound

This paper cites If a label is necessary, keep it short and functional.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation If a label is necessary, keep it short and functional

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.924230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:7fb9ec1140e2e214d9065106934ee80929f4d567cc1970c6fc98109aa2a880c0

Observation b8f8701e-c889-4ac7-b526-5fa9350cb287 · outbound

This paper cites ** Output constraints:.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation ** Output constraints:

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.922535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:174585b19231345aa332366d1a94053c4815db108d35e00e7b2b154dfb29295e

Observation 1c3e32ad-29bd-4f37-bf94-8f60671af91a · outbound

This paper cites an unresolved cited work.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-07-08T16:15:06.925836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:efedfacf2b4dd19af139ee23665884683fb44d006188ad5fafdd24bbcb90d923

Observation e97886db-926a-49e4-8c38-8a96b4eccb5b · outbound

This paper cites first... then... next... during... finally.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation first... then... next... during... finally

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.927739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:df1be61cfbcc49d1f26339067860401634a11ef55ce8eda03a190634221dc48b

Observation 204ad6d6-e95a-4e3d-ac4a-e059da368e67 · outbound

This paper cites an unresolved cited work.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-07-08T16:15:06.930806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:42f8e8936c11039582e000043e5052cae2d7bb56dddd02a668b75b6bdc8e87a7

Observation f5f3d101-7966-47b9-b9c4-e1475a5d480b · outbound

This paper cites User Prompt: Task: <task> Website: <website> Trajectory: <trajectory> Now generate the operational document sentence:.

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation User Prompt: Task: <task> Website: <website> Trajectory: <trajectory> Now generate the operational document sentence:

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T16:15:06.934323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T16:11:49.725631Z digest=sha256:86bffa97eefa90fee7ca64c1a7481968d41a98524c04304b165de6f9600fb203

Pith citing papers

No inbound Pith citation observations are available.