Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:27.016316Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 4 inbound Pith citation observations for arXiv:2505.24878.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:27.016316Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T14:45:56.506288Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T22:56:19.580749Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2d9246e3-756a-499c-88e0-38d0731381b7 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Flamingo: a Visual Language Model for Few-Shot Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1fde33c-9802-4c66-9479-4a51cca99e45 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Claude 3.7 Sonnet System Card
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8cd90733-1730-4365-a1f1-1da81cf6a2df · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f992786-6b47-46f0-b163-09ba2149e385 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7884ebb-dedb-46c9-b75f-d74ce5e0840c · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92da3802-4361-4744-80a0-7b8dd8df97c1 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7f316e4-5392-451e-afc9-0b28dc1118ec · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Gemini 2.5 pro: Our most intelligent ai model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb064953-eadb-486c-bd05-5b0e3ab59606 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f17ade13-c1bf-4358-9893-8e99c5fef9a9 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Making the v in VQA matter: Elevating the role of image understanding in visual question answering
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ac3070e-20df-47bc-9ca2-4d8c227e0e65 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5223b282-d58a-4cd3-a7d7-12bb6f221b0f · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d5a1ea1-65a1-4371-b8c7-ab92b984136f · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d8a4ee-d7b6-4213-a4f5-a3f1bce89de7 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f02b9289-7554-4051-ba3c-2161f2c635a1 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c9e3463-166b-4613-881a-2dbfec41ae86 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d802964-4990-42c3-a3a9-833eac9fc557 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e5c7f2-04d0-407f-88fc-59ac5af16301 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToRL: Scaling Tool-Integrated RL
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6adc218f-cc4e-4935-9596-8e6def33a041 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Visual Instruction Tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 085bd903-297a-4553-a88e-6c011123e860 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents AgentBench: Evaluating LLMs as Agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 742ae582-6b3c-4111-ac23-744c6737e7d6 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a88e7b80-dd04-41b1-a90f-4c9d997186d1 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cea61a5-de58-4ee5-a8a8-9174f04cd906 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53e93438-613d-47f7-bda1-d00dd6bc3ef8 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bcd00016-d170-4464-abb1-30000aab1cf8 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Browser use: Enable ai to control your browser, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3a0d1e3-f580-453e-beb9-1dfc4044fa32 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents WebGPT: Browser-assisted question-answering with human feedback
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c369baf7-f509-4b3b-9b3e-2e13c5623797 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Deep-captcha: a deep learning based captcha solver for vulnerability assessment, 2020
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cbb6873c-85fa-4b90-b060-8caa652f0848 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Openai o3 and o4-mini system card
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00261942-5dbe-4968-87d8-283ee522ff59 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander M ˛ adry, , et al
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12e7a838-09c1-443a-ac54-c9d83c113092 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29736ada-6fbe-413c-bfc1-7cd7b3432b3c · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Symbols of One-Loop Integrals From Mixed Tate Motives
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01e7ca81-d3a2-4dc8-af91-f8ea9b4dd923 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Autoplan: Automatic planning of interactive decision-making tasks with large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 602fb106-998e-48e3-8e3c-78549edb9631 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToolRL: Reward is All Tool Learning Needs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 774b04e1-77b9-4522-a558-b8866e1da4a4 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Plummer, Liwei Wang, Cristina Cervantes, Juan C
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c791ef3-4db7-4304-a588-3e1aef87e671 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Autogpt: An autonomous gpt-4 experiment
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6aa3af0b-9d82-40a7-9d69-c5b52652698d · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee478595-f1e9-4889-9292-680e2f3b3052 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Partially observable markov decision processes
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a8f005a-2bfa-4f72-b3bb-f717f73a7e5f · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Towards vqa models that can read, 2019
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d533068-f779-403b-b4fd-e6fedf97e8dd · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Phrasecut: Language grounding in images by text-based mask segmentation.ECCV, 2020
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06cac098-554d-4f85-97d7-084e3f74b64f · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents AdaPlanner: Adaptive Planning from Feedback with Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a70e694e-7803-4b19-a508-f58b603a9a32 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a684f4db-af1e-4cbd-9758-0b86697091d1 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Voyager: An Open-Ended Embodied Agent with Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c0b0e93-0804-4c73-af47-4ef6676fbd90 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382, 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeaf88a7-c605-4ead-ae1a-6f075e29978e · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents DeepSeek-V3 Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be8dd6d-a812-4ac5-920a-78dbcd120b33 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bd161c2-6d3c-4a36-99b9-a4db26cae643 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents An illusion of progress? assessing the current state of web agents, 2025
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ceaa361b-e936-4422-af1b-2d0fb9da548e · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Berg, and Tamara L
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24ae8380-6409-4b5b-8208-233f31bdc736 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Survey on evaluation of llm-based agents, 2025
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 162ccc53-e690-435f-88a0-cef2f6fc6705 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a051344-355f-4ef5-83e4-820e08d2ce20 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e1f3d13-c12b-43b4-ae24-d0ae1f1b2938 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents PubLayNet: largest dataset ever for document layout analysis
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7932a883-2bfe-4888-803c-20a7cea739a8 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0ebfb9b-12ba-458b-a62c-90823959b5cb · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a82974c-ce36-409f-a4b8-61390314f9de · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a877dd6-c95b-48b8-94c7-068352c75590 · outbound
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Unresolved cited work
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc9c335b-065f-405c-a134-7afa15c65c37 · inbound
Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f2c9b87-98f2-4dbd-b6e8-d8c6495768d7 · inbound
COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2f85212-f096-43df-936e-dca4295716bc · inbound
HLL: Can Agents Cross Humanity's Last Line of Verification? Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e9173ce-1855-42c5-9deb-ca06122ddbce · inbound
Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.