Pith. sign in

Paper Citation Record · LEDGER

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 4 inbound Pith citation observations for arXiv:2505.24878.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24878 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:27.016316Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:45:56.506288Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:56:19.580749Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved38
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2d9246e3-756a-499c-88e0-38d0731381b7 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.058110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.058110Z digest=sha256:4c09d79aa1ad3b3485cfbbd68a740b04e0f699a48b3a992f79f9eaf0d407bc74

Observation f1fde33c-9802-4c66-9479-4a51cca99e45 · outbound

This paper cites Claude 3.7 Sonnet System Card.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Claude 3.7 Sonnet System Card

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:30.038832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:22.148677Z digest=sha256:eb4f934c0adc187f092826f01fe4e54a8f33b50150edce67e6db54e1b3e151fc

Observation 8cd90733-1730-4365-a1f1-1da81cf6a2df · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.240512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.240512Z digest=sha256:a454321063a069fd091bbdd65f2c2f1b8fce190e75b232dad8a9d7f1e7d7be6d

Observation 8f992786-6b47-46f0-b163-09ba2149e385 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.413093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.413093Z digest=sha256:db5231c9bec78f9010621617b3449b5c654ebac737664c60efca09bc22f99df9

Observation a7884ebb-dedb-46c9-b75f-d74ce5e0840c · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.513582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.513582Z digest=sha256:33b48581f91dadc7c7c35da6a61a0da7dfe86e10ab9b2b26b777e01a0eb6db10

Observation 92da3802-4361-4744-80a0-7b8dd8df97c1 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.644201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.644201Z digest=sha256:c8306f078851770251db62f93bc75210f50fdf03f94d8f9a07cea478577800b3

Observation c7f316e4-5392-451e-afc9-0b28dc1118ec · outbound

This paper cites Gemini 2.5 pro: Our most intelligent ai model.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Gemini 2.5 pro: Our most intelligent ai model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.915407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:22.722596Z digest=sha256:f7c1f9d5a3aad5546bd02e486e9b32dbe78223ba7b64a39dd9dd5015ebed5884

Observation cb064953-eadb-486c-bd05-5b0e3ab59606 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.878232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.878232Z digest=sha256:1c3b6fd7aa08580f8a95acc8039f4017c39ecb5b49aae550b1153a2c507d283d

Observation f17ade13-c1bf-4358-9893-8e99c5fef9a9 · outbound

This paper cites Making the v in VQA matter: Elevating the role of image understanding in visual question answering.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Making the v in VQA matter: Elevating the role of image understanding in visual question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.588992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:22.979257Z digest=sha256:e9bfac94f989b86610187fa28f1255d87b39cf789edc6c8257660086f658e0e7

Observation 4ac3070e-20df-47bc-9ca2-4d8c227e0e65 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.076403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.076403Z digest=sha256:64e2de608c8c3a8a4c83a91de1086ab2efd4ea6af8ba3c979f2cba6a438b654b

Observation 5223b282-d58a-4cd3-a7d7-12bb6f221b0f · outbound

This paper cites PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.180088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.180088Z digest=sha256:5608a021bd97561e58d67ae3fa8af596fa30154161a51f700ea06e8854727750

Observation 4d5a1ea1-65a1-4371-b8c7-ab92b984136f · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.288396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.288396Z digest=sha256:ae1d786f7430aaa3162705bff50548c5dd971b878ffa7b16dc5cf780176471f0

Observation e3d8a4ee-d7b6-4213-a4f5-a3f1bce89de7 · outbound

This paper cites FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.358620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.358620Z digest=sha256:576668af8bf9bc3be6d663c7f803d84e3f48f53d7a76e300d44416b40dbe5a3a

Observation f02b9289-7554-4051-ba3c-2161f2c635a1 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.446173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.446173Z digest=sha256:6cf7bb8b704ff10a66db0f2fd8becb2961e55d615afe5359c002bb0000c83cef

Observation 7c9e3463-166b-4613-881a-2dbfec41ae86 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.507875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.507875Z digest=sha256:1a8afa4461a01888aca5798854ddf86bd234f4adb6d5ed19ffce51283402f577

Observation 4d802964-4990-42c3-a3a9-833eac9fc557 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.619007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.619007Z digest=sha256:8f6a17e23ed31c29ca160a7a6d999f22e81ddc026042ea637b518231d7324894

Observation f7e5c7f2-04d0-407f-88fc-59ac5af16301 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.665196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.665196Z digest=sha256:53d51c6c8228fe5be333f8a30b61b91aee3f2d9e7c920946c1f5d5582ba90699

Observation 6adc218f-cc4e-4935-9596-8e6def33a041 · outbound

This paper cites Visual Instruction Tuning.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Visual Instruction Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.747374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.747374Z digest=sha256:e07c83bf71b21e0216f7b66d0a6d6606d04f46b5fa7caa7d2849dc4a029382cd

Observation 085bd903-297a-4553-a88e-6c011123e860 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents AgentBench: Evaluating LLMs as Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.821960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.821960Z digest=sha256:203711d7f6378726d19464416ab971df10fa8af8ba2e898c4520eb4889f7a2a5

Observation 742ae582-6b3c-4111-ac23-744c6737e7d6 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.897548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.897548Z digest=sha256:3fc3ea4af788763aa19eb447e8074306be670c96ad91861d24f1cd06c791a443

Observation a88e7b80-dd04-41b1-a90f-4c9d997186d1 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:24.009766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:24.009766Z digest=sha256:b6017d821a7d90791b484d8e6df8af19e0b34baaff3c9d1b573927a84b727fc4

Observation 8cea61a5-de58-4ee5-a8a8-9174f04cd906 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.419922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.090500Z digest=sha256:319ba97ea55d1292bc10584d21dd970c729054aad7e860c4bd52352cea65712d

Observation 53e93438-613d-47f7-bda1-d00dd6bc3ef8 · outbound

This paper cites an unresolved cited work.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:16:29.297288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.177550Z digest=sha256:79f83e6cadabb25b204e6093a835467c4f7489456c4dcb9b3f65ee527e07501b

Observation bcd00016-d170-4464-abb1-30000aab1cf8 · outbound

This paper cites Browser use: Enable ai to control your browser, 2024.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Browser use: Enable ai to control your browser, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.192433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.252099Z digest=sha256:30c786f24f8852c98b18f08cf5a59ef20503b5a390e60bcc9d6c26e49e87c635

Observation f3a0d1e3-f580-453e-beb9-1dfc4044fa32 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents WebGPT: Browser-assisted question-answering with human feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:24.378965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:24.378965Z digest=sha256:3b3db75c274da1eafa60fd5e5c3cfae67bded8accbabfd024cc2eb5d9e6cf6a7

Observation c369baf7-f509-4b3b-9b3e-2e13c5623797 · outbound

This paper cites Deep-captcha: a deep learning based captcha solver for vulnerability assessment, 2020.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Deep-captcha: a deep learning based captcha solver for vulnerability assessment, 2020

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.084204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.467672Z digest=sha256:970ac342ce81a154455e8555f57f607c3e6a6439a30908f91b62293c4e670b87

Observation cbb6873c-85fa-4b90-b060-8caa652f0848 · outbound

This paper cites Openai o3 and o4-mini system card.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Openai o3 and o4-mini system card

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.977538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.581003Z digest=sha256:c13b0e291e6ccca1f2938f94643cf5c4890e8fa57fe0cc9ef92c0ccc91c78140

Observation 00261942-5dbe-4968-87d8-283ee522ff59 · outbound

This paper cites Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander M ˛ adry, , et al.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander M ˛ adry, , et al

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.674104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.772851Z digest=sha256:8e47acc7f70513053cffa87aef9144f2e559bd4b1101dd3878580143a309a2bc

Observation 12e7a838-09c1-443a-ac54-c9d83c113092 · outbound

This paper cites an unresolved cited work.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:16:28.841033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.679222Z digest=sha256:ce1690e5dfb505f9e23c0db862566d76a5da4c7e91a59eb9667edccfa51f964e

Observation 29736ada-6fbe-413c-bfc1-7cd7b3432b3c · outbound

This paper cites Symbols of One-Loop Integrals From Mixed Tate Motives.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Symbols of One-Loop Integrals From Mixed Tate Motives

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:24.914755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:24.914755Z digest=sha256:6ca69d1b4f67b0b4cf783cd1c614c28ab6d3a40ad3ed28ecde4202c9b2cd15c8

Observation 01e7ca81-d3a2-4dc8-af91-f8ea9b4dd923 · outbound

This paper cites Autoplan: Automatic planning of interactive decision-making tasks with large language models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Autoplan: Automatic planning of interactive decision-making tasks with large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.504648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.857751Z digest=sha256:2a8b14d56e8f14b6fd490333f7b8f1a5c6c0cfae13fe2b27f96c52470f420db0

Observation 602fb106-998e-48e3-8e3c-78549edb9631 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToolRL: Reward is All Tool Learning Needs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.104200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.104200Z digest=sha256:ca32060fe7845443b19167ea248d4702366adf723390c6d47e6c78bd045f84d3

Observation 774b04e1-77b9-4522-a558-b8866e1da4a4 · outbound

This paper cites Plummer, Liwei Wang, Cristina Cervantes, Juan C.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Plummer, Liwei Wang, Cristina Cervantes, Juan C

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.339740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:25.011371Z digest=sha256:5204d2e5e60466b490c7e350c6032cff532970989c979841d6aeb06721a02f55

Observation 9c791ef3-4db7-4304-a588-3e1aef87e671 · outbound

This paper cites Autogpt: An autonomous gpt-4 experiment.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Autogpt: An autonomous gpt-4 experiment

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.185113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:25.302039Z digest=sha256:42e3dcf070a1be44f213f4e094dde336b4c375e8e60acf1215b6a8b871f878b8

Observation 6aa3af0b-9d82-40a7-9d69-c5b52652698d · outbound

This paper cites StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.216062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.216062Z digest=sha256:33a85ae65998069fd9b2501fc0152c146901b60c77beba14f8f32b774751117d

Observation ee478595-f1e9-4889-9292-680e2f3b3052 · outbound

This paper cites Partially observable markov decision processes.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Partially observable markov decision processes

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.490284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.490284Z digest=sha256:19e2def3dedde5287ed73e4a96f7d9d525a9cb4ec013156faaad29fca5ade304

Observation 9a8f005a-2bfa-4f72-b3bb-f717f73a7e5f · outbound

This paper cites Towards vqa models that can read, 2019.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Towards vqa models that can read, 2019

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.390603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.390603Z digest=sha256:7fa4e7523feaa9feaac822cf62b86d6f376c620979fd9dbcede8770b5970bd29

Observation 1d533068-f779-403b-b4fd-e6fedf97e8dd · outbound

This paper cites Phrasecut: Language grounding in images by text-based mask segmentation.ECCV, 2020.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Phrasecut: Language grounding in images by text-based mask segmentation.ECCV, 2020

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.038738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:25.671645Z digest=sha256:932015698518e49f2f391f7f9531bb00f15de2561b5c430f98afd625f850708a

Observation 06cac098-554d-4f85-97d7-084e3f74b64f · outbound

This paper cites AdaPlanner: Adaptive Planning from Feedback with Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents AdaPlanner: Adaptive Planning from Feedback with Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.568370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.568370Z digest=sha256:a194c04c21e073e62bd16bdfcc987092a727666f527db1f02dd7f484863e5561

Observation a70e694e-7803-4b19-a508-f58b603a9a32 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.853689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.853689Z digest=sha256:45e03b356bd250f90365f43f50d1c7856e09fd43da19512bc72a797589ef5314

Observation a684f4db-af1e-4cbd-9758-0b86697091d1 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.770961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.770961Z digest=sha256:a8108f42ab1e803e2030f6a7a0a437f7d90c87cd416d90ddd8ea4ebdb304fee2

Observation 7c0b0e93-0804-4c73-af47-4ef6676fbd90 · outbound

This paper cites An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382, 2025.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.066956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.066956Z digest=sha256:94ff05cd5590903392d301477e19c357dd62810d0f0835b914e5fcec6614ccfd

Observation aeaf88a7-c605-4ead-ae1a-6f075e29978e · outbound

This paper cites DeepSeek-V3 Technical Report.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents DeepSeek-V3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.977120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.977120Z digest=sha256:baba76ef724622efa32ee28b5f216f0d5e63f5acbe5f7d31bd1128f3b46a441b

Observation 3be8dd6d-a812-4ac5-920a-78dbcd120b33 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.288484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.288484Z digest=sha256:853e17630b64d023410136aec04265c8eedeb75af477d776cb000d33c5504329

Observation 1bd161c2-6d3c-4a36-99b9-a4db26cae643 · outbound

This paper cites An illusion of progress? assessing the current state of web agents, 2025.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents An illusion of progress? assessing the current state of web agents, 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:27.859356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:26.170273Z digest=sha256:40c694c5b7fa192a183f3299bd1bc8c12fd6faac43b6adaa1f73da966e3810e2

Observation ceaa361b-e936-4422-af1b-2d0fb9da548e · outbound

This paper cites Berg, and Tamara L.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Berg, and Tamara L

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:27.494034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:26.499271Z digest=sha256:88bade6478b10834b85993a8392091653d934d72089f65abb69b63b224ee627d

Observation 24ae8380-6409-4b5b-8208-233f31bdc736 · outbound

This paper cites Survey on evaluation of llm-based agents, 2025.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Survey on evaluation of llm-based agents, 2025

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:27.702921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:26.402569Z digest=sha256:cebdeadeb7c8fea66ec6b2e005209465adb48d1f7bb7a342b80176036bd9659f

Observation 162ccc53-e690-435f-88a0-cef2f6fc6705 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.677705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.677705Z digest=sha256:d179bdc410653f279c2d3b18a6c928c85d8b14ce444b42f361c573e2f60f04c8

Observation 7a051344-355f-4ef5-83e4-820e08d2ce20 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.604788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.604788Z digest=sha256:ba71f113d05ad4007cec95101da5924bbec631b379ec4cfee2f88a0046ea31c3

Observation 5e1f3d13-c12b-43b4-ae24-d0ae1f1b2938 · outbound

This paper cites PubLayNet: largest dataset ever for document layout analysis.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents PubLayNet: largest dataset ever for document layout analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.842977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.842977Z digest=sha256:94019c0977264e653cf6faba0acaa8cba2553883f85b0cade4082e21449431f3

Observation 7932a883-2bfe-4888-803c-20a7cea739a8 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.754466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.754466Z digest=sha256:8e9310a66a6ddbccb5ee0f12146653e97aac2f5ef415c2b48d35a6efd70fd7cd

Observation b0ebfb9b-12ba-458b-a62c-90823959b5cb · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:27.016316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:27.016316Z digest=sha256:f6928449c0a0feb2b00397f1704290e7768bdb54d058aa9be19a86981c5b7c88

Observation 6a82974c-ce36-409f-a4b8-61390314f9de · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.935337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.935337Z digest=sha256:9bb052f8411b8fc639efe8c5330162df5bc3c2f3eee0c5cd782a9c46e7cf6005

Observation 3a877dd6-c95b-48b8-94c7-068352c75590 · outbound

This paper cites an unresolved cited work.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Unresolved cited work

Reference 2025

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T12:16:29.739225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:22.800470Z digest=sha256:fa311882dae8c249bda3adff548ab9f2ac14870143eb88260fdcfa5dadf508dd

Pith citing papers

Observation dc9c335b-065f-405c-a134-7afa15c65c37 · inbound

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges cites this paper.

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:42:22.039877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T03:42:10.703369Z digest=sha256:b5f53673344711a3dcb209df4039bdafb9a1e6f930514267d0849a87d8bce5e7

Observation 7f2c9b87-98f2-4dbd-b6e8-d8c6495768d7 · inbound

COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers cites this paper.

COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:23:57.983899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T03:23:03.198463Z digest=sha256:2fedb171f09af70f00e4209af1e109b1f6cef9e35b7260a18fb47004dbb3155a

Observation b2f85212-f096-43df-936e-dca4295716bc · inbound

HLL: Can Agents Cross Humanity's Last Line of Verification? cites this paper.

HLL: Can Agents Cross Humanity's Last Line of Verification? Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.583281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:57:57.218669Z digest=sha256:10f4a26d90dc7d770bafd1b8601f817352103b7a467c3c6a286fb7abc653dda3

Observation 1e9173ce-1855-42c5-9deb-ca06122ddbce · inbound

Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents cites this paper.

Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T14:45:56.506288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:45:56.506288Z digest=sha256:74338ad12a08f435471377bcce3a49c8356c7e97956721a111ddd9e2ae3b8b8f