Pith. sign in

Paper Citation Record · LEDGER

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 4 inbound Pith citation observations for arXiv:2505.24878.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24878 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:27.016316Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:45:56.506288Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:56:19.580749Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved38
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2d9246e3-756a-499c-88e0-38d0731381b7 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.058110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.058110Z digest=sha256:0a3f3e65a768e6e75607ed2decb04516a358d3c6e1518346a40fef42e4910348

Observation f1fde33c-9802-4c66-9479-4a51cca99e45 · outbound

This paper cites Claude 3.7 Sonnet System Card.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Claude 3.7 Sonnet System Card

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:30.038832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:22.148677Z digest=sha256:d79d6cc7207b503767e24817b03f37f3c9d68f7b259f05369e3a8df1ec0801bc

Observation 8cd90733-1730-4365-a1f1-1da81cf6a2df · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.240512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.240512Z digest=sha256:e73a3b9d2a71faaf4a4c4a363a732e791bd9b162b4e32a82c6a83a18ccda97ea

Observation 8f992786-6b47-46f0-b163-09ba2149e385 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.413093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.413093Z digest=sha256:bf0734b2f4d8184fe5081e8cd5d283265c5b83318644ac649789c5661fe4dba9

Observation a7884ebb-dedb-46c9-b75f-d74ce5e0840c · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.513582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.513582Z digest=sha256:8c98a148d323112d5b960bc6dcc56a9652893ebaf6c4008e92f2c3e67373d894

Observation 92da3802-4361-4744-80a0-7b8dd8df97c1 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.644201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.644201Z digest=sha256:d18207c31720c07d25634dfad99b4d903703f6345cfa516d1e493c4af116d579

Observation c7f316e4-5392-451e-afc9-0b28dc1118ec · outbound

This paper cites Gemini 2.5 pro: Our most intelligent ai model.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Gemini 2.5 pro: Our most intelligent ai model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.915407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:22.722596Z digest=sha256:ed92100cd2917e99dbff4041cc0005cb8e2f5e974169ac55469e6897e6ad55f9

Observation cb064953-eadb-486c-bd05-5b0e3ab59606 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.878232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.878232Z digest=sha256:78164149774c116fd63e460462980effad507838979b8bad8cdc058cfc556ad1

Observation f17ade13-c1bf-4358-9893-8e99c5fef9a9 · outbound

This paper cites Making the v in VQA matter: Elevating the role of image understanding in visual question answering.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Making the v in VQA matter: Elevating the role of image understanding in visual question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.588992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:22.979257Z digest=sha256:61611b377f260862b49412882630e48acf3a35dceaf55d1c31b47e6436885406

Observation 4ac3070e-20df-47bc-9ca2-4d8c227e0e65 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.076403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.076403Z digest=sha256:e8a216deea3962b8b38b5b10a9e27e2516916f27276ababcb80fb7a85c0c564a

Observation 5223b282-d58a-4cd3-a7d7-12bb6f221b0f · outbound

This paper cites PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.180088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.180088Z digest=sha256:ef01caaac282aded039d9729e3b8b5584f1a685dd9474934c2d24fa780dbf055

Observation 4d5a1ea1-65a1-4371-b8c7-ab92b984136f · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.288396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.288396Z digest=sha256:c17fde8b7ca1ee4b89e6ad2313247a4c8cbe9e9a5e14563feb941711d56b73f8

Observation e3d8a4ee-d7b6-4213-a4f5-a3f1bce89de7 · outbound

This paper cites FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.358620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.358620Z digest=sha256:fdbe58f33f4d232d3d252fea7519df6fdee568affbf91e5fc1a304971840c9e7

Observation f02b9289-7554-4051-ba3c-2161f2c635a1 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.446173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.446173Z digest=sha256:1389fa1a2a299a9d546b5d580b08d30c15dd591b29544f1c673450736eb7bbc0

Observation 7c9e3463-166b-4613-881a-2dbfec41ae86 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.507875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.507875Z digest=sha256:4706a2a8bfbded4228c50aa22c84c7ba33946dc1eaea1dd78fe1e6275490db09

Observation 4d802964-4990-42c3-a3a9-833eac9fc557 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.619007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.619007Z digest=sha256:4217274b040a17ac6ce5e32922923c2e61c0703f92a1c9223fbda1d2d8978336

Observation f7e5c7f2-04d0-407f-88fc-59ac5af16301 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.665196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.665196Z digest=sha256:d739c7f109fadbaa96c0f6ca58e226091c6ce54a569dc653fd9de811cb04dbdb

Observation 6adc218f-cc4e-4935-9596-8e6def33a041 · outbound

This paper cites Visual Instruction Tuning.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Visual Instruction Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.747374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.747374Z digest=sha256:a3526567537b86db0ddaa600cb4f46674b4ad01c121d61728e56a2c034212ffa

Observation 085bd903-297a-4553-a88e-6c011123e860 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents AgentBench: Evaluating LLMs as Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.821960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.821960Z digest=sha256:e5ead61b4f7efb5b0efaf5a7428256bef2b2bfb69e81564a41a4ef3e471c836d

Observation 742ae582-6b3c-4111-ac23-744c6737e7d6 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.897548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.897548Z digest=sha256:7c5ace97c43f492cec003b3c7d495bcacb8839f42a805842990d7ec5994bb6a5

Observation a88e7b80-dd04-41b1-a90f-4c9d997186d1 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:24.009766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:24.009766Z digest=sha256:d9d958f8903a8863351dbf3582fee4f52e46cf54b06e9a9f45d886a09d109cdd

Observation 8cea61a5-de58-4ee5-a8a8-9174f04cd906 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.419922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.090500Z digest=sha256:5a6cd83761baa3f77b248e1f235a6df4d0a53172242271133896af330c947df4

Observation 53e93438-613d-47f7-bda1-d00dd6bc3ef8 · outbound

This paper cites an unresolved cited work.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:16:29.297288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.177550Z digest=sha256:5b9ed34cfb8f4408df4b22faf6c91a715b803050db5fa15885efede2d03a9b66

Observation bcd00016-d170-4464-abb1-30000aab1cf8 · outbound

This paper cites Browser use: Enable ai to control your browser, 2024.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Browser use: Enable ai to control your browser, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.192433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.252099Z digest=sha256:cc9849b73733e3ab8d2b9ccb5233c62309176dfbcb3b65d576f534c321e100e4

Observation f3a0d1e3-f580-453e-beb9-1dfc4044fa32 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents WebGPT: Browser-assisted question-answering with human feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:24.378965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:24.378965Z digest=sha256:38dd7020a4b0eaa6a5240fc4250cd6ad22a3aa9e0ffdf38a0aec8f2fb46586d7

Observation c369baf7-f509-4b3b-9b3e-2e13c5623797 · outbound

This paper cites Deep-captcha: a deep learning based captcha solver for vulnerability assessment, 2020.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Deep-captcha: a deep learning based captcha solver for vulnerability assessment, 2020

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.084204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.467672Z digest=sha256:901effa180518cfd10f17ef1d0be634be572180c119486448560ceea871e8106

Observation cbb6873c-85fa-4b90-b060-8caa652f0848 · outbound

This paper cites Openai o3 and o4-mini system card.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Openai o3 and o4-mini system card

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.977538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.581003Z digest=sha256:296117735220f55e741ffdeb739830565c7804ec9d0ff09368ae218cf21e81ce

Observation 00261942-5dbe-4968-87d8-283ee522ff59 · outbound

This paper cites Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander M ˛ adry, , et al.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander M ˛ adry, , et al

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.674104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.772851Z digest=sha256:6cd5334c007d01a606120ce2d2c7c67c9645a804bc724bf8bb610bb8058453ad

Observation 12e7a838-09c1-443a-ac54-c9d83c113092 · outbound

This paper cites an unresolved cited work.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:16:28.841033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.679222Z digest=sha256:3bbcd67ccb786d1b556e6002705f55ea333d67218e6b4d8f0100eb16f9451952

Observation 29736ada-6fbe-413c-bfc1-7cd7b3432b3c · outbound

This paper cites Symbols of One-Loop Integrals From Mixed Tate Motives.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Symbols of One-Loop Integrals From Mixed Tate Motives

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:24.914755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:24.914755Z digest=sha256:4f39e82401c0c8463129bf017e0e06c5b7666d022acb1be43449f12c3e15fd91

Observation 01e7ca81-d3a2-4dc8-af91-f8ea9b4dd923 · outbound

This paper cites Autoplan: Automatic planning of interactive decision-making tasks with large language models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Autoplan: Automatic planning of interactive decision-making tasks with large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.504648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:24.857751Z digest=sha256:9b0ddb3a77c0d7879c6454a23ac8f2abe47e38e1e1a7ff446e0438b7e82567f7

Observation 602fb106-998e-48e3-8e3c-78549edb9631 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToolRL: Reward is All Tool Learning Needs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.104200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.104200Z digest=sha256:da1cbfe703eda0b937628f40f66da142133915932349f907f1d6c9ec93bba32e

Observation 774b04e1-77b9-4522-a558-b8866e1da4a4 · outbound

This paper cites Plummer, Liwei Wang, Cristina Cervantes, Juan C.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Plummer, Liwei Wang, Cristina Cervantes, Juan C

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.339740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:25.011371Z digest=sha256:9f1b221840e410cc147d04a9917e2707c2f5fd1f710003c40db65c5086ee20b3

Observation 9c791ef3-4db7-4304-a588-3e1aef87e671 · outbound

This paper cites Autogpt: An autonomous gpt-4 experiment.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Autogpt: An autonomous gpt-4 experiment

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.185113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:25.302039Z digest=sha256:3ab6f8f4d2f10ee7869d7c8d2408f786087ff95c4bb64f943f0c1305e60b869f

Observation 6aa3af0b-9d82-40a7-9d69-c5b52652698d · outbound

This paper cites StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.216062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.216062Z digest=sha256:ca535c9b49e6b681856f5823c3b3f2710ce80ef6bdc540d8a37bc5062af3ca2b

Observation ee478595-f1e9-4889-9292-680e2f3b3052 · outbound

This paper cites Partially observable markov decision processes.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Partially observable markov decision processes

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.490284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.490284Z digest=sha256:82a3e08dbc6c7a18a48ac4fba84a3852e5437010e9adf7f5d58a870c2edc5e7c

Observation 9a8f005a-2bfa-4f72-b3bb-f717f73a7e5f · outbound

This paper cites Towards vqa models that can read, 2019.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Towards vqa models that can read, 2019

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.390603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.390603Z digest=sha256:677510dc161ff192b3c0e8c8bab7c843aa52485dcde351f2c5c3ff62aaa5ac5f

Observation 1d533068-f779-403b-b4fd-e6fedf97e8dd · outbound

This paper cites Phrasecut: Language grounding in images by text-based mask segmentation.ECCV, 2020.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Phrasecut: Language grounding in images by text-based mask segmentation.ECCV, 2020

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.038738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:25.671645Z digest=sha256:cb8d71828e8419023260198f0fd9b9aac753ba3af50ffa269eef8581b5823ca7

Observation 06cac098-554d-4f85-97d7-084e3f74b64f · outbound

This paper cites AdaPlanner: Adaptive Planning from Feedback with Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents AdaPlanner: Adaptive Planning from Feedback with Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.568370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.568370Z digest=sha256:b0301e466d2f3fbebd0b60b6d936804208cb3668b94bb197407cec5ba9deb832

Observation a70e694e-7803-4b19-a508-f58b603a9a32 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.853689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.853689Z digest=sha256:9f813266623f35236820020b2d434e8edecc6a223bcf01dbfa5a903dd8d1c708

Observation a684f4db-af1e-4cbd-9758-0b86697091d1 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.770961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.770961Z digest=sha256:41447b3124320d079e6efdac2f2464389f4f76d5ca79ba0ce1a20a93040cd063

Observation 7c0b0e93-0804-4c73-af47-4ef6676fbd90 · outbound

This paper cites An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382, 2025.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.066956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.066956Z digest=sha256:b14a4fe437cc653600f3140dd2f66122098b9b7c5fa23896fa607a599bbf6585

Observation aeaf88a7-c605-4ead-ae1a-6f075e29978e · outbound

This paper cites DeepSeek-V3 Technical Report.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents DeepSeek-V3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.977120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.977120Z digest=sha256:3058e94cc7e374b330ac77e8d02c5a0e16fee042e9268f5fd819d82cf6dd067e

Observation 3be8dd6d-a812-4ac5-920a-78dbcd120b33 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.288484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.288484Z digest=sha256:cdcdc860da68c85552476ed01438a515409fe9ef6d6b0c8ba0c075fc82271a6d

Observation 1bd161c2-6d3c-4a36-99b9-a4db26cae643 · outbound

This paper cites An illusion of progress? assessing the current state of web agents, 2025.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents An illusion of progress? assessing the current state of web agents, 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:27.859356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:26.170273Z digest=sha256:3bd8d032374d74a8400412a8e7c4b3326fa8ebbd70f8c89b7e61a15b2902f2c1

Observation ceaa361b-e936-4422-af1b-2d0fb9da548e · outbound

This paper cites Berg, and Tamara L.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Berg, and Tamara L

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:27.494034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:26.499271Z digest=sha256:d6eb9ba4a851a309d581662ae2c83cdb2207d658bec86baddff76aed870d74c8

Observation 24ae8380-6409-4b5b-8208-233f31bdc736 · outbound

This paper cites Survey on evaluation of llm-based agents, 2025.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Survey on evaluation of llm-based agents, 2025

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:27.702921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:26.402569Z digest=sha256:61239a0855d9bd7bbf7822236f45e79cda8c81c012636c4b94a6c8bc434fb37f

Observation 162ccc53-e690-435f-88a0-cef2f6fc6705 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.677705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.677705Z digest=sha256:c738454161df4227fcf86510dbe057fcbba3a4596726c63f57df18e115a15c11

Observation 7a051344-355f-4ef5-83e4-820e08d2ce20 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.604788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.604788Z digest=sha256:77b18d05033e5610f0904da53b7b5001ee46fc9d62036f8f6e35696ac75f78f3

Observation 5e1f3d13-c12b-43b4-ae24-d0ae1f1b2938 · outbound

This paper cites PubLayNet: largest dataset ever for document layout analysis.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents PubLayNet: largest dataset ever for document layout analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.842977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.842977Z digest=sha256:a5fc9e6aff76d5f914acadc1fd2cf9454621126f10732beeb1da3b4228374e40

Observation 7932a883-2bfe-4888-803c-20a7cea739a8 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.754466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.754466Z digest=sha256:91657e4d19776e8aa23a7a838cecb9bc565f828f8fa0874ef00aba799846f5f3

Observation b0ebfb9b-12ba-458b-a62c-90823959b5cb · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:27.016316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:27.016316Z digest=sha256:c712ebf065fa688cc2a1fac8515ee6d9f12ce7825ca2a3c7e129fd34ce7b4680

Observation 6a82974c-ce36-409f-a4b8-61390314f9de · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.935337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.935337Z digest=sha256:c673eb2f073716eb6f3d45df58e3a7a51b1faae4d061c4159163f25aa4cf55a0

Observation 3a877dd6-c95b-48b8-94c7-068352c75590 · outbound

This paper cites an unresolved cited work.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Unresolved cited work

Reference 2025

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T12:16:29.739225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:16:22.800470Z digest=sha256:4399c6e436cd64a3acf72cbaed0b546a177b74b723b5b7b7bdcb6633f2baf69a

Pith citing papers

Observation dc9c335b-065f-405c-a134-7afa15c65c37 · inbound

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges cites this paper.

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:42:22.039877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T03:42:10.703369Z digest=sha256:844cf3a46b571c5ef0e75e1f1eb8cad1d9b6329520583d2f71a039dda6f9c04b

Observation 7f2c9b87-98f2-4dbd-b6e8-d8c6495768d7 · inbound

COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers cites this paper.

COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:23:57.983899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T03:23:03.198463Z digest=sha256:cf8f68c54e0b3376df91f14c93fd0294e01613c249b781c0dabb2d2ddee4d1e7

Observation b2f85212-f096-43df-936e-dca4295716bc · inbound

HLL: Can Agents Cross Humanity's Last Line of Verification? cites this paper.

HLL: Can Agents Cross Humanity's Last Line of Verification? Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.583281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:57:57.218669Z digest=sha256:04c16e31565bc91f4b0c2e52f342188bf8a641d7a55f4117d28ae97935f931b7

Observation 1e9173ce-1855-42c5-9deb-ca06122ddbce · inbound

Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents cites this paper.

Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T14:45:56.506288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:45:56.506288Z digest=sha256:e8360c168150a2899970534cf0bbe63a9619f1396765d4a2ebcf430e747a80d3