Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-09T21:09:26.626179Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2604.21304.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-09T21:09:26.626179Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T10:54:54.558241Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-20T10:58:14.288775Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 46a2641c-806f-4a23-b03f-205b14c721a5 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 53f19e1b-54ad-423b-9bf5-7344e63cc7e6 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2023 , html =
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9b8cd3de-22c9-4b72-ad36-7dfb638359c6 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8d3f7822-cf79-4da0-8722-21ed503d1f11 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 141a2fe6-05f1-41b7-a45e-28777a61810a · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2024 , eprint=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation de68e949-ce2f-4575-8fcd-514dc73e75ac · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs PubMedQA: A Dataset for Biomedical Research Question Answering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ba0d5b4d-8fb4-47c4-b1a1-c53e8a2ad49f · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs PaperQA: Retrieval-Augmented Generative Agent for Scientific Research
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9627780f-6b84-47fb-9eb6-742b4115e5a6 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 97540a3c-8421-4ff0-b5c5-af206c1edb1a · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs M 3 S ci QA : A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1bbcf13c-0eae-4fc1-bb61-f7abe1ed8b08 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c83a5bf8-4e61-41f9-a3e1-473d765c17a0 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c87c2a54-0e94-4598-ad95-5329d53fba35 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs D oc A gent: An Agentic Framework for Multi-Modal Long-Context Document Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dddb219b-9982-4589-84bc-68285ca652b4 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs S ci DQA : A Deep Reading Comprehension Dataset over Scientific Papers
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7a087a80-fbb0-4dc4-bba6-28c9cdb6ee1b · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs SQuAI: Scientific Question-Answering with Multi-Agent Retrieval-Augmented Generation , year =
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f8b7f3c2-0d07-44dc-bc8e-0a86ef2e81e7 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Docgenome: An open large- scale scientific document benchmark for training and test- ing multi-modal large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9daba706-bfd3-44d2-8565-3dcfa0187c17 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Scientific Reports , year=
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a88082b0-4c86-428e-a4da-9762d5ca55bd · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 69e8a263-64be-4f79-bfa5-f5b5647826be · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs SCITAT : A Question Answering Benchmark for Scientific Tables and Text Covering Diverse Reasoning Types
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 324c5c4e-ba18-4801-9ccd-bcb865bdcedc · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Proceedings of the 40th International Conference on Machine Learning , year=
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5820d47d-927e-4e28-a6fb-293ed3c936ca · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3059b9e6-9fb6-492c-963b-4528b8fde27d · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs P eer QA : A Scientific Question Answering Dataset from Peer Reviews
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dba1fc4b-d086-4ecb-a445-aaac7d6b904b · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs ISBN 979-8-89176-251-0
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2c4e9284-b6d0-4bbf-8207-ab4a1b2bc4df · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Fact or Fiction: Verifying Scientific Claims
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a9eded5a-044c-4e29-bb8b-331e3e0f2ac2 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs EMNLP , year=
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 95203c37-29c6-478e-8b70-f3461e284fcc · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs URLhttps://doi.org/10.18653/v1/2021.naacl-main.365
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 13223644-feef-4d76-b88b-7c48004fc467 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs AutoGen: Enabling Next-Gen
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c13644c4-dab9-4cc8-b749-ad1a7a4ed836 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs author =
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 79ee0db5-9365-4c98-a3d1-f88b4464971b · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs The eleventh international conference on learning representations , year=
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 93592e2b-a5ca-40a5-b6a9-0cfb096db338 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2e6c59d4-9189-428e-afda-a33d41fe395c · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2024 , eprint=
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation de21906c-8242-4011-adba-e7db423ad342 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b859a663-b6a6-4845-b078-59c79ab53be4 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2026 , eprint=
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2de3d62f-593c-4328-96c7-f759253ea115 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 375ff912-ec44-478f-9746-10334a915519 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2026 , eprint=
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b4e7bb77-ddf9-4e57-a38e-24e8ba0d08e8 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Proceedings of the AAAI Conference on Artificial Intelligence , author=
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d7dafc78-4151-4b30-9695-6d0710688840 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2026 , eprint=
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9cf19d4d-6853-4292-834c-9567e21fb707 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2cd4d8e2-4436-4c4c-aab0-2ef90973c9ce · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2026 , eprint=
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 97f65ef6-38da-4fb3-93b8-3a9df5748bb8 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 720c5091-6c6f-4f3b-846b-51229b15e22d · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8aaba336-5b8b-4677-a11a-b744ac728d83 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs and Staar, Peter , title =
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c4b8fe5e-2089-456c-b0f5-2e50111db920 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d328c1f4-6490-49b2-913e-6f76a26ac216 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs On the Use of ArXiv as a Dataset
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b22cd762-7425-41e5-832d-719e8d6f5e86 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Zhang, Z
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2cf12307-a70e-41da-8af0-c36a875a6a96 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e0b67d88-4b95-46b9-81db-bd425bb1f31b · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Claude-3 Model Card , volume=
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 19c2f941-5081-4f00-aa73-a530f813ca46 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs GPT-4o System Card
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b300b491-7e79-43d1-8661-b5b01430ec3d · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2024 , eprint=
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fc93264e-8eb1-472a-8ad0-28004753b011 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 645186cf-2f21-4da5-aab7-bc197ad03471 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 147d3036-ecba-49d6-90b8-1c6e02dbb9d1 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e290d9db-733b-40ee-8fdb-81ae5b8bca54 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Large language models are not fair evaluators
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fe7a4b0b-4404-4615-bc13-f361870dc69c · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs and Zhang, Hao and Gonzalez, Joseph E
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 476c9413-af3f-44dd-a3cf-b2ff7249e77e · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs PDFFigures 2.0: Mining figures from research papers , year=
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 54237905-2490-435c-9ece-12740a36ad87 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Toolformer: language models can teach themselves to use tools , year =
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d4e29a83-00d9-4ab1-9afb-6b6e764ff6ed · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Proceedings of the 40th International Conference on Machine Learning , articleno =
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7c88c490-5df4-437a-93f2-94aca270da30 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 336d0b29-a025-4777-971b-bc0ff62b5260 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs PeerQA: A Scientific Question Answering Dataset from Peer Reviews
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 582b4fdc-f9e7-4a14-a6e1-6f9ada859f4e · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Advances in Neural Information Processing Systems , volume=
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c2588cef-a207-4d07-90f0-18f1591a9b24 · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2024 , url=
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7e5c9dc7-c8d1-44fc-b453-7e8875b1cb5e · outbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs S ci VQA 2025: Overview of the First Scientific Visual Question Answering Shared Task
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e4f3cd5c-4d9c-4be2-81ed-dac18813f480 · inbound
Code as Agent Harness PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs
Reference 208
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.