Pith. sign in

Paper Citation Record · LEDGER

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs

As of 5 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2604.21304.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.21304 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T21:09:26.626179Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T10:54:54.558241Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-20T10:58:14.288775Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact13
  • verified fuzzy34
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch13

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46a2641c-806f-4a23-b03f-205b14c721a5 · outbound

This paper cites MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:41:53.825825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:f184f39232b83c2724b8464cf606672e343a5f9716884cfd045de98b6107a234

Observation 53f19e1b-54ad-423b-9bf5-7344e63cc7e6 · outbound

This paper cites 2023 , html =.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2023 , html =

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.513154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:b8c6227256a48f856c01754795b078e66da64e87b03a4f1221d79071516dcf2f

Observation 9b8cd3de-22c9-4b72-ad36-7dfb638359c6 · outbound

This paper cites MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T07:31:08.404728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:6c5d92ed4d6a3fc2a54c8a997e6e83bb0dc2f1010ef15902f530b4910e0a6369

Observation 8d3f7822-cf79-4da0-8722-21ed503d1f11 · outbound

This paper cites Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.428262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:6ed178db9a1c06a2306e14c060ff653917756be03d4c2fcb71d96fb91bc29a9e

Observation 141a2fe6-05f1-41b7-a45e-28777a61810a · outbound

This paper cites 2024 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2024 , eprint=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.414059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:03483e4adb351277c9aa57db849f71dc641ba988ac6ce1fce451fdf61083476e

Observation de68e949-ce2f-4575-8fcd-514dc73e75ac · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 6

Resolution
verified exact
doi, observed 2026-05-09T21:13:24.035725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:f81c729176ec5bf19a3cab83aa50574733e6234b846f3e02f8a507469afaf833

Observation ba0d5b4d-8fb4-47c4-b1a1-c53e8a2ad49f · outbound

This paper cites PaperQA: Retrieval-Augmented Generative Agent for Scientific Research.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs PaperQA: Retrieval-Augmented Generative Agent for Scientific Research

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:54.269387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:73bee63b321767a8787ddfb85716b20db00f261002bb2c13641ba852eb84a3f0

Observation 9627780f-6b84-47fb-9eb6-742b4115e5a6 · outbound

This paper cites 2025 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.455450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:b1391834f9bac82d92074bddd9cd6a63618f8c81d4b5f4bd1d1eb2c9f6de6544

Observation 97540a3c-8421-4ff0-b5c5-af206c1edb1a · outbound

This paper cites M 3 S ci QA : A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs M 3 S ci QA : A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models

Reference 9

Resolution
verified exact
doi, observed 2026-05-09T21:13:24.021481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:d67da5f69c41e782efe590e648885bb04406d8ce1a3c9f9dbf1b5caec7f5c528

Observation 1bbcf13c-0eae-4fc1-bb61-f7abe1ed8b08 · outbound

This paper cites M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:54.539181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:2ad4de7dd912356955ec18e15ec7dba2b897bb9b1abc6392a90e3fb3b9c2c1cf

Observation c83a5bf8-4e61-41f9-a3e1-473d765c17a0 · outbound

This paper cites 2025 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.488671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:f3badbf70807a58cf2f3b166b28e784766afe7d97a8c0a7de9821de60bcfe5d7

Observation c87c2a54-0e94-4598-ad95-5329d53fba35 · outbound

This paper cites D oc A gent: An Agentic Framework for Multi-Modal Long-Context Document Understanding.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs D oc A gent: An Agentic Framework for Multi-Modal Long-Context Document Understanding

Reference 12

Resolution
verified exact
doi, observed 2026-05-09T21:13:24.023616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:78ef70241e79f02d36fda335be4fa25aef7603c43d2ad49487d11e1d6e719a68

Observation dddb219b-9982-4589-84bc-68285ca652b4 · outbound

This paper cites S ci DQA : A Deep Reading Comprehension Dataset over Scientific Papers.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs S ci DQA : A Deep Reading Comprehension Dataset over Scientific Papers

Reference 13

Resolution
verified exact
doi, observed 2026-05-09T21:13:24.025776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:de5ba5d3a220fcab3c308dc8606423f519a5a8723147c20f934a0bfc048cf5b9

Observation 7a087a80-fbb0-4dc4-bba6-28c9cdb6ee1b · outbound

This paper cites SQuAI: Scientific Question-Answering with Multi-Agent Retrieval-Augmented Generation , year =.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs SQuAI: Scientific Question-Answering with Multi-Agent Retrieval-Augmented Generation , year =

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.424590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:437c4297350c577a245770c5bb88d4cfd3a461f1e48253e8c93a4c7dcd33968e

Observation f8b7f3c2-0d07-44dc-bc8e-0a86ef2e81e7 · outbound

This paper cites Docgenome: An open large- scale scientific document benchmark for training and test- ing multi-modal large language models.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Docgenome: An open large- scale scientific document benchmark for training and test- ing multi-modal large language models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:54.909380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:0957288ca3e9da23ea7bc03786f563eac5919d1985950c2ea6e9dd1116da09ba

Observation 9daba706-bfd3-44d2-8565-3dcfa0187c17 · outbound

This paper cites Scientific Reports , year=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Scientific Reports , year=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.406668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:bf6231c7f5356f4e55ed7623f1f44726cfab9cc344888f135cf6a0b87a36bb77

Observation a88082b0-4c86-428e-a4da-9762d5ca55bd · outbound

This paper cites 2025 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.417757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:f2de6f796ae1281e586061b6e6b91fb68207177d7287a681f7cc4ed0c84f1b92

Observation 69e8a263-64be-4f79-bfa5-f5b5647826be · outbound

This paper cites SCITAT : A Question Answering Benchmark for Scientific Tables and Text Covering Diverse Reasoning Types.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs SCITAT : A Question Answering Benchmark for Scientific Tables and Text Covering Diverse Reasoning Types

Reference 18

Resolution
verified exact
doi, observed 2026-05-09T21:13:24.032760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:685c5933c8735e52bfe2d812774e0ce9aabed13891279fa2c1b30df4dd2dcb94

Observation 324c5c4e-ba18-4801-9ccd-bcb865bdcedc · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning , year=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Proceedings of the 40th International Conference on Machine Learning , year=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.436691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:0b0070a72f77e780a49669bda0ae934622b6c7f6a8b63d15856d0f480107decc

Observation 5820d47d-927e-4e28-a6fb-293ed3c936ca · outbound

This paper cites 2025 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.440732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:351d9622d0979b9f003c862dfc3e04608592be7b96f15d72450296334401db9a

Observation 3059b9e6-9fb6-492c-963b-4528b8fde27d · outbound

This paper cites P eer QA : A Scientific Question Answering Dataset from Peer Reviews.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs P eer QA : A Scientific Question Answering Dataset from Peer Reviews

Reference 21

Resolution
verified exact
doi, observed 2026-05-09T21:13:24.053015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:5fb560ca2f83fd1c81cea8f67c49ca288f20cbeb6b35c5f22b3f1d0ef3df4bc8

Observation dba1fc4b-d086-4ecb-a445-aaac7d6b904b · outbound

This paper cites ISBN 979-8-89176-251-0.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs ISBN 979-8-89176-251-0

Reference 22

Resolution
metadata mismatch
doi, observed 2026-05-09T21:13:24.050553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:21f1d4d5007bdb75fa6c2c05f402950a621ccbc197a7da582109eabce9f5641c

Observation 2c4e9284-b6d0-4bbf-8207-ab4a1b2bc4df · outbound

This paper cites Fact or Fiction: Verifying Scientific Claims.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Fact or Fiction: Verifying Scientific Claims

Reference 23

Resolution
verified exact
doi, observed 2026-05-09T21:13:24.055658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:e092186611eb79f64ccc6e0d0e4e3fb8b9c4bcf039025a1523e5d3b7dc9a87cb

Observation a9eded5a-044c-4e29-bb8b-331e3e0f2ac2 · outbound

This paper cites EMNLP , year=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs EMNLP , year=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.459762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:36844276dda21ab83762b3567a2d2934c12447afdedb33251358872bcd3db0d5

Observation 95203c37-29c6-478e-8b70-f3461e284fcc · outbound

This paper cites URLhttps://doi.org/10.18653/v1/2021.naacl-main.365.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs URLhttps://doi.org/10.18653/v1/2021.naacl-main.365

Reference 25

Resolution
verified exact
doi, observed 2026-05-09T21:13:24.048285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:71204291cc1b17f62835d63b8f7f3f32b64a413e31b3e180607d13811a5edc48

Observation 13223644-feef-4d76-b88b-7c48004fc467 · outbound

This paper cites AutoGen: Enabling Next-Gen.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs AutoGen: Enabling Next-Gen

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.470297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:2b7019ee6770cd85fc277e4813d7e2370a53bb999a7c384c312a86d1180f7b32

Observation c13644c4-dab9-4cc8-b749-ad1a7a4ed836 · outbound

This paper cites author =.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs author =

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.474545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:eaad7e4470766b4f11a7011cbf38f113feedc565c3f13a366fd32f118919d4ae

Observation 79ee0db5-9365-4c98-a3d1-f88b4464971b · outbound

This paper cites The eleventh international conference on learning representations , year=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs The eleventh international conference on learning representations , year=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.499408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:40600179c98bc4b29fcb50cf7d09cd9548a8f671040f4542cd1c2406feea1f49

Observation 93592e2b-a5ca-40a5-b6a9-0cfb096db338 · outbound

This paper cites 2025 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.509527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:8030541aa8a14028d403e71214e63cbb3cb317457b2b006710ed05351f17e39d

Observation 2e6c59d4-9189-428e-afda-a33d41fe395c · outbound

This paper cites 2024 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2024 , eprint=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.492050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:493741224e3da58fdba9c25c0b2a31d0eeebbecd20be7ccf8d6d521a90c67f8f

Observation de21906c-8242-4011-adba-e7db423ad342 · outbound

This paper cites 2025 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.495536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:f74370cd3298e6c7220b5454c9403033a1939e79a57de932db865fb4f502c476

Observation b859a663-b6a6-4845-b078-59c79ab53be4 · outbound

This paper cites 2026 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2026 , eprint=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.502517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:74f41f4cd98f777e8442bb162a231d12f718d92a9f80caae35881fff6cb4b599

Observation 2de3d62f-593c-4328-96c7-f759253ea115 · outbound

This paper cites 2025 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.485592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:567f7962d785735120bce3a15fe3f9b5347cef43deaa5fbe5dd82fe6dc366c96

Observation 375ff912-ec44-478f-9746-10334a915519 · outbound

This paper cites 2026 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2026 , eprint=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.466486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:a357f87ee40637e2aef1b66717af5d8c5a7e42bf7d4c6f93d793900659e81a98

Observation b4e7bb77-ddf9-4e57-a38e-24e8ba0d08e8 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , author=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Proceedings of the AAAI Conference on Artificial Intelligence , author=

Reference 35

Resolution
verified exact
doi, observed 2026-05-09T21:13:24.040842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:070495082f88538559efe6c104cab96443cdfd62646979fae0f6dc53fd6ecfad

Observation d7dafc78-4151-4b30-9695-6d0710688840 · outbound

This paper cites 2026 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2026 , eprint=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.463015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:90a45b95d73cca376f0bb2227c86a84ff839d5575d57a98410a6297f6294be99

Observation 9cf19d4d-6853-4292-834c-9567e21fb707 · outbound

This paper cites 2025 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.478242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:51cde0941edaf80be3fb11e1b0c5125a5563245bbb8bc89242d05808783e3819

Observation 2cd4d8e2-4436-4c4c-aab0-2ef90973c9ce · outbound

This paper cites 2026 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2026 , eprint=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.481900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:3f3bdf765c67141fe27f4be4a9efeb8e261a888293cd3cb019c600ff6ab8fe46

Observation 97f65ef6-38da-4fb3-93b8-3a9df5748bb8 · outbound

This paper cites 2025 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.448732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:cc8f017a073b0c8fde529e5d9a1e8264b62b1781b9dbd9b350d8325da886e0f4

Observation 720c5091-6c6f-4f3b-846b-51229b15e22d · outbound

This paper cites an unresolved cited work.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-23T16:18:11.396395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:5d3a00bb9572040e90e58d4c5ca8696ff35a2ee54fc21ade818cb0275c48dff0

Observation 8aaba336-5b8b-4677-a11a-b744ac728d83 · outbound

This paper cites and Staar, Peter , title =.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs and Staar, Peter , title =

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T21:13:24.046083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:921faf67432107af7e7435a02e425856262284a72960782139f2ba199a95c58b

Observation c4b8fe5e-2089-456c-b0f5-2e50111db920 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:41:55.154858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:4f16525d948621300f240f0f6692c6d101fb8aea575bfe4561bf5978a3a5e394

Observation d328c1f4-6490-49b2-913e-6f76a26ac216 · outbound

This paper cites On the Use of ArXiv as a Dataset.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs On the Use of ArXiv as a Dataset

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:54.765786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:4f802fa9057a351b33fc3e7546d7f04503e884a8f1804349987cdc980c915371

Observation b22cd762-7425-41e5-832d-719e8d6f5e86 · outbound

This paper cites Zhang, Z.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Zhang, Z

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:53.935668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:bb8284864d539d8e435c9798280110bb05a5c315759150147eb3279c0e150363

Observation 2cf12307-a70e-41da-8af0-c36a875a6a96 · outbound

This paper cites Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:54.443885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:31f4b0d9f9fc014ff0419669a4f1be3cf1c0c1a458c24daa0f684a66219e6e5f

Observation e0b67d88-4b95-46b9-81db-bd425bb1f31b · outbound

This paper cites Claude-3 Model Card , volume=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Claude-3 Model Card , volume=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.392661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:1a33fa82485e7a4a0de170382792d0748f8a4d8f429c4a24cbed305f95371748

Observation 19c2f941-5081-4f00-aa73-a530f813ca46 · outbound

This paper cites GPT-4o System Card.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs GPT-4o System Card

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:41:54.099113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:96f8ce637e242c960396a3b5397e232f9301470c09641e905344e0a869424848

Observation b300b491-7e79-43d1-8661-b5b01430ec3d · outbound

This paper cites 2024 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2024 , eprint=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.399919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:2e5c3339f0211f36a5d51a123e90c3838e0dfdd0d717cffcb753ba6bb77716cb

Observation fc93264e-8eb1-472a-8ad0-28004753b011 · outbound

This paper cites 2025 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.410559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:a2fa60e68bff2fdccc8097d5b02d88108ec076b9434b843230b2d316f554cb94

Observation 645186cf-2f21-4da5-aab7-bc197ad03471 · outbound

This paper cites 2025 , eprint=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2025 , eprint=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.388191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:8d5a2caa1787b7718a99cd02e7168aea2c21e119f16aef1c3d43f520b3097d67

Observation 147d3036-ecba-49d6-90b8-1c6e02dbb9d1 · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing

Reference 51

Resolution
metadata mismatch
doi, observed 2026-05-09T21:13:24.038294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:487e8256cd81c2d58ac9855137175ca3a272eaef06cd219532f1808be5eb123f

Observation e290d9db-733b-40ee-8fdb-81ae5b8bca54 · outbound

This paper cites Large language models are not fair evaluators.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Large language models are not fair evaluators

Reference 52

Resolution
verified exact
doi, observed 2026-05-09T21:13:24.042721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:f3cc2bdaa62c15f0df55206ff6fe42767b5048cdcb555352feb786fe6e1b9c8f

Observation fe7a4b0b-4404-4615-bc13-f361870dc69c · outbound

This paper cites and Zhang, Hao and Gonzalez, Joseph E.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs and Zhang, Hao and Gonzalez, Joseph E

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.403377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:fd849d3f976a59e0929fca0f6d3187fba2d48f89f65af21575a7d6cad69c72a6

Observation 476c9413-af3f-44dd-a3cf-b2ff7249e77e · outbound

This paper cites PDFFigures 2.0: Mining figures from research papers , year=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs PDFFigures 2.0: Mining figures from research papers , year=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.420797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:39f355869d7db789b0fff70dcae7f219a27da26c5e650b1b3e113f5569143919

Observation 54237905-2490-435c-9ece-12740a36ad87 · outbound

This paper cites Toolformer: language models can teach themselves to use tools , year =.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Toolformer: language models can teach themselves to use tools , year =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.432520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:baeab38d8f0be3db51bec5a4778cc9c0c968df134c3d2019d64742bbb2585ec8

Observation d4e29a83-00d9-4ab1-9afb-6b6e764ff6ed · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning , articleno =.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Proceedings of the 40th International Conference on Machine Learning , articleno =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.444614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:d348f9231e9eb679defb4696ff0e0986f23b9b2016e0d5c48db6e40d821ca279

Observation 7c88c490-5df4-437a-93f2-94aca270da30 · outbound

This paper cites Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models

Reference 57

Resolution
verified exact
doi, observed 2026-05-09T21:13:24.030615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:320e759fbadf3455a1ac106767b65b76eb12aff7c62fe4e4a1ab7bdf42a5843e

Observation 336d0b29-a025-4777-971b-bc0ff62b5260 · outbound

This paper cites PeerQA: A Scientific Question Answering Dataset from Peer Reviews.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs PeerQA: A Scientific Question Answering Dataset from Peer Reviews

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:55.266227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:7e2c93a9b95a1aa434b0f694c487cebc4183b8c3e979d83983d3f3a6bf23397e

Observation 582b4fdc-f9e7-4a14-a6e1-6f9ada859f4e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Advances in Neural Information Processing Systems , volume=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.452164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:72f8ddadda14e069fbc3317c0c27aa4f55b3043dfc4890026d292806027d1c51

Observation c2588cef-a207-4d07-90f0-18f1591a9b24 · outbound

This paper cites 2024 , url=.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs 2024 , url=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T16:18:11.506335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:2d98e1e35ba5e738f043b49ac4326bd039511c2a85dcd714df143a776882f3ad

Observation 7e5c9dc7-c8d1-44fc-b453-7e8875b1cb5e · outbound

This paper cites S ci VQA 2025: Overview of the First Scientific Visual Question Answering Shared Task.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs S ci VQA 2025: Overview of the First Scientific Visual Question Answering Shared Task

Reference 61

Resolution
verified exact
doi, observed 2026-05-09T21:13:24.028015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:7a5be684d133fd9fb23c5867c989bd95f29b4332d9b8fd9c1c7b7c036f0524cc

Pith citing papers

Observation e4f3cd5c-4d9c-4be2-81ed-dac18813f480 · inbound

Code as Agent Harness cites this paper.

Code as Agent Harness PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs

Reference 208

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:58:14.290698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T10:54:54.558241Z digest=sha256:ffdc6a8a7539e51a1ebbf998b9fcf02fc3117fdd2e2fb594e32aa5b63473b287