Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:41:13.655335Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2507.13405.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:41:13.655335Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T04:01:48.251493Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T18:33:50.200070Z
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9ecf877d-c8c4-4b1a-8c1f-99a4d8bd3e9c · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39f49f9f-7fe3-4f23-afa7-7c5ebae49050 · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 284bbb27-812a-4401-b4af-d444e8cf7cde · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bad059c7-b2ee-4e1c-ab95-768d91d5a10d · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1720286c-ee3f-4044-bbb9-3943b9585fb8 · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66eaf9cf-7123-4b91-bd8d-f881b07038ea · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aeac34aa-ad93-4f30-a0da-a40b4b5066b8 · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13202050-d695-432d-b7a4-79bdccf119ae · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark CrowdHuman: A Benchmark for Detecting Human in a Crowd
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3215973-2793-40bc-82db-132191999531 · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Visual Entailment: A Novel Task for Fine-Grained Image Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49b1b18b-a872-47ff-8655-2e2f8a63c8c1 · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cb4cf62-8367-4010-a1de-5c6103ea2db4 · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0c63bf0-9da3-404b-a112-d41aa3f3a32e · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09201af8-cbf8-45e1-bfca-4b3e0b1d79d0 · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark When in doubt, use more conserva- tive qualifiers
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28ff9e58-16a6-4a6a-beb9-2db731eed007 · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark COREVQA requires models to perform multi-step verification by decomposing complex claims and meticu- lously verifying each component against visual evidence
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94ee7c5c-6f97-4031-9c8b-b618516ea93f · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01a7d63e-f7dd-4c42-8484-267f692b0260 · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c717dd3-edca-4384-b167-0a529f19c289 · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark R., Bashir, S
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d466147-a3f3-4582-81cc-cf2bfb5d755d · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e4749f-3741-47c9-840e-643562c7864e · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b305033a-7bc9-47a0-86da-c716cee83a57 · outbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4340bbb5-e9ee-4e65-a263-2a16a363cfbe · inbound
Scaling Mobile Chaos Testing with AI-Driven Test Execution COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b651095a-580d-412d-a407-a5cb8419b22e · inbound
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.