Pith. sign in

Paper Citation Record · LEDGER

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

As of 7 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2507.13405.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13405 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:41:13.655335Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:01:48.251493Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:33:50.200070Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ecf877d-c8c4-4b1a-8c1f-99a4d8bd3e9c · outbound

This paper cites GPT-4 Technical Report.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.166689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.166689Z digest=sha256:d8834bff82ea6ac783f99e24342f757527267fb0c00d8e91f85ebac52e042afb

Observation 39f49f9f-7fe3-4f23-afa7-7c5ebae49050 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.409602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.409602Z digest=sha256:a60ee52845a4f4bcb51595d98f62fd70214fbc4378d1efeea4f926fb04d51289

Observation 284bbb27-812a-4401-b4af-d444e8cf7cde · outbound

This paper cites NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.639417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.639417Z digest=sha256:2e1b2880b9d95142ca5d2daa8fc5e4001636d582129450a4056f8fa3997c04a3

Observation bad059c7-b2ee-4e1c-ab95-768d91d5a10d · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.716870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.716870Z digest=sha256:2ab9ef00dc6706b24f8a80e8482b309bbe9388f3ca2cb8d646c7d58949a5ec0a

Observation 1720286c-ee3f-4044-bbb9-3943b9585fb8 · outbound

This paper cites an unresolved cited work.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:41:14.859133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:41:12.809200Z digest=sha256:6e9c7071c6eefc256656d9e92ab6c8f9a0928ac0fcc1a5700a287ccc46f5adb1

Observation 66eaf9cf-7123-4b91-bd8d-f881b07038ea · outbound

This paper cites NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:41:14.160149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:41:12.884118Z digest=sha256:0a259ef46f6281991f9ba54d558711870a0f715f3f910d6292c7691718673e00

Observation aeac34aa-ad93-4f30-a0da-a40b4b5066b8 · outbound

This paper cites VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.007547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.007547Z digest=sha256:61b925f59b014fad5841aa10fb126ae87e045d7508f262baa68b6c7244bb8557

Observation 13202050-d695-432d-b7a4-79bdccf119ae · outbound

This paper cites CrowdHuman: A Benchmark for Detecting Human in a Crowd.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark CrowdHuman: A Benchmark for Detecting Human in a Crowd

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.138030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.138030Z digest=sha256:573ab8e12b0e77e2cba8c8a5089090ffa104fbb16358bc55fe22f66eacd25c7b

Observation f3215973-2793-40bc-82db-132191999531 · outbound

This paper cites Visual Entailment: A Novel Task for Fine-Grained Image Understanding.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Visual Entailment: A Novel Task for Fine-Grained Image Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.278562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.278562Z digest=sha256:460ba75d276e9d16f4297bda53f24be3a138b647e5bb726c8eee31e45f5e0216

Observation 49b1b18b-a872-47ff-8655-2e2f8a63c8c1 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.357316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.357316Z digest=sha256:09cb1ab61705ba8b40fa1cf794fec07ccc2a3281a977efdda4351880f316e435

Observation 3cb4cf62-8367-4010-a1de-5c6103ea2db4 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.436572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.436572Z digest=sha256:f3ec273b97aa93ed9951c07d315a33d979f4f03d83c1b8d0a38d09f3de79de13

Observation f0c63bf0-9da3-404b-a112-d41aa3f3a32e · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.513227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.513227Z digest=sha256:1a5a12763acd023efb068f1df33266e5b96b78467576f1fa88b0dfe03da32cdb

Observation 09201af8-cbf8-45e1-bfca-4b3e0b1d79d0 · outbound

This paper cites When in doubt, use more conserva- tive qualifiers.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark When in doubt, use more conserva- tive qualifiers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:41:14.620596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:41:13.590449Z digest=sha256:0dd6ea1013fca5deb892d63caa692b6f1be4892592880722a9c16c7ee02caf13

Observation 28ff9e58-16a6-4a6a-beb9-2db731eed007 · outbound

This paper cites COREVQA requires models to perform multi-step verification by decomposing complex claims and meticu- lously verifying each component against visual evidence.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark COREVQA requires models to perform multi-step verification by decomposing complex claims and meticu- lously verifying each component against visual evidence

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:41:14.459868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:41:13.655335Z digest=sha256:6dc6d5d7f12a344ec2c79b0ede5c2123bf0fde3a40a7b88a151652b2dbe22de9

Observation 94ee7c5c-6f97-4031-9c8b-b618516ea93f · outbound

This paper cites Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.234673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.234673Z digest=sha256:4c6982730e7180e6a0b81b6e100a7e9310370ca9975db20526f583d4ddf5b3ba

Observation 01a7d63e-f7dd-4c42-8484-267f692b0260 · outbound

This paper cites M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.203791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.203791Z digest=sha256:92c4da66b9f2781b9688c266bc55e08e12f86a960c067d3e368dec0df7789a45

Observation 8c717dd3-edca-4384-b167-0a529f19c289 · outbound

This paper cites R., Bashir, S.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark R., Bashir, S

Reference 2021

Resolution
verified exact
raw_fallback, observed 2026-08-06T16:41:13.982818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:41:13.076746Z digest=sha256:bc0b76071b7070d66238590f16628b1be7ab287546b9662c7c108022d36b82af

Observation 3d466147-a3f3-4582-81cc-cf2bfb5d755d · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.553911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.553911Z digest=sha256:12daf57477674e744ad76aee43275d557d4b263e44bb0ada7c815d1d2ba94f8b

Observation 97e4749f-3741-47c9-840e-643562c7864e · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.321982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.321982Z digest=sha256:bc23749b8be3c339a863ab281664be6cbea65f7a1ad10db2a6d50529cabda04c

Observation b305033a-7bc9-47a0-86da-c716cee83a57 · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:12.474596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:12.474596Z digest=sha256:a0d2178fe7e0697a932f520191254efdbed9c8411fa825d8eecb840b1963a0ab

Pith citing papers

Observation 4340bbb5-e9ee-4e65-a263-2a16a363cfbe · inbound

Scaling Mobile Chaos Testing with AI-Driven Test Execution cites this paper.

Scaling Mobile Chaos Testing with AI-Driven Test Execution COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:01:48.251493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:01:48.251493Z digest=sha256:bbc332f4e445dded647c38f6f56437dbae79a577f560ca5dba7b329b8daf85d8

Observation b651095a-580d-412d-a407-a5cb8419b22e · inbound

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling cites this paper.

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.201757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:32:41.348880Z digest=sha256:997acbfc16df58c8f7218e261023c6f912d46c49f239f8347529dc462ca43271