Pith. sign in

Paper Citation Record · LEDGER

VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2207.00221.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2207.00221 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:19:45.886327Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T20:44:09.047810Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f76ee1ca-8be6-4b4d-8d62-c1543f2210fa · inbound

Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs cites this paper.

Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T16:19:45.886327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:19:45.886327Z digest=sha256:dddd7754b1166ce29c4cbf693784d13c4c4076ebdb060f81da1bd35c2b1148f2

Observation a418555b-8f7d-4cb4-8f58-b38d546dc182 · inbound

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models cites this paper.

Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:18.657659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:17:18.657659Z digest=sha256:b683748a9a1767d6087d6f863d34bc5c8ae1890e40d7e0e21cf5d8c6bd05dd55

Observation 06c6fe06-7548-4ce7-b2f1-1f808c4ebe90 · inbound

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models cites this paper.

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:06.258515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:17:06.258515Z digest=sha256:6df48cc4b00383610a1a916dd53e7d8de898c71398de528964a2fba09c6f2c87

Observation 2df23896-35c7-4108-b3f5-1628aab206ea · inbound

Causal Graphical Models for Vision-Language Compositional Understanding cites this paper.

Causal Graphical Models for Vision-Language Compositional Understanding VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.372152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.372152Z digest=sha256:c5ff500d0a3ef7fabd833ed8cb17770b8040c274aba655de3945b07759461005

Observation 7b581a1c-981d-4577-bf25-65524f03fa8f · inbound

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models cites this paper.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:41.725304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:41.725304Z digest=sha256:aba15d95172d88196ee67e3212cead8667afaea9bb555f058a1fd7cdda08b698

Observation e405d9a0-45e7-4a0e-a115-2c5228d7f9f2 · inbound

What's Missing in Vision-Language Models? Probing Their Struggles with Causal Order Reasoning cites this paper.

What's Missing in Vision-Language Models? Probing Their Struggles with Causal Order Reasoning VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:00:18.753401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:00:18.753401Z digest=sha256:6166892312d1a2ea30ca072a44dacdc2fb99d7292934de2b0d26ec55f1984251

Observation 42821a0d-6d10-4218-a568-1379fe734f45 · inbound

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks cites this paper.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.157557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.157557Z digest=sha256:2312d19bea07f1d0f5f60f9d5ed45f95dc1bf8c571f18109b695df0d63337641

Observation e0bd20ed-6479-47b0-bb65-0026b8d5f9d2 · inbound

CF-VLM:CounterFactual Vision-Language Fine-tuning cites this paper.

CF-VLM:CounterFactual Vision-Language Fine-tuning VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.140504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.140504Z digest=sha256:fc1b1490ceab2e48d31df0b8a72bf79a7c9a823f2dc27ccdfbd0ca1faebacee7

Observation 3ccfcbb2-0c3b-4b86-ac09-df84f94e9a8b · inbound

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation cites this paper.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.649660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.649660Z digest=sha256:24eba36bad0b7ebf6b1312ddb7318bf34bf468f3f9f6757eed8663a4f8d50f88

Observation b54f18df-cf1d-496b-8186-365b28cf9c0a · inbound

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets cites this paper.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:53.011960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:53.011960Z digest=sha256:0e790d37d64ec9f8cea796c8ccc275ad1737f21c4a22b54536ba30c7267deb1e

Observation 6460b670-f6d7-44a2-9ea5-af64b6d2838c · inbound

Can VLMs Reason Robustly? A Neuro-Symbolic Investigation cites this paper.

Can VLMs Reason Robustly? A Neuro-Symbolic Investigation VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T19:17:12.798847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:17:12.798847Z digest=sha256:cb6edaea07820d98d66b485bb1d6b5ff0cc422d3ab2d14a60c37d779f667054d

Observation a32aa503-64ed-4cca-ac3d-733737cd2430 · inbound

All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding cites this paper.

All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:31:03.860342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:26:55.369840Z digest=sha256:72a4ef71f56cb57bf65ba03e273fb59987583ef67ba5088ec0d109ab6d59f6d5

Observation 15c11f3d-0ed5-49cb-b08a-6ea5281eac64 · inbound

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding cites this paper.

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:04.985463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:24:57.169737Z digest=sha256:255669ac44bf9dc50b4c8c7c06ab7dddbffd4fd78bb958b026a434b2b3626320

Observation 53fbf125-c7e9-4ec0-84c5-e10536b15f1c · inbound

DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models cites this paper.

DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:36:36.527792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T06:34:29.555022Z digest=sha256:d795149cdd6de3393dcb214422ac921c5b2e51f61efe560b00d1c0ed0f935317

Observation b19db82b-644c-4afe-bfda-2e8a4cad5b6b · inbound

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality cites this paper.

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:31.484644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T07:03:50.311891Z digest=sha256:08d0768f1445114e122e1b2180b48a2823ef42521ed369ba50b7ca8dfb77b40e

Observation 32673b3f-e221-4c90-a8fa-5f2088c1bd37 · inbound

HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding cites this paper.

HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:17:07.117700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-02T15:15:24.784685Z digest=sha256:3c9edeb30f147fffc8d2d51dc69e3ea42bd03c3f7ee044905faf5c675c7ff871

Observation b4105ae4-4d75-4c85-b84d-a6f71f50aedc · inbound

Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models cites this paper.

Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T20:44:09.056579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-07T20:38:37.820535Z digest=sha256:8573f55e2289bd74472928f43eb641c42f35da95ed1dd10753e013575e67ef44

Observation a3a95dcf-3f05-4d4f-8d53-d081e05c16cb · inbound

Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models cites this paper.

Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T16:16:24.318000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:16:24.318000Z digest=sha256:0840f22aeada193de4dea442d8a0e83028844c7a92bc0c7fdea7ebf08b6377c8

Observation 6a600bcf-6fe0-42fc-be01-71f7d9b3e9e3 · inbound

Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models cites this paper.

Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T08:32:59.699718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:32:59.699718Z digest=sha256:c593273c26cdee0fcb9b128caa55f9906634e9c3a166b206c4200a991b27d88f

Observation d8815c50-dd81-4662-b334-57cfb1ad0e28 · inbound

Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models cites this paper.

Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:22.367179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:22.367179Z digest=sha256:9a79fb0062e3996285fa4eaf745f1ad121847487c14d968a16d55e25047634da

Observation e4d9758e-34a0-445b-9616-93d4ca47a694 · inbound

EEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits cites this paper.

EEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T00:14:11.250291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:14:11.250291Z digest=sha256:7b64a838009ee14565da8643b8cfe010facdd6166aac0f913ed6ba47cb97c68b