Pith. sign in

Paper Citation Record · LEDGER

Revisiting the Role of Language Priors in Vision-Language Models

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2306.01879.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.01879 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:31:56.440129Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T07:31:13.950276Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32744962-ab80-4aa2-9641-9069a427bd09 · inbound

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features cites this paper.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Revisiting the Role of Language Priors in Vision-Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.711881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.711881Z digest=sha256:f346359510f0a05b954a172a4fbf9fb7537687d0eae6757f00941a597e25a2b8

Observation eb2297da-7066-4777-b5ff-46b0a8b2543e · inbound

Explainability for Vision Foundation Models: A Survey cites this paper.

Explainability for Vision Foundation Models: A Survey Revisiting the Role of Language Priors in Vision-Language Models

Reference 247

Resolution
unresolved
no resolver link, observed 2026-08-10T17:26:35.957183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:26:35.957183Z digest=sha256:ffc0f29442584bdb214bae1485022151ebe2334cb22cc300a13b79cff1f17ea6

Observation 4e3b89c1-6997-4504-8905-d1072be5caae · inbound

Towards Understanding Camera Motions in Any Video cites this paper.

Towards Understanding Camera Motions in Any Video Revisiting the Role of Language Priors in Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.440129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.440129Z digest=sha256:a91aa5240c61eb059e9d90b4f1a7bc9ab337e0fd43910b1ddd2775a841a2195e

Observation 04b7d7ec-d4f9-402a-ae47-7cdbe860d0e6 · inbound

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models cites this paper.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Revisiting the Role of Language Priors in Vision-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:37.515954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:37.515954Z digest=sha256:3faa49285846a3a15a3abe1534c0f2d4a3df431bf03c6a56abb92c04df78e29d

Observation ecbdd126-2c17-4950-8c3c-f65d8f2cea17 · inbound

ARGUS: Hallucination and Omission Evaluation in Video-LLMs cites this paper.

ARGUS: Hallucination and Omission Evaluation in Video-LLMs Revisiting the Role of Language Priors in Vision-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:39.695176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:39.695176Z digest=sha256:7c225eba2395115a7bf1d8bdd1103f9252bf61ef4598db600e6c472500e0a9fa

Observation 3cea00c3-9b4f-4d96-9dd4-cab80c24a988 · inbound

Trade-offs in Image Generation: How Do Different Dimensions Interact? cites this paper.

Trade-offs in Image Generation: How Do Different Dimensions Interact? Revisiting the Role of Language Priors in Vision-Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:01.756106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:01.756106Z digest=sha256:ccf9c69ae2d7414b753019586820e4e6c255a93cf39fee08e803e1d875c32f09

Observation da03bf5f-cf85-4386-a535-6bc73e606300 · inbound

TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models cites this paper.

TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models Revisiting the Role of Language Priors in Vision-Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T13:52:15.074817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:52:15.074817Z digest=sha256:5737ecf9b8abafd78629d96ae796693f46960e52f8e6b074a94cb7e3b6d7680e

Observation ed0e9454-7441-4148-89dd-c43c9464f95d · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Revisiting the Role of Language Priors in Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.507451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:cd1ba80ff93665cbcb2963edd25cd942efae0e2e1ae9412911a00f84cbe0be25

Observation 27270a0a-9306-4d62-a338-67bc4726689e · inbound

SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals cites this paper.

SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals Revisiting the Role of Language Priors in Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:31:13.953141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-22T07:29:49.950084Z digest=sha256:0d7b6a99a00e219ded55a28052535527c65c156ee18b7e09f73d10d9733ae196

Observation db7c25f4-b9aa-461a-8b8f-5d7f5e3f644f · inbound

Prior Bias in Vision Language Models on UML Diagram Interpretation cites this paper.

Prior Bias in Vision Language Models on UML Diagram Interpretation Revisiting the Role of Language Priors in Vision-Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T06:36:12.761601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T06:36:12.761601Z digest=sha256:672992e47c82895f68a7ae36fc54d06ab04a91be8a363c2e770febcd572e4f13

Observation d20ccc5c-43f6-487e-aaf1-5b1d8fa8e22a · inbound

Scalable Visual Pretraining for Language Intelligence cites this paper.

Scalable Visual Pretraining for Language Intelligence Revisiting the Role of Language Priors in Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T07:33:08.314581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:33:08.314581Z digest=sha256:62fd789719059ff16daa433e93a6bc2479ee72d98227bf2461d8bbdc97e46ac0

Observation 02261088-172b-430b-bb3d-1085ad446991 · inbound

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO cites this paper.

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO Revisiting the Role of Language Priors in Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:44:46.262146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:44:46.262146Z digest=sha256:d8fb1e5fe4d54c061e6dc7d23f32ff2659fc82ad799c4a2f99a9d94c0a5dac14