Pith. sign in

Paper Citation Record · LEDGER

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes

As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2601.07737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.07737 v2

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:04:30.771165Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 021886da-65a4-42f4-8d74-d938d1062f25 · outbound

This paper cites GPT-4 Technical Report.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:28.175587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:28.175587Z digest=sha256:ab9dfbefcc0fcf9c8dffc06f1e02fa9402b7ffbac186c8695dcccbfa7cf2a811

Observation 1e342688-fd15-40f3-bab2-544540154321 · outbound

This paper cites Visual Instruction Tuning.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Visual Instruction Tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:28.297023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:28.297023Z digest=sha256:fff4ebc6db7a4cd703aef69198da70929f49e0d4646f02477e0ecbdf0d6dbfff

Observation b78909c9-3a91-4cd1-a5e3-1d2e17e38b2b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:28.475209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:28.475209Z digest=sha256:1f3c76e97ac661b70f61fdb24a6799d426486e56737b31c457d7d5076641a37c

Observation da682945-5c05-4646-9e52-6e78e68c7b56 · outbound

This paper cites Describing Common Human Visual Actions in Images.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Describing Common Human Visual Actions in Images

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:28.612236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:28.612236Z digest=sha256:2c34d52c41dfbfeaa825ce49cc370e1b6573c7215190ac6d5a511bd3b7a670ac

Observation 503c734b-774a-4358-90c9-0bd897f3e98f · outbound

This paper cites Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:28.748779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:28.748779Z digest=sha256:4678266827e91eff5fcd5befb66768bff517cd98d37e3f2d769b28e20d246f8e

Observation 6644e92c-2524-4f1d-87c8-f18661e1a298 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes High-Resolution Image Synthesis with Latent Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:28.945250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:28.945250Z digest=sha256:b0d6b53f65d34fff62ed9b9ce37daf5064ed13677897bd13fab994d30e6e5867

Observation fd3235ef-1b24-4ada-a739-6c168f2accf7 · outbound

This paper cites Attention Is All You Need.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Attention Is All You Need

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:29.114666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:29.114666Z digest=sha256:19bd8d4000fbe8c2845922c33b313e119e97e41a9daa009be3576e37fc1b2d1b

Observation 0baf72b6-474a-4389-b53f-3e8334787282 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Learning Transferable Visual Models From Natural Language Supervision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:29.281991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:29.281991Z digest=sha256:41c5a579e85ea44f44f1750e6d17b13b3e95ba637cc3504dff4eb1070803e917

Observation 44d7a4e3-1e43-4a98-9cee-66ffd9274589 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:29.529985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:29.529985Z digest=sha256:708b692769450a51e94b5c1c1411e0f2c10f21f1cfee657ede250711cd0cd0cb

Observation b2f53301-a414-4608-9463-8e19be9d1b80 · outbound

This paper cites Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) network.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) network

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:29.665318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:29.665318Z digest=sha256:ee11b96cfbf514238c19ef27c2f60af5dd5e9e5af3b0b36eb2de332cee33c7f2

Observation 6b673d08-da35-4992-9893-a9ac9311b4d3 · outbound

This paper cites RWKV-CLIP: A Robust Vision-Language Representation Learner.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes RWKV-CLIP: A Robust Vision-Language Representation Learner

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:29.801204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:29.801204Z digest=sha256:4049f49069e3592cf2f798a83ee3b66fd11e50ae0f7510edeb1efa2cc4fe0f48

Observation cfa90d08-115c-4ace-b284-cab866ec017f · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes LoRA: Low-Rank Adaptation of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:29.971959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:29.971959Z digest=sha256:e1dcb948577d469b349b6e0394f5d5ea1c24e08ad9eaa985df8462d7e95ad37a

Observation 4dd06a10-af46-46b4-9459-abc9637fa29e · outbound

This paper cites 315VerbNet: Capturing English Verb Behav- ior, Meaning, and Usage.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes 315VerbNet: Capturing English Verb Behav- ior, Meaning, and Usage

Reference 13

Resolution
malformed identifier
no resolver link, observed 2026-08-03T11:04:30.088504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:30.088504Z digest=sha256:5893708746493a07ea10e8150b61531c423129485582ab696d3a83270300e83a

Observation e58732ef-dbf8-453b-88c9-8bad01ef3eb5 · outbound

This paper cites Qwen2 Technical Report.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Qwen2 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:30.236155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:30.236155Z digest=sha256:854b1c20fef1eb32baf4138ea3083bce5f820cff8ac55177cf13c2d2dfc578ba

Observation fbc43bae-0fc1-48d0-adae-ac3e9101a6b0 · outbound

This paper cites Language Models are Few-Shot Learners.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Language Models are Few-Shot Learners

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:30.536424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:30.536424Z digest=sha256:54331d4ed0bf4b2d21d61bda8861381864735118e4e316833b2f2fe5d164b024

Observation ea1cfb8f-249f-479e-a28e-2e2d4c46f626 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:30.660635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:30.660635Z digest=sha256:4e9f08f863c9a29ed620cdd370c7bcc96659e59ddd8f497d1dd2a4cbb950e158

Observation 0dd2b232-20e5-42ab-95c6-669904522db1 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:30.771165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:30.771165Z digest=sha256:ed495245e8625d9447157c7f971e04208f9aab8298c04410f73ebe1b66ef2745

Observation 0d40bf03-abe5-4f0c-bf74-05ec45ea1029 · outbound

This paper cites Qwen2 Technical Report.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Qwen2 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:30.394667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:30.394667Z digest=sha256:cf0fccc59bcb07cd264b49e22eb077e32db87cbee899c99156036995c15e1635

Pith citing papers

No inbound Pith citation observations are available.