Pith. sign in

Paper Citation Record · LEDGER

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes

As of 17 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2601.07737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.07737 v2

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:04:30.771165Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 021886da-65a4-42f4-8d74-d938d1062f25 · outbound

This paper cites GPT-4 Technical Report.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:28.175587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:28.175587Z digest=sha256:98bd522bc966581a7aca0e9210e6ca309ae2edb07c165e9e4297775ce470cdeb

Observation 1e342688-fd15-40f3-bab2-544540154321 · outbound

This paper cites Visual Instruction Tuning.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Visual Instruction Tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:28.297023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:28.297023Z digest=sha256:9aedf718f37e07971ff5da75a9dd7ccfc984466555f4b0765b54c70767d4302c

Observation b78909c9-3a91-4cd1-a5e3-1d2e17e38b2b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:28.475209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:28.475209Z digest=sha256:064c2ea30cc2dc103bb61a40e10f0fdf9ae7b4935a30a295e4268caf9fb21ce7

Observation da682945-5c05-4646-9e52-6e78e68c7b56 · outbound

This paper cites Describing Common Human Visual Actions in Images.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Describing Common Human Visual Actions in Images

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:28.612236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:28.612236Z digest=sha256:57e8088437b7e2e6a7fab606ec7876db2a75255f661b7a56c58f3a95241c6a5a

Observation 503c734b-774a-4358-90c9-0bd897f3e98f · outbound

This paper cites Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:28.748779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:28.748779Z digest=sha256:3045afad6b67ded49f4cfd9d5ec9091c2dde0f89ba7074db9064d64f5709bcd2

Observation 6644e92c-2524-4f1d-87c8-f18661e1a298 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes High-Resolution Image Synthesis with Latent Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:28.945250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:28.945250Z digest=sha256:ab7540e7afe4aaf537ad81c044f8c1b1fde3ab2b5b03b5bca1a587fee1129746

Observation fd3235ef-1b24-4ada-a739-6c168f2accf7 · outbound

This paper cites Attention Is All You Need.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Attention Is All You Need

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:29.114666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:29.114666Z digest=sha256:eb9ade8f2d6e499a89bff7316abb01a2efa05c7a58e9886aba5c1dba1b6bfbe4

Observation 0baf72b6-474a-4389-b53f-3e8334787282 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Learning Transferable Visual Models From Natural Language Supervision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:29.281991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:29.281991Z digest=sha256:7921ee79bc244c87b124eab3b8560cd31b0f272babbfa14f9bb30816d3567733

Observation 44d7a4e3-1e43-4a98-9cee-66ffd9274589 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:29.529985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:29.529985Z digest=sha256:87f9d647f4a9213c0973532893a50cd806f7bfd5e213b54b51db1b8dd6f99d64

Observation b2f53301-a414-4608-9463-8e19be9d1b80 · outbound

This paper cites Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) network.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) network

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:29.665318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:29.665318Z digest=sha256:ee444197d5c5e59891a3ae1e86fe4c6c4d9268dceedb51bd3fcf84f3d2e827ef

Observation 6b673d08-da35-4992-9893-a9ac9311b4d3 · outbound

This paper cites RWKV-CLIP: A Robust Vision-Language Representation Learner.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes RWKV-CLIP: A Robust Vision-Language Representation Learner

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:29.801204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:29.801204Z digest=sha256:bbde8d9fc17e62158966f8b2d67d80174db734b4b3ab83814816770a060d86c2

Observation cfa90d08-115c-4ace-b284-cab866ec017f · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes LoRA: Low-Rank Adaptation of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:29.971959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:29.971959Z digest=sha256:b9a575550216650af38ce78e66055c55ccb0c2c39319dd598bde9716311d4ec6

Observation 4dd06a10-af46-46b4-9459-abc9637fa29e · outbound

This paper cites 315VerbNet: Capturing English Verb Behav- ior, Meaning, and Usage.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes 315VerbNet: Capturing English Verb Behav- ior, Meaning, and Usage

Reference 13

Resolution
malformed identifier
no resolver link, observed 2026-08-03T11:04:30.088504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:30.088504Z digest=sha256:cb2367d81f203a75bbeb7748a8c149d6b18b5148828b309bb588b1d6207f6948

Observation e58732ef-dbf8-453b-88c9-8bad01ef3eb5 · outbound

This paper cites Qwen2 Technical Report.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Qwen2 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:30.236155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:30.236155Z digest=sha256:627a9affb2791655fd6313f76792e5c88f0a28027b17bca29edc132f1a3b46ac

Observation fbc43bae-0fc1-48d0-adae-ac3e9101a6b0 · outbound

This paper cites Language Models are Few-Shot Learners.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Language Models are Few-Shot Learners

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:30.536424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:30.536424Z digest=sha256:30169f6cdafe78fea9c09b9451ae0bb6a85be58596d8af952716d6bd7ebce4d9

Observation ea1cfb8f-249f-479e-a28e-2e2d4c46f626 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:30.660635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:30.660635Z digest=sha256:efc33e6852fe54fb4d122433abe09763b8f6162093a758a0968163e3be53df6c

Observation 0dd2b232-20e5-42ab-95c6-669904522db1 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:30.771165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:30.771165Z digest=sha256:4d93ee7c2cd96ef6138a7c1633c05032a207b7681123a5877ac2617e05863c0a

Observation 0d40bf03-abe5-4f0c-bf74-05ec45ea1029 · outbound

This paper cites Qwen2 Technical Report.

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes Qwen2 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:30.394667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:30.394667Z digest=sha256:c4c1b7ecacb114538da8297d9f12a07e2a112c41fce9309cd6fec3724a6c7b11

Pith citing papers

No inbound Pith citation observations are available.