Pith. sign in

Paper Citation Record · LEDGER

How Do Vision-Language Models Process Conflicting Information Across Modalities?

As of 24 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 8 inbound Pith citation observations for arXiv:2507.01790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01790 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:47:07.903561Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:07:46.242486Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T17:15:52.016694Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 81f33fad-f608-4169-ab62-9fcec184c70b · outbound

This paper cites Multimodal biomedical ai.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Multimodal biomedical ai

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:11.775843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:03.167374Z digest=sha256:0fee303952a8f1ef2938f904f9fb509cf60a5e02a44836176ee81bd740d89406

Observation 6d595cf1-56ec-47b9-ba72-ddeadc186849 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Understanding intermediate layers using linear classifier probes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:03.208420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:03.208420Z digest=sha256:1e35ef77f40ee9f7d1d735f01c1686a269d1f35c298cfc113c66a18210491d31

Observation 8c09864a-642d-4096-b634-7f795a35f720 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:11.681952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:03.286867Z digest=sha256:7ef71d9a41416292721dd3e61a8b94021b09f1c648ee29070a02f0b0a33cb9d2

Observation 26f38637-60c2-403e-8948-5ecb129deaba · outbound

This paper cites Qwen2.5-VL Technical Report.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:03.368750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:03.368750Z digest=sha256:914e0cfd6deaad3bdc6a48e96bfe04870e88b0416322caba48e97ec75a3c8aa4

Observation 04f238e6-90fb-493e-80c5-fe5d5fd89292 · outbound

This paper cites Probing classifiers: Promises, shortcomings, and advances.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Probing classifiers: Promises, shortcomings, and advances

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:03.435165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:03.435165Z digest=sha256:05de77714d8bbbeb280b0902a53a91b7e396118a20da1ce04044acf1a1e4d925

Observation a8ff95b2-42ca-4a9b-a95a-f968dcc63c28 · outbound

This paper cites On the robustness of large multimodal models against image adversarial attacks.

How Do Vision-Language Models Process Conflicting Information Across Modalities? On the robustness of large multimodal models against image adversarial attacks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:11.672456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:03.567710Z digest=sha256:c02c7619208d877a7b300e179c7bef63e9c0c322e2fa5f591eff4b821cf953ed

Observation 5dd82d92-11e5-4409-87d0-506b45b03509 · outbound

This paper cites an unresolved cited work.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:47:11.661002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:03.705815Z digest=sha256:50858f5c1cd9bd3f6f02d3e8b25a2e134c3f152e9e7ef084d69d3ccf2641623c

Observation 921bb0d0-f044-467b-b493-b92e6524b6c4 · outbound

This paper cites Words or Vision: Do Vision-Language Models Have Blind Faith in Text?.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Words or Vision: Do Vision-Language Models Have Blind Faith in Text?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:03.900892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:03.900892Z digest=sha256:2b73fcb0f7ba82aa1a2b2140bcd2ee498ec97f66883f8e8f79a6111832712737

Observation f7efe53c-d129-4982-91bd-484f48087a92 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Imagenet: A large-scale hierarchical image database

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:04.064194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:04.064194Z digest=sha256:7003802451c825b77b98e553045b15fda360a464056df0a563e4ab09f2b736ca

Observation 1b330963-4990-4baa-a8ad-3333f6fcf127 · outbound

This paper cites The pascal visual object classes (voc) challenge.

How Do Vision-Language Models Process Conflicting Information Across Modalities? The pascal visual object classes (voc) challenge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:04.232791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:04.232791Z digest=sha256:513c5c5fc37b8459897f5a4e79737ae4fd35c6e0a35e1f764f1c72de975d0317

Observation d8e1b6d8-03f9-410f-b75b-59f6b38387d5 · outbound

This paper cites Pixels versus priors: Controlling knowledge priors in vision-language models through visual counterfacts.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Pixels versus priors: Controlling knowledge priors in vision-language models through visual counterfacts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:04.394039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:04.394039Z digest=sha256:3fb932dda06c3dc06f04d84e1d806382e1229df2874f01458d07a831446b8b16

Observation c58d2bf2-5ac0-4aac-9135-1e7583da58b1 · outbound

This paper cites What do VLM s NOTICE ? a mechanistic interpretability pipeline for G aussian-noise-free text-image corruption and evaluation.

How Do Vision-Language Models Process Conflicting Information Across Modalities? What do VLM s NOTICE ? a mechanistic interpretability pipeline for G aussian-noise-free text-image corruption and evaluation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:11.373646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:04.559268Z digest=sha256:27de4b946f68cec3f4c1a2c07eb9b633406919951fd6d23d694a16533c89fb1d

Observation 219fdcb1-a9dc-46e2-894a-0bfa23f61067 · outbound

This paper cites How does GPT -2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model.

How Do Vision-Language Models Process Conflicting Information Across Modalities? How does GPT -2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:11.040573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:04.726134Z digest=sha256:8bb4c53b0ef4edd8686099f99ca135c39773ef33daea23656a4c74775d72154c

Observation eacdb6e6-0dc5-4f5a-a378-24aacb8a8d4d · outbound

This paper cites an unresolved cited work.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:47:10.677191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:04.858328Z digest=sha256:72cdb7b05dde0da8a6997b5bb96df78caf680e5f5e6fb5fd1cf22fd79a8ad4a3

Observation 5f5bd9f7-d638-4adb-a967-910e90828056 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Adam: A Method for Stochastic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:05.021127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:05.021127Z digest=sha256:71ddfeee0242f400d9e4b045ebe361e9da416822fed7d279817262a628348e0b

Observation 9f6745f4-7f2e-40e3-9df7-942bb196b3cc · outbound

This paper cites Learning multiple layers of features from tiny images.(2009), 2009.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Learning multiple layers of features from tiny images.(2009), 2009

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:05.200094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:05.200094Z digest=sha256:c06acafe7c8586befdf69c7732ca815ad275571bc551912fa5064de670d4be4c

Observation 948663c9-e9f1-4b37-8be8-10e5ca903714 · outbound

This paper cites Beyond the doors of perception: Vision transformers represent relations between objects.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Beyond the doors of perception: Vision transformers represent relations between objects

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:10.367160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:05.331504Z digest=sha256:09966e09ab8ec96658043137e66db35c4475c862b086da0042f14f09d2edb073

Observation c6a79849-c175-4e02-a595-9753ee31a037 · outbound

This paper cites LL a VA -onevision: Easy visual task transfer.

How Do Vision-Language Models Process Conflicting Information Across Modalities? LL a VA -onevision: Easy visual task transfer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:10.174208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:05.533823Z digest=sha256:a647786e3c75f7a99b46ec032b2601431daffc0788563032d9403a878bbf769a

Observation a7aab324-ef49-4635-a7d8-1afb4b707de1 · outbound

This paper cites Inference-time intervention: Eliciting truthful answers from a language model.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Inference-time intervention: Eliciting truthful answers from a language model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:05.693814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:05.693814Z digest=sha256:fa6920fd20877427f4b709e00442ca711cec85416e104bf68fc066dbe9efb483

Observation 88497041-3c36-4e81-98c0-491e01213a2f · outbound

This paper cites Benchmarking Multi-modal Semantic Segmentation under Sensor Failures: Missing and Noisy Modality Robustness.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Benchmarking Multi-modal Semantic Segmentation under Sensor Failures: Missing and Noisy Modality Robustness

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:47:08.457560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:05.868757Z digest=sha256:7b332004a990cd2f8fb716004722f74cbec6622aa12174f5f1491ed46e5db19b

Observation 1a037fbd-c560-40f1-aafa-e548c41e9a26 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Improved baselines with visual instruction tuning, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:06.023717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:06.023717Z digest=sha256:f7c96795e0361cc8500d3ea117a0726818425c00ad317f61a4c0188a076123f9

Observation 49b6407e-6b08-4795-83df-bbc8955a76a8 · outbound

This paper cites The quest for the right mediator: A history, survey, and theoretical grounding of causal interpretability.

How Do Vision-Language Models Process Conflicting Information Across Modalities? The quest for the right mediator: A history, survey, and theoretical grounding of causal interpretability

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:06.133530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:06.133530Z digest=sha256:e10bbebe282ecb97d7dd15fbbe3ec1d211b19dbbdfc028eb83e7e738a5ef2069

Observation 835729aa-e31d-4519-90a5-7c51c4c1ff2d · outbound

This paper cites Towards interpreting visual information processing in vision-language models.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Towards interpreting visual information processing in vision-language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:09.966918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:06.237627Z digest=sha256:9e35133bf1e8c347da4ffc5ca8b6eb25e05f6eebf322c0f329f2e5ce449eed32

Observation 78824f13-9509-4b94-8d80-070aee0ff6c5 · outbound

This paper cites Same task, different circuits: Disentangling modality-specific mechanisms in vlms.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Same task, different circuits: Disentangling modality-specific mechanisms in vlms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:06.342412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:06.342412Z digest=sha256:05f89bca2ef52386cb54b3bd3598ec4c54222d0d547cef3cd72bf834ed67f387

Observation 4045827d-8ad6-4910-aef7-c948e23bd086 · outbound

This paper cites Introducing operator.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Introducing operator

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:09.786996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:06.444748Z digest=sha256:f6eda85635c7483f1331b51186eedc639ac973284a5aef4b77104d2d7c7de70a

Observation a68d5511-6e85-4683-8fb3-0e6847d566ba · outbound

This paper cites GPT-4 Technical Report.

How Do Vision-Language Models Process Conflicting Information Across Modalities? GPT-4 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:06.548025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:06.548025Z digest=sha256:62fe4009f2dfefd3de5734039e29652050a6fa3f7555ef328758f079fcf933b0

Observation f206f484-5183-40d2-a623-9fcab1b725fd · outbound

This paper cites Interpreting the linear structure of vision-language model embedding spaces.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Interpreting the linear structure of vision-language model embedding spaces

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:06.627050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:06.627050Z digest=sha256:9b17f357297caa9d598c50bd588b42d6bb92b0309a08237fd78facbba96e30c3

Observation bd9f27a9-779d-4f5d-9fa9-6b5c4f9df9a6 · outbound

This paper cites V -measure: A conditional entropy-based external cluster evaluation measure.

How Do Vision-Language Models Process Conflicting Information Across Modalities? V -measure: A conditional entropy-based external cluster evaluation measure

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:09.607422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:06.718852Z digest=sha256:1e673d9f9cf3dd545a472efb371138a89ff0bcd59367d9728b8941588b720a37

Observation 490fa9ea-1e47-4873-b46e-c7600c1a4bab · outbound

This paper cites On the Adversarial Robustness of Multi-Modal Foundation Models.

How Do Vision-Language Models Process Conflicting Information Across Modalities? On the Adversarial Robustness of Multi-Modal Foundation Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:06.822297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:06.822297Z digest=sha256:e658cecd652f5a44c22e6b28fcf3af6d5522bc6d55f2c9c1d8fb2d1fa418d66e

Observation 27494d1a-1258-4ee3-8986-4321e9c8df62 · outbound

This paper cites What do you learn from context? probing for sentence structure in contextualized word representations.

How Do Vision-Language Models Process Conflicting Information Across Modalities? What do you learn from context? probing for sentence structure in contextualized word representations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:09.418278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:06.925496Z digest=sha256:0aa9147de3017e966a988d570d94a8316e13f56a805ba07d7e99676c8e86dfa7

Observation 13cb4d14-85a6-4c25-978d-617ffba7998b · outbound

This paper cites Bert rediscovers the classical nlp pipeline.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Bert rediscovers the classical nlp pipeline

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:09.236606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:07.030149Z digest=sha256:7ebf7715885cb8a08f2137763ff45bce76ef8e53f2be932ccf302f89e25f540f

Observation 355d1c63-4f5c-4747-8764-fd6dccae3cb5 · outbound

This paper cites Investigating gender bias in language models using causal mediation analysis.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Investigating gender bias in language models using causal mediation analysis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:07.129570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:07.129570Z digest=sha256:dc7cd2fee6ac6dbff88374a849522f8e9867260a3aa3ccb470f32ea5252cbb88

Observation 2dec1286-d704-46a3-a389-0369d48bd845 · outbound

This paper cites Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:07.202331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:07.202331Z digest=sha256:8b87569dd63e15ad7bf5a6f13c2b215294b50b4826b06378757c02b305bfaa9b

Observation 7c22cf30-9bc8-4c25-b8de-56ac146da63a · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

How Do Vision-Language Models Process Conflicting Information Across Modalities? The caltech-ucsd birds-200-2011 dataset

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:07.294072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:07.294072Z digest=sha256:9cc6364daede55a1facf6b64920f81f3f80f716389941630cab2f9abc888f804

Observation 85216b34-3275-4f8e-8904-190d27a6f832 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:07.394735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:07.394735Z digest=sha256:db86e2f827cac046ee7cbfa4353c44c37beb16d1dcfc5944295cae4fec10d01e

Observation be17bcbc-6a94-4caf-9ef8-f2d0e315a1ea · outbound

This paper cites Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:07.466049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:07.466049Z digest=sha256:03e1755cc3a0127043030a1bf0bde499871b1817baba66aad09e4e4a9935d3f9

Observation cce0599b-a114-4793-bb89-cf434f2a65d5 · outbound

This paper cites Characterizing mechanisms for factual recall in language models.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Characterizing mechanisms for factual recall in language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:07.551656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:07.551656Z digest=sha256:3c038dc4238bdbd88a40876611a25c8236beb0f921595a94d71f20853f58d152

Observation 22a2a9d2-601a-437d-925d-a220b9229005 · outbound

This paper cites Does vision-and-language pretraining improve lexical grounding? In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4357--4366, 2021.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Does vision-and-language pretraining improve lexical grounding? In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4357--4366, 2021

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:09.052513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:07.648337Z digest=sha256:6232c04a5c1999d61682929193a752118fe2e014ce5b3dd9ba476c0e384bf074

Observation 1027564e-6fa4-4942-bd0d-549b07549037 · outbound

This paper cites Emergence of abstract state representations in embodied sequence modeling.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Emergence of abstract state representations in embodied sequence modeling

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:08.751769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T20:47:07.710731Z digest=sha256:7be1fb8429c11005c5005176d11c3e85023e0152a84f56ff89b817381aac2b88

Observation 658b2ef7-f1d6-4d56-b06a-25407e037b18 · outbound

This paper cites Calling a spade a heart: Gaslighting multimodal large language models via negation, 2025.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Calling a spade a heart: Gaslighting multimodal large language models via negation, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:07.809459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:07.809459Z digest=sha256:0a2a31c2e4172a415620cdc33ea9d6a7507bcb19c3d4e4bdf9bca681cc2b5701

Observation 9c47deb8-34ed-426f-b729-1e8ce3c23b4d · outbound

This paper cites Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:07.903561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:07.903561Z digest=sha256:77b419c860c54c7acb2e2a4ea58b66f31a4e47625075df58025b8e9261157763

Pith citing papers

Observation f9767cf6-0900-4952-9bfb-3c0ca9bc9db0 · inbound

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability cites this paper.

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability How Do Vision-Language Models Process Conflicting Information Across Modalities?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T01:00:26.919995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:00:26.919995Z digest=sha256:407d79be700a7cc18108c3b1e46f30297903b0fba3d0e68f51001efbe3f3d0d1

Observation 0c9086df-9a57-439d-a560-1ad63b475312 · inbound

TRANSPORTER: Transferring Visual Semantics from VLM Manifolds cites this paper.

TRANSPORTER: Transferring Visual Semantics from VLM Manifolds How Do Vision-Language Models Process Conflicting Information Across Modalities?

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:01:34.739667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T05:59:14.478279Z digest=sha256:4e76c341a47d68145736210629c278bdc261df775d88fcdf797509a6699f4b2c

Observation 524773b5-315c-4eac-8bf8-a010888435df · inbound

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models cites this paper.

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models How Do Vision-Language Models Process Conflicting Information Across Modalities?

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.710872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T07:02:02.752466Z digest=sha256:5d593d02871ea22bc60ffc7e905e039a2d467814374259bde71a1c33b07f4a36

Observation 854e9a80-1a6a-48de-8edb-a3c3607b15ab · inbound

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models cites this paper.

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models How Do Vision-Language Models Process Conflicting Information Across Modalities?

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:15:52.018576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T03:53:04.984303Z digest=sha256:6e3f3c398ea56fa162d75b68164117bdf5d03baf0d7446be4535e6318a542481

Observation cf1015e8-e448-412d-a1fd-3f9e5624e96b · inbound

ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision Language Models cites this paper.

ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision Language Models How Do Vision-Language Models Process Conflicting Information Across Modalities?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:14:21.583728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T07:06:36.599771Z digest=sha256:19af5cb41088baab41e800fddc72d036e1a9a7a3fc9e03388acbb740b36d5bb3

Observation 218f5833-c688-4f6d-b7e9-af96fcff062a · inbound

Attending to Multimodal Generation One Token at a Time cites this paper.

Attending to Multimodal Generation One Token at a Time How Do Vision-Language Models Process Conflicting Information Across Modalities?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:43360bab56d9dd874b99ce4be0a67ec22c855336cc72071ac26eeaedb759ac97

Observation a0acd5ca-061f-4b96-8016-9cd40cbb5726 · inbound

Linguistic Context Recodes Visual Representations in Vision-Language Models cites this paper.

Linguistic Context Recodes Visual Representations in Vision-Language Models How Do Vision-Language Models Process Conflicting Information Across Modalities?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T01:41:52.790018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:41:52.790018Z digest=sha256:0404984e0c002d71f4f7d2eefc3da2a843da2a58ebf823cfa4447a872b1aafeb

Observation c362de38-7834-40c6-885b-21f210e261b2 · inbound

PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models cites this paper.

PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models How Do Vision-Language Models Process Conflicting Information Across Modalities?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:07:46.242486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:07:46.242486Z digest=sha256:3825a02500a0794e7ee96add6011137be1de27c522f17761c7028a1d12d61c59