Pith. sign in

Paper Citation Record · LEDGER

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs

As of 6 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 2 inbound Pith citation observations for arXiv:2512.08923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.08923 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T23:52:13.813452Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:46:24.074055Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T00:46:24.164407Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact19
  • verified fuzzy29
  • unresolved7
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8110ebdf-c7d2-484e-9085-aa3bab10f518 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.823660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:158c4b63e6d6100e759c38b314ff7d9b41968c365f2c295987e6e46cbdcb1ac6

Observation dbd07447-2ccc-46d0-8aa9-74421aaa251e · outbound

This paper cites Phi-4 Technical Report.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Phi-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.806385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:96b650db04cf30c1df6998b0c097bd540383063bafb2639cf12c2bbdb60f925a

Observation 629b82f1-6678-4283-88aa-6f857f9fd5d9 · outbound

This paper cites GPT-4 Technical Report.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs GPT-4 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.801800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:4b611eefca077c41ab8d137c8956f032e9c323531f62e043b7e13d61fc372adf

Observation 6d08b4ea-d40d-40ad-b6fc-47877c3b6597 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.689256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:31116a5f8693e624c840f2df27e2af751db091a7289fb5227609b1dfc76949fa

Observation 82fd1ed4-f002-4700-bb0b-9d9a2288c4ea · outbound

This paper cites Vision-Language Models Struggle to Align Entities across Modalities.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Vision-Language Models Struggle to Align Entities across Modalities

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:53:42.846348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:1aa77e152fb4a5dcab07cadc217cfdedf2918c6358a1394a07d374056a8a4db3

Observation 9ec85214-d38a-4720-8d5f-89e6031c39db · outbound

This paper cites Claude haiku 4.5.https : / / www.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Claude haiku 4.5.https : / / www

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.680031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:d057f9a615610b7d11309e6dc721f9d902609017b54f4063f69b8f2be816217f

Observation a032a83c-12d9-41c7-a11c-72323f28eaee · outbound

This paper cites Qwen2.5-VL Technical Report.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Qwen2.5-VL Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.870999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:9d9fe1e92b8caf6adbe65e6e715a82de1c1f1bd94a1984609fcf45dffea6f55e

Observation e10d9367-0423-4a63-9e85-fb27745fe9c2 · outbound

This paper cites OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:53:42.853563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:fd5189ddd5c4802a13320392800f5c8acf22d3a665763a44090508899fbcfcf8

Observation c74bdb3d-134f-4358-a7fe-9ec764820783 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.849602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:21730101f6b8acc4df9140df583e8ad52694dc33a7adad2b5813bbae86f56f20

Observation 9fd7b350-d676-42e6-97f5-e846835ed321 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.790193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:90822a3e4f04207770844deb5dbdc299abb7c3016061eb6f42214dde4eeafdcb

Observation bb41fc09-05d4-4189-ae06-247272359f38 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Imagenet: A large-scale hierarchical image database

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.779247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:1315bfd098734d9b21d7d3233dc1f2cb27e7fbb99db54da4fe69284030777092

Observation 04488c18-9fbe-4051-a41e-7d9cb246797c · outbound

This paper cites Mitigate the gap: In- vestigating approaches for improving cross-modal alignment in clip.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Mitigate the gap: In- vestigating approaches for improving cross-modal alignment in clip

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.783049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:148746f88a4f806afc47abd219a9d2f1ec0813a57ae98064260b3d4f5d3a3c40

Observation b66000c2-f3a1-4559-a26c-e897629f440e · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.786978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:33ccff850b4e5c5465400341e4e87ba825224ed082a644d12f0255db7681385c

Observation c41ac8c6-9930-4b05-b1c8-a82d8499863a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Measuring Massive Multitask Language Understanding

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.867680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:9d74dc2c657e68b734eace71b69d4616ef7c002f1ece09768b534943865605be

Observation dc49d84c-e77d-4ca1-aa50-02ac446c14f8 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.793654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:7efba582034c3f67e373b22eb8518ca8c50599b33df8872fc4e8987e835f6fd7

Observation d54af141-9264-4dcd-9f07-babb89050309 · outbound

This paper cites GPT-4o System Card.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs GPT-4o System Card

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.856747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:d484eea0dddef4e7aaa41b82aecc166afbb31147f4c608101e9dd5c26ec7a6b5

Observation 2791da44-24dd-4661-a9f7-10b0f6fbfd46 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Gonzalez, Hao Zhang, and Ion Stoica

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.774116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:5a1ebe0073bda50d76d6e3b8bc9329a2d2ba65e94f5dcdf896eecd09f4447215

Observation 731d9b5f-7839-4f7e-8515-90899bb1b679 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.Advances in Neural Information Processing Sys- tems, 35:17612–17625.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.Advances in Neural Information Processing Sys- tems, 35:17612–17625

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.765623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:8d467bc9eb5cb84619ff3dfb55fb7ee5f7e1dfeb41e9dd4e2ca37da298860858

Observation b1db986c-ea88-4467-894c-9afd7736f253 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.769727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:d0e433f9a66d0c05918319a1a2e8391ef3d75f06f92d37c06729d16ae1b7bd56

Observation 62d9f306-c088-4f27-91fc-b3923ce4fd38 · outbound

This paper cites Improved baselines with visual instruction tuning.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Improved baselines with visual instruction tuning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.753682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:9e8f629c931e2b61f4d3922290a12978f9914da6614e68af4502b0e53b54c406

Observation b6178d22-fd8e-45bb-b607-186af747abe3 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs MMBench: Is Your Multi-modal Model an All-around Player?

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.815533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:7e6ccf33c57c6aa85d936a735c32b422b14d91e6f3626617211dd95e85cd845e

Observation aa6a5384-6833-4f69-857d-98ce2052ccac · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.835334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:6418312f298c589578c3e21969f0d405328286095c3cd1c664f35c51d6fd13b8

Observation 85b3e2d9-801d-4184-863f-19dc87ade8b0 · outbound

This paper cites Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.748547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:dff6437ead39db9842f1b1a65215b5961952a8a1609efe9de09150e3a6129bc7

Observation 8dd87484-e0cd-4ec2-89d8-c410751296e7 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Docvqa: A dataset for vqa on document images

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.757062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:4b8334df833bed63a104de73be1511ec89e93acd8dc61a0d6a0bac419c1c6d20

Observation 93f68678-e4ad-493c-9245-c3aedf3d2aa3 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.827500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:5af1c77e31973d0cbbbdbd59bcb647e01b82079f53ba9b89b74ba8afd49dc849

Observation ce62c00c-a3b6-4228-ad23-c9ccc56b3db7 · outbound

This paper cites Mistral-small-3.1-24b-instruct.https:// huggingface.co/mistralai/Mistral- Small- 3.1-24B-Instruct-2503.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Mistral-small-3.1-24b-instruct.https:// huggingface.co/mistralai/Mistral- Small- 3.1-24B-Instruct-2503

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.761259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:9c23f8cd6953f420223aac9a064f4487d8be2d170899c3c4ee3d61c516d4d926

Observation 197a3225-651a-47ab-bb69-71a1647c11b1 · outbound

This paper cites Gpt-5 mini.https://openai.com.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Gpt-5 mini.https://openai.com

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.737507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:6db388ad318e37ee8e14ba8a68692c829d764e136af2b0200f0e41c2f5cf52c7

Observation a0f24f34-b1f2-4247-ae93-73c32ff85795 · outbound

This paper cites Kakade, and Stephanie Gil.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Kakade, and Stephanie Gil

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.740729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:11d356eae24786a82e89776a914cbd950e4359f6371b47ac8e5f14f4e188f9b6

Observation 196592b0-8581-4318-80dd-f6c9647a0572 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.743674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:2d635e77dc52db069ba64004f6c58d08990535616394b5c2707f8d749829d559

Observation d86cdffc-8199-47d3-95ca-41dc8fb298b5 · outbound

This paper cites Privacy-aware visual language models.arXiv e-prints, pages arXiv–2405.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Privacy-aware visual language models.arXiv e-prints, pages arXiv–2405

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.734622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:728bc8b9a65e0a5ad43b778233f5e9863518e9bcdb6a8c4a80b59c12c604cf32

Observation 3b3bbf27-a853-45e7-a520-f34aff2159a4 · outbound

This paper cites Large vision-language model alignment and misalignment: A survey through the lens of explainability.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Large vision-language model alignment and misalignment: A survey through the lens of explainability

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:53:42.831307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:4d83657089d6ca91cb269627c43a0d2a081d0adf8bf84509b746f9567f6df1d4

Observation 54906976-e569-47dd-ab12-621841c58412 · outbound

This paper cites Implicit multimodal alignment: On the generalization of frozen llms to multi- modal inputs.Advances in Neural Information Processing Systems, 37:130848–130886.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Implicit multimodal alignment: On the generalization of frozen llms to multi- modal inputs.Advances in Neural Information Processing Systems, 37:130848–130886

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.726472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:14a31a0e3c0ed3c22d6e78d1af499548748ba2e57208975e75a7d53a51fac910

Observation 60b280d5-0c3d-4b31-938d-4cc16d19dcfc · outbound

This paper cites Can vlms actually see and read? a survey on modality collapse in vision-language models.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Can vlms actually see and read? a survey on modality collapse in vision-language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.728845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:3bc2791d2b30c59bba722a7c9cfe334554d115e43d3cf551bb8658f0d9b72332

Observation 5843b2d5-4f6c-44ef-8e0d-bd96c7c5bd29 · outbound

This paper cites Towards vqa models that can read.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Towards vqa models that can read

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.731293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:54273b495c6e57c984973c0fdb642295ab3177d8f9f6e001de2fabf7a13bc939

Observation d4b85dbf-9948-48fd-b1ce-e644443b3680 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.863583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:c24e20aa8b56b7b0b604b04a75b5f96aa709d54060bab53be8c82c5cb3366461

Observation 7fe26bc2-6141-4b98-856f-f0f32a9bfad4 · outbound

This paper cites Gemma 3 Technical Report.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Gemma 3 Technical Report

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.810041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:d3fc0e5151bb31433bf93d44f63189a3651830e286c845e588ab124028c8466a

Observation e47361dd-c980-4304-a62c-686b6f60c25d · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Sys- tems, 37:95266–95290.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Sys- tems, 37:95266–95290

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.724208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:7c51adc16e4f66276db289ea88b58b4dd0d08a8560b50b13dd31a63e5dd1adcf

Observation 0c09722c-09c7-4516-8c4b-8a5fc25217ba · outbound

This paper cites DeepSeek-OCR: Contexts Optical Compression.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs DeepSeek-OCR: Contexts Optical Compression

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.843074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:bbdf0cf3e9697f4a6b7b4e8e6e7de100589eee46882311d3250c26c56e550819

Observation f325dff1-7aee-4b84-a507-45dde8614b54 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.819295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:580b1ed2ed2ab702aaa5f108a42f3794a2840c0bd0d6540ab75affbbb6b8c8fa

Observation ae593e76-deed-40ef-a2f0-8995678414e6 · outbound

This paper cites Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:53:42.860433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:6728cc2e73ba38565bf39b8df6a8b63a80356957bf32599f2cd450eccfcf7b6e

Observation 16645c6a-fe7d-4e50-9e8f-7dd4d34e5d90 · outbound

This paper cites A survey on multimodal large language models.National Science Review, 11(12): nwae403.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs A survey on multimodal large language models.National Science Review, 11(12): nwae403

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.717680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:7b404da48682c8b4f1e8d28a73cb03784d11cb0925a3463c9af51d1f3e5e36fd

Observation 4afca1b8-36af-4872-9d3e-b756c4d311e3 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.CVPR.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.CVPR

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.713011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:f3e33b6e63f370e4cf59eb744ed6fcf1784ebd399e7958be188661a4d2ee9c1a

Observation 589a76f0-0df6-47e9-bee1-2a261110dfcf · outbound

This paper cites Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.693262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:5a1967da9dc1df922ed0e8f202e282118e89d1db9167c84f2680229763e887aa

Observation 4fa82398-d1d5-474a-9a91-4008e3f7a5f1 · outbound

This paper cites an unresolved cited work.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-16T23:53:43.710916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:49f67da67342a7cfd2f72ca7054a6cc4c23e79e535e51e6ae9fc6e1ded93d968

Observation d1440028-f5c9-4e3b-98b2-809c01b43c95 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.839388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:52d514ea8b370d31a8e767c5c3ba055d56dce8f66790e649429b56d4122db3f1

Observation 8d104479-2ce9-4e3e-89d7-c282b2ac5f93 · outbound

This paper cites 2A + B = 15.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs 2A + B = 15

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.714896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:f284a9db48728a412ac472a6feef2abf567fe32f2ddfdc07b7cab2ef1bfcb2fd

Observation 9083f459-afaf-4966-a411-c16d97baf504 · outbound

This paper cites Open-source models follow vLLM’s recommended con- figurations for optimal performance.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Open-source models follow vLLM’s recommended con- figurations for optimal performance

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.719732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:c4a9430cb7e4e30e04fff33ae278d7ae33226862b4b4740517dac9b54ca1bdb0

Observation aff17873-d202-4236-82ae-b308bed27214 · outbound

This paper cites an unresolved cited work.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-16T23:53:43.721801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:d95c6742a2985ea8a538dabe576dfcb337bdb8d140c11e31d16e7977cfb343e8

Observation 4919310a-3e2d-4e33-858b-144c73884cff · outbound

This paper cites Please transcribe now.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Please transcribe now

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.706121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:e7ecf856d23ab48c9d5d98c6b848aaf8f85657e245558fd33969c105768480aa

Observation ff4b2e3c-aa2d-42d9-8cbf-f04827f9239d · outbound

This paper cites an unresolved cited work.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-16T23:53:43.708346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:11e198192a9277e90c09683423905bca181eb4d0d43bab7b91c23a4c86239a2a

Observation 80e681cc-4f07-46ab-80f8-db1826b653c9 · outbound

This paper cites an unresolved cited work.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-16T23:53:43.703977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:b458af8958917518c325fe04ea5832910a75e78188bfbdec51a93eb4bbd0a590

Observation 84c410b5-148c-4dcd-90e0-d5930986959e · outbound

This paper cites an unresolved cited work.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-16T23:53:43.699022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:ea7e84d5baefdeee77a5e6a0b5630789c587f2dae6cb0c33f7dd2d07e3ffd276

Observation 6e6ee886-1cbe-46f2-b96b-1ca12f7d06c0 · outbound

This paper cites an unresolved cited work.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-16T23:53:43.695382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:a203ac18b796df642faac7b3cdab1327743ca566a7114d55dd703b4ac28f567b

Observation b205d216-89bc-4f31-ad83-31c9d435137e · outbound

This paper cites an unresolved cited work.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-16T23:53:43.691091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:dffc2445447e708d6a364fd0955b204ecda3092b7913ff197c5c934cc6eab7d0

Observation 5169d7d2-0314-47b0-87fe-0124493ff0a6 · outbound

This paper cites Format your output like so: (1) 2a + 3b = 10 (2) a + 3c = 30 (3) 2b + 5c = ? Please transcribe now.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Format your output like so: (1) 2a + 3b = 10 (2) a + 3c = 30 (3) 2b + 5c = ? Please transcribe now

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T23:53:43.701503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:c8c9f20ccf76e006f43930b6f83f0130a0e14389caecb49a6c4b14adf6dbb00f

Observation 613873bb-867a-437e-bcfe-216708b10fd5 · outbound

This paper cites OCR-first.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs OCR-first

Reference 58

Resolution
malformed identifier
raw_fallback, observed 2026-05-16T23:53:43.686739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:c910905bc589314e0f3a13e7548a5254817dae6c0fcafcf3aa4ebe596725113a

Observation 86e56011-75ea-4e83-add1-2007859ee779 · outbound

This paper cites Image accuracy, RER consistency and OCR correct scores stratified by resolution (50, 100, 200 DPI) for all questions.

Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Image accuracy, RER consistency and OCR correct scores stratified by resolution (50, 100, 200 DPI) for all questions

Reference 59

Resolution
malformed identifier
raw_fallback, observed 2026-05-16T23:53:43.684161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:52:13.813452Z digest=sha256:75f998b99db39e5816d5c9e9ee03996d046cf362693ae78e0fa9087dad3aec93

Pith citing papers

Observation c707c0d3-f389-4d8f-8197-e927dbff2e72 · inbound

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs cites this paper.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.402402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.402402Z digest=sha256:48856da36c6b589b6d234bc87ab2385236efeb0474e6c8654f9e00dab11b82dc

Observation b17d3675-72d1-4890-90cb-bf5c3a2a02df · inbound

Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations cites this paper.

Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:46:24.168279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-05T00:46:24.074055Z digest=sha256:486e528f2b724927f97ab0937ec40b8965be3a58d70a5fbfadac84a80e13cbc0