Pith. sign in

Paper Citation Record · LEDGER

Cross-modal Information Flow in Multimodal Large Language Models

As of 19 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 7 inbound Pith citation observations for arXiv:2411.18620.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18620 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:06:10.147882Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:44:10.027628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T00:02:17.731733Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 94d4e1d1-55df-4aef-a79d-d5fdbb77f07b · outbound

This paper cites https : / / www.

Cross-modal Information Flow in Multimodal Large Language Models https : / / www

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.836713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:09.948464Z digest=sha256:989f39234e7274f03d8ac65304276d827311de9d3065cf2ebf9d8166266bbc46

Observation b93005f4-16e5-4816-8d8c-3f3374ffa350 · outbound

This paper cites https: //huggingface.co/lmms- lab/llama3- llava- next-8b, 2024.

Cross-modal Information Flow in Multimodal Large Language Models https: //huggingface.co/lmms- lab/llama3- llava- next-8b, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.823396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:09.953444Z digest=sha256:1326c41eaf0e21a5a9fa869e24b842facaca463aec7635cd7b97551c49c73cda

Observation 656b0e57-6fb0-4515-80e9-c2ccddac5ba5 · outbound

This paper cites Vl-interpret: An interactive visualization tool for interpreting vision-language transformers.

Cross-modal Information Flow in Multimodal Large Language Models Vl-interpret: An interactive visualization tool for interpreting vision-language transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:09.957745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:09.957745Z digest=sha256:3f04e697e94c4153e6575912421e3ea0b1bcb00a343e79444a10bec5223e7fc2

Observation d663816d-610a-4413-b229-5146f384283c · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Cross-modal Information Flow in Multimodal Large Language Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:09.962225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:09.962225Z digest=sha256:8a22788f2ebdeaa2f487a6664b577f0b2fc0a17432d21e52c741e79794a15cef

Observation 35af6b75-8de0-45d1-a119-2ea0f693bbd5 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond,.

Cross-modal Information Flow in Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.801689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:09.966787Z digest=sha256:64939e0ef53326f643d02d08ae70d0d7a2a43bb9ac764e6baf90b4e0d1eb0b85

Observation 8b5ca4c8-7f59-4eb3-a0c7-10259dca8f40 · outbound

This paper cites Understanding Information Storage and Transfer in Multi-modal Large Language Models.

Cross-modal Information Flow in Multimodal Large Language Models Understanding Information Storage and Transfer in Multi-modal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:09.975089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:09.975089Z digest=sha256:87da270dd52c211b47cd06e690583db2863adce59be217fffb4375b15410f851

Observation 5679eee8-5e35-46a9-8d91-13f3d1c77669 · outbound

This paper cites Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models.

Cross-modal Information Flow in Multimodal Large Language Models Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:06:10.458456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:09.979415Z digest=sha256:30afef017a60af78190b2d37f8f2d0271dd498210932da08a06e62abfc0f9068

Observation a4c46df5-8b62-4585-9e3b-cf128bfcc9dc · outbound

This paper cites Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers.

Cross-modal Information Flow in Multimodal Large Language Models Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:06:10.438330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:09.983791Z digest=sha256:731bc0c68a9d4148434ce805690011dd2dff6a7b37b2ce09105585d908f4d4c2

Observation ac4c82cb-e752-4536-b798-a8e0067ceac1 · outbound

This paper cites Probing multimodal embeddings for linguistic properties: the visual-semantic case.

Cross-modal Information Flow in Multimodal Large Language Models Probing multimodal embeddings for linguistic properties: the visual-semantic case

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.788524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:09.988420Z digest=sha256:3e9d5e5dd5895044de0396cbe92b7615c2b7206309e200f141003dbf3a59512a

Observation 4bfa2db5-488c-44fb-baad-05e8fa16563f · outbound

This paper cites Knowledge Neurons in Pretrained Transformers.

Cross-modal Information Flow in Multimodal Large Language Models Knowledge Neurons in Pretrained Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:09.992993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:09.992993Z digest=sha256:d6be9c2d7103746672673060ccfca5b34a65bae9af7b133e74d3e5480e197d58

Observation c359de28-3689-433e-a264-2f747cfb7911 · outbound

This paper cites InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,.

Cross-modal Information Flow in Multimodal Large Language Models InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:09.997576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:09.997576Z digest=sha256:c04176cee930767214cee8662453835ead776b65c65a95c4156df23bbdae14d2

Observation 8ddc5ead-d7ef-44f1-805f-914c96ea8c55 · outbound

This paper cites Analyzing Transformers in Embedding Space.

Cross-modal Information Flow in Multimodal Large Language Models Analyzing Transformers in Embedding Space

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.006430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.006430Z digest=sha256:fc7d4f3b4c8d05de285141e5928ca9c4bacbbd2df7852f13ebc29b3cc4a4e5a5

Observation 73934708-6a89-4f42-919c-c82c3fb29a1e · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Cross-modal Information Flow in Multimodal Large Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.001598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.001598Z digest=sha256:970ffd6ab8eb36ff7649435ffb681d7d13b4da8c16ea2689450ea8feae3258e0

Observation 506a9b9d-a8d7-4a18-986f-01af989092a7 · outbound

This paper cites The Llama 3 Herd of Models.

Cross-modal Information Flow in Multimodal Large Language Models The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.014632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.014632Z digest=sha256:ef89e9dbfbbaac582a8fe4a5adfb309d15d4f8eec461faa1f196110b7e1544af

Observation 9743b524-6887-4e46-b4df-1f50d94e4b79 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Cross-modal Information Flow in Multimodal Large Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.010600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.010600Z digest=sha256:56ab75866c415b2b209aaae4728e5c58a907b4944e182b3cd3b576106164508f

Observation 9b3243e9-8b72-4c70-9301-003728575019 · outbound

This paper cites Eva: Exploring the limits of masked visual representa- tion learning at scale.

Cross-modal Information Flow in Multimodal Large Language Models Eva: Exploring the limits of masked visual representa- tion learning at scale

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.752851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.022505Z digest=sha256:10e233fb5fcb1f6ac64e2828950f977bfbcacb2783d414e6f527bf69af89ad7b

Observation fc5d2625-70ea-493e-9483-d3f5209f3c9b · outbound

This paper cites A mathematical framework for transformer circuits.

Cross-modal Information Flow in Multimodal Large Language Models A mathematical framework for transformer circuits

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.765687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.018553Z digest=sha256:0bf33ccf5d5348b2435c52c176bdd530556e853d9c3bebc95379c58b4042aa62

Observation bc3efc44-c022-4ca5-82b5-ef1278bfe58c · outbound

This paper cites Transformer feed-forward layers are key-value memories.

Cross-modal Information Flow in Multimodal Large Language Models Transformer feed-forward layers are key-value memories

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.739359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.030663Z digest=sha256:7982aeaf772f47fdf4aa34096c23bce899c608c26d2155e3f5a2fd40aa770621

Observation d646a383-9468-4e36-af97-baa081b32688 · outbound

This paper cites Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers.

Cross-modal Information Flow in Multimodal Large Language Models Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.026636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.026636Z digest=sha256:ee10ffeee12ae53d583637e9dc5bed0bc79317e2495de1a2ddfa348ed2a24658

Observation 1ceb7cdd-dd6e-4f22-884e-130efcc4e62c · outbound

This paper cites Probing image- language transformers for verb understanding.

Cross-modal Information Flow in Multimodal Large Language Models Probing image- language transformers for verb understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.726312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.038701Z digest=sha256:8ff4e84f1de41245c11d59b8a3c4f5507c4621552e258451e3c95f24f83054f7

Observation c0887d75-c0cd-4fce-ba23-d1fb115fd837 · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-Regressive Language Models.

Cross-modal Information Flow in Multimodal Large Language Models Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.034653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.034653Z digest=sha256:d5e31ab6f84d3567638fb83b10dae1e5bcd86c09be557efca0af79a238234da4

Observation 24f55189-75cb-47ef-a182-ab2065133cca · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Cross-modal Information Flow in Multimodal Large Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.698375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.046431Z digest=sha256:4f5e6d95356def33b1eebf6008a44df8af5a5f8ef49a25f2f07a0af1dd36ab33

Observation b529b4a9-d90d-4394-9e30-6cac4dd11304 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Cross-modal Information Flow in Multimodal Large Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.713166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.042574Z digest=sha256:7aed7955e334930deab1973d927e8acf0f5891d3d9092a9661b16a1403760351

Observation 095b51f7-cf03-4e9c-b171-196e1198c93f · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Cross-modal Information Flow in Multimodal Large Language Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.054165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.054165Z digest=sha256:e4a2b40aa499ad01cc801b57d7be02ce859dab2759a0cd1b9fd77bbd15a32265

Observation 03c8f297-00e2-4a9e-80f2-7088517572ae · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Uni- fied Vision-Language Understanding and Generation.

Cross-modal Information Flow in Multimodal Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Uni- fied Vision-Language Understanding and Generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.685237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.050256Z digest=sha256:4edd356d66623306cef2eb0a30cd29212aac2b5934c8a0fa01f6869ee05cdf6a

Observation 3d512324-036e-4835-b169-691bff587c95 · outbound

This paper cites Visual Instruction Tuning.

Cross-modal Information Flow in Multimodal Large Language Models Visual Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.062531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.062531Z digest=sha256:493c1f5de9da6086103ba064745129ea0ef46636e583af0dc3e6fede443804d8

Observation 1f2f1f64-2e6c-44a6-9cc1-e04fdfc19dde · outbound

This paper cites Mini-gemini: Mining the potential of multi-modality vision language models, 2024.

Cross-modal Information Flow in Multimodal Large Language Models Mini-gemini: Mining the potential of multi-modality vision language models, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.671695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.058671Z digest=sha256:a734b7d8af738b616a0cc8b866ac1612d23cd51a01def8f7a86cad50017073d8

Observation 7b909667-9ad6-475c-a12b-1769e563eb92 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Cross-modal Information Flow in Multimodal Large Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.642189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.070475Z digest=sha256:23bd4b150789bbbb22576a9c6642a8498bc752599b0b0ccaf229bf80084140df

Observation 2554fcdb-e476-4b9a-9307-a460cbfcd171 · outbound

This paper cites Improved baselines with visual instruction tuning.

Cross-modal Information Flow in Multimodal Large Language Models Improved baselines with visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.656793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.066709Z digest=sha256:44b41312ba9c47d530659f3306bb7611e44a75e259e3ee05f0109a54f3f9d5b9

Observation 43bfb556-49b8-4768-93a6-1771ed119ad4 · outbound

This paper cites Locating and editing factual associations in GPT.

Cross-modal Information Flow in Multimodal Large Language Models Locating and editing factual associations in GPT

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.614732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.078402Z digest=sha256:4b349d1cdae21942176a80c851c9fb26979d34ac74de578b0df2c4fcdac0f0fa

Observation 897b8e6a-1902-476c-820a-b999ced49df6 · outbound

This paper cites Dime: Fine-grained inter- pretations of multimodal models via disentangled local ex- planations.

Cross-modal Information Flow in Multimodal Large Language Models Dime: Fine-grained inter- pretations of multimodal models via disentangled local ex- planations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.628724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.074445Z digest=sha256:000aef52daa8366180748f28c0f09395c96bb5795c663a09464862786f0a2040

Observation df9fd07c-8bed-4973-9504-cd4d08f1ae97 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

Cross-modal Information Flow in Multimodal Large Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.086378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.086378Z digest=sha256:6a281d731ae57f21a9fcff220fac71f4ade7b687f4c3187ab7346dc4e7b0a1fb

Observation 583d1950-009e-4941-8ae9-59e0e38da9e1 · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

Cross-modal Information Flow in Multimodal Large Language Models Progress measures for grokking via mechanistic interpretability

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.082279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.082279Z digest=sha256:29bd9da973db76db5d837f9fc55686d306f4a0550a42283c40dfa88ad11dadd6

Observation 131e2a97-032c-4d66-b3f7-5af3468968ef · outbound

This paper cites Towards Vision-Language Mechanistic Interpretabil- ity: A Causal Tracing Tool for BLIP.

Cross-modal Information Flow in Multimodal Large Language Models Towards Vision-Language Mechanistic Interpretabil- ity: A Causal Tracing Tool for BLIP

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.586304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.094722Z digest=sha256:4ac1096f427b7b3688c3253df4b80f4dfe75efbbad02a12f173488f03ca1fcce

Observation d8d59e14-b39e-467d-9856-5b30d0efe9d4 · outbound

This paper cites Mechanistic interpretability, variables, and the importance of interpretable bases.

Cross-modal Information Flow in Multimodal Large Language Models Mechanistic interpretability, variables, and the importance of interpretable bases

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.600490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.090735Z digest=sha256:6d03ffb8cc8f3db5c9c236e81bdea0647b8f42718c21d5a7fec1b3fc9c67b7f3

Observation 4df14011-c962-45a4-8808-e5d70d67ecf7 · outbound

This paper cites Are Vision-Language Transform- ers Learning Multimodal Representations? A probing per- spective.

Cross-modal Information Flow in Multimodal Large Language Models Are Vision-Language Transform- ers Learning Multimodal Representations? A probing per- spective

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.564761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.102413Z digest=sha256:3ad8d5dff2ce693fb9e49e4408fb18b52fcbf76c9072427d919141466dcdd883

Observation d62d5ceb-a28a-46be-be91-5003edef89a5 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Cross-modal Information Flow in Multimodal Large Language Models Learning transferable visual models from natural language supervi- sion

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.098826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.098826Z digest=sha256:57637ed9938d5bbb990f6d24a1d96bf96cededb0cdcf596ed710079da0f7ce00

Observation 918cbea9-5af6-48d8-8ab0-361f6fec4e68 · outbound

This paper cites LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models.

Cross-modal Information Flow in Multimodal Large Language Models LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.110272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.110272Z digest=sha256:7950852b546faefe9cd98b62c5b2ce99ad1d179d971bb70427ba0c164538175f

Observation 6e240497-9f54-4a25-acaf-0bc3eb56ca71 · outbound

This paper cites Multimodal Neurons in Pre- trained Text-Only Transformers.

Cross-modal Information Flow in Multimodal Large Language Models Multimodal Neurons in Pre- trained Text-Only Transformers

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.550805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.106441Z digest=sha256:903a742934065f3443d49d0a0a64d9ee63f2450aa1388dd30949e48816fc98bb

Observation 5c458eaa-781c-488f-b47a-84b191c24b86 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Cross-modal Information Flow in Multimodal Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.118405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.118405Z digest=sha256:7ddfa39561fc8d99f3a863e4ee75a4cfc9321b50a080c1fee3f50081c862d6c4

Observation d985c5aa-b104-416c-8262-bb2e933a553b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Cross-modal Information Flow in Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.536182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.114376Z digest=sha256:e4f5026d06681055440ba4cbdd8550812decb3df3e1a0a16459c707fbb97aa41

Observation dbc8cd6a-4d42-44a7-b9ce-0ab0fcb4713f · outbound

This paper cites Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning.

Cross-modal Information Flow in Multimodal Large Language Models Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.126476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.126476Z digest=sha256:012d87e49015ca8a6322a227c1f44bb013725e4cbc75ecaf4ac88b3e33aa0392

Observation eb6cd9b0-6eba-4555-9d87-cb8cf918d49d · outbound

This paper cites Attention is All you Need.

Cross-modal Information Flow in Multimodal Large Language Models Attention is All you Need

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.122548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.122548Z digest=sha256:fea713d763969fe640d11bc7336cef60670e4f5c61a106529fc6544789ff4fa2

Observation 454bc649-dece-44a8-897c-b51279e05c9c · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Cross-modal Information Flow in Multimodal Large Language Models OPT: Open Pre-trained Transformer Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.134669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.134669Z digest=sha256:6872ffbc039ac6ec8cb6b379fef0cc6dc86c3542f8697fa62bd11bc0b53143be

Observation 1b5e16aa-6ce4-453b-b9d6-fbc87cb9fb93 · outbound

This paper cites Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models.

Cross-modal Information Flow in Multimodal Large Language Models Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.130548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.130548Z digest=sha256:c9b448b5712b33ed98630af576f045af74d275c6e7d5e070cea0f6fd2db33cce

Observation 9393027a-a1a3-4aaf-be1e-faff84f16657 · outbound

This paper cites The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?.

Cross-modal Information Flow in Multimodal Large Language Models The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.142953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.142953Z digest=sha256:592c04e8976e19b1579adcf0a54aab5071ad8eb60866b1649048033b90353930

Observation c13a0267-c980-4bde-b20f-2cdeacd98c95 · outbound

This paper cites From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks.

Cross-modal Information Flow in Multimodal Large Language Models From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.138708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.138708Z digest=sha256:28d8d4230cf7e08a86241f2008abbfb37252d1ed6b5d8b595b21328f2d529c9a

Observation f7589b04-ade1-42b5-8c32-67636639326b · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Cross-modal Information Flow in Multimodal Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.513689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T11:06:10.147882Z digest=sha256:e18a756d4cf14607dda1b2326f0360a225564ee7336fec4c6af6699fa8dda49c

Observation 0cca1222-5a9e-48da-a720-c98c617057bc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Cross-modal Information Flow in Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:09.970888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:09.970888Z digest=sha256:9c30cfa124de4c64e803c446ba9fa4c7e237897c24a6eccf49a244ff21d2e0b3

Pith citing papers

Observation fbc8c65d-34f7-4862-ad73-5e0c6f2c7509 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models Cross-modal Information Flow in Multimodal Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.735126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:49f4879c467e2d7220056f17fd1a667ff39b405d23a69f1f3b8792705da171a5

Observation d595b4d8-664f-4f4c-a21c-947feaa1fcd5 · inbound

Interpreting Social Bias in LVLMs via Information Flow Analysis and Multi-Round Dialogue Evaluation cites this paper.

Interpreting Social Bias in LVLMs via Information Flow Analysis and Multi-Round Dialogue Evaluation Cross-modal Information Flow in Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:10.027628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:10.027628Z digest=sha256:d22476e51306397b59c40907fbc88303e0921287b784ca4f61a2dd59acd52627

Observation e90e38c3-498d-4f73-a74f-53f56c143979 · inbound

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation cites this paper.

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation Cross-modal Information Flow in Multimodal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:26.857624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:53:26.857624Z digest=sha256:90854687c209f54f50b00dc05a4b80aa7307d73f77bd04300d99054f73226a19

Observation 03c0985e-2934-4751-9bbb-32103dfcca6e · inbound

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models cites this paper.

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models Cross-modal Information Flow in Multimodal Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:53.735239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:53.735239Z digest=sha256:7fbb6a7b6346756be3d7eb3bdd6ab35177799a84cea4c2166879055e2d2812fd

Observation 75c6598e-4096-4d67-8213-229f86126856 · inbound

Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models cites this paper.

Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models Cross-modal Information Flow in Multimodal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:40:05.544824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:40:05.544824Z digest=sha256:c500695d25a2a31a9bbe42ddc2345e2709aa287fa30384b809c902b7d6013828

Observation 37595e36-29d1-423e-8221-6cd810049313 · inbound

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models cites this paper.

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models Cross-modal Information Flow in Multimodal Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T16:30:12.654347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:30:12.654347Z digest=sha256:0b549756b8f2bb8251b4a3fa7c69691736eec0ea2398d590ca3ab0e79bb7f996

Observation 9f9da4c6-c19e-49d8-91e9-b939a7f244c4 · inbound

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination cites this paper.

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination Cross-modal Information Flow in Multimodal Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:28:05.500293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T06:25:43.369455Z digest=sha256:cc11ec6bfe1a8cf99381431a9ec3335176e051b03e8b215b4b8eedf6023c0a64