Pith. sign in

Paper Citation Record · LEDGER

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

As of 9 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 3 inbound Pith citation observations for arXiv:2505.17316.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17316 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:09.476244Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T04:35:49.701061Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation fadc0ca7-aeda-4650-b84f-c888ac9a3105 · outbound

This paper cites Visual instruction tuning,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Visual instruction tuning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:14.433783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:04.097733Z digest=sha256:c32da8faca6c381ab109c39e6383234235bfc12f1cc22b15bd95cb3c34e2f5b9

Observation 29de9b4f-7a6c-45b0-ba66-65863a79d134 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Improved baselines with visual instruction tuning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:14.277054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:04.208405Z digest=sha256:72931dc5495911cab48adf7d7b03b264a2447b3991134f98c59ab39bfacccaa4

Observation 208045f1-7904-41c4-a316-77bcd68edc2c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:04.379623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:04.379623Z digest=sha256:6f814bd7260fca44347f2d577afe87dbd163986abd078f795eb7b93a09dc37b4

Observation 70584995-12c7-48a0-89ce-68f725df6f8d · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:04.539750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:04.539750Z digest=sha256:e80503c92e04d35e0e0ea51e1746cd2fcc93b88b66805b7cc6a2a6ef6a6e1eb0

Observation 5db061ba-bb92-4291-80e2-ec1c3ab061f2 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:04.677406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:04.677406Z digest=sha256:1c3ff31bbfa7435f1f723c20f69c2cf88504dfee4d7f4c69a8b8dc5a30c28198

Observation 98609cdd-071d-48fc-a888-967810b5d0dd · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Flamingo: a visual language model for few-shot learning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:04.770528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:04.770528Z digest=sha256:2b1d6cf69ecd43a3d481feed16cd21b3db6cc20c8a84d407b2d4778f7bb513df

Observation bc4869a1-3879-4476-9c1f-272048cb17a6 · outbound

This paper cites Language is not all you need: Aligning perception with language models,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Language is not all you need: Aligning perception with language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:14.117267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:04.917798Z digest=sha256:c5e0ddae72f667004e9cf11b2e4e563d967aaca7a458e32183fd59c876fd3c27

Observation 47117c81-1b56-4a5b-aae9-9bfb30599547 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:05.094450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:05.094450Z digest=sha256:8425fd6fb1bed54c9f638fa956192cb6bb8e9c0ef6027b632e0c55fc0f8d9b1c

Observation 09fdddc4-ddbc-4e47-8c00-b13a5032e0aa · outbound

This paper cites Kimi k1.5: Scaling reinforcement learning with llms,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Kimi k1.5: Scaling reinforcement learning with llms,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:13.964943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:05.186005Z digest=sha256:b6805cecddf8d90b7cafda0c4eeaba0cfd4a9ae8e05dfbb9f80686e0697bab21

Observation 61a44add-9445-4537-aaa6-470852bb6cc7 · outbound

This paper cites Multimodal transformer with multi-view visual representation for image captioning,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Multimodal transformer with multi-view visual representation for image captioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:13.839155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:05.270372Z digest=sha256:ef7e51923715cbe687193b878d71115fdc9c87062909379938dbe023c91e67a7

Observation 12569a1d-b954-4dba-951a-aa8085cbd105 · outbound

This paper cites Vqa: Visual question answering,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Vqa: Visual question answering,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:05.384694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:05.384694Z digest=sha256:b020e62716d7858ffbbacdc49fdadfdc0f5094fdaf9d8fa340c6a0c9f3903a7a

Observation 8bb64434-aec4-491f-8c0d-89bf212f3a5b · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:05.476390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:05.476390Z digest=sha256:983cfbeee6b4bfa79bba0ad9e2f2ca9b42229fac5ea9182f15ac56e66b968a2e

Observation cc12dd15-7ac6-46de-a797-54b4ce1bddf1 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:05.597599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:05.597599Z digest=sha256:93f4c609f11bdade88e906b0f1914af334bd9566faa695cf64eb6097dffe441d

Observation 53446d3c-8f1f-4f37-b759-51d558bf7049 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:05.690123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:05.690123Z digest=sha256:a8f45060be88392f1d84a62da2d99e399050c22ad5cfb3bf26f29924e9562886

Observation 29156912-f77a-44c8-ac02-3ef445607f04 · outbound

This paper cites Glamm: Pixel grounding large multimodal model,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Glamm: Pixel grounding large multimodal model,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:13.659627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:05.769581Z digest=sha256:2071a02701256caf693525d66f357b859bb7b6b40061cc48190c7a27a28da64c

Observation 099610fd-b9dd-44de-855d-e0271fb59288 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Hallucination of Multimodal Large Language Models: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:05.858115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:05.858115Z digest=sha256:b199928c84c0b707d8050a2563de1b5028b2b9522fc22ad9eb80b366ee1b82a7

Observation d751b048-6e58-4c5e-9b9d-f1acb8e9589a · outbound

This paper cites Honeybee: Locality-enhanced projector for multi- modal llm,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Honeybee: Locality-enhanced projector for multi- modal llm,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:13.495833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:06.013664Z digest=sha256:b4834ae7e28c35408ae6576ee6fa1a2459b142e966efeb4f234e3ee6a3b7d94b

Observation 3bcaf127-cac2-4f33-aa77-d8a119ec9186 · outbound

This paper cites The Platonic Representation Hypothesis.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models The Platonic Representation Hypothesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.128909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.128909Z digest=sha256:2806ed597cc6fac57d94bcd1babc574b3ac335177da24299c2b2d0221c45437d

Observation 29f11f75-9d76-409a-8d86-51f6192b84f6 · outbound

This paper cites von Neumann,Mathematische Grundlagen der Quantenmechanik.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models von Neumann,Mathematische Grundlagen der Quantenmechanik

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:13.345684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:06.209608Z digest=sha256:758e74d2dd8fc331f65aa34a45dc60dd1fca35831ac88516795432838d0b9ea8

Observation 1cd1b476-c3c0-49af-a2cf-202ca427a365 · outbound

This paper cites Linear algebraic structure of word senses, with applications to polysemy,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Linear algebraic structure of word senses, with applications to polysemy,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:13.022576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:06.426208Z digest=sha256:49f7637f720c8a6bdc720c38af2df9fb8489bb09a2d5f5531140961188595456

Observation ced2b0b2-19d7-456f-875c-eeb48feff724 · outbound

This paper cites Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.502588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.502588Z digest=sha256:0cbbedf54439d1bfdd11473bdc69748a6d828dad89337ea6a48f24c1e8fd1f1d

Observation ba7c44c2-883e-48e9-96ee-9444c5df2e11 · outbound

This paper cites Signal recovery from random measurements via orthogonal matching pursuit,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Signal recovery from random measurements via orthogonal matching pursuit,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:12.858259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:06.619408Z digest=sha256:818549e5207581652eae08b8c15bdb43415d8c272cc52fb937406499240cf02c

Observation bc3cf9d0-2f91-4990-82cc-56bd5650f420 · outbound

This paper cites Recognize anything: A strong image tagging model,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Recognize anything: A strong image tagging model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:12.717265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:06.690111Z digest=sha256:b07264712da3d28d4eb4ebd162ba374ea44657195bc762715c644af7bcee8edb

Observation 8fedc27c-4a1b-4068-8304-468498f53fef · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:12.505247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:06.753385Z digest=sha256:8caba9678e05d09820417430bb9648693e3f9425f40e49300a3c596ef1d98723

Observation 69cb3100-d9ea-4117-95a9-8016afa821b8 · outbound

This paper cites Segment anything,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Segment anything,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.817193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.817193Z digest=sha256:0558b637475086da71caaf472212512bb534565078bf593dceba8c0a2be89329

Observation f5cf6810-da43-4cd8-81d0-acd914e552a2 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.857116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.857116Z digest=sha256:3cf883a59002ae286060dd6aa4412670e0c0a1d4d55cdbeeadeb3ea15fc8c4ac

Observation 26a3b850-f28d-414e-a021-fcf65bd986ea · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.924738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.924738Z digest=sha256:69d383feba57d6f2ef173c47b23948d23083f76648830ce3f3388a9df4e401c9

Observation a61da63f-e94a-4644-8fe6-02badd35fdaa · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.984546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.984546Z digest=sha256:30c90649f91ca14ce36496feaeed964b156d60574cc76aa9d884251534a74373

Observation a07487ff-fff9-47ec-8cc0-69e5657edb1b · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.035989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.035989Z digest=sha256:ec40aa79aea37dfc2d2818e8fecc44e6c534b0b5482799b8270118092c2e7e38

Observation 294bd963-4843-4895-a63b-ea8567de710b · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.161010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.161010Z digest=sha256:d2f417850b14a3762740e812e888b7d22d348666d95f7b539ca5c73c744be8f6

Observation 410e58b5-1c87-450d-9d58-7086b04f6ba4 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.253612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.253612Z digest=sha256:ce53a4b6d06127283167d397113454aa23bc356a16b7f984be87a0b3ab6aface

Observation f78e1c9d-00dd-436e-bc06-0b9b921d9ed9 · outbound

This paper cites Law of vision representation in mllms,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Law of vision representation in mllms,

Reference 32

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:52:09.875483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:07.363562Z digest=sha256:76151b521d14c06743865f937070f8499c951288aefce6895fb659757f46ecef

Observation 58d0ae74-b066-4fa9-8c35-3430164cc0ef · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Learning transferable visual models from natural language supervi- sion,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:12.324474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:07.459692Z digest=sha256:146b0b0f3512c6469aa640c7e900805622dacfe19fa1eeaf6bc00cce06feb975

Observation afd9b67a-ee7f-4db6-a5d8-cecc6819e3b2 · outbound

This paper cites Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.553797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.553797Z digest=sha256:89feb395036c119f39725ade89fd2673c92c7bc0dd339adfc2a1f397c0e3f523

Observation d5c0e1f4-08fa-4fae-a4d5-f324eebe1557 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.702315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.702315Z digest=sha256:3b682fbea19ddf64a3d223c4b6aa0e183ff33a51c2a85613c49b222cc774c005

Observation 258e03ae-c032-4490-886e-02438a42f2c5 · outbound

This paper cites Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.801312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.801312Z digest=sha256:9950ca509eec333da2e8b4dd2ace912b63197891570b21dfb0c1189672d608ab

Observation 15f7dcb8-de70-4ef8-a31d-a938af74c5c5 · outbound

This paper cites SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.914884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.914884Z digest=sha256:aafdebe22d4432b95e220baf886e8bfa99cd385c446757af1e30a810332d974a

Observation ffb7bfab-4a76-425f-9254-89705a2460c8 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Honeybee: Locality-enhanced projector for multimodal llm,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:12.176356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:08.003998Z digest=sha256:852a39c2d98211c8c935bc206df73f6d09faa5d6ba2d9652c594f78a2c163642

Observation a4e8b62b-1cb5-45b0-aee7-596158711fb4 · outbound

This paper cites Microsoft coco: Common objects in context,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Microsoft coco: Common objects in context,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:08.137009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:08.137009Z digest=sha256:d8772f3944e1000075ab73b6cb104c2df24a44948422c81e4dff893a2d95c84f

Observation 66925e60-7f15-4ea4-b05c-42d93478d3d6 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Referitgame: Referring to objects in photographs of natural scenes,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:11.991080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:08.278205Z digest=sha256:c6a32a201ce3e395366d42d9c7837c209bb4559e73fb637faaea7402cfc80521

Observation 766fa7c7-55c0-4cef-81ca-9e6e07e5deed · outbound

This paper cites Generation and comprehension of unambiguous object descriptions,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Generation and comprehension of unambiguous object descriptions,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:11.792342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:08.425913Z digest=sha256:a8c03d4ab11921c938ed2bed83d8ed3951a67b414c22e4019ab84e783e7e3373

Observation 519e912a-9cd7-48be-99a7-28ddc9d0707c · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:08.552384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:08.552384Z digest=sha256:7b2b7a74c6b05d7a330564a79b6285027f96baa840c6de8931c31d8d3be4f45a

Observation 94b6ad84-73b1-435c-aefb-38348b595dfe · outbound

This paper cites Introducing idefics: An open reproduction of state-of-the-art visual language model,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Introducing idefics: An open reproduction of state-of-the-art visual language model,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:11.567486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:08.630284Z digest=sha256:9dace4b3b8db1dc68f6fe211c33783664a5166e81ffaba9fe994a48e8362356c

Observation aebb9424-4cd8-4cc6-9db1-d85de7550613 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:08.711601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:08.711601Z digest=sha256:e7d46bc4b044e418fe1a28aa62c7ef178e0a7cab51d37f203abd00b2a907fc41

Observation 3ee107fa-bfd8-4446-bb2a-7f31b967b7e0 · outbound

This paper cites Instruction Tuning with GPT-4.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Instruction Tuning with GPT-4

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:08.800402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:08.800402Z digest=sha256:7bc88588d9f632275905590ddc36d273b7aa8f1fa19e6e1cc3969ef6979500a0

Observation da6a0ac9-3e8b-4782-b358-5e15b6c0d4cf · outbound

This paper cites The Llama 3 Herd of Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models The Llama 3 Herd of Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:08.905998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:08.905998Z digest=sha256:ecc1789d9717eee32cad315441825315bc4d30bca84e938184c610c873d48f45

Observation 55246c22-573a-4a8c-a38f-94e1528d6620 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:11.305627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:09.039392Z digest=sha256:2d6a3105135a5193197cdb9907563ec656aed9e86292b4bbd6672f7f5adcfdad

Observation 717aab0c-5678-414b-b5a8-09c0685b4fdd · outbound

This paper cites Towards vqa models that can read,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Towards vqa models that can read,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:11.043224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:09.193236Z digest=sha256:d9c2890fe11f87a33f8f6090e552af7bbaf1a3057107b05253c377de83874130

Observation b5e92fd3-23ec-473a-ba67-f160ea063043 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:10.787542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:09.291633Z digest=sha256:36b21699888617420a292a0e1aca59adda6e25ee96465d6b6b003f500504b712

Observation 9143f5af-b801-45d4-b5fa-ced2989a23a3 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Ocr-vqa: Visual question answering by reading text in images,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:10.559675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:09.394927Z digest=sha256:41819fe0d28f5c557e5058892140f5a4082e0182bc5a769ed313329c2ce03dc7

Observation 894462a2-6627-4fe4-a7f7-ad0abbf5ec84 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:10.298358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:09.476244Z digest=sha256:da64ec41af00d48f19ab6c1c48f8d85eb98a704e0e984346cc55d7c404ca1593

Observation 81ffb6bc-c061-4c86-96d1-c38e50e59701 · outbound

This paper cites an unresolved cited work.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Unresolved cited work

Reference 1932

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:13.179783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:52:06.306580Z digest=sha256:fcce91c9640bcd5614f86c434cf001e7e775a4a0293a3e23c5b4823fb3119de5

Pith citing papers

Observation 016138de-0e18-46f9-a76b-1f2029ae70c9 · inbound

Latent Denoising Improves Visual Alignment in Large Multimodal Models cites this paper.

Latent Denoising Improves Visual Alignment in Large Multimodal Models Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:09:26.689929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T23:07:54.806529Z digest=sha256:d9515d8ed7b45b4eb81cc28d79752edb923d12643fe0c3ffb04f95a927b5829c

Observation 9c925f13-3f9d-4834-b9d5-bd95ebda16bf · inbound

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media cites this paper.

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:08:20.436198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T14:05:14.737146Z digest=sha256:26192978943d223709f12bf414f1e020328a958358d0cf626d39618ae8aa1d1f

Observation 5774306e-2e1f-44d5-b230-98bb897ca4f1 · inbound

Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning cites this paper.

Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T04:35:49.701061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:35:49.701061Z digest=sha256:c505415e5fe0a2364d654da29e9152c9f0c05dfaa82d70b8cb63b60fbfb9c084