Pith. sign in

Paper Citation Record · LEDGER

Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2410.07167.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.07167 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:09:09.883745Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T19:45:00.682859Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 874e5f9d-2eff-4ffd-9b1c-7bc2eaa1a9ea · inbound

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models cites this paper.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.603097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.603097Z digest=sha256:bfb5ff53f13a1b422e8f5c67f0d550445a0dbb93deb0ede6edc21641cbb20635

Observation 97833a46-8d04-47f4-9ce4-7e47936c38dd · inbound

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion cites this paper.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.456702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.456702Z digest=sha256:8e89a1486ea6fc04a7f6172738ea2d999f01055de2f0c48307c23ef8b726a0bb

Observation e2d55118-93a7-4eff-ab3e-8831eb90cd91 · inbound

A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models cites this paper.

A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:09:09.883745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:09:09.883745Z digest=sha256:580274b646366bbbb341d9f0d9cf9fecbed3930870fc643c26ce29fc373054db

Observation d6970a0a-68eb-4b8b-8841-49d7ff231925 · inbound

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM cites this paper.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.527741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.527741Z digest=sha256:d55aa527d5373da15c58b122c31a897f88325aca54eec29d9e5b548f11279ad9

Observation afd9b67a-ee7f-4db6-a5d8-cecc6819e3b2 · inbound

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models cites this paper.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.553797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.553797Z digest=sha256:fc77e6d1ff349bd0a6edab11bcd9c27d4a0c88af96526b05facb5634ee1340ed

Observation c11fcb7d-a043-4238-9988-5fe90ad14551 · inbound

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap cites this paper.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.984510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.984510Z digest=sha256:aed0af496ace23f1232ebc6e870afa6e5c5cdf21f52105d58458f22bac1a731f

Observation 538615cf-305a-4137-9b6d-9e3a569fbaa3 · inbound

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing cites this paper.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.433181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.433181Z digest=sha256:18c3633357319c18a07aa6882807e53f005b8439d9895a30fb0284a2ad6906a3

Observation ea38c2f4-6165-4719-8c34-37e37b7bb1a4 · inbound

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models cites this paper.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.939167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.939167Z digest=sha256:58dcd152ca36d6f18a8fa247c4c736c98556c1ba1c7d33b5f084d4b4d03523ce

Observation 37d341dd-0cd8-44ac-9657-7fd031ebfc28 · inbound

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation cites this paper.

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:33:19.172520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T13:31:16.012419Z digest=sha256:9cb257db53ac854071a29bc4ef3548a1eb402185750a997c63a7979b2cc26fde

Observation 06f3f885-7a39-4372-90a5-42c8deca1fbc · inbound

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation cites this paper.

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:45:00.684556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T19:42:41.238072Z digest=sha256:0ca5ef04a90e6dd0e6e700c32c6a1bfc01d688e339ef39391fd97529ed3c8900

Observation eed17356-21fa-4a12-90e4-9976229c8bf9 · inbound

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion cites this paper.

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.119706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T08:20:05.081980Z digest=sha256:952c7be371aa3cff3c3098a3aa82cb2c2b4087a5a5dea096db841310a0d5d165