Pith. sign in

Paper Citation Record · LEDGER

X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2311.18799.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.18799 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:05:44.701187Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:41.279523Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f02d1db-a0d8-40b6-b7b2-d537c641e7b9 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:44:53.605377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:89c4cc75c7ede650e77492e44f22f29060170b5274806aee01a88759fda19d25

Observation af7ff0c7-ffa8-4aa4-919c-5a1ac56f26e4 · inbound

Modality-Inconsistent Continual Learning of Multimodal Large Language Models cites this paper.

Modality-Inconsistent Continual Learning of Multimodal Large Language Models X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:52:40.170982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T06:50:12.315919Z digest=sha256:05539cf0532a9689f8e34fc0984317bd8a5aac554583a2e3af5636ad8e7a16a2

Observation e3f1422e-6964-464e-9a20-bf4909b6da1c · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.174028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:348594ed6a6b5449a400af6d2b9a8e4702fd9d684ccdacca1d1c59e543f70e1c

Observation cfebbb1a-da40-42bf-b345-d6b3fb941bfb · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.404419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:11e4ed98403c84d8bbb9213c427319e1ed1d8392a9bcd21c2c3164e02b5a655b

Observation 191b7aef-7ad7-4a45-8d33-809fc4b8aa07 · inbound

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes cites this paper.

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:44.701187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:44.701187Z digest=sha256:60e65d1580eb75f7baff7038f763a0604bb1beea3aea3997b16e4e0a3f7e63cc

Observation bee4125a-b824-4a19-a701-e96c733a7aa2 · inbound

Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization cites this paper.

Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:35.610779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:35.610779Z digest=sha256:8a9258053e961f848169dcfe2844226761317fae5f847e2abaf6a837bea5e839

Observation 1fdfb7a6-5a81-47c0-a436-b1e74a46124e · inbound

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs cites this paper.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.328981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:bdbb6b649431ca02909e204e1b4542b0ceea613922da6fa0adc032a7478b72dc

Observation acb3a1a4-d757-4d89-a3fd-b84f2b0363c0 · inbound

Closed-Form Spectral Regularization for Multi-Task Model Merging cites this paper.

Closed-Form Spectral Regularization for Multi-Task Model Merging X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.382888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:40:00.510742Z digest=sha256:98b76b1e459f70a4873982d4280b9d104118c34cf9c8ed01d859949e67c0a0f5

Observation 04d39992-225f-4a31-be9b-a27bbdc21052 · inbound

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales cites this paper.

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:41.281286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:09:01.026544Z digest=sha256:c8cc582b4ff5d45fe9d632e08c6bc7f8eaead4539c98635bf5fbc5cbd6931154