Pith. sign in

Paper Citation Record · LEDGER

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.04788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04788 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:36:56.237018Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 426e823d-bf67-4157-8c7e-5a3c82d199f3 · outbound

This paper cites GPT-4 Technical Report.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.180523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.180523Z digest=sha256:4d136481e57a83d5ae2aa783d2fde07a9bfe340db6932fe4f1e2b0e350b42cf3

Observation eded0f9a-9741-4aa3-b944-1b04c6b765ce · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.191276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.191276Z digest=sha256:2f0745b8aa9da9471105bf930de25c9db60fda769d2488636074047f2b18dd7c

Observation 500868d0-b501-4c3c-965b-4714d8c7625b · outbound

This paper cites InIEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 15120–15130.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques InIEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 15120–15130

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:36:56.754264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:36:56.196880Z digest=sha256:bcf074d134f981fffae679c6131ec2e41b1c8dd7852961963636907b4106780a

Observation 935f29c8-819d-44e0-9ed8-0b5acf5b550a · outbound

This paper cites Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.201686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.201686Z digest=sha256:9d810142f0b4d1fd1fcd66136baf42fcebd7d0642c6c4d249dc06f6460657d04

Observation 7d6180a7-6901-41b0-a5dc-fbc53b65ccf8 · outbound

This paper cites Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.212001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.212001Z digest=sha256:e8a176efbdc3c452e48b0043b3b6d1eb566e558fe018e08441ea64f866bb5182

Observation 0f0e7a96-b62c-437c-9140-88118de914b4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques Gemini: A Family of Highly Capable Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.216528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.216528Z digest=sha256:e10dfa411213229b07816bcd598201284cae66c30b57221ff85442264f69a180

Observation 3514720f-5be1-4c11-a1c7-5f55dd0e539f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques LLaMA: Open and Efficient Foundation Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.221847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.221847Z digest=sha256:3b487395d730e3daf5a738dbb6d9cd6584a6c4b17bf192de676a4ed36de3edd3

Observation 21ca545f-01e0-475d-80f4-ac9c32598554 · outbound

This paper cites Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, and Jifeng Dai.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, and Jifeng Dai

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.226890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.226890Z digest=sha256:5de10edd508145b74e7657771924ec6b95ed591f9ed0f2597ce6a91b6975af9e

Observation 67d5af40-f86a-4030-a61a-90de85db71b9 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.232186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.232186Z digest=sha256:f2063735e3d995b75e82a5a35083c6e0148582c881f73cd787e6390ac4802328

Observation 8d7b984d-99c0-441e-ba9a-dbb4663cdf44 · outbound

This paper cites an unresolved cited work.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:36:56.738577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:36:56.237018Z digest=sha256:1aca321aeb3b68f9065c96fc3c9d89e46b7ceb77c7eb2271f40b719a3e6f78e1

Observation 48eff9f7-c9bc-4a51-b885-8cc48c7e8ec1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.173390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.173390Z digest=sha256:7fcd728fc88be06939209f3d9fbaced1285413a6543e3eb30760f5871ebcfed3

Observation 3e11c5f6-141d-4a64-b401-f753d12fb7db · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.185722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.185722Z digest=sha256:94cd9ab6e14dc7ce3fe83ba90710e7bc5c1aa4beda54d323f89d87e308d16b59

Observation d8a8afa8-b18b-4ba7-93ce-417e3287ed56 · outbound

This paper cites How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model.

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:56.207237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:56.207237Z digest=sha256:d02e89c6cd79920b15821e641f25ffda0656408d3b9931f03948c759220a6ab2

Pith citing papers

No inbound Pith citation observations are available.