Pith. sign in

Paper Citation Record · LEDGER

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities

As of 19 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 2 inbound Pith citation observations for arXiv:2412.11694.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11694 v3

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:43:41.879700Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:53.953513Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T06:00:27.821608Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b095287b-cdc7-4af5-8189-c1dbf3fd4c9a · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.772627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.772627Z digest=sha256:0ffd2adb44d65df4b62a746449c3ff9278aa551f6e96e513239726f1af48212f

Observation 6015470d-8b7d-42e3-b7b4-cca96d21a289 · outbound

This paper cites Textbooks Are All You Need.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Textbooks Are All You Need

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.791396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.791396Z digest=sha256:e514d00e6b8eeee4ebca8a9951d62b164a177437065bb6e7ef9e77315f79db6f

Observation 6cfa550e-ead2-42fa-a8d2-ac8358eb7bfd · outbound

This paper cites LLMs Meet Multimodal Generation and Editing: A Survey.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities LLMs Meet Multimodal Generation and Editing: A Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.797243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.797243Z digest=sha256:f114f8f2f8704e6371824505239c8edac1e2b815fd8ec87648bbaa4ee4db3dcc

Observation 0bb9eb8e-87e8-4f5a-8150-19ea002bba58 · outbound

This paper cites Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman K.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman K

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:43:42.195622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:43:41.800929Z digest=sha256:9c975904aeee0aa40ba3372253e36ea2a2f5a7a4097c157b687e3f2d78b99c75

Observation d5aa68bd-f30b-4e65-858d-7977993ceef8 · outbound

This paper cites Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.818671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.818671Z digest=sha256:04d9663ef6e6ee923cf25d35bc51944d0709ad444dbaf5b3068569db3b7e591b

Observation 8172c3f8-a721-4b5e-8508-f550d531a49e · outbound

This paper cites Zero-Shot Text-to-Image Generation.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Zero-Shot Text-to-Image Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.822155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.822155Z digest=sha256:f9fb71ba3fa22ef03025879cc0953b385f215b04e17c135a29a570b3428074af

Observation 3ad8012f-d2b8-4168-b533-1dbf434585cb · outbound

This paper cites In 2021 IEEE/CVF International Conference on Com- puter Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021, pages 12159–12168.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities In 2021 IEEE/CVF International Conference on Com- puter Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021, pages 12159–12168

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:43:42.177216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:43:41.825123Z digest=sha256:803c1133887de2ca779ad352bf960959493cfffddfbf9864b6a25bff8c59945c

Observation 33b339dd-fd64-463a-9ed4-2f4dcbc55f24 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.827742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.827742Z digest=sha256:24d1bd426afb52ce0ac32cc105946043f44c5bcddffbe44727257e20602ea082

Observation 936a8bfd-94c7-4b20-bbb5-a12fe8f43a86 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.830371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.830371Z digest=sha256:7a462760f6d5cc10aab24465118653779974ebe41350f648ebbcdccf30abce38

Observation 63acddee-8ffa-40b1-b0b9-95846f211d63 · outbound

This paper cites Improving Image Captioning with Better Use of Captions.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Improving Image Captioning with Better Use of Captions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.833630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.833630Z digest=sha256:c9eb1bacbba9923a61709155affa3535ad506e83c93491cf44d7a794a55c069e

Observation 45d4ccea-e6ad-4a11-beed-c4faca960180 · outbound

This paper cites Audio-Visual LLM for Video Understanding.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Audio-Visual LLM for Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.836545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.836545Z digest=sha256:82e3a1d277f822e42743ae768d7e1d77690bfc033066e087ca1c20947d16c204

Observation fc8d910e-f776-45d3-b739-4fd24cac4519 · outbound

This paper cites SNAC: Multi-Scale Neural Audio Codec.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities SNAC: Multi-Scale Neural Audio Codec

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.839513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.839513Z digest=sha256:339c0d404bf9dfcded9d3816063b18353c7847920e05ff2b972eb6397f57a274

Observation 354bc140-7518-4a8d-9e2f-28348bb231a1 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.848739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.848739Z digest=sha256:0a381d9d191dd97fbe5987d98408191a2bd22c81db7aa1e4d733aebb5a0a508f

Observation e8354e47-abb4-42aa-9939-458d6491b53a · outbound

This paper cites In IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 16354–16366.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities In IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 16354–16366

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:43:42.157717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:43:41.859995Z digest=sha256:8001fdd0a7ec59de3fdd8588c324ec75037b92abc0c00e4dd8ff45b913680076

Observation c85f9204-f626-49d7-bbc9-ce286fc6e683 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.863299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.863299Z digest=sha256:73f981321e5b234b50c16ac99af828f142e7e3f2bf4bb138edf90ac9f08927c1

Observation 5d2f0c67-bdcd-4a91-a1e2-b65f44880e02 · outbound

This paper cites In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 9637–.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 9637–

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:43:42.145503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:43:41.866664Z digest=sha256:09422a1318ce9e3cd739c9220c4ad2fcbc6c0102c45316e893dc3d624b44ab19

Observation de52a44d-27fa-441d-a937-1ec31f5bc019 · outbound

This paper cites an unresolved cited work.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:43:42.134703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:43:41.873468Z digest=sha256:46dffd5f6fbc6bf9fd025144985210457d76d5b2d5053fcff318fccb293e25ec

Observation feadd281-1466-4ae1-96e5-880b312e6605 · outbound

This paper cites an unresolved cited work.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:43:42.124556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:43:41.876686Z digest=sha256:6f7a3c91c1fe1a5aa9a69596bf6e64052c49d3cf17fedd2eab78f09b5921e445

Observation e71699d6-ca2e-4f4a-830b-6e0c309a4701 · outbound

This paper cites Video generation models include Zeroscope (Cerspense, 2023), VideoFusion (Luo et al., 2023c), VideoCrafter (Chen et al., 2023b), and ModelScope (Wang et al., 2023a).

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Video generation models include Zeroscope (Cerspense, 2023), VideoFusion (Luo et al., 2023c), VideoCrafter (Chen et al., 2023b), and ModelScope (Wang et al., 2023a)

Reference 33

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T14:43:42.114567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:43:41.879700Z digest=sha256:f082475e6f35b7e114c6c4d2d803396e084484fd5b7022f93656f54451297168

Observation feaf6af2-0313-457d-b292-c6f11801bf44 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 576

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.842628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.842628Z digest=sha256:c32aa20710c49f5354d564348ddf9e515d0d78145c6b90ee3224668e75f627b2

Observation 795a0c97-bf2a-43b1-9a13-ebaa2a2f75fb · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 1464

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.852286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.852286Z digest=sha256:c491c81c49a2eb3b8805284ab7fa088b01d2a505e50510527542be1c2e778ccb

Observation f3237d09-3b9d-46ec-9dc0-50da29f93d73 · outbound

This paper cites In Advances in Neural In- formation Processing Systems 24: 25th Annual Con- ference on Neural Information Processing Systems.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities In Advances in Neural In- formation Processing Systems 24: 25th Annual Con- ference on Neural Information Processing Systems

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:43:42.186658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:43:41.815337Z digest=sha256:3aa16fbf826fcc2ceac6a3cc228a385b174f0a088433962235a997596e192d92

Observation 223f260c-c00c-4627-9ef4-1277e5a1070e · outbound

This paper cites IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.811825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.811825Z digest=sha256:fc1e913f41276ad3f8009b4f6c7184677aeb38f31dc048075ea7d79eccaccb66

Observation 8be402c9-2cd2-47a7-843b-f576b414bafe · outbound

This paper cites Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.803976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.803976Z digest=sha256:82b3bb56e7197df97eb148a6a2d2815c50bf4e4a84f2fcf26853524457f9380d

Observation 9daa699d-1640-4a67-8399-4500d6a02ed0 · outbound

This paper cites A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.856076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.856076Z digest=sha256:75d9e84559bcae534137bb30ef6efbcec9700c8ec40a6db73b988d86d66bbefc

Observation f9b8c00e-b17d-4aef-9056-860bc1537b8d · outbound

This paper cites Jukebox: A Generative Model for Music.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Jukebox: A Generative Model for Music

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.780304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.780304Z digest=sha256:a2ab9b7555e1bb3d031081d2e463bcd7fa317f18e14e4b240c3c35d9d6c509f7

Observation e14dedcb-cd1d-49fe-9acf-919d218b88b4 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.768579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.768579Z digest=sha256:b520912313de3d323cc3aa4ad500baf22aee7e62fd49e0b40e97bb265b32f1fd

Observation e724bc1e-fd5e-4edf-b490-cb9ac53cd8b7 · outbound

This paper cites an unresolved cited work.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:43:42.204784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:43:41.787668Z digest=sha256:00a2333e6d515fb6328a099c90a6a740eae426847c0facb3ce3d9b9e0a40586a

Observation 74a0feec-5154-4dfc-a652-0bf5f0cc89f2 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.776532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.776532Z digest=sha256:738657f5a5faf5cefd9e049811297fc75e90debe2db817dcabc79f71a7239c7e

Observation b897068a-707c-4498-8b53-e96c1b68727e · outbound

This paper cites LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.783734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.783734Z digest=sha256:ce131d8ba235a9b612f20fcdcd81b4451f0dfa342e885e878b3f5a0bdf9e3f94

Observation 3b2a0404-5fe8-453c-8c4e-3c059110fc3b · outbound

This paper cites Licai Sun, Zheng Lian, Bin Liu, and Jianhua Tao.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities Licai Sun, Zheng Lian, Bin Liu, and Jianhua Tao

Reference 6121

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:43:42.167871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:43:41.845738Z digest=sha256:4b2bff78757d9531587b9b70a376897fc9d76f278366257ec32af516cb5ac885

Observation 3e45a59d-d3cf-4ddd-8bd8-dfee93308a3d · outbound

This paper cites The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 8298

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.807896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.807896Z digest=sha256:0db7a2a0c605b1bb3626ebb61d501f318d1044815697e6229d82d38cc37c0d72

Observation ce5c3553-7261-4844-a9e9-5a631ee13d56 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities OPT: Open Pre-trained Transformer Language Models

Reference 9662

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:41.869741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:41.869741Z digest=sha256:70d5016568f331128799ea9677fe1d15a5e45ac55f0edd5f877046c70db2def0

Pith citing papers

Observation df313483-fe6a-4964-a8f7-6a58ae9e0d70 · inbound

Unified Multimodal Understanding via Byte-Pair Visual Encoding cites this paper.

Unified Multimodal Understanding via Byte-Pair Visual Encoding From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:53.953513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:53.953513Z digest=sha256:84548f55bb92ccdce6ad0740719f60fb6426725a44e2e27e74e3287ccb71e74d

Observation 6a7c9cf0-32e4-4045-b016-693b24c42168 · inbound

Sample-efficient Integration of New Modalities into Large Language Models cites this paper.

Sample-efficient Integration of New Modalities into Large Language Models From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-05T06:00:27.825437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T06:00:27.446056Z digest=sha256:4584351618d3412a07dbab0da8ffdc0ff45ab7ba46c40938faeb06dbe344302a