Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T03:18:51.582340Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 26 inbound Pith citation observations for arXiv:2311.04257.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T03:18:51.582340Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:55:55.637449Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T06:39:37.569021Z
77 of 77 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dabf9d8a-4d5a-42e2-b1ba-c0a29fdd05f0 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration http://sharegpt.com
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7ffc64e1-5a3f-42c7-a32c-3669cca1f861 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4ac22e38-c165-4d77-b5f3-35bbba7c2251 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 51e9787a-644c-486e-9945-12d980ea1d67 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Layer Normalization
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd4efb21-6969-4288-8197-2b6727b2a138 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 46cc1ed0-9c66-4c19-8723-c62ad851a954 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Language Models are Few-Shot Learners
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1cc48520-f9e2-4231-9bfb-ad9d3a3d0c03 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Coyo- 700m: Image-text pair dataset
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b788be39-6370-448f-8d45-518299d7e79e · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration End-to- end object detection with transformers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f7a8dbf-f0ed-495c-9f10-12fa286256be · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aa5f582f-dcfa-46cf-ada1-b57716b280ed · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 045c4426-7361-4bd1-bdc4-7d51ba13ce1b · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration PaLI: A Jointly-Scaled Multilingual Language-Image Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2a5f5f0e-1200-4f1e-95e3-ccd6186fafa1 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aa1e4a5a-7b0b-4bc6-a4f7-d685307adbf0 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bf1ec431-8207-4d45-97b4-7adc78b71842 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Opencompass: A univer- sal evaluation platform for foundation models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c4ec7c7f-4f2c-4ac1-848d-4188eba549c3 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation adf7d802-bdb9-4848-88d6-6120231ba215 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Xia, Mehdi S
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c95f173e-1bca-4237-879c-5b610b3c123f · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50537838-c88e-44dd-add6-32810a4947d5 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration DataComp: In search of the next generation of multimodal datasets
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 531e3cd3-b9be-42ca-b882-bac1a200d757 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f444185-6efc-40b1-b169-19812663166e · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8bf39e90-255e-427a-ba7a-98a6cb1d969f · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 376a0f8a-d07d-4bf2-9d9a-e6680b461bd9 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Measuring Massive Multitask Language Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 54258956-4631-49d7-83ad-fb6969eace15 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Language Is Not All You Need: Aligning Perception with Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc357717-0205-49cc-80cc-eb69a2f89d42 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7138ca41-f18e-4299-8197-30f5e168234d · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1ef21c4c-19b1-427a-80c1-9421962abb50 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Visual genome: Connecting language and vision using crowdsourced dense imageannotations
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc7a416a-d0d6-4b29-8305-a111b9e8910b · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Masked Vision and Language Modeling for Multi-modal Representation Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c7c84476-4eee-446d-8602-cc2768e882de · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 38daaa6e-8ee2-4656-b93c-8e1be25435fb · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation de6f8bcd-3f79-4abd-a14c-0f2f7af7bd6d · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation edddc23b-c9ff-4a87-b594-1a214319f62f · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 48b1f484-3639-4ead-a4b8-01f3bb01bddc · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration VideoChat: Chat-Centric Video Understanding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45ca166c-de2a-4486-a74b-f5882a991822 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Evaluating Object Hallucination in Large Vision-Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8dbd5756-be1f-45f7-b9cc-d292468be12c · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b1de4368-e754-4b37-8ba0-98fde650054e · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Microsoft coco: Common objects in context
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b27c7cbc-14db-4736-95db-3e4986caa806 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e24a6a25-5838-4fab-aeb3-8cd7f8a95a70 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Improved Baselines with Visual Instruction Tuning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4a7681a3-6995-44ea-ab55-5b2fb58eecfc · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Visual Instruction Tuning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation beda4e00-d371-4df3-addf-64e0a44886ba · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MMBench: Is Your Multi-modal Model an All-around Player?
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c3a4e4df-863c-4a38-acf7-45d60b07b916 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Fixingweightdecayregu- larization in adam
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae304f91-df2e-44eb-85c5-d36eacc2c498 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc3ae5ba-d5ed-47e0-a352-6eebeda33c17 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 71cd6aa0-cb97-401e-89a3-5d24420e1974 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ebf9cce4-8f62-46fa-9ecc-4efda528a7f4 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Ocr-vqa: Visual question answering by reading text in images
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 77041ce3-75a5-477c-ba05-a8b120a3c703 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration 5, 13, 17
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 672c078b-d93c-44f2-a26a-0d824906bbdb · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Gpt-4v(ision) system card
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65d53aca-4246-4221-9440-c247fc70004f · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration GPT-4 Technical Report
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 363859b2-683e-4927-b3d5-1f50505575ea · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7df40d85-0076-4fc6-af2a-847eba804395 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Learning transferable visual models from natural language supervi- sion
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 624e37af-4de2-4e27-bc0a-ea21285d6f6c · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Laion-5b: Anopenlarge-scaledatasetfortraining next generation image-text models.Advances in Neural In- formation Processing Systems, 35:25278–25294
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf33827b-ab92-4308-aa38-b74c81478bfb · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration A-okvqa: Abench- mark for visual question answering using world knowledge
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 175c5a8c-d7bf-42f8-a43a-cb8bf1cae27f · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration 5, 13, 17
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8e4012cb-1099-48ea-83fb-5d86b182abd7 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration GLU Variants Improve Transformer
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b0cd216-b39c-457e-90d1-776d007642af · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 263cd0dd-6105-4614-a500-a3b7ae773355 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Textcaps: a dataset for image caption- ingwithreadingcomprehension
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3570651b-bb5e-4146-b65d-de388019a2cd · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8c92b47f-a279-4e35-9317-1968abfce620 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 601cfc42-d13f-4588-ad3d-701803007823 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Hashimoto
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2f105fe-74aa-4fa0-bd08-295baa97feb6 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration LLaMA: Open and Efficient Foundation Language Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f0a0de5-3a8a-403a-9a69-65e6a3f1558e · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81fd797d-0d47-40e5-97a1-46b2a468fb81 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration GIT: A Generative Image-to-text Transformer for Vision and Language
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f42ccc44-ad91-4fa8-a89e-cc5147a91bb8 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5fe46ad5-5219-465a-b98a-f6ff0b8df7b4 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 36b968f5-7243-47ad-9beb-eb6f63ceb33e · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Video question answer- ing via gradually refined attention over appearance and mo- tion
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5663e2a-2c52-41d4-97fa-d8c0c55f2202 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration In InternationalConferenceon Machine Learning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a78790e-29d0-4296-8fb6-9ab2139f53c0 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Zero-shot video question answering via frozen bidirectional language models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4b8cd34-6934-476d-8aab-f6f555e1387c · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c2a20675-9658-414a-915f-2d96836e3a41 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Ureader: Universal ocr-free visually-situated language understandingwithmultimodallargelanguagemodel
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c1184676-be13-44ee-b55d-c037f1989874 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Hitea: Hierarchical temporal- aware video-language pre-training
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15a88951-f765-42ba-bef5-0210290d1b26 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5705a2da-3131-4c87-9d8f-d21d7b23303d · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Modeling context in referring expres- sions
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 33796428-5124-40ac-bd43-d23f3ef608aa · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc4a8f39-1acd-4589-9407-0614558899a3 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8a1aeacf-975c-4c5d-9467-ad30e2b5fedc · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration SVIT: Scaling up Visual Instruction Tuning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7a39f36c-0833-4151-9d18-3f9fad54b0ec · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration P Xing, Hao Zhang, Joseph E
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9644ed94-4617-459d-9b84-777600b65e97 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation caccece5-ae84-4be2-ba8b-24fb0dc06ac3 · outbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c720735b-edf2-4085-87db-a7b5e3bcc50f · inbound
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 86566b5b-222c-4cfb-8475-21b98c56d9f4 · inbound
MMBench: Is Your Multi-modal Model an All-around Player? mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 176fbb6b-ee30-4a1c-80a2-e6bf6a6721d7 · inbound
CogVLM: Visual Expert for Pretrained Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 77838a10-3067-4695-ab75-e857e2e36946 · inbound
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 74aa7ba5-e508-4027-8609-6a373889c352 · inbound
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15bd591b-ad1d-46be-b116-789496e35c97 · inbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5d51d752-3267-4779-abed-cab4a8de5409 · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 724b22dc-be35-4f0b-929c-96a9e7d3ea53 · inbound
Hallucination of Multimodal Large Language Models: A Survey mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 189
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4232900f-91e2-49fc-9527-e3c9a2d30a05 · inbound
Detecting and Evaluating Medical Hallucinations in Large Vision Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d364d906-c976-4e03-99ba-c89be89abf89 · inbound
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0297dbd4-e743-4d32-8608-ca401d71eec5 · inbound
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 03b48de3-3b9a-4b9f-a672-fbc1e5b13f38 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 276
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b0f5dcfb-043e-4bf3-9dd1-6181e967d880 · inbound
Qwen2.5-VL Technical Report mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1265ab54-13b9-47eb-b30b-ae6965dc4041 · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 137
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a28beb4c-f76f-4ec7-a2c7-82d2cb6444bf · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 166
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98d6f1f5-70c9-4e02-80d5-f8a463baf381 · inbound
Parameter-Efficient Adaptation of mPLUG-Owl2 via Pixel-Level Visual Prompts for NR-IQA mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 249f0c18-328e-4706-a150-e7c42a0b3163 · inbound
MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 71c0aba1-8f21-4af8-b5e2-ca6021598b25 · inbound
From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b7aec3d1-05ed-4b01-934e-e0a5e2ed6045 · inbound
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 28979578-b0f6-4c53-bb45-ade7405486a7 · inbound
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3349ccc7-1b8b-4a25-9b6a-e76b8a020c3d · inbound
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 08253312-f120-4863-b5ed-e4ae3fdcce76 · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 234
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c8fe25bd-83ba-40eb-b1ec-c57bb2d6184c · inbound
Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4e66300e-43c7-4670-bcf5-881d882097a8 · inbound
Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1aaf88b7-bedb-42e7-8b3a-51431da31c21 · inbound
Qwen-Audio-VAE Technical Report mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 181
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abbab7a4-eef6-4abe-8d74-a2efe25f6549 · inbound
Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.