Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-24T09:02:31.211260Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 100 inbound Pith citation observations for arXiv:2304.14178.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-24T09:02:31.211260Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:24:22.627249Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T06:15:00.866473Z
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c966cff0-1189-4175-b608-21e0d3cb2f22 · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Flamingo: a Visual Language Model for Few-Shot Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ec80962-ecf4-4b0d-ab78-4a128e7c5401 · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ce41ffec-77dc-4295-9b45-32ed5be9ae1b · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality PaLM: Scaling Language Modeling with Pathways
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 31c7845c-ccb1-462c-a422-7c12891b1e07 · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Scaling Instruction-Finetuned Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9e89c9d5-617c-40b2-8dd6-4055bd7caac9 · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality PaLM-E: An Embodied Multimodal Language Model
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1a325725-06b9-4be4-b2ca-1766b6c1bc84 · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2e18a3f9-186c-4bc8-b2f8-cd0cf5a4288c · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Visual Instruction Tuning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 350f9dfe-3a83-4c0d-a890-cb095211731b · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality GPT-4 Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 64cc3e3f-973d-405a-a234-24816a57a3b4 · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Training language models to follow instructions with human feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8db6f3b4-7956-4623-a776-c4e649a9554d · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c25a1df7-1ab1-4632-80d6-320c4f4b13db · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bdf32398-5d4c-470f-a001-a008a26c0be8 · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f7883de6-9577-4d88-a56c-ef3b9b8c50b3 · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality LLaMA: Open and Efficient Foundation Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 477faf47-d840-4a0b-a645-1d5be12e8edf · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3ddb50c2-298c-4961-b890-67743b597329 · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e2490bfd-d056-4063-8440-6ec0db0165a3 · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 71a320d5-2ef4-4aa6-a654-0a2a5620c7ba · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 10d8118a-1f5e-44b4-ae73-3d4294e0bd8c · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 28c91947-cc6a-49f5-8f2a-b9382edf2624 · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e1a8c733-b4fb-417e-8c6c-509b6f930665 · outbound
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 76944607-a24b-4725-a4f2-3624f6bcd79a · inbound
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c3394408-acdc-493f-a1e5-9115a84c87b1 · inbound
Evaluating Object Hallucination in Large Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b1aa540c-3827-4de9-bdb3-5caf9b589431 · inbound
A Survey on Multimodal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f8962a29-202a-4945-9ccc-b4868541b0ca · inbound
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 153230d6-ba90-4b15-9dc5-2fa0aa48f79c · inbound
MMBench: Is Your Multi-modal Model an All-around Player? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6e22e92b-a58a-4578-a0df-435bbc5a0475 · inbound
A Comprehensive Overview of Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 78427e1f-bc1a-40d0-8b56-b633d9b8f0dd · inbound
SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5c0198fe-1053-450a-b3e8-9b545abb631d · inbound
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 45bac780-abc3-4ac0-b7a8-807648c83609 · inbound
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff276869-2731-4eec-967e-10de06a5b520 · inbound
Aligning Large Multimodal Models with Factually Augmented RLHF mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4a56a9e4-1bf6-4ad4-999a-014ee602264e · inbound
The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 146
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ce33ff3c-5afd-4b5f-a530-f9ae1e4c4163 · inbound
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dbef0d99-48d9-45d8-8792-bd82c02c9789 · inbound
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cd581fc1-a0c9-4997-9dd8-5fb8592363cf · inbound
Improved Baselines with Visual Instruction Tuning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 870f95de-9c13-47aa-ae07-365bd2148e3c · inbound
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4c4fd9db-3542-404d-b6c1-39921caeb41f · inbound
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 15a88951-f765-42ba-bef5-0210290d1b26 · inbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 88b2bf15-26bd-46fa-9b0f-4474af0248a8 · inbound
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4358022d-95c1-45e0-ab53-c662ce04dc35 · inbound
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f8fd3119-352a-4101-9d27-eaa6f1784838 · inbound
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1827c624-c1f7-4e8f-a474-d7470f1bfd18 · inbound
An Embodied Generalist Agent in 3D World mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6370e879-f16f-40d3-9cd9-0a591343968a · inbound
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 54661be6-1b54-4002-b669-da2f5e91851d · inbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 612cebfb-9871-416e-b028-61034679cc2d · inbound
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 25c59aa5-2992-4d29-85d6-5b5b6a0230c6 · inbound
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3c38bf03-f69f-42d8-87d8-fd248438338d · inbound
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 29f21e54-a2f2-4286-9570-f57881e28d40 · inbound
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e779d337-483a-43dd-9583-b1117bc6f29d · inbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f4ede3bb-da84-41ef-a5aa-b6de47c2bc2b · inbound
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 179
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c66988cc-63eb-421a-aa47-81bf253b1b29 · inbound
TempCompass: Do Video LLMs Really Understand Videos? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ffea0252-1c39-4f0f-961d-a817d10e3b08 · inbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4acff808-ec0f-4638-b458-70db110dbda5 · inbound
Generative Models and Connected and Automated Vehicles: A Survey in Exploring the Intersection of Transportation and AI mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fe117760-de67-47c6-86df-6ba21ed3c357 · inbound
Are We on the Right Way for Evaluating Large Vision-Language Models? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0db715e4-c6ce-4d06-935a-ecb9cfbe959a · inbound
Hallucination of Multimodal Large Language Models: A Survey mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 188
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 43f69a16-3446-4395-b26d-a6b62830dca0 · inbound
VideoPhy: Evaluating Physical Commonsense for Video Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2fca4937-5447-4c41-af00-8ddddb1a404f · inbound
MLVU: Benchmarking Multi-task Long Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5c109a0-9114-462a-bf25-31dd5942c855 · inbound
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b7089374-dc6c-4598-9382-6d94c5bedff9 · inbound
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 14aa7ede-5221-4739-83d3-02b5cd04641a · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 261
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a9bdbae0-afc3-4e40-ad65-3f37fe8d9869 · inbound
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e64a57d2-caf1-41e1-b96a-46707c56b9e6 · inbound
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dca7e261-89ab-422a-802c-6b3872e4f7e1 · inbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3bb13275-a751-4d4e-a841-41ade18e559f · inbound
Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 584107d8-04e9-44bc-8705-32a18bcb5e03 · inbound
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 59dd7a6c-167a-454e-84d6-48c4e8adf649 · inbound
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5932dfb7-3179-4cc2-a915-3303f017e8d2 · inbound
When Large Vision-Language Models Meet Person Re-Identification mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fa1e2b3a-1aca-45ce-bfec-f549f0c2e0ff · inbound
Open-Sora Plan: Open-Source Large Video Generation Model mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e3eb39e2-a421-42ba-801a-86a9868077ed · inbound
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0558c787-2bf6-4a06-9024-203621515970 · inbound
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9e60d85c-99e4-4327-829d-5946a85e52b5 · inbound
LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a2f37d5-f00c-4f4e-8bec-9a88201beeee · inbound
Q-Agent: Quality-Driven Chain-of-Thought Image Restoration Agent through Robust Multimodal Large Language Model mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d9029727-9535-4217-990a-46e0d006d43e · inbound
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee130942-ea7a-4237-81b9-e3f9d0b59eb5 · inbound
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bc6634c1-5380-4f0c-bbe8-421a3de6dab4 · inbound
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 575aeedd-6a0d-48fb-a16f-3af8b3d3eacb · inbound
ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 656eb6dd-2e8a-4b3e-a877-200dba689f01 · inbound
ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6f18ac0c-2eff-4346-8fd9-94fb847f3974 · inbound
Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0981b908-739f-4a9f-aa52-001c61bf3bca · inbound
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e59d0de6-02af-4d0f-98a8-6243b17f1a14 · inbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cdc6667-9167-471e-ba50-2e699eb78d92 · inbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 92116bc7-0121-4aed-a4df-7aac4060b207 · inbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8217daf5-0d1c-46b2-ac43-aa2968335a53 · inbound
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60f37dde-6ba9-43ff-88da-749d9fe23481 · inbound
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6521beb8-f74b-476a-a313-c4b183aaa3e8 · inbound
Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67790b3b-d1f7-4a77-842c-a0b3b64531b6 · inbound
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ab547cc-cda1-41e6-b117-eba8546c1d2b · inbound
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa6d3f82-8639-42fa-8dd9-fcd97d668a2b · inbound
DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4561c020-80f0-475e-b953-e97ae17bd251 · inbound
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 603e4b0d-d928-43ff-96ac-d27b9c253815 · inbound
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a23dda5a-a2fb-4b65-a15d-ee527e89e665 · inbound
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7a846a2b-d46c-4c2f-b105-de90b785b8a6 · inbound
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d44ef7c9-740a-4b14-84b9-bbba089d64d9 · inbound
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 37c41524-cf77-4589-b7b7-e8f28b166e66 · inbound
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a0a9d8e0-8560-441b-9d16-5054bd9e2c72 · inbound
Latent Denoising Improves Visual Alignment in Large Multimodal Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ec175278-467b-4f7c-94d3-07433fceaa42 · inbound
ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 72aaa357-9d65-4b0f-b511-d34c966df4e4 · inbound
ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a82033d-3cda-4f74-b2da-a22d1a1497ea · inbound
AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0ca23b14-1dcf-4089-ae60-7c869c778804 · inbound
ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a774b6b9-dc84-436d-85a5-29e93bdb4f59 · inbound
LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 36c3ac5b-1ba3-4d49-a10e-b561e49192a7 · inbound
Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a01fdfc8-680c-41b3-bd5d-22b4f8366ac9 · inbound
Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c5e2dafd-2067-41b0-a964-7d8a0c76c25b · inbound
AffectVerse: Emotional World Models for Multimodal Affective Computing mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9c0c9dae-d5f3-4b24-a83f-7ccb7dd9480a · inbound
Mitigating Hallucinations in Large Vision-Language Models via Causal Route Gating mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7a0f665d-f0f7-49d3-9cc5-f925a4ff2880 · inbound
Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7547257a-0f44-45c3-8c3e-867ea7e2cf4a · inbound
Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ffa5f6a3-f059-4790-a24d-e68400af9c00 · inbound
The Hidden Power of Scaling Factor in LoRA Optimization mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 19d3cfb6-64a2-42d2-b9c7-9d2bd0607996 · inbound
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e7d40f5e-62f0-4007-ace4-21dc05b5f166 · inbound
NEST: Narrative Event Structures in Time for Long Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 288
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 18c50c91-7666-4f9d-8b9e-b539045cb3d6 · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cb54ec4c-5023-406a-897a-55cc01b83f35 · inbound
An LMM for Precisely Grounding Elements in Documents mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7e892e0e-68fe-4e70-862d-bb51d0142525 · inbound
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 23fa663c-7f5f-495b-b8ff-27fcefa72f8e · inbound
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 502db605-70bb-49ac-99d3-9aada034aebc · inbound
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 143b3b6d-b76f-414e-84ef-d5e764626aae · inbound
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f55624f8-0cc2-48c3-89cd-6e9cd757d53e · inbound
See Only When Needed: Context-Aware Attention Intervention for Mitigating Hallucinations in LVLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3b280dc7-d374-4a0c-a64f-c8c249e4708d · inbound
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5a096d5c-d394-4091-8171-91a6caf37252 · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 151
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78fb47b7-a639-4a07-93fc-ec7dfdae28ce · inbound
MentalThink: Shaping Thoughts in Mental SVG World mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 206
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2d6a28b-f309-4d26-ae2c-d32521d46a3a · inbound
Qwen-Audio-VAE Technical Report mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 141a615b-57f2-4749-a930-356f8a16468b · inbound
Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.