Pith. sign in

Paper Citation Record · LEDGER

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

As of 4 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 100 inbound Pith citation observations for arXiv:2304.14178.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.14178 v3

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-24T09:02:31.211260Z

measured 120 of 120 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 100 of 104 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:24:22.627249Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact19
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c966cff0-1189-4175-b608-21e0d3cb2f22 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.076622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:9b743bf7a219596a63faa1e67650c2d1504531d952f7fa134eae7b65c22dddc0

Observation 8ec80962-ecf4-4b0d-ab78-4a128e7c5401 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.032403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:17d2fef1a5544ddac4c61bff4fbc3a77f47235bd2b65e86cafb682068c2dd677

Observation ce41ffec-77dc-4295-9b45-32ed5be9ae1b · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality PaLM: Scaling Language Modeling with Pathways

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.080433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:b7a61db5f076e1525160e6314e81322946239e9c3677679114ef2445ff36dc2d

Observation 31c7845c-ccb1-462c-a422-7c12891b1e07 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Scaling Instruction-Finetuned Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.083930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:18b81d2c8c9e53f9dc6b52e01432a99385adec57731afc8a278047e3cc85284c

Observation 9e89c9d5-617c-40b2-8dd6-4055bd7caac9 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality PaLM-E: An Embodied Multimodal Language Model

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.068093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:d25e64254deed4e90304daf670053e4040500c5420a8db6face21c7758a3b285

Observation 1a325725-06b9-4be4-b2ca-1766b6c1bc84 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.087475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:00cc60c2cb219d9770b765d841b755260f464d4f79442c0bd913a494b1fbbc73

Observation 2e18a3f9-186c-4bc8-b2f8-cd0cf5a4288c · outbound

This paper cites Visual Instruction Tuning.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Visual Instruction Tuning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.028188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:b7f84856f43fd68f26242f6c21da2717ff4e6333afca29cadcf58522dcf5bbdd

Observation 350f9dfe-3a83-4c0d-a890-cb095211731b · outbound

This paper cites GPT-4 Technical Report.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality GPT-4 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.024208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:5b42bf2453d853709e65e3c36f079af081306ff1ea5f1226e4facdca4ee718a5

Observation 64cc3e3f-973d-405a-a234-24816a57a3b4 · outbound

This paper cites Training language models to follow instructions with human feedback.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Training language models to follow instructions with human feedback

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.048146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:e54e0b26135414d53c3528c40e12ca109e8a08167c9cf3fa3d59a666615f8c54

Observation 8db6f3b4-7956-4623-a776-c4e649a9554d · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.040147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:ee299708ac23a05631b188aebc77b7389afbd5f754805cc5e26da6be9626d719

Observation c25a1df7-1ab1-4632-80d6-320c4f4b13db · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.036222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:fa39a75a59f44e88db04a89f679fed08e93fcce5266954a44677c567ea9a4264

Observation bdf32398-5d4c-470f-a001-a008a26c0be8 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.044052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:49459736f6a48b335ffdc8c3bf8803334b7c252dbd69080249c35ed6929e0983

Observation f7883de6-9577-4d88-a56c-ef3b9b8c50b3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality LLaMA: Open and Efficient Foundation Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.097534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:31ceb7f9f99baa5b95134f8d1648c49c7e64627eb9b8ae21ad4ff95a59239886

Observation 477faf47-d840-4a0b-a645-1d5be12e8edf · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.064182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:bd6df469f39be30a157d29d9a4daeb45bdb51f6b3f34fb790cc95efa680413f6

Observation 3ddb50c2-298c-4961-b890-67743b597329 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T09:04:14.794976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:f21b41c08a14133e898bc003b97797366b24338607e73bcab311ebdfc03122bf

Observation e2490bfd-d056-4063-8440-6ec0db0165a3 · outbound

This paper cites Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:04:15.052649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:a36ac8617b3f6dfce2723f315024e3e1a9b3dd522a4f393739016cf7d6cfb251

Observation 71a320d5-2ef4-4aa6-a654-0a2a5620c7ba · outbound

This paper cites mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:04:15.057623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:6207b502cc1cc05c9bc9722255523e7e5c9425825dd5242489bf1a0d7bc689ff

Observation 10d8118a-1f5e-44b4-ae73-3d4294e0bd8c · outbound

This paper cites HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:04:15.071854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:7c7e477412571a2dde0c9a973cfaf358ca8109997f9f63d1a681045d88241d57

Observation 28c91947-cc6a-49f5-8f2a-b9382edf2624 · outbound

This paper cites HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:04:14.790817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:f2ccd6fe8d89599280d0083c9234c87e9906e0cce4a6dc599001147f3cc238e7

Observation e1a8c733-b4fb-417e-8c6c-509b6f930665 · outbound

This paper cites Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:04:15.092474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:7219160562fbc03f6c5dfb8937b92d68eca4cde505a76614ba308cf908d94cfa

Pith citing papers

Observation 76944607-a24b-4725-a4f2-3624f6bcd79a · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-14T23:07:42.918442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:dfafddcc9c48975129cfffa8a69cd005ca0e56c14092a6f7fc93279ed86b5e6c

Observation c3394408-acdc-493f-a1e5-9115a84c87b1 · inbound

Evaluating Object Hallucination in Large Vision-Language Models cites this paper.

Evaluating Object Hallucination in Large Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:44:09.893263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T13:44:09.626361Z digest=sha256:c7f8b02df49eedb6e75f5524a864941e2dea24beab2c58c81cbac7eba12fd25e

Observation b1aa540c-3827-4de9-bdb3-5caf9b589431 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:56:42.609524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:b5b455dc820f36051873b93b7c748700cb26da36c75b3ab0d4a5701e403bbfaf

Observation f8962a29-202a-4945-9ccc-b4868541b0ca · inbound

Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning cites this paper.

Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-14T17:34:56.988590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T17:34:56.836034Z digest=sha256:5e3fe62b849820313371e89ae09fa22583e69d7b26297b63cbc3be8db9257b0b

Observation 153230d6-ba90-4b15-9dc5-2fa0aa48f79c · inbound

MMBench: Is Your Multi-modal Model an All-around Player? cites this paper.

MMBench: Is Your Multi-modal Model an All-around Player? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:20:53.880932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T17:20:53.687692Z digest=sha256:e717cf94595fa3eaad133daee548e302591cd9d9da3ed071d92658424cd560f8

Observation 6e22e92b-a58a-4578-a0df-435bbc5a0475 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:28:39.264298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:6ec60532f5fe21e29e2aac41e7d70c550f4c43a56b1b2068c89ac1003b19c415

Observation 78427e1f-bc1a-40d0-8b56-b633d9b8f0dd · inbound

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension cites this paper.

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:59:50.618607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T16:59:50.495335Z digest=sha256:c6309c33e86d1b9ec2a5b5729f9c05b186ddd6bb495b33e0b28ac9bddea0ad23

Observation 5c0198fe-1053-450a-b3e8-9b545abb631d · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:52:01.336399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:d2d2602db9ddac62f2e19d8b19a9edafbb8d0992101a4a5310e158befed73405

Observation 45bac780-abc3-4ac0-b7a8-807648c83609 · inbound

Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models cites this paper.

Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T14:21:16.530157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T14:21:16.453610Z digest=sha256:5968bb7a1e17e9c922f40824696b33f2ea4fd39af19967d6bf4a65c9ddfb4376

Observation ff276869-2731-4eec-967e-10de06a5b520 · inbound

Aligning Large Multimodal Models with Factually Augmented RLHF cites this paper.

Aligning Large Multimodal Models with Factually Augmented RLHF mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:58:17.771669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:58:17.699042Z digest=sha256:90a9200451b3de7519b20b0e797f4728d6b4741ec153784a598f2f323f277f09

Observation 4a56a9e4-1bf6-4ad4-999a-014ee602264e · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 146

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:26:06.485947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:351f9fbf97d4ba94a646ced302da502fe5077647a6bf279a1ae014ac354686fa

Observation ce33ff3c-5afd-4b5f-a530-f9ae1e4c4163 · inbound

Analyzing and Mitigating Object Hallucination in Large Vision-Language Models cites this paper.

Analyzing and Mitigating Object Hallucination in Large Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:46:52.820083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T22:46:52.791128Z digest=sha256:979e08942d9b63d5b49eafbd1707e71a7ebef14d3bafcec6eb9c7bc35c2cfdaf

Observation dbef0d99-48d9-45d8-8792-bd82c02c9789 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 91

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T03:27:59.111880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:6630e21db839f135e30090c2cadacccf8587545e52f01f986bc6c838e73b45d0

Observation cd581fc1-a0c9-4997-9dd8-5fb8592363cf · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:11:33.968118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:c172f668c3f21aa0f50be1410de19b82ba4b1b42b529fd04222d1db65ea4fc5e

Observation 870f95de-9c13-47aa-ae07-365bd2148e3c · inbound

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning cites this paper.

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:13:09.002100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T07:13:08.867745Z digest=sha256:77ac3f0c0a69a03fa38d662ccb7b171a55a498ade20530eef3546f6836920d84

Observation 4c4fd9db-3542-404d-b6c1-39921caeb41f · inbound

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models cites this paper.

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:22:04.174548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T01:22:04.035994Z digest=sha256:9c9476383725ab171e26e1238ca06ef05c652633a48844bddd75e97c69190c2f

Observation 15a88951-f765-42ba-bef5-0210290d1b26 · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.778085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:c6f12608ec1d917875821fd8a86be61487725b2991639a7e67cc725f5f8630f5

Observation 88b2bf15-26bd-46fa-9b0f-4474af0248a8 · inbound

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models cites this paper.

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:03:26.796658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:03:26.723464Z digest=sha256:aa6f5f2d320ddac36503c4e2b0bc14f9dffc67f72e013770dedd426d47e3cb60

Observation 4358022d-95c1-45e0-ab53-c662ce04dc35 · inbound

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection cites this paper.

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 79

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:08:01.386169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T18:08:01.166072Z digest=sha256:d645cbd9083252786412f9b6ec2b2defc0d8e6547e5ed2690c222a4a2e403818

Observation f8fd3119-352a-4101-9d27-eaa6f1784838 · inbound

ShareGPT4V: Improving Large Multi-Modal Models with Better Captions cites this paper.

ShareGPT4V: Improving Large Multi-Modal Models with Better Captions mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:12.985200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T17:08:12.727773Z digest=sha256:54950af8eae79174948956aa92ae57da3d24e5d5d093ddf99c6d2c5cb365d769

Observation 1827c624-c1f7-4e8f-a474-d7470f1bfd18 · inbound

An Embodied Generalist Agent in 3D World cites this paper.

An Embodied Generalist Agent in 3D World mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:22:18.658537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T14:22:18.606817Z digest=sha256:68e3a7f9c49a8f0de6f1265018b6ef0ae5a9c980f3072bd4b14ce8edac6f48b1

Observation 6370e879-f16f-40d3-9cd9-0a591343968a · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:37:41.704687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:af608351e871af386a097d83976452ee2353b117fe9761d4fe2e7677f531ab7c

Observation 54661be6-1b54-4002-b669-da2f5e91851d · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.123593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ed1b2ca4a8cf4e7799263d4be4036e08e5450d56fcd65acdfd26a33bcf395faa

Observation 612cebfb-9871-416e-b028-61034679cc2d · inbound

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices cites this paper.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 127

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.171663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:73c4963042263e7be64adf84d0d8729122d21dace18401ff7f5504e41c8ca287

Observation 25c59aa5-2992-4d29-85d6-5b5b6a0230c6 · inbound

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents cites this paper.

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 105

Resolution
verified exact
local_arxiv, observed 2026-05-17T10:09:46.632276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T10:09:46.447508Z digest=sha256:5d7ff2af0c0bf6bd9eeb0153bf54e5a305ce7299956497e6dc4a3f4622cde6d2

Observation 3c38bf03-f69f-42d8-87d8-fd248438338d · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:33:30.430063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:b8522a6bae3565aa76c684faa6db9beaec49e2cecc0cd0f415067c4b8166739b

Observation 29f21e54-a2f2-4286-9570-f57881e28d40 · inbound

Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception cites this paper.

Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:19:28.009730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T00:19:27.965902Z digest=sha256:cf94dcb069eea6f39b0c81e10a26dfe4e8e92442c67510d6a1336ee6605d4199

Observation e779d337-483a-43dd-9583-b1117bc6f29d · inbound

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model cites this paper.

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:27:52.125475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T15:27:51.839171Z digest=sha256:13b4df9b65d6821959ee9db1f655c074b1c5298c77bd66156e7fb351324279ed

Observation f4ede3bb-da84-41ef-a5aa-b6de47c2bc2b · inbound

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning cites this paper.

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 179

Resolution
verified exact
local_arxiv, observed 2026-05-17T10:58:53.480487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T10:58:53.215887Z digest=sha256:599a1f096ee51aca4efe65d3643a03e1fcad280fad5449890d9c16eaf858dbfd

Observation c66988cc-63eb-421a-aa47-81bf253b1b29 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 129

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:46:16.862645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:85a86cfc476943644b2c85c982c718b02ba7ccde6694618e6b47ca70cbbeda70

Observation ffea0252-1c39-4f0f-961d-a817d10e3b08 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 124

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.359145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:849a7a8318c5d7153df798ae9beddb65e806a0272b31c0cd60865a489ddadb61

Observation 4acff808-ec0f-4638-b458-70db110dbda5 · inbound

Generative Models and Connected and Automated Vehicles: A Survey in Exploring the Intersection of Transportation and AI cites this paper.

Generative Models and Connected and Automated Vehicles: A Survey in Exploring the Intersection of Transportation and AI mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-05-24T02:48:47.383923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T02:48:10.934475Z digest=sha256:07aae00e016533c1d94f0d2a3d234afce8883dcc52b4c24105141ed81b78a7bc

Observation fe117760-de67-47c6-86df-6ba21ed3c357 · inbound

Are We on the Right Way for Evaluating Large Vision-Language Models? cites this paper.

Are We on the Right Way for Evaluating Large Vision-Language Models? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.416241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T19:41:44.263663Z digest=sha256:0c7b54931d8627419016eb42c5d967a89ccbe6c4ff6b59179ad61e8c56d1e1d2

Observation 0db715e4-c6ce-4d06-935a-ecb9cfbe959a · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 188

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.243343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:8afb477a886e69473969934063c5a41751daaf9e871cf820bb8af71cc28385b7

Observation 43f69a16-3446-4395-b26d-a6b62830dca0 · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 114

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:34:37.739196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:8501ff81ec175a0d79671bf913a5f3ab45a71cf98b1bb6fc241618a45c407c5e

Observation 2fca4937-5447-4c41-af00-8ddddb1a404f · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:55:26.482600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:127f92e04ac65c0d60ae39786dc5759e139a28c206f247ed52281639404980d2

Observation d5c109a0-9114-462a-bf25-31dd5942c855 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.704805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:d27a78a4ff72b9616f0169521b61804ef47c9994cba1daf78df813e66aa393c1

Observation b7089374-dc6c-4598-9382-6d94c5bedff9 · inbound

We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? cites this paper.

We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:55:41.004026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:55:40.808698Z digest=sha256:dbb64419ced49e93911af19f2be59021eadf95b888e508645224c9c60411459a

Observation 14aa7ede-5221-4739-83d3-02b5cd04641a · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 261

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:20:36.560997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:05852150f94243ee11ee03db7c2337e7ee86208bf2e1aa66af3390897e234c8d

Observation a9bdbae0-afc3-4e40-ad65-3f37fe8d9869 · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:59:32.734287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:cfb83926468c4a1b99d6c2883ad5e9fd263eba7e639feeda19b3555600353467

Observation e64a57d2-caf1-41e1-b96a-46707c56b9e6 · inbound

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark cites this paper.

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:51:48.441410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T00:51:48.163349Z digest=sha256:317c0d515034b1d203cafa0900302d6e945fcd60db79bfe8d9c942afdeb25b94

Observation dca7e261-89ab-422a-802c-6b3872e4f7e1 · inbound

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference cites this paper.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:58:32.359216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:cc71d1c186372db3eae6c6066e884b4859ef18bff26b8e37393e0aafcb3d6b58

Observation 3bb13275-a751-4d4e-a841-41ade18e559f · inbound

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark cites this paper.

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:05:47.864373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T20:03:38.336841Z digest=sha256:8ab58042db9acd637571a68248bd87390595830e8c8a74a4fd0809ca1003d373

Observation 584107d8-04e9-44bc-8705-32a18bcb5e03 · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:53:33.714961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:c4be5137b9309d5ad41e8d5bceb9b1c2ed57a44cebea2de4a0fcb5329bda0df9

Observation 59dd7a6c-167a-454e-84d6-48c4e8adf649 · inbound

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs cites this paper.

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-23T16:58:12.147855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T16:57:12.821916Z digest=sha256:0448992fbac41548aaeecdef010044b1705dccd7ab6b04ebe1ea56e967724fc9

Observation 5932dfb7-3179-4cc2-a915-3303f017e8d2 · inbound

When Large Vision-Language Models Meet Person Re-Identification cites this paper.

When Large Vision-Language Models Meet Person Re-Identification mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-23T16:58:11.858648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T16:58:02.417053Z digest=sha256:8d54ab9f2f021dd41afd5764608530aa693e4ab04a2b4ddc44c6a6c3a66ca4e4

Observation fa1e2b3a-1aca-45ce-bfec-f549f0c2e0ff · inbound

Open-Sora Plan: Open-Source Large Video Generation Model cites this paper.

Open-Sora Plan: Open-Source Large Video Generation Model mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:42:45.159395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T08:38:27.946746Z digest=sha256:4b0b90e09a5d91ba141c34e83b4750171a1f51911663354148c0e7585a3113a4

Observation e3eb39e2-a421-42ba-801a-86a9868077ed · inbound

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks cites this paper.

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:05:29.213504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T07:05:08.716223Z digest=sha256:a29db0d162ca3a2e7f4d5f2480210908099a767f69488ebcd67e120a955dd0c5

Observation 0558c787-2bf6-4a06-9024-203621515970 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:02:37.657152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:bbd0e2a97b535be30410637d9fe9790dc435e823f2820b62f17d7a1861c84089

Observation 9e60d85c-99e4-4327-829d-5946a85e52b5 · inbound

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL cites this paper.

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:15:46.338926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T15:15:46.255296Z digest=sha256:3e4b50abc894b992cda2471a2ed4df23d3ee4db0984ea1492d743b19f5966ff4

Observation 8a2f37d5-f00c-4f4e-8bec-9a88201beeee · inbound

Q-Agent: Quality-Driven Chain-of-Thought Image Restoration Agent through Robust Multimodal Large Language Model cites this paper.

Q-Agent: Quality-Driven Chain-of-Thought Image Restoration Agent through Robust Multimodal Large Language Model mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:22:09.793387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T21:17:15.694349Z digest=sha256:991f7cb9f64ce751d02792b3a99587b794460c9b917d3ee54b8689ffe58888c9

Observation d9029727-9535-4217-990a-46e0d006d43e · inbound

Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration cites this paper.

Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T12:52:18.070105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T12:48:44.324236Z digest=sha256:dba0774069b678d45344f34b407f2001b179693d65f95eba48cd8882e1689406

Observation ee130942-ea7a-4237-81b9-e3f9d0b59eb5 · inbound

Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding cites this paper.

Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:43:02.172931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T09:42:14.171289Z digest=sha256:8c2be059fc7f3c68edf280e2a06f701a1bf06e7f4b31c50192c42587f2f2e162

Observation bc6634c1-5380-4f0c-bbe8-421a3de6dab4 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:42:04.355425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:783c451466c68f97b1c8e3e4210c661b3bdcbbc81e57897c661c0c6dbc4b7dcf

Observation 575aeedd-6a0d-48fb-a16f-3af8b3d3eacb · inbound

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models cites this paper.

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:26:33.887861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T15:23:13.318310Z digest=sha256:cdb4fcdbae17f231200118049e4ba401bb8e5ee4e578a975f1953b7b51373d3a

Observation 656eb6dd-2e8a-4b3e-a877-200dba689f01 · inbound

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models cites this paper.

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:30:39.467615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T21:28:59.898029Z digest=sha256:a9341178906dedc320ff257adc848b08db86a9f0ed864803fd35fa9e86fe04f3

Observation 6f18ac0c-2eff-4346-8fd9-94fb847f3974 · inbound

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM cites this paper.

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T11:24:22.627249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:24:22.627249Z digest=sha256:fd8cc66d74aff10c3830a52947eb9fb0980758a38fcd5ca03f6cd09874a4b52b

Observation 0981b908-739f-4a9f-aa52-001c61bf3bca · inbound

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model cites this paper.

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:18:44.550010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:18:44.550010Z digest=sha256:536ed723eedaba770eec88c89f4a2cc02f6b24da98bd22d24d0ab813825b024a

Observation e59d0de6-02af-4d0f-98a8-6243b17f1a14 · inbound

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs cites this paper.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:08.346069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:08.346069Z digest=sha256:85c633d186fc32d497ae2c08355b96b2f862cb48aa49f3703d1b1eabcfb8255e

Observation 0cdc6667-9167-471e-ba50-2e699eb78d92 · inbound

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents cites this paper.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:11:29.873760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:09:31.161760Z digest=sha256:0c054bd460f2b8e620bced90736148e42a241b012e4d91bd10a61f5339615721

Observation 92116bc7-0121-4aed-a4df-7aac4060b207 · inbound

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents cites this paper.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:48.397328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:48.397328Z digest=sha256:6f9b125d7eb64ae849fc5520c0e0b6387216edf5d676fe03734906d8de45fff5

Observation 8217daf5-0d1c-46b2-ac43-aa2968335a53 · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T13:02:00.686182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:02:00.686182Z digest=sha256:6f2a7ad90fbcb0546d13c7c8528181a1a6eb886ec335be67ac7e3ba992976249

Observation 60f37dde-6ba9-43ff-88da-749d9fe23481 · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T06:36:33.949676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:36:33.949676Z digest=sha256:b5c680abfcbc53c90c935cd142b407a859910c56d6117b6d6ceb327774b84f01

Observation 6521beb8-f74b-476a-a313-c4b183aaa3e8 · inbound

Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation cites this paper.

Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:30.854898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:30.854898Z digest=sha256:ad30aab26d393eee92edaa1811163261039f7266d3934cc64d53b346bbe9f295

Observation 67790b3b-d1f7-4a77-842c-a0b3b64531b6 · inbound

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion cites this paper.

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-15T13:43:40.241796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:43:40.241796Z digest=sha256:ff151e9bf731789ff4090df2b7437ab4af262bd71d8a3fea9c86ddb8dd331109

Observation 6ab547cc-cda1-41e6-b117-eba8546c1d2b · inbound

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning cites this paper.

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-14T22:18:56.115559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:18:56.115559Z digest=sha256:44af8e41c36d763dd53dff1a69099dcbc9b654a3e9258db51884997e42f0f725

Observation fa6d3f82-8639-42fa-8dd9-fcd97d668a2b · inbound

DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization cites this paper.

DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T11:10:02.352378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T11:05:07.336827Z digest=sha256:fd074fc70f0ea63bd09673bfee8a0d2d10570efa32ef8ed13404d9adab3f770e

Observation 4561c020-80f0-475e-b953-e97ae17bd251 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:08:04.297157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:d137a0a102bcbed212743c09ed92fe7150f006ba15a4bb3bb967150f1c3563f7

Observation 603e4b0d-d928-43ff-96ac-d27b9c253815 · inbound

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM cites this paper.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.463336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:56e734a94fdfc199625b99b7fc587f5200b09f6dab72ed2585d62b42ba2d72d9

Observation a23dda5a-a2fb-4b65-a15d-ee527e89e665 · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.188098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:53c2126b2d224dc97609f9ef9500c291d4eb95d19da7a2c970e966e7e9452e87

Observation 7a846a2b-d46c-4c2f-b105-de90b785b8a6 · inbound

Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding cites this paper.

Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:26:01.395330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:29:49.681399Z digest=sha256:316599c1991e01711db5152cd128940e2ff7462fdf724dbf8ebaae85a3a31cf6

Observation d44ef7c9-740a-4b14-84b9-bbba089d64d9 · inbound

R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs cites this paper.

R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:04.422982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:28:14.039910Z digest=sha256:0be851661e2f9fd65e885a5454b40f90a1d9fcedc6ca6a6268e5100392c6d35f

Observation 37c41524-cf77-4589-b7b7-e8f28b166e66 · inbound

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models cites this paper.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.832543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:278cce3f1374b1ec23942b73614511a8700c71a501cee888ddb6dbdb4d1edf4f

Observation a0a9d8e0-8560-441b-9d16-5054bd9e2c72 · inbound

Latent Denoising Improves Visual Alignment in Large Multimodal Models cites this paper.

Latent Denoising Improves Visual Alignment in Large Multimodal Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:09:26.730864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T23:07:54.806529Z digest=sha256:29ab3447e1d3d1e6560844cdfee7fdb5072d2efd41b157777c2384fe7f478d9c

Observation ec175278-467b-4f7c-94d3-07433fceaa42 · inbound

ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models cites this paper.

ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:51:11.023925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T10:53:27.265815Z digest=sha256:9e9b5dd9c618d21cf4993ef3800cf685f855c96281317b0177d4c93ec61e2c16

Observation 72aaa357-9d65-4b0f-b511-d34c966df4e4 · inbound

ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models cites this paper.

ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:43.477913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:43.477913Z digest=sha256:cdf1c4e41137fddc99a3a99168b4969e95ea15b1105389c91d32ce773d2e5812

Observation 4a82033d-3cda-4f74-b2da-a22d1a1497ea · inbound

AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition cites this paper.

AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:01:11.833187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T07:25:42.870223Z digest=sha256:4613f900f92d58ea9929fa0b44bf6dfb21c3831b0e82fd0e403ea727e7116d3a

Observation 0ca23b14-1dcf-4089-ae60-7c869c778804 · inbound

ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning cites this paper.

ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:59.717874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:51:42.304051Z digest=sha256:9f589ac55aa3f975c70a68c197eb6eba894f5e8aba9fc2eaa8e7c95fb691672a

Observation a774b6b9-dc84-436d-85a5-29e93bdb4f59 · inbound

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? cites this paper.

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:16.343060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:06:45.858231Z digest=sha256:cc334cfbb1f0397b3b41404e69cd8a4b88a9633332fc77f96001759cc2bdfc80

Observation 36c3ac5b-1ba3-4d49-a10e-b561e49192a7 · inbound

Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement cites this paper.

Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:47:21.423502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T05:44:05.093926Z digest=sha256:bef92aff1cd103d3b3cdd7504300aefa3a1b06ec69cb14a237ca22308813169e

Observation a01fdfc8-680c-41b3-bd5d-22b4f8366ac9 · inbound

Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models cites this paper.

Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:23:44.400235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T21:21:11.388050Z digest=sha256:7974c3d4cdb987238aef088b5ff8ac13c51c8a0e5ff411070ac386cff3fe447c

Observation c5e2dafd-2067-41b0-a964-7d8a0c76c25b · inbound

AffectVerse: Emotional World Models for Multimodal Affective Computing cites this paper.

AffectVerse: Emotional World Models for Multimodal Affective Computing mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:48:05.774694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T06:46:33.612905Z digest=sha256:f161aa67963fa15bc7628c6c2eeedfc68433824c30fdf85fd75374749e8f8ae6

Observation 9c0c9dae-d5f3-4b24-a83f-7ccb7dd9480a · inbound

Mitigating Hallucinations in Large Vision-Language Models via Causal Route Gating cites this paper.

Mitigating Hallucinations in Large Vision-Language Models via Causal Route Gating mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T17:24:56.634573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T17:19:45.341282Z digest=sha256:3f6c2ecc90671ae550ac6270b1cdebdb59c6c06d9f15f72c279bc12350aa1b44

Observation 7a0f665d-f0f7-49d3-9cc5-f925a4ff2880 · inbound

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models cites this paper.

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:16:34.738320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T10:10:38.261635Z digest=sha256:cf9e9f92c55d2b406172a5b7114a7efc00af128dbfc25212bee6e5cab973337f

Observation 7547257a-0f44-45c3-8c3e-867ea7e2cf4a · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 123

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T09:17:48.386689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:44941727dbc37fcbf697f0b6193af85e61f10e4c3637109b9cea547d1286511e

Observation ffa5f6a3-f059-4790-a24d-e68400af9c00 · inbound

The Hidden Power of Scaling Factor in LoRA Optimization cites this paper.

The Hidden Power of Scaling Factor in LoRA Optimization mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:08:21.974063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T07:14:08.479610Z digest=sha256:11132ad75c17514bd83dc93649bedd5c4d49a00a3234ee39ab074d8aeb563bdc

Observation 19d3cfb6-64a2-42d2-b9c7-9d2bd0607996 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T17:18:43.806230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:c24048dea50aa9c374dd6195c2a764e0e7a7212c23d46032b39722866998c0ad

Observation e7d40f5e-62f0-4007-ace4-21dc05b5f166 · inbound

NEST: Narrative Event Structures in Time for Long Video Understanding cites this paper.

NEST: Narrative Event Structures in Time for Long Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 288

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:31.187726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T17:57:55.366051Z digest=sha256:8e76efa7521f0b03bcc2f7e4568484baee9a4a12ba305399b9999afc26e9fd5b

Observation 18c50c91-7666-4f9d-8b9e-b539045cb3d6 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 150

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:39:37.691004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:e8a957622942537b18799481334d0dbef0209350d374f1571861b2c29c2818d6

Observation cb54ec4c-5023-406a-897a-55cc01b83f35 · inbound

An LMM for Precisely Grounding Elements in Documents cites this paper.

An LMM for Precisely Grounding Elements in Documents mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:09:54.720233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:51:50.043216Z digest=sha256:9d88e0500a042211f338a3655607b226935bd84a2b30b079c80a5ba40de95a1a

Observation 7e892e0e-68fe-4e70-862d-bb51d0142525 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 95

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T19:20:06.429472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:b6c2db15c5b09483933833ac4f2ea97382c7dff0e98710a307daf0b7cc06240d

Observation 23fa663c-7f5f-495b-b8ff-27fcefa72f8e · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T10:17:44.329479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:17:44.329479Z digest=sha256:b9279d03f876fe1488b2d2843039bb4bd41c490066951cf74b579e0161fc08bf

Observation 502db605-70bb-49ac-99d3-9aada034aebc · inbound

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference cites this paper.

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T14:09:53.838517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T04:24:38.917137Z digest=sha256:0fd977df20656f0195d5368a4ae9e1ff63b1da55fbfa362f587c6cd937507e15

Observation 143b3b6d-b76f-414e-84ef-d5e764626aae · inbound

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy cites this paper.

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:51.443168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T05:06:09.216428Z digest=sha256:a160090e911514d696b5bb78d3e2ec3f8cb8a759f1bf4fd8ef7075ba2d089c4d

Observation f55624f8-0cc2-48c3-89cd-6e9cd757d53e · inbound

See Only When Needed: Context-Aware Attention Intervention for Mitigating Hallucinations in LVLMs cites this paper.

See Only When Needed: Context-Aware Attention Intervention for Mitigating Hallucinations in LVLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:54:20.794831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T06:47:01.711336Z digest=sha256:f4fea618ab1e6b20bbb3584e072f245244a433e255540bb9308436605c98698c

Observation 3b280dc7-d374-4a0c-a64f-c8c249e4708d · inbound

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning cites this paper.

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:48:32.415187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-03T14:47:35.377391Z digest=sha256:5c7f6f6bf97a41d501d283b9af8a436b696e59995a9297b4b16e43c70e33f79d

Observation 5a096d5c-d394-4091-8171-91a6caf37252 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 151

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:db5f97ff23d58a55391f6b68ed5df479c28bb5fc6f55f4947ba1aba3deb85a86

Observation 78fb47b7-a639-4a07-93fc-ec7dfdae28ce · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 206

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:6dd2f63fa6d3c0d9bd3a67d551a93d3586c292b3a9aa28efe8f6a1502ed90216

Observation c2d6a28b-f309-4d26-ae2c-d32521d46a3a · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:ade6908fbcdc4393ac3f698a49fce574e0e2c33a1dc0e99b350906164c33ce99

Observation 141a615b-57f2-4749-a930-356f8a16468b · inbound

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions cites this paper.

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T01:58:01.554825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:58:01.554825Z digest=sha256:8cae8d6f541a1ab8e8465b41face06b8804e1e1a978decb2956b6e0208e32cf4