Pith. sign in

Paper Citation Record · LEDGER

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

As of 5 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 26 inbound Pith citation observations for arXiv:2311.04257.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.04257 v2

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T03:18:51.582340Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:55:55.637449Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T06:39:37.569021Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact40
  • verified fuzzy32
  • unresolved0
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dabf9d8a-4d5a-42e2-b1ba-c0a29fdd05f0 · outbound

This paper cites http://sharegpt.com.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration http://sharegpt.com

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.887484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:118dc77ac5fb07af1605b064a374154d86670fe431118eb784a06c5639612a1e

Observation 7ffc64e1-5a3f-42c7-a32c-3669cca1f861 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.849222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:6551b29f778ffdee7b1cce1985eb8781ee3f283a9f535628af755046604481a2

Observation 4ac22e38-c165-4d77-b5f3-35bbba7c2251 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.806210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:f065bac0dee4b62293b528f13934ed4e7b35f03a1f93896a6e37de3eff3b95c8

Observation 51e9787a-644c-486e-9945-12d980ea1d67 · outbound

This paper cites Layer Normalization.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Layer Normalization

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.810253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:cfa70bc9ea20a7d062e5732dcf99f4adc3dd89a8fcaa5ca019661388d06c9565

Observation cd4efb21-6969-4288-8197-2b6727b2a138 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.814142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:18668be8636bf0e312dd760ec3fc68f52ad7d55a50154423c8847b21e7b8fb81

Observation 46cc1ed0-9c66-4c19-8723-c62ad851a954 · outbound

This paper cites Language Models are Few-Shot Learners.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Language Models are Few-Shot Learners

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.818322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:ea27da0fa0f512759d7961768aafd84ebeaf9a0293bf4913f9c968217d1cf8c3

Observation 1cc48520-f9e2-4231-9bfb-ad9d3a3d0c03 · outbound

This paper cites Coyo- 700m: Image-text pair dataset.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Coyo- 700m: Image-text pair dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.868050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:cc415e0931864e4a1f66f25db407d822787c8815ba2e4e572279e146f1771b51

Observation b788be39-6370-448f-8d45-518299d7e79e · outbound

This paper cites End-to- end object detection with transformers.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration End-to- end object detection with transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.871325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:2266fb0f695292d44b51beb3ca76df7192ac10465715b0274e7f79db533cd95e

Observation 8f7a8dbf-f0ed-495c-9f10-12fa286256be · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.874884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:4ab5f1e6bd7eea226bc408ad047a1f75f83fe1bd154d85583b624f9505725c4c

Observation aa5f582f-dcfa-46cf-ada1-b57716b280ed · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.822093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:3e632b50dfd938024735f1279b77ae195087c6c6d1ffb973c014cc02bf639673

Observation 045c4426-7361-4bd1-bdc4-7d51ba13ce1b · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.826291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:8ce312c353fbd02ff8e0f7fbe795c02d31a30bc48781c6917b5978988dd250e0

Observation 2a5f5f0e-1200-4f1e-95e3-ccd6186fafa1 · outbound

This paper cites Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.884663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:4c455dd4893e769fef628b7e81baa3dafff2c7836afc5e6b196e3c9d2c861625

Observation aa1e4a5a-7b0b-4bc6-a4f7-d685307adbf0 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.830880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:66eb08e00777481d255a022ce8cc4afd388d7249af612407be2d45237f7de81e

Observation bf1ec431-8207-4d45-97b4-7adc78b71842 · outbound

This paper cites Opencompass: A univer- sal evaluation platform for foundation models.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Opencompass: A univer- sal evaluation platform for foundation models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.890389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:ca93b61739330ffb45e4e628e6a2e51e8192f8c03aa732571e3b92749b4e727d

Observation c4ec7c7f-4f2c-4ac1-848d-4188eba549c3 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.835167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:b844dec26d36da20bc21a32ad273e46aed4c151f21d8a6e1f6740da8f86abccb

Observation adf7d802-bdb9-4848-88d6-6120231ba215 · outbound

This paper cites Xia, Mehdi S.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Xia, Mehdi S

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.896393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:019cfa9342f5f92d9a69380c6fddfc80520174333bf8dc35eed2136824dd077a

Observation c95f173e-1bca-4237-879c-5b610b3c123f · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.839531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:8eec32d37becc8bfac8ca2a724f82947bd908397150adcd416af22fcc1c1c0b9

Observation 50537838-c88e-44dd-add6-32810a4947d5 · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration DataComp: In search of the next generation of multimodal datasets

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.844203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:a64f94c50e3c3d871bcf237783e1698b1bb9bef4fb813170acf97e9e7b0b631b

Observation 531e3cd3-b9be-42ca-b882-bac1a200d757 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T03:18:51.639903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:383873045fba0ca2e7ea1b3070f8e4ceb7e391ff34d610b2beafb65c58af7032

Observation 8f444185-6efc-40b1-b169-19812663166e · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.645922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:f46a6255d139171f0db5f83b6518594b6c62e617d74358d448349467c0acce54

Observation 8bf39e90-255e-427a-ba7a-98a6cb1d969f · outbound

This paper cites Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.910903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:410fee72fa6538640d94052d51893cdce7c6b88352d349b4267b26eef9e839c7

Observation 376a0f8a-d07d-4bf2-9d9a-e6680b461bd9 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Measuring Massive Multitask Language Understanding

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.652730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:24f3677d1e293b53405c5123ad80aee775a855dd074f15d83ee7d0af828427a4

Observation 54258956-4631-49d7-83ad-fb6969eace15 · outbound

This paper cites Language Is Not All You Need: Aligning Perception with Language Models.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Language Is Not All You Need: Aligning Perception with Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.658614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:ebb22815387f250ebf9a701ef9aff4dc9bfaeac40181987a27db2f159aa719ae

Observation dc357717-0205-49cc-80cc-eb69a2f89d42 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.918823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:65107def08cabd4e9ad932d737069444456dab807422427cda822b6e3a930ae8

Observation 7138ca41-f18e-4299-8197-30f5e168234d · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.921195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:3fda957b886f016e88d225e87a4d8f107c1022b998347f0bba813d3b78b8b06f

Observation 1ef21c4c-19b1-427a-80c1-9421962abb50 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense imageannotations.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Visual genome: Connecting language and vision using crowdsourced dense imageannotations

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.923967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:84d149332b98801acd174929423f7df0c52430d51c7f55ac19632e932887e6fe

Observation fc7a416a-d0d6-4b29-8305-a111b9e8910b · outbound

This paper cites Masked Vision and Language Modeling for Multi-modal Representation Learning.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:18:51.664930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:50fe1ae7ae6543c8e9d1e9162d7419f3a213dd9cbc34ab6db41e4e2cc3ea1b03

Observation c7c84476-4eee-446d-8602-cc2768e882de · outbound

This paper cites Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.928895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:ba9144eb842568b237a9ad7b235f337145547a935e21bb727283d241e45f3797

Observation 38daaa6e-8ee2-4656-b93c-8e1be25435fb · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.671246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:1623d7ee18a4f04fc348169d24273c855f69e1c3b4f7898655afb5956ca508be

Observation de6f8bcd-3f79-4abd-a14c-0f2f7af7bd6d · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.676627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:e1966d3e5ba8e59d1e3c7219331e83abaa667e12e97ff779fa9ec86bbfdfcf17

Observation edddc23b-c9ff-4a87-b594-1a214319f62f · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.682240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:7b76b8c5619b432345db6809bcfa4280e9f43a48103d70d93b80da2afc9c8b98

Observation 48b1f484-3639-4ead-a4b8-01f3bb01bddc · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration VideoChat: Chat-Centric Video Understanding

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.688069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:b56e39a73031c526fe1b3fd1c99c17053602b3b622f3c2dddd67da1039187899

Observation 45ca166c-de2a-4486-a74b-f5882a991822 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Evaluating Object Hallucination in Large Vision-Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.693529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:606cb34ecdd54203b539904f9bc498a9642643056ce010fac3ae986752f86155

Observation 8dbd5756-be1f-45f7-b9cc-d292468be12c · outbound

This paper cites Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.945711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:cf84c11dcf141b251a6932f0e5655fb76f0f6ee03b5646de510c52a31f95d032

Observation b1de4368-e754-4b37-8ba0-98fde650054e · outbound

This paper cites Microsoft coco: Common objects in context.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Microsoft coco: Common objects in context

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.853136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:7d013ac3525a8546c650ae0313c9dd58e09efbb075b54e4e8647b116503ebceb

Observation b27c7cbc-14db-4736-95db-3e4986caa806 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.700114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:105dcda9db0832e8c71696b0ff11a4a3c3faaeffcf7d3c2326db9a82fcec740f

Observation e24a6a25-5838-4fab-aeb3-8cd7f8a95a70 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Improved Baselines with Visual Instruction Tuning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.705895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:7044023f8cff2bf7717de1ac9b410013513ee873736998c8d2dcd0fcbd38687a

Observation 4a7681a3-6995-44ea-ab55-5b2fb58eecfc · outbound

This paper cites Visual Instruction Tuning.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Visual Instruction Tuning

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.712380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:3d090889228403d078d1a48c7c5e056156dc9f9deb20256b398efbd1cfa0ccde

Observation beda4e00-d371-4df3-addf-64e0a44886ba · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MMBench: Is Your Multi-modal Model an All-around Player?

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.717146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:fe96fb81d6d0b378b9ca211a534fb05bae10573925f1efe3f1dc2925cb07e824

Observation c3a4e4df-863c-4a38-acf7-45d60b07b916 · outbound

This paper cites Fixingweightdecayregu- larization in adam.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Fixingweightdecayregu- larization in adam

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.881298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:131479f166648d5db1bbe46f7258df541228b5456abd8aa851eb20328a2c1d45

Observation ae304f91-df2e-44eb-85c5-d36eacc2c498 · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.723093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:e760274bbcf7f03ee6e5e4823aa678511b517db518be83d2876cb5a752f07d43

Observation cc3ae5ba-d5ed-47e0-a352-6eebeda33c17 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.728083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:c81a4ab1f16a49c9f910be1ecc8ec0d3af667cab3ba5e46ccd416191ec2dedcf

Observation 71cd6aa0-cb97-401e-89a3-5d24420e1974 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.899248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:ea82041e779629232ebcb74c5a14f943d952755165ecae5b9ecd7b23117b55d5

Observation ebf9cce4-8f62-46fa-9ecc-4efda528a7f4 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Ocr-vqa: Visual question answering by reading text in images

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.903199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:39b920771c5665158d8eb368728e1233d9350673af57d4c809d1548c21b2d40f

Observation 77041ce3-75a5-477c-ba05-a8b120a3c703 · outbound

This paper cites 5, 13, 17.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration 5, 13, 17

Reference 45

Resolution
parse uncertain
raw_fallback, observed 2026-05-18T03:18:51.905783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:f2d22608b70be46d814a945e60b50ea2f369a7423ea254e077cdd6728ca52d4e

Observation 672c078b-d93c-44f2-a26a-0d824906bbdb · outbound

This paper cites Gpt-4v(ision) system card.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Gpt-4v(ision) system card

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.908520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:617370629962399f721b040302575edef4e47f740c2cde7f0b5a44d0e9d0a706

Observation 65d53aca-4246-4221-9440-c247fc70004f · outbound

This paper cites GPT-4 Technical Report.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration GPT-4 Technical Report

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.732444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:fa0a127de052bf39db21639f92e14aea21f6bfa1a3d63d2593bc525aecf12b18

Observation 363859b2-683e-4927-b3d5-1f50505575ea · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.736822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:5b918621e31e9bf2394b630d877e771b162621ac6fa888fb4b16c2ab677b6399

Observation 7df40d85-0076-4fc6-af2a-847eba804395 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Learning transferable visual models from natural language supervi- sion

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.926547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:76b828e4d8cea32c62d3ffa571431da229c4b4865d3ea6cc3d6793ea93db165e

Observation 624e37af-4de2-4e27-bc0a-ea21285d6f6c · outbound

This paper cites Laion-5b: Anopenlarge-scaledatasetfortraining next generation image-text models.Advances in Neural In- formation Processing Systems, 35:25278–25294.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Laion-5b: Anopenlarge-scaledatasetfortraining next generation image-text models.Advances in Neural In- formation Processing Systems, 35:25278–25294

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.931904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:d334070edf2769d1a78d08c7f3c2eff918db4399d36a0feaa31e54d8ced9d866

Observation cf33827b-ab92-4308-aa38-b74c81478bfb · outbound

This paper cites A-okvqa: Abench- mark for visual question answering using world knowledge.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration A-okvqa: Abench- mark for visual question answering using world knowledge

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.934741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:2a9542e901677a5096a861d42c09f7b0aca0c77ab0249d2e279c7be148b77fe5

Observation 175c5a8c-d7bf-42f8-a43a-cb8bf1cae27f · outbound

This paper cites 5, 13, 17.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration 5, 13, 17

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.937146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:80c89485f9b3cb8840d439e09f97d3a56620a96a22eea78dcbc41bf771726d71

Observation 8e4012cb-1099-48ea-83fb-5d86b182abd7 · outbound

This paper cites GLU Variants Improve Transformer.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration GLU Variants Improve Transformer

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.740513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:fb637d9b4c8233cb8c71cdf7acf62e67a685693e69610f160e72055901426376

Observation 3b0cd216-b39c-457e-90d1-776d007642af · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.744387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:31fc25cb38b5bdc6ef88908518989e77e18132ce69588537eea4afebc2014047

Observation 263cd0dd-6105-4614-a500-a3b7ae773355 · outbound

This paper cites Textcaps: a dataset for image caption- ingwithreadingcomprehension.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Textcaps: a dataset for image caption- ingwithreadingcomprehension

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.857737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:d048b2a0c41fe7965e02ea3b2a311cfef1b0217de4c254542a4741759d0b6328

Observation 3570651b-bb5e-4146-b65d-de388019a2cd · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.748972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:69ac6125b4295aad4bdb5793c92bbabd25517a994b599fa17c5adf9632b4b91a

Observation 8c92b47f-a279-4e35-9317-1968abfce620 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.753187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:1185bdaec2a56f0d5d56c805f93c90a06d56f1de95a4b0f324f6dafa3fe49f3d

Observation 601cfc42-d13f-4588-ad3d-701803007823 · outbound

This paper cites Hashimoto.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Hashimoto

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.878373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:a8a5a0cec789ecbed0ae0590c59447824112d6a3d5078f663daec6f1e24ddab1

Observation e2f105fe-74aa-4fa0-bd08-295baa97feb6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration LLaMA: Open and Efficient Foundation Language Models

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.634100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:72392be7d43eafecfd196afb155d48862498ce3c5869114fee5fd9bc610b922b

Observation 9f0a0de5-3a8a-403a-9a69-65e6a3f1558e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T03:18:51.757298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:f7953ba8690b0c49c281f71333cb9c043d36910362c3a3a3058b09cbac520b25

Observation 81fd797d-0d47-40e5-97a1-46b2a468fb81 · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.761087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:d87d8d3813394f7673b5ef13b66f1e85704edb1c1271bc01c86f8d8c90da7246

Observation f42ccc44-ad91-4fa8-a89e-cc5147a91bb8 · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.765863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:6a279c994f9cc2d0f9be5fcabe836a54e75f200f5567c908efbe60c4bdb3a103

Observation 5fe46ad5-5219-465a-b98a-f6ff0b8df7b4 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.769928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:dccc37f54d287bf0a8af861aab25b8f4c55ed22ed1f6dd9d4cc21b3e2c87c213

Observation 36b968f5-7243-47ad-9beb-eb6f63ceb33e · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.942415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:8dbf46fd2626a9f214d416cd9986625d13a49dbbbdbb90abf4a8adae53a6da5a

Observation c5663e2a-2c52-41d4-97fa-d8c0c55f2202 · outbound

This paper cites In InternationalConferenceon Machine Learning.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration In InternationalConferenceon Machine Learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.861000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:cc0bf1b587a5cfdcb60f0496cbc5d3f77080bd90c3d3355946d07ab949db45c4

Observation 0a78790e-29d0-4296-8fb6-9ab2139f53c0 · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Zero-shot video question answering via frozen bidirectional language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.865165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:f8670e199a72e185d562fee7a6fea4caa2200480a02d71852f0166cac8e23db9

Observation d4b8cd34-6934-476d-8aab-f6f555e1387c · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.774332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:599a44881a9b76e4d1d6a33e6ae8999790d7a19910ac244f2a79658c550897d5

Observation c2a20675-9658-414a-915f-2d96836e3a41 · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understandingwithmultimodallargelanguagemodel.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Ureader: Universal ocr-free visually-situated language understandingwithmultimodallargelanguagemodel

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.913586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:cbd6b7107092758190bf424a5f6ef11982f841193ecbe899685388fba839c086

Observation c1184676-be13-44ee-b55d-c037f1989874 · outbound

This paper cites Hitea: Hierarchical temporal- aware video-language pre-training.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Hitea: Hierarchical temporal- aware video-language pre-training

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.916523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:1de0925b864d4d65a80bdfb57620cab5cba175e236e39609a7ad2e47be2652ce

Observation 15a88951-f765-42ba-bef5-0210290d1b26 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.778085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:d3c1be845701f03e0a44b40ccc6355cc0aac7ae1f93a1e9f355aa29930fd5ae8

Observation 5705a2da-3131-4c87-9d8f-d21d7b23303d · outbound

This paper cites Modeling context in referring expres- sions.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Modeling context in referring expres- sions

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.892981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:5b1b96ce56e22a76830608ba7a226036f74cda595536f1ce7f6bc9edbe7b3bb3

Observation 33796428-5124-40ac-bd43-d23f3ef608aa · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.782964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:4564a64d4f174c702026a445c29d062cf831e76de38523c2263d10664e333d97

Observation dc4a8f39-1acd-4589-9407-0614558899a3 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.787911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:c695b36f7d373f252ec2bc85a0350ed8975c3d0b527d5a12c6bb550cf30cfd1f

Observation 8a1aeacf-975c-4c5d-9467-ad30e2b5fedc · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration SVIT: Scaling up Visual Instruction Tuning

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.793866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:11bf74db373ac8bafc10449f63bd93b3623c280d4b8f78f4f963f763636fc71f

Observation 7a39f36c-0833-4151-9d18-3f9fad54b0ec · outbound

This paper cites P Xing, Hao Zhang, Joseph E.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration P Xing, Hao Zhang, Joseph E

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T03:18:51.939744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:8d76f42d3e7702149bf6ea2796b6eb2f8596087f44ebabe37da4809be0d7e107

Observation 9644ed94-4617-459d-9b84-777600b65e97 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.798370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:a348eda61cbfe7d9df3f259cb299ad63e389ef694abe88eca29e9ebd3935fc42

Observation caccece5-ae84-4be2-ba8b-24fb0dc06ac3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 77

Resolution
malformed identifier
local_arxiv, observed 2026-05-18T03:18:51.802150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:dc05888c6aaf67ec7c58caee8a456dd304b2a045f6ea6c4649bf184c36c335f0

Pith citing papers

Observation c720735b-edf2-4085-87db-a7b5e3bcc50f · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:a2cf3cef56d4bf60651f4be07e58577d3a722190ca8c658cd16255b8b90044e3

Observation 86566b5b-222c-4cfb-8475-21b98c56d9f4 · inbound

MMBench: Is Your Multi-modal Model an All-around Player? cites this paper.

MMBench: Is Your Multi-modal Model an All-around Player? mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T17:20:53.687692Z digest=sha256:068580dcc3bd29743ab736da7270a7ced6ec68896c277efdc2c062b302f7e297

Observation 176fbb6b-ee30-4a1c-80a2-e6bf6a6721d7 · inbound

CogVLM: Visual Expert for Pretrained Language Models cites this paper.

CogVLM: Visual Expert for Pretrained Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:1d656c56bc1b8e87ea35a1bb1d8295940d861870a0a8e3df9dd8193aad2b8898

Observation 77838a10-3067-4695-ab75-e857e2e36946 · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:4c2224c7c63af40196d1043a7521b83244a4552b214d8ce5ba3dbb48946be233

Observation 74aa7ba5-e508-4027-8609-6a373889c352 · inbound

Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception cites this paper.

Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:19:27.965902Z digest=sha256:f7349b86ea634daa670a3426d3f3b8c9170ba4ea45c45c293bf9a4ec5870a819

Observation 15bd591b-ad1d-46be-b116-789496e35c97 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:8e6fc01f90ac0a73f772026795f922e5c3c4df3c56372e0e8ca96e7bebf4a833

Observation 5d51d752-3267-4779-abed-cab4a8de5409 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:f720afddb38d233ebb7af49c81fee1cde43d4a96a9b79e5152baddb9fa1b92fa

Observation 724b22dc-be35-4f0b-929c-96a9e7d3ea53 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 189

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:06aae43ee69e1ce5adb1707259742dd7ebe1c86f80873e2379699d0c9c16bc4d

Observation 4232900f-91e2-49fc-9527-e3c9a2d30a05 · inbound

Detecting and Evaluating Medical Hallucinations in Large Vision Language Models cites this paper.

Detecting and Evaluating Medical Hallucinations in Large Vision Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-23T23:58:39.595999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:55:57.103971Z digest=sha256:d94b4db11e3e5fb294071042175c0e31dbf687c6698489b3a10ed5cbf36978e4

Observation d364d906-c976-4e03-99ba-c89be89abf89 · inbound

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark cites this paper.

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T00:51:48.163349Z digest=sha256:b4aac010d75122aff72f44ad035f5c9b35b1d1c7a7fb50ca329c0db2cd1f19e8

Observation 0297dbd4-e743-4d32-8608-ca401d71eec5 · inbound

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs cites this paper.

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-23T16:58:12.153327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T16:57:12.821916Z digest=sha256:0bd44788eba0274b8261e8728ae40b9f8fdee86ea296b790d0f45427f8ac46f0

Observation 03b48de3-3b9a-4b9f-a672-fbc1e5b13f38 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:afee8ba41cb6847cfade72f3fb8943653e3d1e419cbb024c1405b72d600a0573

Observation b0f5dcfb-043e-4bf3-9dd1-6181e967d880 · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:25:19.057831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:6ef22dfe62d6ab4c90e2c9b69b13a5026ecbe07615263f8b9844ec22567ab269

Observation 1265ab54-13b9-47eb-b30b-ae6965dc4041 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:8d05c4ac09990d77bd0e5cc85caa3498359f9204cf0fc943b27a386af795f1ff

Observation a28beb4c-f76f-4ec7-a2c7-82d2cb6444bf · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:e45ce908670c97e77540e705624bf14606b9981d75a3befcaa109f2dd414cd8b

Observation 98d6f1f5-70c9-4e02-80d5-f8a463baf381 · inbound

Parameter-Efficient Adaptation of mPLUG-Owl2 via Pixel-Level Visual Prompts for NR-IQA cites this paper.

Parameter-Efficient Adaptation of mPLUG-Owl2 via Pixel-Level Visual Prompts for NR-IQA mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:55:55.637449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:55:55.637449Z digest=sha256:a183bc9373d4e73b3b7ee2f4b46a3ed8823933c4838941e947f54f51d3091d8e

Observation 249f0c18-328e-4706-a150-e7c42a0b3163 · inbound

MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models cites this paper.

MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:51:31.394779Z digest=sha256:2c19545b9cfda8f1faa882d39fe745da36a5ed2205df0c9c2ed7276a5e43fe61

Observation 71c0aba1-8f21-4af8-b5e2-ca6021598b25 · inbound

From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception cites this paper.

From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:25:46.990772Z digest=sha256:ec65c57429a3ac7ddebca74b6571f438c9efe9e51101de5d7efbdd820e41522e

Observation b7aec3d1-05ed-4b01-934e-e0a5e2ed6045 · inbound

Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation cites this paper.

Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 136

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T00:33:39.960170Z digest=sha256:1a7a0097d42df76363f5bda28aab5651ecc8b4245928cf0054801f6c80b63dec

Observation 28979578-b0f6-4c53-bb45-ade7405486a7 · inbound

Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing cites this paper.

Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T04:50:08.866969Z digest=sha256:06d8bfc5341ef5504e2fe9e961c8928523fce082a7823936419df2fec23632c4

Observation 3349ccc7-1b8b-4a25-9b6a-e76b8a020c3d · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 168

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T17:18:44.000141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:bdd0ce40e4be5fa3da292ba36420a3567e8a526a8b42665558930008df84705b

Observation 08253312-f120-4863-b5ed-e4ae3fdcce76 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 234

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:39:37.570172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:2a5f50204821a5011fc41b6ad8768a56ba85a5debd0b37fa9cbae89a867237aa

Observation c8fe25bd-83ba-40eb-b1ec-c57bb2d6184c · inbound

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation cites this paper.

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:14:18.740368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T06:13:11.013395Z digest=sha256:5324bafe2f50bfa2ff43a5216edf0c4ac76978df5c7dbee5902da098bed6ed84

Observation 4e66300e-43c7-4670-bcf5-881d882097a8 · inbound

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation cites this paper.

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:05:28.622861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T07:01:26.911110Z digest=sha256:46f2f87a771d6a73e433374bb41aff5eb72396fdb739ecc1a353afd396549d43

Observation 1aaf88b7-bedb-42e7-8b3a-51431da31c21 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 181

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:a356769bc2fe6ff8e69db6b86ac439cbeb1063d6c3f8357571a4d4bce51daf55

Observation abbab7a4-eef6-4abe-8d74-a2efe25f6549 · inbound

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution cites this paper.

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T12:40:56.954792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:40:56.954792Z digest=sha256:78451660ecc476f11721d837582abcfdbb25aeeede93dcb75acf0c64ed177448