Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T06:20:36.235304Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 238 outbound references and 55 inbound Pith citation observations for arXiv:2408.04840.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T06:20:36.235304Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:25:01.561578Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T03:19:31.849614Z
100 of 238 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d152c559-f91b-46f6-83ba-2551ba297dd3 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Scaling Learning Algorithms Towards
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 00ec488e-6ab0-4a15-a07a-7098679feaf7 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models and Osindero, Simon and Teh, Yee Whye , journal =
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a2ad1236-f8a6-488b-b6f1-10d7dad77009 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2016 , publisher=
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 88700bb6-3084-4318-9425-781e7a465ddc · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1c2852a6-e767-40dd-90b9-a368384dc822 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 94ade3e2-91ae-4a81-8725-7d45ecf2d541 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 46833eb5-afb4-4337-8c04-6031360c0806 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a70f8a3-f81a-487c-99e5-9bc5df9289bb · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2023 , url=
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 753a7c4b-2dbe-4119-9bfc-4f2e1d7a3ca5 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f14e61f1-7c5a-40c1-a6e1-cc2918ad9c63 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3e935220-6d7a-4359-9129-931142f84d12 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d26736a-cca4-4234-9531-4ece34d31204 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 421a475f-44f5-4c04-ae7e-dbac6f0d756e · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec1cfc84-7bf1-473b-b461-2c28a5ba1a12 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3183f8bd-6770-4ead-b564-3a2a86d224ed · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation feae5519-5a67-47ea-9151-94055b25e7b9 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c639c225-477e-42a1-af18-42444ea4011a · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e9277c1-f2b2-412f-bcd0-450571db335c · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2239683f-121c-44ed-bc35-8fe1192c9202 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e3d8ce9d-ee23-481c-b8ab-4e454a72f78d · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models International Conference on Machine Learning , year=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3ee6392-4914-41c0-a32c-5793cdbd2768 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 15550b04-61ab-458a-b25b-6ecde2b8a29b · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2c4051a9-1a87-41fd-91a3-323ab187e8f8 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models International Conference on Machine Learning , year=
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 618385b2-e032-4a55-8429-5103f9c4efb3 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Advances in Neural Information Processing Systems , volume=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3343d62f-8f14-4329-9194-3f02a33b645f · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8280fb1e-c6eb-47c7-8150-d422a1d9cef2 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models GIT: A Generative Image-to-text Transformer for Vision and Language
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e95bc2cb-2ed3-4f09-9ea4-a17ae1c7323f · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models PaLI: A Jointly-Scaled Multilingual Language-Image Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a99aa96-9325-405c-9b22-6efa8d54da14 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6b8f8874-c2ce-4656-8578-bcd40876a4e9 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b68542f5-6f08-47b1-ae87-63610d78ca0a · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Advances in Neural Information Processing Systems , volume=
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d3c1fb7a-e4b3-4b61-a5a9-e3ea664d1690 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Advances in neural information processing systems , volume=
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 60c055f4-18a6-4f37-82f7-f6087c6bb65d · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 173a94b3-db98-48ef-93d6-96fc96bc7b4d · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Advances in Neural Information Processing Systems , volume=
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a442456b-ee74-40b2-b22b-aa57346c2b3b · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Retrieval-Augmented Multimodal Language Modeling
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f3bf0015-a5bd-4cc5-b815-455984531663 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fa2f54f5-b5d8-4d21-a342-ac3040be2aac · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models NeurIPS , year =
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3954fdc-5170-456a-9141-5e821f801377 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 61211925-b1ef-4910-bb97-21e40b3b16c2 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies , pages=
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c9587af5-780b-40ea-82de-3e86e6d28b66 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3972d5a9-9d91-4a5c-b1ef-47d199c8f63e · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2f4d5ca9-07c4-4ec6-9ca1-ad65d68f4174 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 459734e9-3e66-47a7-acf6-9c60bbe997a1 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 62173ff7-ba5e-47a1-b44b-738f8e06a365 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1fdd638d-6af2-4dca-a9b2-b360ba23d48c · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models European conference on computer vision , pages=
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5e39adac-8c17-4a81-a131-ee9321220ea2 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models GLU Variants Improve Transformer
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 47ae2a54-4c68-4fe3-b7b7-8ebf33bfda54 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f08ae79-f840-43a5-9061-abb794005b2d · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2023 , eprint=
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6368b160-0a54-4b3a-95d9-fa3da198c1b1 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE conference on computer vision and pattern recognition , pages=
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ed11c71-f922-4ce9-839c-e6f1bc8540e0 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d1a283ce-0177-4dc8-a8da-bf383082fd9a · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models SVIT: Scaling up Visual Instruction Tuning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e2317f9b-8599-4c53-8e34-11678bff397f · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models International conference on machine learning , pages=
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 87898014-2e2e-4545-b3f2-469bd9d4a018 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Hashimoto , title =
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ffb8abba-46ec-46d0-bee7-cee713c7fcb6 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2023 , publisher =
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 16dd2632-0ae1-4f92-983c-b1fa3a700b35 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6276443a-ba48-4c65-a2aa-85382af71546 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5e0b7c72-b03d-47e5-913d-e571be8c4224 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part IV 14 , pages=
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ea24bf46-f79e-4581-8d74-e8261f784f22 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2e455416-bbef-4168-ba29-0d1a98080757 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6fc3f590-22c1-4606-a51a-c12b083a6bca · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Measuring Massive Multitask Language Understanding
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 60b36141-770e-44b0-8135-0e137ef45a1e · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc643947-5dc4-4ac4-b4fc-a6db1271f161 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1a08c6b7-3c0b-491d-ac6c-dfa147dfaa7d · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0c51f8da-d84c-40c9-a7f9-ca643263d9f1 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the 25th ACM international conference on Multimedia , pages=
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a97c1128-9332-49b0-b54d-25c07fda91ec · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE conference on computer vision and pattern recognition , pages=
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bd1fb5d2-f1ab-4256-aeb0-4db9e777d852 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2023 , eprint=
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2364c3d2-e62b-4eb6-b278-d5b5eb496c17 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF international conference on computer vision , pages=
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aad7096f-a958-4ecd-aa90-85009d76a08f · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models The 2023 Conference on Empirical Methods in Natural Language Processing , year=
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 73483d60-7ded-49e4-9ff1-bc6852fde0c2 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 51412131-f3b4-4537-a344-7e3287006371 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation efe76cd9-42a9-44cf-87eb-0afb48a74dff · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 111472cc-437a-42d0-8acf-13d335379d19 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Advances in Neural Information Processing Systems , volume=
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dabce7bf-65da-4c23-8192-94447874a840 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2022 , howpublished =
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b3db8ade-94a2-4f95-8e14-eba49653e079 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 , pages=
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation baf3f318-bcd0-4af5-8ffc-ac0bb0681d35 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Making the
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e1b41dfd-6553-4f96-9e69-5dfc736d5adc · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2019 international conference on document analysis and recognition (ICDAR) , pages=
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fb3aafd5-6a05-491a-aeab-84cfdc61d901 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f786beee-39e0-48f5-a98a-56abbe7797b1 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7280a6c9-5301-46e3-8c2b-a664cfe3307a · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part II 16 , pages=
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0e2108c9-3e57-4226-b915-b74f688589f2 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/cvf conference on computer vision and pattern recognition , pages=
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9fc45383-eca4-41c8-b4b4-555684d65f86 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models European Conference on Computer Vision , pages=
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 27a7f066-5dac-44d7-8fd9-15d65a60fc2e · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models International journal of computer vision , volume=
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9d0133a3-e6f3-47de-834e-3c41370938d0 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14 , pages=
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4f0c569c-2947-40d3-b5a5-39c76f571b76 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 79617ae0-2c61-4ab2-84d6-3190b00c92c6 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Unresolved cited work
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 55cf0588-1a63-43f3-a548-c69e5be1de70 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Masked Vision and Language Modeling for Multi-modal Representation Learning
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 75f46906-3c13-4aca-a331-f74e13f0615c · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6ae7e215-c939-4ac4-9d4c-1e01b59e066c · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Layer Normalization
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3c2de6fa-cf34-4987-ac84-5884a3a36d71 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 755bcfab-8b96-441c-98ba-e66b641fc83a · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a15194ec-c396-4ffe-b937-900484ea5728 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models International Conference on Machine Learning , pages=
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a6ad327f-dae7-4ff0-bbb7-30cc3804efa1 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a3829c38-823c-4691-beb7-0f3f0998ebda · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Exploring Diverse In-Context Configurations for Image Captioning
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e66d72d4-c06f-42cb-83c9-5fdd9237b91b · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models How to Configure Good In-Context Sequence for Visual Question Answering
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e1525ec3-824b-4709-8235-9f7ad85d1a7d · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0c0549af-ce0f-4f67-9497-b5d0257fbf9f · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Manipulating the Label Space for In-Context Classification
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a297d3f5-7a6d-45a9-9dc4-69131e4c092d · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation faa4d1e7-e021-4d15-9c8d-f08cfd1dae1b · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Gemini: A Family of Highly Capable Multimodal Models
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 04cca4ad-d10f-4e47-966c-7749221d48db · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ee5ae40d-22fc-4835-b1ae-e9ab01209402 · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2023 , url=
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b524ec34-9c4f-4248-bbcb-32bc0f777fee · outbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=
Reference 120
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ab603d58-ef63-4ccd-abb8-47741ba195c2 · inbound
LVBench: An Extreme Long Video Understanding Benchmark mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3608fb53-3b6e-426d-953a-e9f524d786fa · inbound
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a47649b8-0f06-49d9-97a9-8dddce700d69 · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 143
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9710c9e3-ac62-487f-8521-303ae1e425ac · inbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4a0c88b2-7a11-4f3c-9bc8-1d33ab1fbebb · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3079ff5c-85cf-4863-9f30-621935021e31 · inbound
DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0668538a-cf7b-40e9-a8cd-75764d2ee628 · inbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5f2bf78-fec8-4178-8f43-4d55feb0d08e · inbound
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 45b905aa-47e5-4ff6-b4e2-799dc5540401 · inbound
HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 559a4317-2615-4f68-9f49-785dc8100937 · inbound
Toward Rich Video Human-Motion2D Generation mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed19ce15-88c0-4f53-b8aa-b72b8f086db0 · inbound
Visual hallucination detection in large vision-language models via evidential conflict mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f970a53b-403c-422c-932f-5ba58f3ec871 · inbound
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f156745b-ab03-44ca-becf-967ee861c55b · inbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3d8d813-aa58-4a52-9185-6336b493b57c · inbound
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 053c4317-01a0-4b80-b2ee-b2bdaf1627d4 · inbound
Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7633f68b-5142-46fb-be42-0dcf7ba2cfcf · inbound
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c631496-3b3d-46f0-b4de-2b6e88a6785a · inbound
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1949ae06-50bd-4c89-a926-48883de00917 · inbound
A Survey on Video Temporal Grounding with Multimodal Large Language Model mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0913bc48-a0bf-460e-a289-d19558f33f32 · inbound
Aesthetic Image Captioning with Saliency Enhanced MLLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee04e35f-2e5d-4c03-a0ba-e946184a1057 · inbound
CAViAR: Critic-Augmented Video Agentic Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58f36d98-be96-4120-a13d-9d454a7a221c · inbound
Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fdfa497-be9c-43d1-956c-9b4e82d77196 · inbound
TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies? mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a095ebf5-ebae-4d08-af40-75a0eb04acbe · inbound
ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ae7cd08-9b32-4089-b065-796ae06b34bd · inbound
StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 501df437-b224-4a3c-ab3a-14689e3b59b8 · inbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3a92804-3732-4cd5-8eae-fb0ad171bbf2 · inbound
Towards Sparse Video Understanding and Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f85148e9-91de-480b-902b-788e24b75114 · inbound
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bdc85bc-4d4f-4bed-bba9-55ff53be74ec · inbound
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6fdfb52-2834-45a6-a785-66b9afb3be91 · inbound
Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ffa1a70f-95ba-49aa-b788-acc41fcdbb5d · inbound
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 23f62cc7-2ed3-42eb-a396-fdceca16c285 · inbound
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 927cbe4f-7be5-4c77-b890-50ba9565aaff · inbound
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8f4f5e7d-70bf-4450-9772-9bf3754028aa · inbound
EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9597da25-df58-4660-b2b9-c90cdd08763e · inbound
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1bbecdb9-12e7-4b4c-b4b5-05754e2e008e · inbound
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9148bc60-ff90-4f24-81f2-c5265907d705 · inbound
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7792c837-fc3c-4a1f-9ede-36a6bdadac0f · inbound
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2ffc2cab-dc13-4d21-937f-9643240ad600 · inbound
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 403222ab-1573-4d5c-81bc-2845bbe8592c · inbound
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c2c5ce59-729d-4f4a-8236-b15ecf5596c9 · inbound
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 951f65a9-c13b-4aa4-986f-4bcf0e2b1742 · inbound
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 87c95789-6104-4b55-820f-e537c2d8b029 · inbound
CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 964e3731-771f-435f-8a57-93418b89f1c2 · inbound
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63407566-6e03-4694-b112-aadf20a542e6 · inbound
Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 139
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31aa1e9b-fa69-43f1-90f2-2efef04227e7 · inbound
Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 894ca6ba-4f74-45fa-8363-550f21fe0e9f · inbound
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 193
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b079ded5-78d6-4ce9-ad7c-f9b5973ba6f1 · inbound
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 169
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 890decec-bec6-430e-9fd1-c85b2dbc2cca · inbound
The Hidden Evolution of Disguised Visual Context inside the VLM mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f301847e-620a-4a7f-87ba-f8462e2c1386 · inbound
Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 36ae4100-0c8d-4190-a485-ebe27b5e2fd8 · inbound
QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e2bfbc0b-1aa7-4e6e-ad9d-1eb2e65a6d43 · inbound
EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bec75ba-1f6f-470b-b336-41f21e2fad57 · inbound
Qwen-Audio-VAE Technical Report mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 182
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b1d9cc7-3b82-4c1c-8637-bccc5f2dcb28 · inbound
Personalized Image Aesthetic Assessment via Preference-rich Sample Mining and Cohort Merging mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23b177a1-d927-4d62-913f-afada9d78050 · inbound
MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d78038-614a-4f76-b6ee-4e875293a278 · inbound
3DGSI-Assessor: A Large-Scale Dataset and An LMM-based Method for 3D Gaussian Splatting Image Quality Assessment mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.