Pith. sign in

Paper Citation Record · LEDGER

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

As of 7 August 2026, this Paper Citation Record lists 100 of 238 outbound references and 55 inbound Pith citation observations for arXiv:2408.04840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.04840 v2

Coverage vector

measured 100 of 238 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T06:20:36.235304Z

measured 155 of 155 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 55 of 55 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:25:01.561578Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T03:19:31.849614Z

Reference resolution

100 of 238 outbound references displayed

  • verified exact6
  • verified fuzzy70
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch21

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d152c559-f91b-46f6-83ba-2551ba297dd3 · outbound

This paper cites Scaling Learning Algorithms Towards.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Scaling Learning Algorithms Towards

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.700677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:7caf38c6403df89a44ef2aaa23271a698970afda753c321feac7ca32de286575

Observation 00ec488e-6ab0-4a15-a07a-7098679feaf7 · outbound

This paper cites and Osindero, Simon and Teh, Yee Whye , journal =.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models and Osindero, Simon and Teh, Yee Whye , journal =

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.798740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:2e5366ee8d1082207e0361a11bb1274810c6a5bae8e6b8b572baacf234b4088a

Observation a2ad1236-f8a6-488b-b6f1-10d7dad77009 · outbound

This paper cites 2016 , publisher=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2016 , publisher=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.801465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:c7db27f75e0c4f253ae7334d5f96f858b0a68e5cfb37107a5586bf03cb911f29

Observation 88700bb6-3084-4318-9425-781e7a465ddc · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.804054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:f649725c6f3f2b200d7014143e9fa22b8049f1274571d17a61c065e60b804ffe

Observation 1c2852a6-e767-40dd-90b9-a368384dc822 · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.308851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:ccb2a59db5caa128da2f0fb568e95fe638ea21c78b16e70247fc33083c34f780

Observation 94ade3e2-91ae-4a81-8725-7d45ecf2d541 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.806512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:9816f8c0eca7ed9f8775e008c74cfdb7c5417030b5a0b6900f75d6acc3a03020

Observation 46833eb5-afb4-4337-8c04-6031360c0806 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.810565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:c493181f5b9097a6350f9ea311d5873a786c43a3fa10c226ffc4613973f90828

Observation 3a70f8a3-f81a-487c-99e5-9bc5df9289bb · outbound

This paper cites 2023 , url=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2023 , url=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.813234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:6952bdcecd020df503cd49626188236ca3c8b116dd5edd4a6ab93fc79e514d1b

Observation 753a7c4b-2dbe-4119-9bfc-4f2e1d7a3ca5 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.816796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:efad646f4990e232a1b5007cfce41dce88810d489eefa6dee7aeac24d17c5a6a

Observation f14e61f1-7c5a-40c1-a6e1-cc2918ad9c63 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.818910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:c1f192354769144c921af17ea4e315e47f6b73d0fa1d58f436f7cb0d6b997a06

Observation 3e935220-6d7a-4359-9129-931142f84d12 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.820819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:4525cd691a0a02f8ffa11368b1ed8a6888ddd0d3bad837c8cfc5de64ed9d1b3a

Observation 2d26736a-cca4-4234-9531-4ece34d31204 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.822855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:3bc79ea96836f6fc6cc3093d539cecc1261cf9762e6ce7866b36439628eaac5d

Observation 421a475f-44f5-4c04-ae7e-dbac6f0d756e · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.825014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:96b8b751bccfd4fe7c3e43b03dfcb37ce1d65e080a5fa0259737f6ecabb56246

Observation ec1cfc84-7bf1-473b-b461-2c28a5ba1a12 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.827390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:5aa7dfde313cc90e8cde26dc26d7d15485e181f6d1b5936de6ed591c94875c6d

Observation 3183f8bd-6770-4ead-b564-3a2a86d224ed · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.829394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:08bdc4dabb398d8181a4de078ed2f07069abfbf15abc02eb5c0739987e117152

Observation feae5519-5a67-47ea-9151-94055b25e7b9 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.831381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:8f373362c99f9627a6cdcfecea725a9119f8879f1a2b919c4ea595c747a635c9

Observation c639c225-477e-42a1-af18-42444ea4011a · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.833781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:977e912fe65b7cc4f6a6ceb93e69c47a54afc5319aa0d5c048605ff5930de13c

Observation 8e9277c1-f2b2-412f-bcd0-450571db335c · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.835831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:b87cba1776eb4f6744fca7fdade6def389d6309e855c1f52bbd8a97d57c9c17a

Observation 2239683f-121c-44ed-bc35-8fe1192c9202 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.837718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:60e6281babf992fffe27700d21606402b01da68f77d872dab2f75ca39f348472

Observation e3d8ce9d-ee23-481c-b8ab-4e454a72f78d · outbound

This paper cites International Conference on Machine Learning , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models International Conference on Machine Learning , year=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.839697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:a4e7f1e259cbc0cb7a075c4a675e204ad525e5c5e51894ae3bab2648304c6236

Observation c3ee6392-4914-41c0-a32c-5793cdbd2768 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.841798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:81fc4f8b6109a2e1da4454b59b9255d834a84654e0628044ec89418318e226db

Observation 15550b04-61ab-458a-b25b-6ecde2b8a29b · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.332513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:a9c36c3dc381a67fd52be6735c54efeb2e6f9cf440ff9330c61c746d0f6da651

Observation 2c4051a9-1a87-41fd-91a3-323ab187e8f8 · outbound

This paper cites International Conference on Machine Learning , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models International Conference on Machine Learning , year=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.844003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:dd1f0952d90930c87147272d94231a19a457669270103fa67b3bed4b00c7f4f6

Observation 618385b2-e032-4a55-8429-5103f9c4efb3 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Advances in Neural Information Processing Systems , volume=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.845997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:7788ef16f2baaebba3b98ffa32a3d30853a2d94568c5fd82f5c06b93eadd3268

Observation 3343d62f-8f14-4329-9194-3f02a33b645f · outbound

This paper cites an unresolved cited work.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-20T06:20:36.847987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:cc926c83dfd4dc86d145e3e6e16943cc532e1e15ea5d15c1b7e26f263cf25188

Observation 8280fb1e-c6eb-47c7-8150-d422a1d9cef2 · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.491574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:82991dd693c758e6809e9b9637bd0ee7feb15fad9fd2cdffea966b3400d5f994

Observation e95bc2cb-2ed3-4f09-9ea4-a17ae1c7323f · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.568176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:fc784420d07711d86c6d51c79ef8d64e186ad54d4a88b88909299baaff416610

Observation 8a99aa96-9325-405c-9b22-6efa8d54da14 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.849804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:a337739802fdd3886b86da5599ffaff42e5f2d95c2626872a136968185c5a547

Observation 6b8f8874-c2ce-4656-8578-bcd40876a4e9 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.851737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:849d616a69dbaa810fc09e7e0208df7dd2666b3f3a881d555394985daef5ff2a

Observation b68542f5-6f08-47b1-ae87-63610d78ca0a · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Advances in Neural Information Processing Systems , volume=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.854197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:193c883e45a50d144da3967d8bc7fc1dca9de1ee804962a13c5f7c50158b13e9

Observation d3c1fb7a-e4b3-4b61-a5a9-e3ea664d1690 · outbound

This paper cites Advances in neural information processing systems , volume=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Advances in neural information processing systems , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.856187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:59231c32c09a1eb3cfb6ad8a4a18398c4a84c459798ee6caae214e4caa572865

Observation 60c055f4-18a6-4f37-82f7-f6087c6bb65d · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.865122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:6e5196d27c9e38651bf7fa6f54d8b27c8c7d28681539b7748cee40bd4ea263f6

Observation 173a94b3-db98-48ef-93d6-96fc96bc7b4d · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Advances in Neural Information Processing Systems , volume=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.867260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:6265cc27d4427b8cc2dae5216ea31e8f0b35da7606ab3ba02b0a03bdfc5e8867

Observation a442456b-ee74-40b2-b22b-aa57346c2b3b · outbound

This paper cites Retrieval-Augmented Multimodal Language Modeling.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Retrieval-Augmented Multimodal Language Modeling

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.381184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:9ea756ca2c2e9b43e8feb781cd8c2fee60446d1be9af73aefb7fdbe65011769a

Observation f3bf0015-a5bd-4cc5-b815-455984531663 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.869372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:b1e924d491ae54bf9d4191b20bc6ed014c124b5fd7d0dde21afec99e623247aa

Observation fa2f54f5-b5d8-4d21-a342-ac3040be2aac · outbound

This paper cites NeurIPS , year =.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models NeurIPS , year =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.871669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:5209e4e8c0f64f4f6ee5573f8328997087fdecd10747492a117f377fb72a02b5

Observation c3954fdc-5170-456a-9141-5e821f801377 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.873716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:451e1fa51e46d66b31874f990d8c641c410c7da5a6150bc3f4b19d721e775242

Observation 61211925-b1ef-4910-bb97-21e40b3b16c2 · outbound

This paper cites Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies , pages=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.875856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:ccab3561f31bcb84666735fe7ca23b148260500bc2388f927d0284b5f201e4ea

Observation c9587af5-780b-40ea-82de-3e86e6d28b66 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.877751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:e8e9a9a45fcb767b72b118b30487aaceff676e5859ea37b080b79ccdca036a88

Observation 3972d5a9-9d91-4a5c-b1ef-47d199c8f63e · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.879732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:319446cc0cc14ad2aed0976f18c5f3f4444374c24e97fc1141a9d01c29370701

Observation 2f4d5ca9-07c4-4ec6-9ca1-ad65d68f4174 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.882031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:9756540b53fcbfde8c2e0c64f0411850a1e5b06d61c6203e382cebf16d8f59d0

Observation 459734e9-3e66-47a7-acf6-9c60bbe997a1 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.883935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:eb97a63de391765d08e9b102023c70a80501c267c5c1e9fffc7e6303a896e0fb

Observation 62173ff7-ba5e-47a1-b44b-738f8e06a365 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.886284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:9273259befa3a61898d1d888ce4ee1dd31863c8bd3a5d2299e311721dc3765a2

Observation 1fdd638d-6af2-4dca-a9b2-b360ba23d48c · outbound

This paper cites European conference on computer vision , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models European conference on computer vision , pages=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.888866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:c2597631185bc90c8a70b914f4924c9464dfc794209f0dbbd791d6c10b1e69f0

Observation 5e39adac-8c17-4a81-a131-ee9321220ea2 · outbound

This paper cites GLU Variants Improve Transformer.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models GLU Variants Improve Transformer

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.397564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:94e88273099989be1950303bf444d59d1e3a23d50112b754ff2886120c981621

Observation 47ae2a54-4c68-4fe3-b7b7-8ebf33bfda54 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:20:36.893584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:23f66bb04fd3c0ac0c83755f638a3017c91bf6c7e2af86036a2b6d781a333276

Observation 3f08ae79-f840-43a5-9061-abb794005b2d · outbound

This paper cites 2023 , eprint=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2023 , eprint=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.064036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:ada33cde8b1de117a19e6bb083e0f8eafe352212f120839cd1ce55ec64c05912

Observation 6368b160-0a54-4b3a-95d9-fa3da198c1b1 · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.046833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:06696fde80e46f92f47ad9b124fa3379fcb85ff47fd71b8748895420996f8c9c

Observation 3ed11c71-f922-4ce9-839c-e6f1bc8540e0 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.356515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:97961225e9dd32e53e148c92a8133cb41d03f427a1c8fec17ea3721a39c0fa9e

Observation d1a283ce-0177-4dc8-a8da-bf383082fd9a · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.368777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:63453c68deeb7867c1d7aa94dc7e9077a9b6c3765850365d70bbf6eebc503c93

Observation e2317f9b-8599-4c53-8e34-11678bff397f · outbound

This paper cites International conference on machine learning , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models International conference on machine learning , pages=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.034944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:c69a166488c8bc6c8c57e75253449802b2e69cb7eb3eb3812de567bf6c9aa8b7

Observation 87898014-2e2e-4545-b3f2-469bd9d4a018 · outbound

This paper cites Hashimoto , title =.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Hashimoto , title =

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.038535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:a23081d7a593cce6e50b8140d1009f4073521ddb9de73c82bf327091e7c247cd

Observation ffb8abba-46ec-46d0-bee7-cee713c7fcb6 · outbound

This paper cites 2023 , publisher =.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2023 , publisher =

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.033138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:8d87168af91ef2e6113bf8477db984176248c57282cc409e2e1e7a16ca3c1bc6

Observation 16dd2632-0ae1-4f92-983c-b1fa3a700b35 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.446419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:07b0270dc5623f5605d4fa31764a20b54a00bcbedfd59d1a9dccabb9c859edce

Observation 6276443a-ba48-4c65-a2aa-85382af71546 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.488698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:90e7bda189bbfeddde1ffe9c96a6b8511b666d21e0653f309dd0961dbffcf596

Observation 5e0b7c72-b03d-47e5-913d-e571be8c4224 · outbound

This paper cites Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part IV 14 , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part IV 14 , pages=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.067495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:0fb115599569b3c52dbd12b7b65e637eb6d841316db855073a4bcf872530832d

Observation ea24bf46-f79e-4581-8d74-e8261f784f22 · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.557416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:96d68bf57fde1a0a6f811a6fa4014778c3df135e71ec7f7047861475f9fffb79

Observation 2e455416-bbef-4168-ba29-0d1a98080757 · outbound

This paper cites ArXiv , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ArXiv , year=

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.029652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:0f644c40d77fd63210a7fbae3a5d2db47d29c656f7b5b3ff9e6b47e3a5a97d2e

Observation 6fc3f590-22c1-4606-a51a-c12b083a6bca · outbound

This paper cites Measuring Massive Multitask Language Understanding.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Measuring Massive Multitask Language Understanding

Reference 76

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.592244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:2d754681c113ce74c2ce7b6b31eee8bd93be6aff6b35aa31e3f4570b0d70efeb

Observation 60b36141-770e-44b0-8135-0e137ef45a1e · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.319194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:b2a7bfce6ba609fe28684873dd89eea9ace421cbaed0f5c399742a15af8aa647

Observation bc643947-5dc4-4ac4-b4fc-a6db1271f161 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.324452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:959729fed2c9a3de53132cf09c04ed848f83902537b08fd3f1f2e83e33d801f7

Observation 1a08c6b7-3c0b-491d-ac6c-dfa147dfaa7d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 79

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.328464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:c1ccbd562ef5a769480b94996b6e0b23e8ed41f283e056ddd76a9e602a8a287f

Observation 0c51f8da-d84c-40c9-a7f9-ca643263d9f1 · outbound

This paper cites Proceedings of the 25th ACM international conference on Multimedia , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the 25th ACM international conference on Multimedia , pages=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.028111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:c002d5bc646de4283a6c79d08ff6d4ca133d03571c1ae70d2258ca7cfad5ba2a

Observation a97c1128-9332-49b0-b54d-25c07fda91ec · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.069417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:dfff1b5bee86b28f795d7a5d99650985fc5083e12090860018474ee59ffb50bf

Observation bd1fb5d2-f1ab-4256-aeb0-4db9e777d852 · outbound

This paper cites 2023 , eprint=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2023 , eprint=

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.018627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:34d249148568f60d52021df4c0e612c3b785aaad56ab5cbd918d1472fa52a69d

Observation 2364c3d2-e62b-4eb6-b278-d5b5eb496c17 · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.020696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:2be55f3790026e545ee836edd599bff266d6ff4a4d2c383a756cb53a0c60d378

Observation aad7096f-a958-4ecd-aa90-85009d76a08f · outbound

This paper cites The 2023 Conference on Empirical Methods in Natural Language Processing , year=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models The 2023 Conference on Empirical Methods in Natural Language Processing , year=

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.012599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:0d391c82af6ea8d7dc93c7dee159183dd213aa056b89ef25dd184b9a514e7467

Observation 73483d60-7ded-49e4-9ff1-bc6852fde0c2 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.016946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:210a561e5ab9ded3522183ebf9b08c6b5b3810ad12f5745661fe8f46f66d89f3

Observation 51412131-f3b4-4537-a344-7e3287006371 · outbound

This paper cites an unresolved cited work.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-05-20T06:23:06.010046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:1d9f51457a97158fbac6dbfc46a4b09c76d34de10c360066d0bec666a5779913

Observation efe76cd9-42a9-44cf-87eb-0afb48a74dff · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.042950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:a8251d4135910e23af3d1011f8b29cebaf3080d755a754d3e8fe496861d4e2f2

Observation 111472cc-437a-42d0-8acf-13d335379d19 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Advances in Neural Information Processing Systems , volume=

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.026404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:95b268c3af9a075124b39be1c3980964f76ded93a3768e162a60bfc63db18432

Observation dabce7bf-65da-4c23-8192-94447874a840 · outbound

This paper cites 2022 , howpublished =.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2022 , howpublished =

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.024462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:4b8ab91d2b3d864209f90aa7e6cbcbccac57fb758cf1e82e386cccc591aecea8

Observation b3db8ade-94a2-4f95-8e14-eba49653e079 · outbound

This paper cites Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 , pages=

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.014593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:bd77de4345ec2b91e3aaabe04c7e8ff7d1119e28ad2bf7acbbd31362cfd27546

Observation baf3f318-bcd0-4af5-8ffc-ac0bb0681d35 · outbound

This paper cites Making the.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Making the

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.078861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:0dfd07936eb30e045a8287b0653251d51d323deb792539aaf7b3e017efa6c243

Observation e1b41dfd-6553-4f96-9e69-5dfc736d5adc · outbound

This paper cites 2019 international conference on document analysis and recognition (ICDAR) , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2019 international conference on document analysis and recognition (ICDAR) , pages=

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.075266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:36ce6993b918db30c61ae0d3b6fe94e9fa8c2e9227837547ce3f06525ec0f44f

Observation fb3aafd5-6a05-491a-aeab-84cfdc61d901 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.036857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:50f149123e6f8cc423500e1f4344965cb3f6177a8a2ea768d8dab9db63ac4a51

Observation f786beee-39e0-48f5-a98a-56abbe7797b1 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.040906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:ad28496fb613ca1f31099b8e7b7691cd4c1dab5743dece7724925ead86412a48

Observation 7280a6c9-5301-46e3-8c2b-a664cfe3307a · outbound

This paper cites Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part II 16 , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part II 16 , pages=

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.048869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:26f082ff0ca8061d279630280752d67c9aa8315a90c085ce596731b3d528520a

Observation 0e2108c9-3e57-4226-b915-b74f688589f2 · outbound

This paper cites Proceedings of the IEEE/cvf conference on computer vision and pattern recognition , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/cvf conference on computer vision and pattern recognition , pages=

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.073532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:45f5b8abf890f677f3036d6db870f7f6fc14e712f2dc35c777e9ae3535958044

Observation 9fc45383-eca4-41c8-b4b4-555684d65f86 · outbound

This paper cites European Conference on Computer Vision , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models European Conference on Computer Vision , pages=

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.050596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:36a1e12495ae9c770f58331996dc2b78b01cfb3d0b28043b55d168328feff64d

Observation 27a7f066-5dac-44d7-8fd9-15d65a60fc2e · outbound

This paper cites International journal of computer vision , volume=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models International journal of computer vision , volume=

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.007756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:a46ffa012612d5f165a1134df0be9e6407f8d93afa50fbae7b8e5d8442ce78cd

Observation 9d0133a3-e6f3-47de-834e-3c41370938d0 · outbound

This paper cites Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14 , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14 , pages=

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.077201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:d4aea1885806548c75a08f99fd372b3413af62266f41ad83577e70010a182700

Observation 4f0c569c-2947-40d3-b5a5-39c76f571b76 · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.388847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:668a0c4cc70c1b543870bb57589b780f36e3dc883f99d50979c1d727dc9e4eb0

Observation 79617ae0-2c61-4ab2-84d6-3190b00c92c6 · outbound

This paper cites an unresolved cited work.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Unresolved cited work

Reference 104

Resolution
unresolved
raw_fallback, observed 2026-05-20T06:23:06.005708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:d1d54101245938ac25012164d8765727db9f228129029158309802e30416122b

Observation 55cf0588-1a63-43f3-a548-c69e5be1de70 · outbound

This paper cites Masked Vision and Language Modeling for Multi-modal Representation Learning.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Masked Vision and Language Modeling for Multi-modal Representation Learning

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.402341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:61e5cdc5a2b4c5628164ced45705e12d019ca1e97cc24bae6ac8d1d98e1fe6a7

Observation 75f46906-3c13-4aca-a331-f74e13f0615c · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 106

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.044909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:b2664537c2d90767e3d7ad4495a66c8bc306c28e6de235075c69de2ed5708abd

Observation 6ae7e215-c939-4ac4-9d4c-1e01b59e066c · outbound

This paper cites Layer Normalization.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Layer Normalization

Reference 107

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.449690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:de6aa35f53fd671f531aaec7611b5c10136b861802c233e6b2ac6dd0545994cf

Observation 3c2de6fa-cf34-4987-ac84-5884a3a36d71 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 108

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.461623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:49ee5177a1330c59afbd878398d48f5d7d9efaea008d1e7192ac40452ece76f7

Observation 755bcfab-8b96-441c-98ba-e66b641fc83a · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 109

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.485479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:a2990f6f6fdbd0eda26e04a16ace361c865ce379796eb82e7b19b47bbe94c187

Observation a15194ec-c396-4ffe-b937-900484ea5728 · outbound

This paper cites International Conference on Machine Learning , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models International Conference on Machine Learning , pages=

Reference 110

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.022490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:ec3fdddee6ca9f617ba8ce77de9410f00385e80f06f22a644f1d87f89b1fe8b8

Observation a6ad327f-dae7-4ff0-bbb7-30cc3804efa1 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 111

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.031290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:b44cffc6efc3d4e30ddaaa2db2d5a7f96881ef6e235a2f5166eb4814893e5f6b

Observation a3829c38-823c-4691-beb7-0f3f0998ebda · outbound

This paper cites Exploring Diverse In-Context Configurations for Image Captioning.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Exploring Diverse In-Context Configurations for Image Captioning

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.495065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:e1f399799868c54fdb5b64b4ba3cd20d4d094a47165a4d9e19ad340dc475cd76

Observation e66d72d4-c06f-42cb-83c9-5fdd9237b91b · outbound

This paper cites How to Configure Good In-Context Sequence for Visual Question Answering.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models How to Configure Good In-Context Sequence for Visual Question Answering

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.501279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:63887514b68a786fe5162cc755d9af95cb05f23180c2c0f7e63db53ae934048d

Observation e1525ec3-824b-4709-8235-9f7ad85d1a7d · outbound

This paper cites Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.515715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:6d719a8943babbf5608a2f294e03ea6fdfa3abe65c247cc32b9c6e01be868246

Observation 0c0549af-ce0f-4f67-9497-b5d0257fbf9f · outbound

This paper cites Manipulating the Label Space for In-Context Classification.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Manipulating the Label Space for In-Context Classification

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.528552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:e5dab1126ba2771581ff9364d8d20553a26f64c0ea6d41bc2a63adfe8ea3ad86

Observation a297d3f5-7a6d-45a9-9dc4-69131e4c092d · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 116

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.535304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:af5b4fe903a4a38ea399e41510044ee70296b29ca89a8b24fd73e32cc4c7d1e1

Observation faa4d1e7-e021-4d15-9c8d-f08cfd1dae1b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 117

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.538521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:b9ff7fd19df457cdcf213c645d4908cf63b9fb25f3093802de5ed225d9737dee

Observation 04cca4ad-d10f-4e47-966c-7749221d48db · outbound

This paper cites Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.548843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:833615c73947cf206fe0e262531f67d3fbb497064a8b0be53c72bd077d1167d2

Observation ee5ae40d-22fc-4835-b1ae-e9ab01209402 · outbound

This paper cites 2023 , url=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 2023 , url=

Reference 119

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.061708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:1e77a19ad2e0a46b4a2c32496d903e99603997f4ddbda9db3d27bb02400dd0bf

Observation b524ec34-9c4f-4248-bbcb-32bc0f777fee · outbound

This paper cites LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=

Reference 120

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:23:06.071464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:679fc70b02a0f9c241561c662ff4dfcdea987b7929ed29f839bb9c9b1dd05ea9

Pith citing papers

Observation ab603d58-ef63-4ccd-abb8-47741ba195c2 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:b7712b943878cfb363119d6879b12b413e79a6c1192f9276ba97c43a1f479e1b

Observation 3608fb53-3b6e-426d-953a-e9f524d786fa · inbound

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs cites this paper.

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-23T16:58:12.099189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T16:57:12.821916Z digest=sha256:fe514ae3cdbd26a3354d37ad71a3e4f64074d14f81230aab5ebc425b51437f22

Observation a47649b8-0f06-49d9-97a9-8dddce700d69 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:595c5369941208a5f914092830d0412107653341340174b9993652e54c580911

Observation 9710c9e3-ac62-487f-8521-303ae1e425ac · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:b2c5e809e4f1ea05daa87a1ad2ed03f6d96a5940f59b965b2f8efdceb17362dd

Observation 4a0c88b2-7a11-4f3c-9bc8-1d33ab1fbebb · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:02daccd1cf87dc17b006f5bde5192b2efaebb8cc33df3b81dedbf5de069ed874

Observation 3079ff5c-85cf-4863-9f30-621935021e31 · inbound

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning cites this paper.

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T14:42:56.565621Z digest=sha256:ee8ae44dd7fb2ec92f0566614c0598f357398a305b433dcde7161d31b77cd733

Observation 0668538a-cf7b-40e9-a8cd-75764d2ee628 · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.824392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:9710d7d98ef19f29b07599d9287771f4b5453418c5a0ffe250ec2aec5a9835b1

Observation a5f2bf78-fec8-4178-8f43-4d55feb0d08e · inbound

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning cites this paper.

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:36:36.687324Z digest=sha256:8e3a7512a9cee6490e4cc48586031cbd3e89b40d65b96b2aea93c90ea563da80

Observation 45b905aa-47e5-4ff6-b4e2-799dc5540401 · inbound

HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding cites this paper.

HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T10:44:01.880405Z digest=sha256:ac2c6a3a6c4053fbb1b537ef14610dd86f2810bfdd22fe7c9f94ff9d4ef499a5

Observation 559a4317-2615-4f68-9f49-785dc8100937 · inbound

Toward Rich Video Human-Motion2D Generation cites this paper.

Toward Rich Video Human-Motion2D Generation mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:01.561578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:25:01.561578Z digest=sha256:08a8d73846f9a64999c72e9163537ca4ef2469d87133b22de60741aff970cfdb

Observation ed19ce15-88c0-4f53-b8aa-b72b8f086db0 · inbound

Visual hallucination detection in large vision-language models via evidential conflict cites this paper.

Visual hallucination detection in large vision-language models via evidential conflict mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:54.884075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:54.884075Z digest=sha256:b0c49c724d2db99b3f1b585b35caeb5a0527e6c836a87ab7c5665f579d5ffbfd

Observation f970a53b-403c-422c-932f-5ba58f3ec871 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:f5e14e05ec31189baad3976f366298e228819bddb276ba980697673a71da255f

Observation f156745b-ab03-44ca-becf-967ee861c55b · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.971505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.971505Z digest=sha256:435574cdcf0212f7d14db0ad1accde3443aecefb24cba80757b1af1f7a598f66

Observation b3d8d813-aa58-4a52-9185-6336b493b57c · inbound

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs cites this paper.

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T12:12:23.706791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:12:23.706791Z digest=sha256:cb7d128c810d612e68433e01217120e99264574608f4f8fc86bf432ce6c590a8

Observation 053c4317-01a0-4b80-b2ee-b2bdaf1627d4 · inbound

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models cites this paper.

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T05:51:16.088023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:51:16.088023Z digest=sha256:03c5a45c9bd74e1c068babc74bae6019e3ebc2ecfed1338e8e72c602185ce1f7

Observation 7633f68b-5142-46fb-be42-0dcf7ba2cfcf · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.097293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.097293Z digest=sha256:079c751bee2aa21b20e5d072bec248907b30dfdd3939d9f55d747cab5c925c50

Observation 1c631496-3b3d-46f0-b4de-2b6e88a6785a · inbound

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models cites this paper.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.952229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.952229Z digest=sha256:4de1329f390810e8ea07099f0283793e72af20937e9f5879d3b24042253241f9

Observation 1949ae06-50bd-4c89-a926-48883de00917 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.804125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.804125Z digest=sha256:f2c388bc3258fa89d31b3d1c8e04d2550d999f8c8e4ecf5f540de66ba0e77e81

Observation 0913bc48-a0bf-460e-a289-d19558f33f32 · inbound

Aesthetic Image Captioning with Saliency Enhanced MLLMs cites this paper.

Aesthetic Image Captioning with Saliency Enhanced MLLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T10:15:50.299374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:15:50.299374Z digest=sha256:4964a7229672b681d97bbe78f629f4543f29e1044ae543a29e4b9492f213f448

Observation ee04e35f-2e5d-4c03-a0ba-e946184a1057 · inbound

CAViAR: Critic-Augmented Video Agentic Reasoning cites this paper.

CAViAR: Critic-Augmented Video Agentic Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:36.085495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:36.085495Z digest=sha256:4630fcfabb5a9520e55386beaca8aeb8f071b907aceb7d6dad52b223f9e1db7e

Observation 58f36d98-be96-4120-a13d-9d454a7a221c · inbound

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis cites this paper.

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:40.837413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:40.837413Z digest=sha256:9c54d70478f0d4f567ce6003341cde3a08ba7f98c068683125d056948dad207d

Observation 5fdfa497-be9c-43d1-956c-9b4e82d77196 · inbound

TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies? cites this paper.

TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies? mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T16:40:16.630602Z digest=sha256:d78e499f6a1ce6a3306d4757cec450c15e3ffabeb6b0cf91807ba267a859f7ce

Observation a095ebf5-ebae-4d08-af40-75a0eb04acbe · inbound

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering cites this paper.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:15.178796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:15.178796Z digest=sha256:693afcea5f6bdf3b2c24d5eca3a6722f939669bf3b8bcf19560476606fdd50b7

Observation 6ae7cd08-9b32-4089-b065-796ae06b34bd · inbound

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback cites this paper.

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T05:26:42.034972Z digest=sha256:137cb056a49287ce4c0df9cf2c33b81030b6dc28d5dac34f22c7bd9e97b9ce7b

Observation 501df437-b224-4a3c-ab3a-14689e3b59b8 · inbound

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs cites this paper.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:08.154037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:08.154037Z digest=sha256:256f0ae4093a5ef1fe6800e4042508020e61ceccf11918a1694d0dde199f0fee

Observation e3a92804-3732-4cd5-8eae-fb0ad171bbf2 · inbound

Towards Sparse Video Understanding and Reasoning cites this paper.

Towards Sparse Video Understanding and Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T23:33:15.573793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:33:15.573793Z digest=sha256:04b74cfc1c34ac9280e263697b8abd70ea8301f3067c451bbe03cca3fc03c536

Observation f85148e9-91de-480b-902b-788e24b75114 · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:17.457027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:17.457027Z digest=sha256:413f33fe554f6b98cdac8e114263d5b432172a4753f711634477a3937ee92e40

Observation 8bdc85bc-4d4f-4bed-bba9-55ff53be74ec · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:e6ceac0eb2a54096daa510284f9a3bab08c7bfba5aef23ad37c72923daed2855

Observation d6fdfb52-2834-45a6-a785-66b9afb3be91 · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:77de59c46207757de1494bf05a67bc383220e433b37153d7c87033bda62c7a84

Observation ffa1a70f-95ba-49aa-b788-acc41fcdbb5d · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:933379970d5eca4baafda0d6acd91bc83c968b8bbb424f1110e081edbe4da096

Observation 23f62cc7-2ed3-42eb-a396-fdceca16c285 · inbound

Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction cites this paper.

Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:16:31.701345Z digest=sha256:1b31a71e5ed3af1369787c6313370ff5c83b9491e76b79a856e2f79c0b87720a

Observation 927cbe4f-7be5-4c77-b890-50ba9565aaff · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:699e0aefa825ab6a163eb9018d46359ea2768d097d27a131455a4adcab873737

Observation 8f4f5e7d-70bf-4450-9772-9bf3754028aa · inbound

EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling cites this paper.

EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T07:00:36.870817Z digest=sha256:a60a16bdcdb44584ff432e078bb82da18361e7588bfbedf0eeafa2e8d1015d46

Observation 9597da25-df58-4660-b2b9-c90cdd08763e · inbound

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding cites this paper.

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:26:01.568507Z digest=sha256:0fe0869cd168d307f786770f566b14ad5067f6a4ba6fe203cb22da75364062cb

Observation 1bbecdb9-12e7-4b4c-b4b5-05754e2e008e · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:48:39.444933Z digest=sha256:fdfeee1d7a0fba50619af5dedd005f5a2b307e0cdb7171b9b4ee222c31e39770

Observation 9148bc60-ff90-4f24-81f2-c5265907d705 · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:57:42.822121Z digest=sha256:3b229d70a6532b0d4f0b58b3849e34507d926067bfb6378c34ba10dc01657473

Observation 7792c837-fc3c-4a1f-9ede-36a6bdadac0f · inbound

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models cites this paper.

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:38:14.420098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T11:38:01.947806Z digest=sha256:a82750a86fa4bee0d84a2ea91e2dd1c05bec65d2c1f1cccd2afde7db9985c27e

Observation 2ffc2cab-dc13-4d21-937f-9643240ad600 · inbound

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models cites this paper.

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.167951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:42:46.422756Z digest=sha256:7a59c75f184a65f3704333dd57af0d02c25647187099b5f2e33a16ccfaa5673b

Observation 403222ab-1573-4d5c-81bc-2845bbe8592c · inbound

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding cites this paper.

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T06:08:12.887394Z digest=sha256:06777914d2827ab121aa7d890c146ab261d68bb75e5ddc5a77cfef2e58400eb1

Observation c2c5ce59-729d-4f4a-8236-b15ecf5596c9 · inbound

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding cites this paper.

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:04:02.793257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:59:57.488107Z digest=sha256:78aadc49754b1582e47348abe9561e303ff1496e431bd61035b3bf471b5e8fb2

Observation 951f65a9-c13b-4aa4-986f-4bcf0e2b1742 · inbound

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding cites this paper.

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:15:00.001853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:09:49.178090Z digest=sha256:456ad4d043149a79259302c159d3f24a071ba132e7408c310a194dee9e2d2959

Observation 87c95789-6104-4b55-820f-e537c2d8b029 · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.894768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:962d2feed2e7f97fdbaf653bcf741bd3becbf9491775f3377898dfe1bac253f8

Observation 964e3731-771f-435f-8a57-93418b89f1c2 · inbound

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning cites this paper.

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.638957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T13:41:44.049230Z digest=sha256:76aa48e8a1f01cfe8d50e7492ce045ddcee70cb71b51284a058cd689b0153e13

Observation 63407566-6e03-4694-b112-aadf20a542e6 · inbound

Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues cites this paper.

Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 139

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T08:26:48.004528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T06:04:28.939248Z digest=sha256:2668f004101d6f987254445e0046ff14b00e12c73d901555c8fd04e3eccf8d0a

Observation 31aa1e9b-fa69-43f1-90f2-2efef04227e7 · inbound

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding cites this paper.

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.051074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T10:04:29.739632Z digest=sha256:4199da5a88829653bdc63eaaaeb0b4a6dce1784ea6bd94d32f139754b1309a99

Observation 894ca6ba-4f74-45fa-8363-550f21fe0e9f · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 193

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:02.993392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:7a477d87ff6c4e344fec313ea0e915dad06e9ae242caebd8dbd5e7f4913fe373

Observation b079ded5-78d6-4ce9-ad7c-f9b5973ba6f1 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 169

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T17:18:43.794018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:7d56e94193436f5f5435d1877e5233978b8fc69083a5ba82f9c40b8ef8a38bab

Observation 890decec-bec6-430e-9fd1-c85b2dbc2cca · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:19:31.851640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:8af2f257a7f605fd1331b050f7e772d2c54ccf1112295cc0a8cda0d694fe02d0

Observation f301847e-620a-4a7f-87ba-f8462e2c1386 · inbound

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs cites this paper.

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:35:47.610810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T01:22:16.176398Z digest=sha256:301f41b2151a15c82542751a1324f356308bcdfbb922acb32748ebaccaed5aa3

Observation 36ae4100-0c8d-4190-a485-ebe27b5e2fd8 · inbound

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding cites this paper.

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T14:17:02.655589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T14:09:54.549499Z digest=sha256:58463ca3d79ddb02e3efef1df150ce17ca2c5599e5bedeb2d28be290fb2488fb

Observation e2bfbc0b-1aa7-4e6e-ad9d-1eb2e65a6d43 · inbound

EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization cites this paper.

EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-11T12:10:52.100410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T12:10:52.100410Z digest=sha256:b4135ed85a29d35bb0e2ff433125d2f2c93fdcfad14c37aa7d4e5b211eed5303

Observation 5bec75ba-1f6f-470b-b336-41f21e2fad57 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 182

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:fa8917a541dc4241ad549be2c177057e28d89d01f72a5a8c3d0860433970fbb6

Observation 4b1d9cc7-3b82-4c1c-8637-bccc5f2dcb28 · inbound

Personalized Image Aesthetic Assessment via Preference-rich Sample Mining and Cohort Merging cites this paper.

Personalized Image Aesthetic Assessment via Preference-rich Sample Mining and Cohort Merging mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T22:28:27.815322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:28:27.815322Z digest=sha256:c7f6496e6710919b381a0b27c4c9c1e0aae284db99130cc920b07a0f2c975045

Observation 23b177a1-d927-4d62-913f-afada9d78050 · inbound

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement cites this paper.

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T08:37:41.845892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:37:41.845892Z digest=sha256:dfc9aa9220d199c8cd3671791d36428033294cef8a1f07bb830c89bdcfd5592c

Observation 13d78038-614a-4f76-b6ee-4e875293a278 · inbound

3DGSI-Assessor: A Large-Scale Dataset and An LMM-based Method for 3D Gaussian Splatting Image Quality Assessment cites this paper.

3DGSI-Assessor: A Large-Scale Dataset and An LMM-based Method for 3D Gaussian Splatting Image Quality Assessment mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T21:53:29.121248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:53:29.121248Z digest=sha256:78acf79569806e6b0dcc83a355feb55937be367e82c2aa17542a2622d814cecb