Pith. sign in

Paper Citation Record · LEDGER

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

As of 20 August 2026, this Paper Citation Record lists 100 of 235 outbound references and 32 inbound Pith citation observations for arXiv:2504.13180.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.13180 v3

Coverage vector

measured 100 of 235 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:36.917814Z

measured 132 of 132 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:35:44.870369Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 235 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 16ca75ab-6316-46f7-87ef-10347eaf7f55 · outbound

This paper cites Visual instruction tuning.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Visual instruction tuning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.433312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.433312Z digest=sha256:b06b7c79971c7103a2fb5f5a77695e629e54d93307c5a7490ec7700cf0effdac

Observation 3a048876-d40d-4740-97ab-02b68cc5099d · outbound

This paper cites Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.439892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.439892Z digest=sha256:c5b49312fa3c65133f5a5a6f278d68372f977444462ed727ee7cfa9b124efc4d

Observation b611945f-069b-4791-9049-12a434cfeff1 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Sharegpt4v: Improving large multi-modal models with better captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.444904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.444904Z digest=sha256:1a623ef8f2512cf9f99035c6c651ea17facad6749f7c85330e4806970cfc9565

Observation 8a525e0e-bb21-4904-b92e-29df26792509 · outbound

This paper cites Finevideo: behind the scenes, 2024.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Finevideo: behind the scenes, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.449728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.449728Z digest=sha256:8be2b839e72aacd2e6e960b71bf7f42b12689bd97548c9afb0ce943b949f4f2b

Observation c492a6bb-b5d3-4bc5-87e5-ac11f57d5a76 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Video instruction tuning with synthetic data, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.454622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.454622Z digest=sha256:bf23ee2f8fd519a02eecba7cc80c2f7017bbd935f8a41e026be8a1ec2082a8b2

Observation c44d3654-2579-4e20-b1e4-fb6f56135dc3 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.459440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.459440Z digest=sha256:6e08d61a40cac83c9a135860f143264d7009483755efe84f05578e20c684e459

Observation 325ab784-ef84-425f-92c9-a3c7369a4249 · outbound

This paper cites EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.465422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.465422Z digest=sha256:0f39b1d655f3070633d4cd1fce4eba90a5b35b5500b8948ba3842e28d8304597

Observation 43398dab-a72d-4441-b1e2-55da1f3df5d2 · outbound

This paper cites HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.470416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.470416Z digest=sha256:8165faab8842e126ec753fd722427fa7ac6d30c105021407eb4323ff2d72c126

Observation 14f6e097-9e1a-4291-9a8a-0ac7eeac1f9e · outbound

This paper cites Where does it exist: Spatio-temporal video grounding for multi-form sentences.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Where does it exist: Spatio-temporal video grounding for multi-form sentences

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.475449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.475449Z digest=sha256:7efe6059ee17f80ff6637cd77c74f6d38592eb00d571842f77f5e1739681d128

Observation d1af8518-3d95-479e-8906-9f88ae1a770f · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.479853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.479853Z digest=sha256:0e66c46a23a430165679ea2c58d3a1fdeb2877a42405e90098428d2ed22a1e1f

Observation f1a6f1ea-4eb4-4391-b086-a6c76dc81130 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.484457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.484457Z digest=sha256:8020ce150a04f9f5490b46fdb24085a93f7567abc6f9d51cb1faba17b5d772cc

Observation 67cfed3d-5642-4c01-a041-1649df3fb22f · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.489108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.489108Z digest=sha256:b4caae8e5d2dba49c456c418fb783e9c757fae39094a0f85fccfac40dc5fe628

Observation 142ca518-f2e4-4386-9f7b-9f524b8f3bb6 · outbound

This paper cites The Llama 3 Herd of Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.493417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.493417Z digest=sha256:f78308c845531b534110087da5923218dddb59e59480bafa1c6f7f919d772ea5

Observation 92330894-b3d3-4abe-b053-4f782f2550dc · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.498151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.498151Z digest=sha256:02fe447b27ad1e7bb8a3e49b636f5faa00e1900d3a78490e48d018de1a236072

Observation 108f75cf-f89d-4de6-8385-6eba40d76e9c · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.502746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.502746Z digest=sha256:5671c20788ff7b2d8de68eae69c8f5382399fca6387a64b780d23feae2542b97

Observation b337e501-00b9-4dd9-b1d4-deac0b119bbe · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Flamingo: a Visual Language Model for Few-Shot Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.507718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.507718Z digest=sha256:5304e3757132f22f0ab1014a41da15288a5e7a84b80113924a0e8a8e33bc3c64

Observation 27fea407-90ae-4867-a924-29b541458a28 · outbound

This paper cites Vila: On pre- training for visual language models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Vila: On pre- training for visual language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.512477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.512477Z digest=sha256:da47162ef21392c6c7f0ce2237e10cbb265adf58a9ebd12763d2b1d25069c5e5

Observation 830a7ec6-e930-45ed-9711-0f60cdbed056 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.517413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.517413Z digest=sha256:75d259197658a19ebec407e80d779bd54f66abec482d1d84c770e190fa47f573

Observation d3988dee-b4ee-4bf1-ad25-ca63bd140d40 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.522967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.522967Z digest=sha256:6d516222b2ff5ec1603e93842b9613286c15ff2b965034040e962628254b734a

Observation 83d3b4f6-dbe0-4ad2-a728-7944bf161e53 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.527507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.527507Z digest=sha256:8b0232552d5d75bc27ff096624248f2ac1d6b7c76e36a87d9d145282dd6514b4

Observation 15a44779-8703-4016-99c9-c504bc4ec177 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding VideoChat: Chat-Centric Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.531929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.531929Z digest=sha256:df9233416ab8a2b6ccabff5dc5a865d75acbb39946559f850aee5cb74b93c5e2

Observation 6aa03cf5-150d-41a0-bfb9-9eb7bdf42f01 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.536792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.536792Z digest=sha256:fcb5569dd03e92dc3ea23727c009164eaf31e4caac87885a594c9e20ee9910c7

Observation aa8098cc-cf6b-4254-ad36-9324668d7b77 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.541708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.541708Z digest=sha256:12701446c6be9876be86f1f1590e2ecc83767add0224b203186b32eab25fd3d1

Observation c9767c1c-77f6-449e-8e2b-03b2b93abe5d · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.546578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.546578Z digest=sha256:f7dcff7f28cbd154b0af1c511cfe229faebbb5fb992a57d0219fb0e5a80129ee

Observation 59e0e807-d82b-402a-842e-9e2b9a188f77 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.552155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.552155Z digest=sha256:620795f233d66edad75e6c232c91d4ee064d5a483be30a6db1ad2365476cb9d1

Observation cf7b1ea1-0e59-4378-bfe5-814968b28754 · outbound

This paper cites Longvlm: Efficient long video understanding via large language models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Longvlm: Efficient long video understanding via large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.557090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.557090Z digest=sha256:c1c4c169dc72fede56b3f740825e6502aa0931efb432b1b40a56307038ba9ce9

Observation 2e779e7f-4f47-466b-9110-9869485177f9 · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.561597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.561597Z digest=sha256:23d5ed1b21ff4cb4b27e3a464741708ebb2043bcbe094eb507a5cfebfb28d8fa

Observation 3c5dbe1b-4619-4bd7-a977-98491a6aab08 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.566251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.566251Z digest=sha256:038eb6c845bdd02ba7c510215646e0a227071ee90144efbc028da6e28594fa31

Observation 1fab2509-b331-4909-a6cc-0558d4851560 · outbound

This paper cites Anymal: An efficient and scalable any-modality augmented language model.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Anymal: An efficient and scalable any-modality augmented language model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.570808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.570808Z digest=sha256:74e92dc883988a646898461fc542790cae063b7057f7a577dd705f2cd05eea39

Observation 000576c0-9e8e-48ab-b6e4-9586afe5f8b4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.575302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.575302Z digest=sha256:0dc3ef0db7a7a2d3637595e54776d16147e4c53e7a924c5b86b54c81d1c33c18

Observation 0c0bb014-6a46-406a-a79a-3156e534898d · outbound

This paper cites Don't Look Twice: Faster Video Transformers with Run-Length Tokenization.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Don't Look Twice: Faster Video Transformers with Run-Length Tokenization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.579952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.579952Z digest=sha256:379ef89ac17c56c1cc35434f71c93122e75deedc3ef015da346a3b39f69755a0

Observation d1862ffb-7313-4922-bce3-aaba6a4eda55 · outbound

This paper cites Gpt-4v(ision) system card, 2023.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Gpt-4v(ision) system card, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.584792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.584792Z digest=sha256:9c9aa8128adc5bd86e51e54b111c753b9d58a85cc8c01d4ac8327a3372564041

Observation 338771fa-3acd-446f-bf72-03e0d55cbdb9 · outbound

This paper cites Gpt-4o system card, 2024.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Gpt-4o system card, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.590068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.590068Z digest=sha256:00d42a91adab0b1f0637d2a1528312948394a23580aa04fabe76599571a0096d

Observation 78f2af48-f529-42b6-9720-0c198ed5e422 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.595350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.595350Z digest=sha256:6e23922550181f3462e2e48451ba2197f5fd3604467b17bf13a45db4d7b98fe1

Observation 940f928e-1b35-4234-b7db-4386b4679e9c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.600023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.600023Z digest=sha256:15a29de53f5c93d4f43fc6f181a17bb192ad73e145f66c70deb847cadf607d2c

Observation 11900aaa-fc6c-4aa7-a5d1-174a2b804596 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding The claude 3 model family: Opus, sonnet, haiku

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.604511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.604511Z digest=sha256:91cdb4bdc7dae7847eeabc4649056ae62062906486e46f4e67b9fb1ab6696270

Observation f7c513db-7dfc-4a03-8fb9-0c44daddaa26 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.609122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.609122Z digest=sha256:6deae271c972c9847b43e02032ffe4d587bc247ccc92c4e778b4fdb83e4c3e5a

Observation 1480b762-335e-4750-93ba-9af59537700a · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.613585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.613585Z digest=sha256:0a13dd39e28c15aa7fe6eddc761db1932a4c49a287306e0cd5b82a81970e1a2a

Observation a66b0859-6aa2-4deb-a6c8-7dbe8c9448bb · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge, 2022.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding A-okvqa: A benchmark for visual question answering using world knowledge, 2022

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.618270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.618270Z digest=sha256:688481e9a5a9d960b6f8e4a6cfdef69753b887b9562d21d6f1cfce846814a301

Observation b8a069c1-855d-4703-92cf-de4b0b2154bd · outbound

This paper cites Vizwiz: nearly real-time answers to visual questions.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Vizwiz: nearly real-time answers to visual questions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.623425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.623425Z digest=sha256:ace4d5d4df7b7626ef7d551c898dbb25fc9ff698e431dbf6b2e3afe9412cc4f6

Observation fec5040a-9dd1-4619-b2ac-7cee39082a1e · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.628596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.628596Z digest=sha256:4628455ca92146667a31f6d8ff0877f42bae83592bbc810e439085c14d51e4c6

Observation f68a8351-1612-4a5a-87ae-24d735c84c96 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.633434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.633434Z digest=sha256:9a372c4d2e566fafca2cebb7f98b58ba21a6d5db99f8b1acb40477b310034d1c

Observation c5d40bcc-f888-44b8-b6f0-c21f3066f810 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.638288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.638288Z digest=sha256:2b37a6e446dc25af866e4c5d26ba6b4b3486b2c14ca93cad15a29ca425f95236

Observation 5b666324-d314-4bcb-9635-c0a985489d67 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Blink: Multimodal large language models can see but not perceive

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.643074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.643074Z digest=sha256:2bc5a6c468583019563f1a371b49a9cacaeeb07aa2f1fee5aeb2dbb920dbdecc

Observation 524002db-2b40-4b4e-9d14-50f16a6f5e8a · outbound

This paper cites Realworldqa benchmark.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Realworldqa benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.648271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.648271Z digest=sha256:a80986354a9dd119e709239848b79775646e7fab7b6e5e535aab11346a5d4e6f

Observation 11cfde4b-e409-404d-a7fc-9fa47b040793 · outbound

This paper cites WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.653521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.653521Z digest=sha256:3e54b24a9c9a60d2ba3ad930c8df7c17af2d3f982bd86741106e59d434f11526

Observation c2c25d4d-1d58-473a-98bd-914753cb556f · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.658561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.658561Z digest=sha256:6eda6f9cc7ea573dba606022ef123f5c31ea1230accd4dda6822ab35d5f9e1fe

Observation 369c0858-dd28-4ac2-9d44-11c325b56693 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.663697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.663697Z digest=sha256:53b0bb9f41ae705f379f6c380945eb58b6921f00baf4e3dbcb6db58cce6c551d

Observation 3f6672c5-1ff3-480b-a74f-fa1c8e916947 · outbound

This paper cites Microsoft coco: Common objects in context.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Microsoft coco: Common objects in context

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.668535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.668535Z digest=sha256:b8028cadd20e503db9bd5aaf67ee986bc03160b6d789d7dd96b47eaa5e565c50

Observation e1506e30-27c7-4c7b-bfb5-cc69f77f826d · outbound

This paper cites Nocaps: Novel object captioning at scale.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Nocaps: Novel object captioning at scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.673236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.673236Z digest=sha256:67086d68df5eb6dd9755d7b9223c15c1cfbff208a6b197994ae00de579111d55

Observation a73ebd77-8a9d-4c79-99f3-48a282097279 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.677990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.677990Z digest=sha256:c86893f18e42e707e2a6929dfeba00f782c9bf3398d4f879dd69d99562cc8cef

Observation c714fac2-7c3c-4f18-bb3e-05cdaa8dee75 · outbound

This paper cites Towards vqa models that can read.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Towards vqa models that can read

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.682652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.682652Z digest=sha256:21ffabdf7b9596dfd9a0beaf363f348de03c95d8d8d1e54736a550adc5feb26f

Observation 0d7a51e6-8026-4ef4-8dc2-1d0433f71800 · outbound

This paper cites an unresolved cited work.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.687155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.687155Z digest=sha256:2e605a78c6e1a4f397f9604e041601fb197e0efa3b2de3911b001b7237c4ef30

Observation 4c43630a-20ec-4d84-a900-3e0e80d9c64f · outbound

This paper cites Advancing Chart Question Answering with Robust Chart Component Recognition.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Advancing Chart Question Answering with Robust Chart Component Recognition

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:19:38.537693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T12:19:36.691634Z digest=sha256:292848f1169de37bd08580388581ca688047d8ee33717f81c283d18e4f2c9cdb

Observation de527b50-6dbc-42b0-b065-316309f9b995 · outbound

This paper cites A diagram is worth a dozen images, 2016.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding A diagram is worth a dozen images, 2016

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.698287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.698287Z digest=sha256:25e1a6bdd4732ec082ad461897a2d6667d07c8380a56e996a2f595d29c40f6e2

Observation 06f64dd8-f3d2-46e3-923f-499ee0a04cf9 · outbound

This paper cites an unresolved cited work.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.702889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.702889Z digest=sha256:4373785be37dc9b49303e528531918309c00c376db159fafe60e31af1b357d66

Observation 9be40560-095c-491e-9eea-0e5f9b3b3599 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.707440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.707440Z digest=sha256:8f446b9452b72f963b212b9420e61ccbdf04153e0d055802807614bfcbbcf925

Observation 150637a6-f450-4ba0-88be-2bbf45db1a25 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.712171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.712171Z digest=sha256:22398e420de6971994420269b23514aefb2181878597cd94c31c027a7e892e3b

Observation 6bd37f51-f0fd-4f8b-8147-8f7e4cecb931 · outbound

This paper cites CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.716754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.716754Z digest=sha256:8179bd53effdeede68567647fa768d2778d54b9ad8ab03f948cb80b64a081b92

Observation aceb6ef3-34d4-49b2-8196-2d1aea9d99f3 · outbound

This paper cites From recognition to cognition: Visual commonsense reasoning.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding From recognition to cognition: Visual commonsense reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.722779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.722779Z digest=sha256:5b34cdc4c9d7925595779877292023d617463c6614385441845753c607327874

Observation b70de624-4752-41e8-b45b-e32d1a8f81e0 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.727449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.727449Z digest=sha256:c245226fe8a179d73df706562594998810960ab6db7ccb7d788a877ee9d78d31

Observation 4246225f-b79b-4929-995e-e7f9123c6bc3 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186, 2025.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.732217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.732217Z digest=sha256:215febed242ddba592c8e5071286e4ad8eac575cbeab25b68ebe59adb2a97f97

Observation 1f3ceeac-ffca-461d-8a67-d4095ac5b3b1 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.737233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.737233Z digest=sha256:83c5bb62315f386a54d8ad47f03f17e29fe6d8b4c46daafa2b9c1c22562a6429

Observation 2ed44cf2-e708-48a7-b235-535dcb997b59 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.742450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.742450Z digest=sha256:6c85457c3e46ecdf3281f671d30bb9fae05237847ea176a687a15005c8015a52

Observation 750f8264-b226-4935-b8ae-ca15a37c60bd · outbound

This paper cites Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.748237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.748237Z digest=sha256:6c42387dbbe71b958382270c206739f69b7b5fcc89478edf3bbb9a9402dfef22

Observation 8f746fb6-72f7-48a5-9544-af59859f16b3 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.752963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.752963Z digest=sha256:b6cedcc8f7edfd60ed904854fede9fa625fb415a14d97024bda02bb45edb3cb1

Observation 80f32b7c-508a-4a05-88fe-ca0109527ca1 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.757707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.757707Z digest=sha256:c5248194da9b01efc6e38a48bab0b48a56b4f8a36562973b21c745f0bf008d5b

Observation 5822dbea-4472-43ad-af67-ebd6fc50fbd3 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Evaluating Object Hallucination in Large Vision-Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.762319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.762319Z digest=sha256:05a9dfe3e86ddef8ee537df73a0af32c04ef0661af41b68fed6032d16f3a9453

Observation a92d91a9-eed8-4b6e-bb55-8b476e87e987 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Next-qa: Next phase of question-answering to explaining temporal actions

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.767843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.767843Z digest=sha256:d96dacd02a29eea6e0ccb1d12eb314f59774832b897c837196ea7d767d14d310

Observation 13d48eb5-8296-4901-9411-78349347a32b · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.773067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.773067Z digest=sha256:4ecf187469334455e30d18e1d1bb1cf9900678fdbe9a97974e25830aa87b7ea0

Observation 9d399e52-fd78-4879-af69-ec2ca6b00792 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Perception test: A diagnostic benchmark for multimodal video models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.777598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.777598Z digest=sha256:d7e4af01796eb34a6d08aeddd23eff69c3a872c8c0eeb29f67f0e41c293bdbc7

Observation b6c9ba23-dfd0-416a-86bc-897eb47a4de9 · outbound

This paper cites Star: A benchmark for situated reasoning in real-world videos.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Star: A benchmark for situated reasoning in real-world videos

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.782318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.782318Z digest=sha256:d8259b313c1ebf3e7be8d31df067e6edb71f4ca0766405a9207c0197293ae933

Observation c749c94c-1f2b-4c7f-9d96-db90b140502c · outbound

This paper cites Tgif-qa: Toward spatio- temporal reasoning in visual question answering.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Tgif-qa: Toward spatio- temporal reasoning in visual question answering

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.786825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.786825Z digest=sha256:f04e31be5f17bf9a0de5234783a6849b3db96654287e8173b3a4a07227386f99

Observation 950d872f-70aa-491e-ab8b-9a6241880bdd · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding TVQA: Localized, Compositional Video Question Answering

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.791298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.791298Z digest=sha256:7eed75fb03f770338ac1b91da2352646656240223464d0f0dca029d693e4b277

Observation a16e8d0c-ca51-4034-9d34-8f68480ce1d5 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.796176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.796176Z digest=sha256:b9eb5968331223b076f77a1cfafe090c4b14f46b5c21ae5cbf060a7a4a5e2650

Observation a8843b2a-2896-4094-b6ea-0a3e5c1062b1 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.801127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.801127Z digest=sha256:7c895b084e2f2510214abe4cf976f1fbea877cd2d877181fb3e30d0710ccf1f1

Observation 99046912-b16a-4197-9d4e-81770bcf636e · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.805570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.805570Z digest=sha256:dc550cf17e8f4972995f5baa58b465b68a1d2eb6e9d3ebe0e5126b64f87a24ee

Observation a6f04011-2e93-4114-a36c-edf06bd2dc7b · outbound

This paper cites Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.810695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.810695Z digest=sha256:22f35ca45055343cee46bf02c6b71d48a43f673dca677b85181ff9f473b9b72b

Observation bab3da26-6402-45d4-aecf-bf17ed42d31b · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.815801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.815801Z digest=sha256:a7b11e444034e8c9652d3db4f396c235c988459054fe849ffaa360e07ae55511

Observation eb4cdb90-95e4-4514-9f5a-af4e447517b5 · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.820475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.820475Z digest=sha256:94e139ab1128aaec4fb409a9425d3ba44472acef9bbeded5f479333922d296c4

Observation e3cb6c5c-ddec-4d68-8ebf-7798c51c6767 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Msr-vtt: A large video description dataset for bridging video and language

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.825402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.825402Z digest=sha256:73d1fb9421e116ce9d0e2f1aea12efcea19ae87c61443c33e4fcfc06f6291b0d

Observation 14823033-ad53-4c92-9094-18d84c93a1ae · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Collecting highly parallel data for paraphrase evaluation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.830109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.830109Z digest=sha256:41d7bc1977f43e4354359783f0210fd4c44e8d7175cbe576f17fc0e3f00560e2

Observation ab77670a-6332-4abf-91d1-a345d88ece21 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Towards automatic learning of procedures from web instructional videos

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.834759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.834759Z digest=sha256:3316f507b59a796e519b096968e8ebe48d163dac499ca3dfa069a73ac96941f8

Observation d279b101-bf4c-4bd8-94ac-18a0610e604d · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Vatex: A large-scale, high-quality multilingual dataset for video-and-language research

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.839341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.839341Z digest=sha256:0e07d4a5eb75121f684a91a0b0fb50f36667098504027334da938018ba25c641

Observation c503fa1c-0df3-4971-b880-a5973c11678d · outbound

This paper cites Dense-captioning events in videos.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Dense-captioning events in videos

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.843968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.843968Z digest=sha256:145092fe01836fc46b9a11095db6b7327c816bfd8cb6d696384a58d067559d3d

Observation 5bc87307-2b8a-4df9-b86e-d9a42f27b391 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.848557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.848557Z digest=sha256:85e5057bb9a5eb4241b28306b0f81beb2d26f8d421cf9ab4bc307c27f38a1dad

Observation 24faaf20-db2e-40d4-8566-e87b52a3e902 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.853647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.853647Z digest=sha256:f39b83f5e38d69cca8184a1741f05a90d354bde240ed47f94afc423ed16a2ea4

Observation 39150a31-3698-4398-b3eb-dc88c042dacf · outbound

This paper cites VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.859384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.859384Z digest=sha256:94e96744bb554a3b352a63c1dd07515bb1b2a4288cba94c045b0bfd7dca06a04

Observation ec2fe659-1e9f-4afc-8c91-c673d2b17594 · outbound

This paper cites Eventhallusion: Diagnosing event hallucinations in video llms.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Eventhallusion: Diagnosing event hallucinations in video llms

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.864178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.864178Z digest=sha256:31725d4352b1f1a71a81cd7edafc1303796a395e068e28c46c343bafd3193ee7

Observation 22f2f325-5907-4713-8906-369d2cdd9d25 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.868641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.868641Z digest=sha256:a4a70bae27ea6f560446592ce752fca76ae39e64925a86bf852f480c570e94ea

Observation 5d9bd916-b723-42d8-add9-51debe881cef · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.873240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.873240Z digest=sha256:2ee9ce718b636bc8d62832e4f47ff8fef3ea3c7071a4a289663b81473014cd5b

Observation 8f7aaed4-1cf8-4b8d-8b51-692555e12207 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.877967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.877967Z digest=sha256:d68268464952a9f7d151b3a1b808d6fa92b8fe043776bfd8abd796dfcd4685fa

Observation da3c861c-1a44-475d-9f38-e90afe065029 · outbound

This paper cites Movieqa: Understanding stories in movies through question-answering.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Movieqa: Understanding stories in movies through question-answering

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.882703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.882703Z digest=sha256:d9450b14533575f01b98d7797c47a82bd97350a34ecb6af45465d2f2b005f7ea

Observation 2c6d55e2-6d24-42ea-b596-0d77650b80a5 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828– 28857, 2025.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828– 28857, 2025

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.887130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.887130Z digest=sha256:a8aa37c2644c67133668e36db4c382f9e9099b06c6a81ba3fa6a91d11a6b046f

Observation b0cc2070-28b6-4b01-a9d2-7f5f2fd686ac · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Moviechat: From dense token to sparse memory for long video understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.892235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.892235Z digest=sha256:06eec7ccd73b1aea8fe098d45b463f9db2f66ac364c486f79898297e0921b667

Observation 80e9632d-b4d5-4aaa-9e94-687a4322bdcf · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.897692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.897692Z digest=sha256:c9a53fe165692705b2e5491b7f5c771096bbed8a15e3e6e12b43b2eec9892505

Observation e4cd5b5c-dcb8-464a-b26e-baa7c1ec1267 · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.902360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.902360Z digest=sha256:c5968aba4838f566b7352fe00eaad29290eed65239a60ee3444d3d990edda673

Observation d565190c-c7ad-4575-ad02-0d6476b5a4a0 · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.907103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.907103Z digest=sha256:a1ca36f6bb2bfc0e69390c088883bcc7c71b03c52afa687e89ab8529ee00ee01

Observation 9658c1c3-82cd-45bc-8251-b58c942b3496 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.911916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.911916Z digest=sha256:1d5771389c7f5229440ac2a3b0383447d2b7c8d7689df60802c459966f8be0ab

Observation 37b008b2-5910-4ab5-a082-1123a2b09afa · outbound

This paper cites TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.917814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.917814Z digest=sha256:eaec2f23fb5818d64a38b2bb8760d7bd6752181c0c81a2a8aaf9507535a7e868

Pith citing papers

Observation f51e9682-09fb-47b3-8924-32619f16dc8f · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:15.817065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:87d359f69c75a1ac5fe102e5cf87304d4e0191e79e09e49afe11550ee3d1e66a

Observation 36bbce6e-2e67-485a-90eb-aae6f5bd0c0e · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.240386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.240386Z digest=sha256:23a251fd51accfd675191a17e2417fb96857dfa42d5a7a17cec9b3431be9e298

Observation feb10324-1938-47ff-a4ff-fc8f6981218f · inbound

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models cites this paper.

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:44.870369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:44.870369Z digest=sha256:5784f13b509f3bc51b07a7a9e85ca37c7688e335f04cc74238a9dba28742f613

Observation ab285409-aff3-450c-b2d9-708aa346803f · inbound

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models cites this paper.

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:59.009056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:59.009056Z digest=sha256:2ca856728b585f1b6f175d720fd6bf4e3e417850de78c4007b447c98abe64cfe

Observation 5481e072-022d-4fb6-974c-a1f73ccbf903 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.751197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:8c5b299c2df53912d9cd0d758800dec2def188d9b79201ea5871cb2a29294968

Observation 254383df-acec-4f27-a082-f36dbfe6d5b6 · inbound

Group Relative Augmentation for Data Efficient Action Detection cites this paper.

Group Relative Augmentation for Data Efficient Action Detection PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:55:37.435979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:55:37.435979Z digest=sha256:1176599820a9ae2e3de8a28f9eec1aa7b6e0d8a76fb16d56998d41e2ae2bc96b

Observation 048c6c3c-bdd1-4fc5-86b3-6ebce32f8450 · inbound

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data cites this paper.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.783270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.783270Z digest=sha256:4476cfbc419b4149727057794bb44dd1fd1b6ae891c4733717c3c5fa562adeed

Observation 9b872af4-c4ba-4404-9fd9-9132e9fe6cae · inbound

SAM 3: Segment Anything with Concepts cites this paper.

SAM 3: Segment Anything with Concepts PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.567661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:4bcccaab55b87607a5a01dc2e6989d1a07202c3f0785441283af48c68814d727

Observation 28b1ff73-f9ff-4cb7-8a2a-d35a64972510 · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T13:01:57.744137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:01:57.744137Z digest=sha256:44010d0f0e7b093229bc9ba59db75377105101b4b56be1e4083d8922b524faaf

Observation c87468a0-646d-48e9-baee-b10aa5b2c2b6 · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T06:36:33.788629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:36:33.788629Z digest=sha256:a507163d92697c2a61be6632b02ed6fad738e6d8dccc0eb915b6103cd210ee9c

Observation 5638f5c9-e25e-49de-bf42-b0e8b754930e · inbound

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding cites this paper.

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:21:29.849841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T04:21:29.526008Z digest=sha256:f7bc549968f199936af4489bf81db99d6f84722dfd3073a48c144fd33f7c0826

Observation 3ab36bd7-1861-4b62-bfd4-6b1b14d6faa5 · inbound

InstrAct: Towards Action-Centric Understanding in Instructional Videos cites this paper.

InstrAct: Towards Action-Centric Understanding in Instructional Videos PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:35:56.793965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:07:49.334104Z digest=sha256:8bf2023604dc911cb9186b3c4cf9e837fcb88c2ec9a9dea26df87c62f60aef6a

Observation 303a81e6-0a18-414c-a1da-9dcd0f3f3662 · inbound

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs cites this paper.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:02.000635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:f74e00fe41a0312dee5346dcb181dccc767ae4d1d89f9a766606e2d3991ef465

Observation 90b75ce0-2758-4e5f-b709-91725c335095 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.390564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:eaf699232289e981462a3e65f521ef38ee5edccfcba840200bd2cea406210f45

Observation f0a3f67e-6673-468f-b957-ee4e6c81beb8 · inbound

Don't Pause! Every prediction matters in a streaming video cites this paper.

Don't Pause! Every prediction matters in a streaming video PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.822135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T04:32:01.379605Z digest=sha256:462b0c106834068e66a1abe6200f1794c97542b04641b0b53a3332f28ed2dffe

Observation d1335f95-8139-4879-9e57-123cf6c50c59 · inbound

ZAYA1-VL-8B Technical Report cites this paper.

ZAYA1-VL-8B Technical Report PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:23.368911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:15:16.607346Z digest=sha256:2478e7e922f0ed7a418faa05666c42bf32c34d598de3d205c52136990e18aa8c

Observation 8542f057-34bb-4891-9e1f-a8d89a7d3f16 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 155

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:28.521335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:3abda8ca59f3d5c1462056abf824747b0c33fdffcfd57a4e5dee54d7afff7b10

Observation a1e13608-5d35-46fa-88cf-ec4009c487b9 · inbound

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly cites this paper.

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:20.656335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T09:20:32.920925Z digest=sha256:2bab41162defb1a20af984327471cfbbd63d2d9c70782cda155c879cfa4e83e0

Observation 883c0a7a-fb6d-484a-9887-2490acff76e2 · inbound

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning cites this paper.

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:46:14.843460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T07:45:56.473188Z digest=sha256:df8a9db3a016bef75aa63900361e3b83b54726cda2de2a8ee734c1f8653ca801

Observation fcf13a7e-7ada-4806-b127-201796c1e010 · inbound

Zamba2-VL Technical Report cites this paper.

Zamba2-VL Technical Report PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.682259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:34:20.970856Z digest=sha256:c113c95ef8988fc50cc17bbebf5b4507ee391a8fa63d9471a6e7b30fa64b047c

Observation df891798-c9bd-4693-95c3-3f58cc7d1c13 · inbound

AdaCodec: A Predictive Visual Code for Video MLLMs cites this paper.

AdaCodec: A Predictive Visual Code for Video MLLMs PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-06-28T15:22:19.555545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T15:20:48.248576Z digest=sha256:27e312db8e2764fa143ccea92733d054cee4ff32481de4dfc36b8adf8c049027

Observation 1765e026-7c55-4e84-8cd2-5ce4315e9981 · inbound

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model cites this paper.

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:12.972644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T21:22:53.792778Z digest=sha256:b6dd1ea77a24b66cccb999f59b512b48ffc8d28f416f260be95849fc900af619

Observation 84eb08e2-7c7b-4151-adec-6a76caa51cbb · inbound

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model cites this paper.

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:28.798366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:34:37.974692Z digest=sha256:de4883a66fd6e5360ee2978b843ed41916230507b167e14fee077a893903786c

Observation 256df914-618f-4b03-9e49-28a8ceef1490 · inbound

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model cites this paper.

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:25.000055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T22:15:14.561596Z digest=sha256:5021026b0255016b3b2353d18da824c5f0878638fb6b1dad96eeceacec1403c9

Observation 73f2b3c5-bdef-455f-ac15-b2fe8c5fe847 · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:59:58.135042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:4b4912a705d32131bf843162ea21eee82d6f04510c94d0da25911714dae2053e

Observation 479012f2-b78e-486c-b73e-fc149cfee7d5 · inbound

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity cites this paper.

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:50:10.224238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-25T21:02:44.441202Z digest=sha256:921b957f04c5c6fbc68af8118373d9fd90d255695b95954ab72d22e11346006d

Observation 9bc7a457-a71c-40a0-b29b-1d5a706ad9e7 · inbound

Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs cites this paper.

Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:50:11.284082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-25T20:58:01.925342Z digest=sha256:1ef87f3d7ce547b5301b6589b14b26377f698e6a7a5ee06265dcdec235c1e2d4

Observation 070e822c-dcc7-4a43-8b9f-8a33109367ae · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.690264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:a295bce127c4ae877664a6f17773f75026e1f683dc1c68c89cd4888f3d4a3f1f

Observation 115430e1-3ed7-4c6d-be2e-554e0056aa28 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.161656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:bb144347bd256ab24afc884ab67f3cfe28177d952624f49af6e48fb0b3f12790

Observation 22c24511-54df-4f63-ba2a-427fe098fd85 · inbound

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context cites this paper.

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:04:21.260794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T06:01:49.904752Z digest=sha256:b9807dd881d9c649057ae395d46589ec9d7c2107fdbcd8ca5b5c7b39f8a66e0f

Observation 0d9a6198-a0f5-4191-9a08-107af9fc74e1 · inbound

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models cites this paper.

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:46:58.655990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T13:44:48.250838Z digest=sha256:3aa07f150375e9c35975cd15f3dd0499884ad95f5288fdde937d51426b633b33

Observation e667046c-a725-4642-b587-c2328b08047c · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:40.660990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:40.660990Z digest=sha256:5b6bac8e0bc40a706dc6eca6295ddb7d9852f63d2a32710db258b492148d165d