Pith. sign in

Paper Citation Record · LEDGER

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

As of 7 August 2026, this Paper Citation Record lists 100 of 127 outbound references and 3 inbound Pith citation observations for arXiv:2511.20272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.20272 v2

Coverage vector

measured 100 of 127 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:23:07.705686Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T16:37:06.384435Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:38:39.619743Z

Reference resolution

100 of 127 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c2785ed-4cd7-488a-8e38-8df42133fc3d · outbound

This paper cites GPT-4 Technical Report.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T20:22:59.834670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:22:59.834670Z digest=sha256:788785f7b491362ab7462b05bd2335a0ed77a6905001ae911f8214ac1932babf

Observation 0e4113ed-1a88-4f11-bed2-153b4d47d61b · outbound

This paper cites On seeing stuff: The perception of materials by humans and machines.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs On seeing stuff: The perception of materials by humans and machines

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:22:59.937363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:22:59.937363Z digest=sha256:ca5ec5140279745a186ca068897409b42e8d03d2203d1fabe8fe6a0278b10d4a

Observation 40d649dd-1e16-41bd-ad37-01207a4f42b5 · outbound

This paper cites Vqa: Visual question answering.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Vqa: Visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.093535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.093535Z digest=sha256:de40dc94569fadecc468f98946a1ff0d2c940996a2225478870c332387cd398d

Observation d19b8ecb-31a9-4742-8a73-f796e7c8d198 · outbound

This paper cites Eye-contact, distance and affiliation.Sociometry, pages 289–304, 1965.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Eye-contact, distance and affiliation.Sociometry, pages 289–304, 1965

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.331240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.331240Z digest=sha256:e9e9cb224986dc6444d926932e83cfb2d1c724f07f2c61d7866c5ff8e719e850

Observation 14fda1d8-9815-4d28-9671-ecacd20eb630 · outbound

This paper cites Qwen2.5-VL Technical Report.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.406527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.406527Z digest=sha256:20856d29699a0884425b9c820fd39530194df4c44409b2134bbd3869174ebf00

Observation 1dae183b-ab4f-4652-9b30-8dcee8737c6d · outbound

This paper cites Representing the existence and the lo- cation of hidden objects: Object permanence in 6-and 8- month-old infants.Cognition, 23(1):21–41, 1986.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Representing the existence and the lo- cation of hidden objects: Object permanence in 6-and 8- month-old infants.Cognition, 23(1):21–41, 1986

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.484396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.484396Z digest=sha256:02bdb770025de713deae48ae7f1581deb57b70c9465fcb746fd74926141d2807

Observation c53a5db6-45f7-4cc0-8170-3321c28fdfa7 · outbound

This paper cites theory of mind.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs theory of mind

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.572585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.572585Z digest=sha256:56791f4c7389c014e7c5976db3bd6eccc3ac49a3dc16549e5b859413a1d679b1

Observation 858ff559-25c1-4076-bb04-fd8184e2384b · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.688451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.688451Z digest=sha256:888c0a72866a78e0bf7bac3a77dc47e074b9ed965fd27be4b9facdef752aedd7

Observation 98860f4a-b3b8-4ce3-91d0-263f87216e4c · outbound

This paper cites IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.771656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.771656Z digest=sha256:0e8d1a93ed1dc7b8b3e23c534e33f43aa13f7a7098734fffb06c00a9f11d8b0f

Observation 0d473e7b-2837-4ab2-b10e-58c29a3aeabb · outbound

This paper cites Routledge, 1995.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Routledge, 1995

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.854703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.854703Z digest=sha256:ed8bc2103f99eb218884f351208eff30835ee864391f3c028908024872ebd7ec

Observation 0c586aa5-8c9d-4252-b923-e4201d3f038e · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Activitynet: A large-scale video benchmark for human activity understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.937721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.937721Z digest=sha256:b66446683cc7c2da8bcf4f916fe900757715cbde602593aa54a5241081e8e711

Observation 4715a1fd-4758-4444-81f4-330e1b44d5c2 · outbound

This paper cites Rextime: A benchmark suite for reasoning-across-time in videos.Advances in Neural In- formation Processing Systems, 37:28662–28673, 2024.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Rextime: A benchmark suite for reasoning-across-time in videos.Advances in Neural In- formation Processing Systems, 37:28662–28673, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.971832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.971832Z digest=sha256:17603e8c7e0b3043a098ace0c12b87488a1325f29206e28c95f221fe8da6a018

Observation f61da4e4-5d2f-474c-8887-1cd84e193873 · outbound

This paper cites Are we on the right way for evaluating large vision-language models?Advances in Neural Infor- mation Processing Systems, 37:27056–27087, 2024.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Are we on the right way for evaluating large vision-language models?Advances in Neural Infor- mation Processing Systems, 37:27056–27087, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.064028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.064028Z digest=sha256:af197d59c9057b2e55ce99517b5a673e051ec397a94dee2330711103b78efe04

Observation 96d07928-9fd3-4a1b-a2cd-89e6c7c79f24 · outbound

This paper cites Video-holmes: Can mllm think like holmes for complex video reasoning?, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Video-holmes: Can mllm think like holmes for complex video reasoning?, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.175032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.175032Z digest=sha256:fb63c040543e5c3b69bbacee262a68373179640aaa0a4052b469b9465bf826d3

Observation ed3b046b-4305-4e84-b01a-dee2201bcd7b · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.290978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.290978Z digest=sha256:eb0f411850b946896f43741481810c239672b0b797710a091657c033ca15ac04

Observation cbe8be57-de9b-4f62-8e83-9a4890fa0284 · outbound

This paper cites Le, Sergey Levine, and Yi Ma.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Le, Sergey Levine, and Yi Ma

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.403982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.403982Z digest=sha256:605960219b585d889271097a0e03be76c99d42eee21cde2e770030ff5069dc7e

Observation f2bc5c7d-9545-4df6-9dde-7ab60c9cf8a1 · outbound

This paper cites Whatever next? predictive brains, situated agents, and the future of cognitive science.Behavioral and brain sciences, 36(3):181–204, 2013.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Whatever next? predictive brains, situated agents, and the future of cognitive science.Behavioral and brain sciences, 36(3):181–204, 2013

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.472386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.472386Z digest=sha256:2329393da2d11c2e934a0f351b04d621b1120a3cfb96348c2bb50651e51a9adf

Observation 7e7f228a-2550-4553-9916-365f8ead4e66 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.531512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.531512Z digest=sha256:20aa90499f01915f0e50c47d92fe5f9ccddacf831293fd2f5b729b5a1c5ef5fe

Observation 5223b74d-761a-450a-9ae5-af8fed1c7b9a · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.579499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.579499Z digest=sha256:7950af9bcd56842874cac3490d75a0dcd45772f334312238bd34853cc729e332

Observation ca6d9e86-e844-4883-b7c6-7f82d5ec0538 · outbound

This paper cites MIT press, 1987.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs MIT press, 1987

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.643970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.643970Z digest=sha256:1109b6a64bf326b1270fe95fc5d15e20692e1d6bbbb646b000ea532aaaf606db

Observation 35bb7093-aff4-45a3-ba58-47c1837078ba · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.727373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.727373Z digest=sha256:6153b63cc9d32cd75a540a7e4b8e323ed4cec0c906808000186eefd6f01825c5

Observation 049c5852-561a-4b6b-9c0a-9ad9ce89d05c · outbound

This paper cites An argument for basic emotions.Cognition & emotion, 6(3-4):169–200, 1992.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs An argument for basic emotions.Cognition & emotion, 6(3-4):169–200, 1992

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.812885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.812885Z digest=sha256:e2809dfd2f276c3d4cf183e2623f8d0053f8fac3450041b8a2e9dec441d6ccb7

Observation 15443549-ad36-4cf8-a9df-f10b57c5db37 · outbound

This paper cites Unreal engine.https : / / www.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Unreal engine.https : / / www

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.869913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.869913Z digest=sha256:8c418cee0987eb7b87ef69fed6b226581676c6bb598409c1044a924d23b4c849

Observation 0141789d-b487-475d-867e-941fac8706f2 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.923190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.923190Z digest=sha256:e0296bc96cce7f4d255d54a5bac99383f372914354cd2b70a2a98b2dd7d5d914

Observation df25d2c5-5169-468f-a084-b69b1fc46b41 · outbound

This paper cites Material perception.Annual review of vision science, 3:365–388, 2017.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Material perception.Annual review of vision science, 3:365–388, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.003342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.003342Z digest=sha256:9afe46d8ebbf03d2ffbd6fae57219759ea97f4d52cf4ac21a0e40316ce6c39ae

Observation 6163625b-6a14-42e8-8f7b-afdba3cc4faf · outbound

This paper cites The free-energy principle: a unified brain the- ory?Nature reviews neuroscience, 11(2):127–138, 2010.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The free-energy principle: a unified brain the- ory?Nature reviews neuroscience, 11(2):127–138, 2010

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.033703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.033703Z digest=sha256:3ece94ac41af0c4ce48e486f8215b8c581960e2354fd4e491dc78179c4a6fe4c

Observation 9948e519-a3c9-4ab7-a4cc-a60d18381849 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.091017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.091017Z digest=sha256:7a4642f438e318ee661b98a724b49eaefac71e5bacb7c9909fa450ddc574f512

Observation 2f5062b3-88b0-4b28-bca1-026698eb8739 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Blink: Multimodal large language models can see but not perceive

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.138277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.138277Z digest=sha256:2beac662149d7ef6b327f8edb7866304bf236d257e46533e2d8001f7615d3425

Observation a30fc0d5-5e85-45b8-9ccc-84d95dcfeaf0 · outbound

This paper cites Atomistic Control in Molecular Beam Epitaxy Growth of Intrinsic Magnetic Topological Insulator MnBi2Te4.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Atomistic Control in Molecular Beam Epitaxy Growth of Intrinsic Magnetic Topological Insulator MnBi2Te4

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.182852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.182852Z digest=sha256:70cf55ab105321ec79422ab1a8710a4cad2e8ddff6e29a6d4abae3d790929bdf

Observation b3b05772-dd96-409b-b6e7-3cbca0d93443 · outbound

This paper cites Houghton Mifflin, 1979.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Houghton Mifflin, 1979

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.249905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.249905Z digest=sha256:773f926c502220dd0f066c9a55f06a553f840036279a97a076131ed0bfdae80a

Observation 37ef0134-653e-4edd-bef5-ff46717a1752 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The” something something” video database for learning and evaluating visual common sense

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.321899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.321899Z digest=sha256:8041c552230a745de1c42b08fd60265fb77a0b64df9876eb093260060b9099f0

Observation f3e67a5c-82eb-49a4-b17f-4932b2dacfb9 · outbound

This paper cites The Llama 3 Herd of Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The Llama 3 Herd of Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.394429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.394429Z digest=sha256:85239bc26744f0d6db38fc592e4d7139a4ccf50b41fa2fb57640c49a14939257

Observation 0cba16a0-eb5e-4666-88cf-e2a04c2e0cf4 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language halluci- nation and visual illusion in large vision-language models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Hallusionbench: an advanced diagnostic suite for entangled language halluci- nation and visual illusion in large vision-language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.480865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.480865Z digest=sha256:3616734bf6a5fb98ced78b375d22c38e65fb4e1bb3d83fcf7ba10cbf242f9667

Observation 6411cd8f-eded-43a2-b78b-3cae95e7e5d4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.549143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.549143Z digest=sha256:b9e1866bb0df94fe11473a63424918316e561cb9265267564959688d18bb316b

Observation dc1af7d2-65d2-48c4-8ed2-51d6db7532a2 · outbound

This paper cites Doubleday, 1966.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Doubleday, 1966

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.624952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.624952Z digest=sha256:dea90a75d15c7805b396ac07c3e6cf0d653fe6ff264bcb63a2f101fa9041f079

Observation 10724494-018f-430d-935e-7415a1293efd · outbound

This paper cites Glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.726794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.726794Z digest=sha256:e45425ceb5cec00af2bc77c760c5da0465668a5bb0bf1d034b7821873d68e5bb

Observation f56e9cb6-14ad-429a-82cd-1483b68685bf · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.778666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.778666Z digest=sha256:6575a0ff9ff8efa9a83c3e0e64b8ea5606ec559475ad9b2dfe074ee7731151c2

Observation 5f8a72b2-eb6f-47ee-9d55-dd7ae9dce488 · outbound

This paper cites GPT-4o System Card.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs GPT-4o System Card

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.877204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.877204Z digest=sha256:75aa85a4e085f26da0c90cf7a2d541d190b3223d8f5c26d4d49a039d8ee997df

Observation 678d0199-f37c-47c2-98c6-88ad754fba18 · outbound

This paper cites an unresolved cited work.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.044818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.044818Z digest=sha256:252ab99771298aed57954856a4761e58ae9b52c7bfa76a4120359260c41c2098

Observation da0fb4f2-cd52-41d2-a5f9-f584920f3438 · outbound

This paper cites From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.179935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.179935Z digest=sha256:a3a98c58248b29e9a8df26cc9c3f982a43aa8fb4aef9c721d78cfad7a5f5a720

Observation f9851049-ad81-41e2-94d3-c00edd3d4bf3 · outbound

This paper cites Towards Social AI: A Survey on Understanding Social Interactions.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Towards Social AI: A Survey on Understanding Social Interactions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.378338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.378338Z digest=sha256:8ea55caa671692f249782321d6a03f108237e30604f52fab26cc238d03c7b60d

Observation d511dc7e-7895-4907-9c2b-c3b73d49c0dd · outbound

This paper cites What is More Likely to Happen Next? Video-and-Language Future Event Prediction.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs What is More Likely to Happen Next? Video-and-Language Future Event Prediction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.468092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.468092Z digest=sha256:90693e978f9a5cf9f69fe74400feacf85050dc51d538febd9d33e4d3da9ebba2

Observation 6f908d3f-9d90-465a-9c6f-a5905c14acff · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Detecting mo- ments and highlights in videos via natural language queries

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.619120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.619120Z digest=sha256:ec59d0de1d705d8d7e5b76cba3ea8c2ee42743455481bf1c39e787252ea5041d

Observation 224dc2ff-e5db-4249-bfec-2fe6ff741b00 · outbound

This paper cites Anisotropic flow, flow fluctuation and flow decorrelation in relativistic heavy-ion collisions: the roles of sub-nucleon structure and shear viscosity.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Anisotropic flow, flow fluctuation and flow decorrelation in relativistic heavy-ion collisions: the roles of sub-nucleon structure and shear viscosity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.810455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.810455Z digest=sha256:c1933f3d5181ac0c8afe17781d74c7a0568808be6900604a909ec3d46cca6cdd

Observation fd73e304-7c6a-4557-b36c-1e76ed112956 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.031544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.031544Z digest=sha256:af091631bfa28156ec333417d67f7cbbc699521540d45e0672002f9ff2cccf55

Observation cf5d0930-c3d0-443f-a63d-92be2ecef23d · outbound

This paper cites From represen- tation to reasoning: Towards both evidence and common- sense reasoning for video question-answering.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs From represen- tation to reasoning: Towards both evidence and common- sense reasoning for video question-answering

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.223426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.223426Z digest=sha256:d4fb3849a83397769d6d8890be538d4886e39c0a882d0af9e62be8feb4961080

Observation f72b758d-fcdf-49d2-b4ba-33bbc058f01c · outbound

This paper cites Mvbench: A comprehensive multi-modal video under- standing benchmark.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mvbench: A comprehensive multi-modal video under- standing benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.327558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.327558Z digest=sha256:083331d2ee58cd260ced3f770926d110faab43e82217f2f92ff101cd4bcca0be

Observation 4fc421a3-3b21-4ed9-bdc0-1e7dec310b92 · outbound

This paper cites Pope: A simple method to hallucination eval- uation in visual question answering.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Pope: A simple method to hallucination eval- uation in visual question answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.422362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.422362Z digest=sha256:02a7fd48c65bdddf341c17c9906f7453517b94b32eb762c9d5b1b3c5ce1656e6

Observation 278caf4c-6ebd-4563-beb2-ea12b21b56cc · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.532245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.532245Z digest=sha256:d57a6ac5458dbb38fccecc0bc41e076c657056038e1c0cffaf410becaea1cd0c

Observation f1a887c0-7353-4b9f-bf56-c24768bc63e2 · outbound

This paper cites Core Knowledge Deficits in Multi-Modal Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Core Knowledge Deficits in Multi-Modal Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.626496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.626496Z digest=sha256:71ca82c168b5f5c586f997f0c902cc3ea19580e154dbe8e9a3a0fc890ac48868

Observation 9ea4dfa3-11d9-4158-9b50-b17e8074ccdc · outbound

This paper cites Self-Rewarding Vision-Language Model via Reasoning Decomposition.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.725008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.725008Z digest=sha256:279da520074439caf6637b65dce71d46b460ba2b207153a168763296a4bd5c7f

Observation a7619fe7-b755-48d1-89b7-3d73dba6fe7b · outbound

This paper cites Explainable Multimodal Emotion Recognition.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Explainable Multimodal Emotion Recognition

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.826214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.826214Z digest=sha256:e5cb33a97bdb8b0861286ba382f321877d8b2058e0870628420e082cb81e1dc9

Observation b4d6e639-687f-48bd-b459-b9d41f707997 · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.926420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.926420Z digest=sha256:85f07ad35da93a6f1085f750f73dc5600695fbd18a1e23ee1c981b46471ecc1e

Observation f150adbc-f5dc-46ac-81f6-0a1fa312d5fb · outbound

This paper cites Microsoft coco: Common objects in context.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Microsoft coco: Common objects in context

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.034079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.034079Z digest=sha256:e9c7eadeeac7bdb6e7510c9d8817821064e83c3c2f80f3475c265c367b3a059d

Observation 9038ad2b-7b01-4e72-b0ee-82aab3cd54fa · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.200466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.200466Z digest=sha256:1766ed487bb03149842a5bfaaf477373998d2d6d27f4a8d35a6450a3cdbbf4b0

Observation c6d62e94-c635-4f87-a6b4-aa202eed3c8b · outbound

This paper cites Generative Physical AI in Vision: A Survey.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Generative Physical AI in Vision: A Survey

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.304254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.304254Z digest=sha256:7bf9a5ad5bf5554a6f68bee2002cf0adff8a64d2a4da8d4a69747c1352f8443b

Observation 4725a73b-cd00-4335-bd38-b554a4a47c85 · outbound

This paper cites Visual instruction tuning.Advances in neural infor- mation processing systems, 36:34892–34916, 2023.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Visual instruction tuning.Advances in neural infor- mation processing systems, 36:34892–34916, 2023

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.407372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.407372Z digest=sha256:576bcea0aca7288e41e7fa73f60ef77c638561d0040fac521fc5410a9051e702

Observation e9aebeb5-ba66-4fb6-b7fc-a4e731d52021 · outbound

This paper cites Improved baselines with visual instruction tuning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Improved baselines with visual instruction tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.487338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.487338Z digest=sha256:da3d1d40dc018735ce7d4f938ad17e31ea6bb1c276558b8a53f3d4cc06a767f8

Observation 4b6cbdf4-27dd-42ad-b025-e4b7e614b806 · outbound

This paper cites Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.553933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.553933Z digest=sha256:d385c43d9a93567bedd969c442f13501b43e08ffb0c0fd1caf2c8f8df18f3e61

Observation 2c559219-0d9f-435a-a826-b1d462cbcf51 · outbound

This paper cites Mmbench: Is your multi- modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mmbench: Is your multi- modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.605866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.605866Z digest=sha256:92b409c11498a38b9c0cf89c12a7569b5c3d76c050409e5aaaac184f430db26a

Observation 75e2c42f-5191-4b9b-bbce-3ebab6794364 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs TempCompass: Do Video LLMs Really Understand Videos?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.644159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.644159Z digest=sha256:418a1079dc7aa83e89256084a256e187e21ff3d7e3c439fbf19ebb7d6553b82e

Observation 0eb26b62-1cfd-4c03-b18b-baa9752989bb · outbound

This paper cites Videoreasonbench: Can mllms perform vision-centric complex video reasoning?arXiv preprint arXiv:2505.23359, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Videoreasonbench: Can mllms perform vision-centric complex video reasoning?arXiv preprint arXiv:2505.23359, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.701446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.701446Z digest=sha256:1b9b437ae9cf3eb101d52c01712037f56a390c5bce4aea3d85ef588916d247f8

Observation 98b90b1a-9759-4ffa-bb92-81e248525fb2 · outbound

This paper cites When thinking drifts: Evidential grounding for robust video reasoning.arXiv preprint arXiv:2510.06077, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs When thinking drifts: Evidential grounding for robust video reasoning.arXiv preprint arXiv:2510.06077, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.749238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.749238Z digest=sha256:1420207a222bef4855c1a9880d3af0ae1e426ae14921a9604bcb8bb7c049b7e1

Observation cf187c95-adac-4098-b1bc-88bddfc558b9 · outbound

This paper cites MIT press, 2010.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs MIT press, 2010

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.792172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.792172Z digest=sha256:523dfa66df8117637ab8635eaf23f63e2fb9f1a6531d90cbfad229b58ae07a42

Observation a349a015-7c69-4313-89ee-8323b3926575 · outbound

This paper cites Ba- sic books, 1988.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Ba- sic books, 1988

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.833758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.833758Z digest=sha256:7ec9a32bc436719d15b453e98e594a4bed1d7156c73f2e038ef52bf516558754

Observation e1d5a2e2-5c32-4bba-a7c1-b626cb290a91 · outbound

This paper cites Clarendon Press, 1978.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Clarendon Press, 1978

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.895137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.895137Z digest=sha256:50b555a97664eafff691ae5e24042b10c05fdd4ea086c791e828eedeb0b7833d

Observation 2e884371-8bcb-4f6f-b24f-2d478ec58b54 · outbound

This paper cites On visual knowledge.Frontiers of Information Technology & Electronic Engineering, 20(8):1021–1025,.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs On visual knowledge.Frontiers of Information Technology & Electronic Engineering, 20(8):1021–1025,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.945973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.945973Z digest=sha256:0bb57ed6e09a9bf80f43a7a57e452da38e4f9e870c307fe64dd636a458ba3cf3

Observation 8baf9970-6bcf-4ac3-a585-5cac0dc0bd69 · outbound

This paper cites Basic Books, 1954.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Basic Books, 1954

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.986585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.986585Z digest=sha256:5ecf2b793a57ad4792454d3de4ab9c336872ae86afa56a32b4d9ba62e86b72d4

Observation 814d2dfb-5841-4200-b9ad-4dbed6af28e2 · outbound

This paper cites Free Press,.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Free Press,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.041280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.041280Z digest=sha256:9e8fa63829081ced1dddf9431143f7abf541f048c8da291f11ed7b91bf85a518

Observation 383f3a06-06dc-41c2-888f-4915c0f5158e · outbound

This paper cites Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1 (4):515–526, 1978.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1 (4):515–526, 1978

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.169383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.169383Z digest=sha256:0362106f97a723197db8acb3f69281a5106bf549ddbcc37d3b5536559e31b971

Observation 8085e3b1-b330-4261-bfe7-ff04d5b2faa8 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Learn- ing transferable visual models from natural language super- vision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.206756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.206756Z digest=sha256:93a89b9e144dabde045d63447fbcde59d21e4e8fb350c28c828a14e385f0a64b

Observation 76b23818-ceb9-458e-a77e-a1b1a0fec841 · outbound

This paper cites Robust speech recognition via large-scale weak supervision, 2022.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Robust speech recognition via large-scale weak supervision, 2022

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.241697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.241697Z digest=sha256:e52b2efe7f293d4761dd83f7c2d44c971beee3a1e06affe9d8b1fe33fec8eb8e

Observation 8e03999a-7e75-490b-8664-0a14bed4c345 · outbound

This paper cites Reka flash 3, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Reka flash 3, 2025

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.292109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.292109Z digest=sha256:b09a0c0e62598a8632775f6b20606be86ee58c0bdf960af48534e6d21cb70b0b

Observation 37092672-20ea-423e-a20d-b2de8430b06b · outbound

This paper cites IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.339842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.339842Z digest=sha256:7ef327a29c3991b43bec0e91f5a85f10506481c351af2b7bbbdb87ddb787f2e6

Observation 0835e414-529d-4a0f-bb76-38dd27af434c · outbound

This paper cites Wellness Insti- tute, Inc., 2001.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Wellness Insti- tute, Inc., 2001

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.400203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.400203Z digest=sha256:47e1851d7cacab778bb3cd5f7843f21f6be2281b93985821c167f1aa19ba681e

Observation f2da25a9-8ede-48f8-89bd-94ad7b8478ef · outbound

This paper cites Lawrence Erlbaum Associates, 1977.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Lawrence Erlbaum Associates, 1977

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.439459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.439459Z digest=sha256:3157b387f620cfda5eea3101b7a065f6a74cd00bed4c14cffb709e2086823be9

Observation 655df557-4eac-465e-882e-bdf5827dc90f · outbound

This paper cites Core dimensions of human material perception.Proceedings of the National Academy of Sci- ences, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Core dimensions of human material perception.Proceedings of the National Academy of Sci- ences, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.482435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.482435Z digest=sha256:dfa470deba00d6e77b4a9e2b2f7d02447870216b66f88077323c6f925c1f1edd

Observation 85483ddc-b72c-445c-a6be-a44e7720ef2f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.535503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.535503Z digest=sha256:6ba3b2575894aa58cbe70552ac94657ca71509679a58e7c17e69621419fbc500

Observation dd66fe63-c1d4-4f88-802a-62595ad9b606 · outbound

This paper cites Towards vqa models that can read.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Towards vqa models that can read

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.578820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.578820Z digest=sha256:3363f08fcea49a788bd1ba2626d1f6271e14bdf4a38a599d4d164cfc95269d8d

Observation 4babefde-5f86-4720-bc87-6e121df70ef2 · outbound

This paper cites A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.635646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.635646Z digest=sha256:6e412e406781abfc6d4c803c83c373ab539dce84d41aa17255058f8267e3e809

Observation bd8084a8-b7ec-405c-a36d-98be1261629c · outbound

This paper cites Origins of knowledge.Psychologi- cal review, 99(4):605, 1992.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Origins of knowledge.Psychologi- cal review, 99(4):605, 1992

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.685266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.685266Z digest=sha256:0ed0492b127141015e403607444c54f6800adc5ab683d765cfffbd197123667a

Observation c5496cb1-740a-4094-994e-73864e367016 · outbound

This paper cites Mimo-vl technical report, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mimo-vl technical report, 2025

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.722348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.722348Z digest=sha256:9fe31fb65db57927e8514f423e9575155874dcef833a2824aa2f0507dc427823

Observation 595fa52c-9b86-407c-876f-a754121fa28b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.763176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.763176Z digest=sha256:ede5f743e2a2bfd8a503cb0cfac0d2b4765e5f450f9add28fa6369149e58960b

Observation 19aaa2ea-9a29-43ce-92bb-673ffdf86ce9 · outbound

This paper cites Qwen2.5: A party of foundation models,.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Qwen2.5: A party of foundation models,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.811188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.811188Z digest=sha256:c4015e14e79e3f4dc1f4a4136e5456cb42d9a876aa22d126823fe460fb613d32

Observation 300bd0a9-c482-4590-9ca2-ee6940c64633 · outbound

This paper cites Trl: Trans- former reinforcement learning.https : / / github.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Trl: Trans- former reinforcement learning.https : / / github

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.869311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.869311Z digest=sha256:9b979f9035f028158965130f8ff66bfa96bd81a30a21618b2e56aeb3676cf343

Observation a632ed77-af1f-40e5-b70b-9ade9057808b · outbound

This paper cites Make Your Training Flexible: Towards Deployment-Efficient Video Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Make Your Training Flexible: Towards Deployment-Efficient Video Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.938062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.938062Z digest=sha256:a1f9d5aac8e21bb54ff384cfff694c8be64911f2bb24ff4c6492012a5c1bf57a

Observation 7e6f7a3f-13d5-4433-b15e-9a32d1ba58a9 · outbound

This paper cites VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.982542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.982542Z digest=sha256:65f19a4b1ee9faa658614ab128ba865e53acebfcdb258429bd6d7ae96ec849a4

Observation 04382e29-c0ac-4578-8504-e3b22ab54b6d · outbound

This paper cites Videorft: Incentivizing video reasoning capabil- ity in mllms via reinforced fine-tuning.arXiv preprint arXiv:2505.12434, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Videorft: Incentivizing video reasoning capabil- ity in mllms via reinforced fine-tuning.arXiv preprint arXiv:2505.12434, 2025

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.036858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.036858Z digest=sha256:08a469636ba47c0ddbb5bcc783e0618687023569f3c3b93d2b033824bde65594

Observation 7976c6c5-d3fa-427d-883a-80ef94acb2d8 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.078012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.078012Z digest=sha256:f422d87509adcf0466598b2bf82fd878c6b35e5a68dc89123abfdc2af2a6a538

Observation 5e3a3b12-e1d4-4821-9fa0-158de14183ab · outbound

This paper cites Visual knowl- edge in the big model era: Retrospect and prospect.Fron- tiers of Information Technology & Electronic Engineering, 26(1):1–19, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Visual knowl- edge in the big model era: Retrospect and prospect.Fron- tiers of Information Technology & Electronic Engineering, 26(1):1–19, 2025

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.123948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.123948Z digest=sha256:498faa9420f1474213bf19f9c960495e194b0538b4d2326bb0633940701714ce

Observation 446e61eb-466a-431c-b28f-f2d23c2e59e7 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding, 2024.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Internvideo2: Scaling foundation models for multimodal video understanding, 2024

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.176103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.176103Z digest=sha256:204549dd2acc0ce989e348bfd6a6fc1a14d76d4d09bb805d363adc56b04b6f9c

Observation cf47c66d-3022-43d2-8fab-2778e12eba31 · outbound

This paper cites Social-iq 2.0 challenge: Benchmarking multimodal social understanding.https : / / github.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Social-iq 2.0 challenge: Benchmarking multimodal social understanding.https : / / github

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.201464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.201464Z digest=sha256:56d1846ca967d5c80809e32074ea5ee736394f38f983b5c864cde0cee71618a0

Observation 4b7087a9-7e98-47b4-92b3-56b13ebd9674 · outbound

This paper cites Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception.Cognition, 13(1):103–128, 1983.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception.Cognition, 13(1):103–128, 1983

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.239523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.239523Z digest=sha256:88bbe8df1ada518f605578ee8677e41b9e436e8524b668194e4e7453a8b77098

Observation 905837be-2b9a-4bfc-bfc4-713a1d332ae9 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.271934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.271934Z digest=sha256:43e9e7a52e5f33fdb3d76bbcfefa85bbecf87fb953456d3cb63c166ffd0a0535

Observation 7cceb949-8946-4862-ab0c-74e261ec11a1 · outbound

This paper cites Visionary-r1: Mitigating shortcuts in vi- sual reasoning with reinforcement learning.arXiv preprint arXiv:2505.14677, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Visionary-r1: Mitigating shortcuts in vi- sual reasoning with reinforcement learning.arXiv preprint arXiv:2505.14677, 2025

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.319389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.319389Z digest=sha256:a07727097475b9124e59b96c07584de8e2fa08f61c712db6bf5639204a8132a1

Observation 42a0454b-2425-4809-989e-d2fa46412140 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Next-qa: Next phase of question-answering to explaining temporal actions

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.349331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.349331Z digest=sha256:87736501a7c4bde1cfa9b1fbb543cf0ccc9f24265ae7899b6967c2885c0566c3

Observation 7330fcfe-4bed-43db-9130-ee96a05e3ba3 · outbound

This paper cites Advancing multi- modal reasoning capabilities of multimodal large language models via visual perception reward.arXiv preprint arXiv:2506.07218, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Advancing multi- modal reasoning capabilities of multimodal large language models via visual perception reward.arXiv preprint arXiv:2506.07218, 2025

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.392482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.392482Z digest=sha256:03f4179e72b3296719bac2b2994b59563c783d73c9be35c8a6d8ac97cee87ce5

Observation 36820249-8bf6-4b53-b020-81b06c3cd21c · outbound

This paper cites Expvid: A benchmark for experi- ment video understanding & reasoning.arXiv preprint arXiv:2510.11606, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Expvid: A benchmark for experi- ment video understanding & reasoning.arXiv preprint arXiv:2510.11606, 2025

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.464600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.464600Z digest=sha256:6ee9930c3fc5474523c0a5dbcaf48956f4cbebabcb9d481163e7c85d9d38c6da

Observation c280d29b-db4e-4c29-98d3-f92a41a34247 · outbound

This paper cites Qwen3 Technical Report.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Qwen3 Technical Report

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.540368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.540368Z digest=sha256:58dea80856ea5a5903385fa2e0f31a19a4c108caf806477b8faa0ac4b786baf6

Observation a88715d9-15f4-484d-b8e2-edd3ccdb777f · outbound

This paper cites Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.705686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.705686Z digest=sha256:31a024371b2fa8633b830432a9d08e7a1ec1e78a9fb34a718bff7c5574005748

Pith citing papers

Observation 243b57be-37fc-49e0-b526-59420390fa47 · inbound

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction cites this paper.

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:00.259742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T02:46:50.373450Z digest=sha256:060f4a7bf5b8bdbb1a74a071358989ce331bec3fa269b8530a1c2c293ba9bab6

Observation 9e3662f1-eca9-4aa7-bc4c-f4783f099257 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

Reference 289

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:00.259742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:a43bc2f06b68a7202a744824e8d3d89e00fe93beb95f9a0f2c6fe8c9e9740d5a

Observation d83de7e4-9f52-4e18-acfc-1e5e5fef7c06 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:00.259742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:334f8a9cc0f77df02b14ea0bf1fc9dfde24c46f7033430a841286073a9972f61