Pith. sign in

Paper Citation Record · LEDGER

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

As of 19 August 2026, this Paper Citation Record lists 100 of 127 outbound references and 3 inbound Pith citation observations for arXiv:2511.20272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.20272 v2

Coverage vector

measured 100 of 127 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:23:07.705686Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T16:37:06.384435Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:38:39.619743Z

Reference resolution

100 of 127 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c2785ed-4cd7-488a-8e38-8df42133fc3d · outbound

This paper cites GPT-4 Technical Report.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T20:22:59.834670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:22:59.834670Z digest=sha256:4bff0f571a318ab9c015e554c117dc77b3767ba04676805c3bc4f0c5f3728de5

Observation 0e4113ed-1a88-4f11-bed2-153b4d47d61b · outbound

This paper cites On seeing stuff: The perception of materials by humans and machines.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs On seeing stuff: The perception of materials by humans and machines

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:22:59.937363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:22:59.937363Z digest=sha256:e672d0853961d675ae54e88dfcf91510c22c8b39ed491db58ec450705b5c3793

Observation 40d649dd-1e16-41bd-ad37-01207a4f42b5 · outbound

This paper cites Vqa: Visual question answering.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Vqa: Visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.093535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.093535Z digest=sha256:f4fe8e1342ecab44281517554d89ecf85950d75abb6115dcf19a13ae441a785d

Observation d19b8ecb-31a9-4742-8a73-f796e7c8d198 · outbound

This paper cites Eye-contact, distance and affiliation.Sociometry, pages 289–304, 1965.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Eye-contact, distance and affiliation.Sociometry, pages 289–304, 1965

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.331240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.331240Z digest=sha256:500c8c5eb359491857b9a7f443d60136bb0b610db0a43e2176610ab5a844ef71

Observation 14fda1d8-9815-4d28-9671-ecacd20eb630 · outbound

This paper cites Qwen2.5-VL Technical Report.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.406527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.406527Z digest=sha256:bf1a5338466780c981fd6f28b1b0ffae6f836a57faa7355860b3eb340860e8a4

Observation 1dae183b-ab4f-4652-9b30-8dcee8737c6d · outbound

This paper cites Representing the existence and the lo- cation of hidden objects: Object permanence in 6-and 8- month-old infants.Cognition, 23(1):21–41, 1986.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Representing the existence and the lo- cation of hidden objects: Object permanence in 6-and 8- month-old infants.Cognition, 23(1):21–41, 1986

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.484396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.484396Z digest=sha256:f7d28eb82243b6c1764446864cedfd2d52391740e41ea79ced67ca26d4d2bfea

Observation c53a5db6-45f7-4cc0-8170-3321c28fdfa7 · outbound

This paper cites theory of mind.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs theory of mind

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.572585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.572585Z digest=sha256:9cb4142c1a3e3a98da1c3217710a5204858cf8eb92935707b8a05534c7116f0e

Observation 858ff559-25c1-4076-bb04-fd8184e2384b · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.688451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.688451Z digest=sha256:d08fed4f6b9e28a7f930355b8c164c41d91dc91b142b4741c04d7f0f4abc8bf9

Observation 98860f4a-b3b8-4ce3-91d0-263f87216e4c · outbound

This paper cites IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.771656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.771656Z digest=sha256:e3fe9edcce31356c3e82e02878c22c1015822de9f55096c75574ec375c52cda0

Observation 0d473e7b-2837-4ab2-b10e-58c29a3aeabb · outbound

This paper cites Routledge, 1995.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Routledge, 1995

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.854703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.854703Z digest=sha256:a741455174ab1e66624f6f45c2c9fc55fb50be65d9b9980e23bd8a878c9aa384

Observation 0c586aa5-8c9d-4252-b923-e4201d3f038e · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Activitynet: A large-scale video benchmark for human activity understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.937721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.937721Z digest=sha256:cc4862aad4fb35a3e94c6864d3c586374a008f0385d6e80996adf66a452179b5

Observation 4715a1fd-4758-4444-81f4-330e1b44d5c2 · outbound

This paper cites Rextime: A benchmark suite for reasoning-across-time in videos.Advances in Neural In- formation Processing Systems, 37:28662–28673, 2024.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Rextime: A benchmark suite for reasoning-across-time in videos.Advances in Neural In- formation Processing Systems, 37:28662–28673, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.971832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.971832Z digest=sha256:3d17aa71e7df1697ab59612761a8628531339815ba37364c8bc5c4cfabc5adb2

Observation f61da4e4-5d2f-474c-8887-1cd84e193873 · outbound

This paper cites Are we on the right way for evaluating large vision-language models?Advances in Neural Infor- mation Processing Systems, 37:27056–27087, 2024.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Are we on the right way for evaluating large vision-language models?Advances in Neural Infor- mation Processing Systems, 37:27056–27087, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.064028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.064028Z digest=sha256:6de64e52779b0665014199fc2b5a937999b6888304d397b557307be9d819e788

Observation 96d07928-9fd3-4a1b-a2cd-89e6c7c79f24 · outbound

This paper cites Video-holmes: Can mllm think like holmes for complex video reasoning?, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Video-holmes: Can mllm think like holmes for complex video reasoning?, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.175032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.175032Z digest=sha256:3f4846065c1645a1d66a43a2a49cb58a3eedf9f2ad3214fe60aede59dfa43b26

Observation ed3b046b-4305-4e84-b01a-dee2201bcd7b · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.290978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.290978Z digest=sha256:825b3923cfdff32f41daadd4d476db34cd9cad996a74dc2cebf5c2ef117a1b41

Observation cbe8be57-de9b-4f62-8e83-9a4890fa0284 · outbound

This paper cites Le, Sergey Levine, and Yi Ma.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Le, Sergey Levine, and Yi Ma

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.403982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.403982Z digest=sha256:c1bad0e831598cf936e1bf272026a102fb3f009a5f45b532cb377e138a1a2b3b

Observation f2bc5c7d-9545-4df6-9dde-7ab60c9cf8a1 · outbound

This paper cites Whatever next? predictive brains, situated agents, and the future of cognitive science.Behavioral and brain sciences, 36(3):181–204, 2013.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Whatever next? predictive brains, situated agents, and the future of cognitive science.Behavioral and brain sciences, 36(3):181–204, 2013

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.472386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.472386Z digest=sha256:559f2dd60e6de02e22fcdb6f8a6a1ffad4a85319b71faaf237e11776bd417f90

Observation 7e7f228a-2550-4553-9916-365f8ead4e66 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.531512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.531512Z digest=sha256:0179f758f23703987eb9f67b68482e7ef533beec01d506235b4cacc3b33cd984

Observation 5223b74d-761a-450a-9ae5-af8fed1c7b9a · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.579499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.579499Z digest=sha256:901e0b797e6807b32e8b12233c9210790283e670347b0771e2fc15e4260213dc

Observation ca6d9e86-e844-4883-b7c6-7f82d5ec0538 · outbound

This paper cites MIT press, 1987.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs MIT press, 1987

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.643970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.643970Z digest=sha256:2340cc2068b0161e91b357031f809d4e677479e1d420cef09f352b8580c76224

Observation 35bb7093-aff4-45a3-ba58-47c1837078ba · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.727373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.727373Z digest=sha256:0ef287252a13d9399c10a028985e41cd1d9d29ed9ffdad694e56ccb29bc73bfa

Observation 049c5852-561a-4b6b-9c0a-9ad9ce89d05c · outbound

This paper cites An argument for basic emotions.Cognition & emotion, 6(3-4):169–200, 1992.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs An argument for basic emotions.Cognition & emotion, 6(3-4):169–200, 1992

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.812885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.812885Z digest=sha256:274bbaf42c15346c7c204cbfb9567cde4b3c55e282617a75783b1e12fd574408

Observation 15443549-ad36-4cf8-a9df-f10b57c5db37 · outbound

This paper cites Unreal engine.https : / / www.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Unreal engine.https : / / www

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.869913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.869913Z digest=sha256:4f1cea8f4fb6cd161329766a20b1cc5253f426d8b97aa285b298cbff9cdeeec2

Observation 0141789d-b487-475d-867e-941fac8706f2 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.923190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.923190Z digest=sha256:085f8c3529b6efeebd66288ac7a6aa98d5ba6d3917d786cf92ffe476adf1b769

Observation df25d2c5-5169-468f-a084-b69b1fc46b41 · outbound

This paper cites Material perception.Annual review of vision science, 3:365–388, 2017.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Material perception.Annual review of vision science, 3:365–388, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.003342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.003342Z digest=sha256:ff7a6f4e645c30d6adcdccccf1545b2065dd0ba0cf51d18662f34f6a7e1353ea

Observation 6163625b-6a14-42e8-8f7b-afdba3cc4faf · outbound

This paper cites The free-energy principle: a unified brain the- ory?Nature reviews neuroscience, 11(2):127–138, 2010.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The free-energy principle: a unified brain the- ory?Nature reviews neuroscience, 11(2):127–138, 2010

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.033703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.033703Z digest=sha256:2a8b72148cd68d83b679b98f073aeab6f3b2f4909f13f03a0791fc2947643acb

Observation 9948e519-a3c9-4ab7-a4cc-a60d18381849 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.091017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.091017Z digest=sha256:2a556660f67275a8e0d63b454941a668804f0f6b605109107733b4d8c12474b2

Observation 2f5062b3-88b0-4b28-bca1-026698eb8739 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Blink: Multimodal large language models can see but not perceive

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.138277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.138277Z digest=sha256:598b1d02b9cfc20952d1fe5bf42000a0653e475ec0eaf950f057e87c179cf50d

Observation a30fc0d5-5e85-45b8-9ccc-84d95dcfeaf0 · outbound

This paper cites Atomistic Control in Molecular Beam Epitaxy Growth of Intrinsic Magnetic Topological Insulator MnBi2Te4.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Atomistic Control in Molecular Beam Epitaxy Growth of Intrinsic Magnetic Topological Insulator MnBi2Te4

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.182852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.182852Z digest=sha256:7b0ade7b986728bf01bfd7390d4feada2f960e3a70f414574eef96dc76dac1dd

Observation b3b05772-dd96-409b-b6e7-3cbca0d93443 · outbound

This paper cites Houghton Mifflin, 1979.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Houghton Mifflin, 1979

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.249905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.249905Z digest=sha256:d9b57bbb24c27acb7d8e1018415a96a2699f1c861c9edb1715cfa61afd8c7421

Observation 37ef0134-653e-4edd-bef5-ff46717a1752 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The” something something” video database for learning and evaluating visual common sense

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.321899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.321899Z digest=sha256:945b7cfbf08583b4a7dd0e7c9a9cf373218a665786e2f4c830ae8399ec0bfc22

Observation f3e67a5c-82eb-49a4-b17f-4932b2dacfb9 · outbound

This paper cites The Llama 3 Herd of Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The Llama 3 Herd of Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.394429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.394429Z digest=sha256:dbbb1467c8478408f52553a3e4dcae0e40deb6aad4ea41ee2a2e7b45aa136795

Observation 0cba16a0-eb5e-4666-88cf-e2a04c2e0cf4 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language halluci- nation and visual illusion in large vision-language models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Hallusionbench: an advanced diagnostic suite for entangled language halluci- nation and visual illusion in large vision-language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.480865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.480865Z digest=sha256:7e77d6b17306ef249d0b7d26f958b9f05e608faaff140cef1c20242aed1148ac

Observation 6411cd8f-eded-43a2-b78b-3cae95e7e5d4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.549143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.549143Z digest=sha256:9e3f131afff79477cd903869fc641989a39ee2d63fd8671ed8c4c0d934ff7c04

Observation dc1af7d2-65d2-48c4-8ed2-51d6db7532a2 · outbound

This paper cites Doubleday, 1966.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Doubleday, 1966

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.624952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.624952Z digest=sha256:7305b97b8902651eea66d973040f1952e597d812debdc720d5f6e95ac16af7e9

Observation 10724494-018f-430d-935e-7415a1293efd · outbound

This paper cites Glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.726794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.726794Z digest=sha256:1249fafec609aaf0f01250a85383925b0b6fd1f60963904e5b472b2259c26f91

Observation f56e9cb6-14ad-429a-82cd-1483b68685bf · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.778666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.778666Z digest=sha256:677c930904850bf0b35a4f68cad279cb96325e3b59b301ac0801ad4bcab6b547

Observation 5f8a72b2-eb6f-47ee-9d55-dd7ae9dce488 · outbound

This paper cites GPT-4o System Card.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs GPT-4o System Card

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.877204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.877204Z digest=sha256:2682d5a98dd357924e101a9834b8a273f8dc29a83c7102dee0cfb5837f2c332d

Observation 678d0199-f37c-47c2-98c6-88ad754fba18 · outbound

This paper cites an unresolved cited work.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.044818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.044818Z digest=sha256:a6a4aebf86ad0e46ba288b8b3df59d6d2296e40c49bef4aa7835e98129449dbc

Observation da0fb4f2-cd52-41d2-a5f9-f584920f3438 · outbound

This paper cites From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.179935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.179935Z digest=sha256:37cb85009a865815ed5b9b15badea8a26b28577e5439114f2e033e73abf91d6c

Observation f9851049-ad81-41e2-94d3-c00edd3d4bf3 · outbound

This paper cites Towards Social AI: A Survey on Understanding Social Interactions.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Towards Social AI: A Survey on Understanding Social Interactions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.378338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.378338Z digest=sha256:aaea31c9abc7310e12ff2502382ad460b78a1e85fc37649125fb659ea9c8229d

Observation d511dc7e-7895-4907-9c2b-c3b73d49c0dd · outbound

This paper cites What is More Likely to Happen Next? Video-and-Language Future Event Prediction.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs What is More Likely to Happen Next? Video-and-Language Future Event Prediction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.468092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.468092Z digest=sha256:0bbb2ce7ed92a40454b5246312251d678b8a2445763fde3e84835142b6aa6ca0

Observation 6f908d3f-9d90-465a-9c6f-a5905c14acff · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Detecting mo- ments and highlights in videos via natural language queries

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.619120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.619120Z digest=sha256:a54224b81e6d8ccbbcba1ff66e4d468cdfb06050d10196d04745c20642d51275

Observation 224dc2ff-e5db-4249-bfec-2fe6ff741b00 · outbound

This paper cites Anisotropic flow, flow fluctuation and flow decorrelation in relativistic heavy-ion collisions: the roles of sub-nucleon structure and shear viscosity.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Anisotropic flow, flow fluctuation and flow decorrelation in relativistic heavy-ion collisions: the roles of sub-nucleon structure and shear viscosity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.810455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.810455Z digest=sha256:ad8543a2a0e682b3960455d8f9bc6ace7e016b353018d690829ddaf98bf27292

Observation fd73e304-7c6a-4557-b36c-1e76ed112956 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.031544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.031544Z digest=sha256:62eed3949e31e70c691eae0a6e52b0b4eb43b406d72a348f559f06b9fb62bff8

Observation cf5d0930-c3d0-443f-a63d-92be2ecef23d · outbound

This paper cites From represen- tation to reasoning: Towards both evidence and common- sense reasoning for video question-answering.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs From represen- tation to reasoning: Towards both evidence and common- sense reasoning for video question-answering

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.223426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.223426Z digest=sha256:69fc7d88464596e1182521a2716f4b17a9993208d92aecb51d9154c023f7815f

Observation f72b758d-fcdf-49d2-b4ba-33bbc058f01c · outbound

This paper cites Mvbench: A comprehensive multi-modal video under- standing benchmark.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mvbench: A comprehensive multi-modal video under- standing benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.327558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.327558Z digest=sha256:64da2d60b9333c176d716114faee7a05842483ce66c01368fdb205d6cc9d04a8

Observation 4fc421a3-3b21-4ed9-bdc0-1e7dec310b92 · outbound

This paper cites Pope: A simple method to hallucination eval- uation in visual question answering.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Pope: A simple method to hallucination eval- uation in visual question answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.422362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.422362Z digest=sha256:0150f1fccd7f35f9ff07c609905199b015d15bfbdbc0f797548dd6ba286746e1

Observation 278caf4c-6ebd-4563-beb2-ea12b21b56cc · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.532245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.532245Z digest=sha256:7b01f4b46e77bae29aec4a98d38f91676496d6a57f25e9df4d9ed1809bfe87df

Observation f1a887c0-7353-4b9f-bf56-c24768bc63e2 · outbound

This paper cites Core Knowledge Deficits in Multi-Modal Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Core Knowledge Deficits in Multi-Modal Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.626496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.626496Z digest=sha256:d8b60cfe104a8183b7fb27d9b804ad91c12f9eed579c85cf30ecec9cf0b91313

Observation 9ea4dfa3-11d9-4158-9b50-b17e8074ccdc · outbound

This paper cites Self-Rewarding Vision-Language Model via Reasoning Decomposition.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.725008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.725008Z digest=sha256:cc425780aec3adc1c3c891e389a4c2e781492a9f9ae8d0147f96b96e3aee502d

Observation a7619fe7-b755-48d1-89b7-3d73dba6fe7b · outbound

This paper cites Explainable Multimodal Emotion Recognition.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Explainable Multimodal Emotion Recognition

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.826214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.826214Z digest=sha256:b566147cfcdc666d63e18e914c5af48f89109c6c9e4aa2e086b7d56d07fba826

Observation b4d6e639-687f-48bd-b459-b9d41f707997 · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.926420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.926420Z digest=sha256:c3098a539b2c515967c38d2b11f3cc7e96734e73f9f97953a10c4dfb69aa557a

Observation f150adbc-f5dc-46ac-81f6-0a1fa312d5fb · outbound

This paper cites Microsoft coco: Common objects in context.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Microsoft coco: Common objects in context

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.034079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.034079Z digest=sha256:5de8fc242e2587c1f1d56ecce578437b56720335773e6b947b4bd760832583f8

Observation 9038ad2b-7b01-4e72-b0ee-82aab3cd54fa · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.200466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.200466Z digest=sha256:1bb8f6cc68e0c2e8e01c7dc5f26227b2a4fc103470feb3f6016ae6aeaa021eb1

Observation c6d62e94-c635-4f87-a6b4-aa202eed3c8b · outbound

This paper cites Generative Physical AI in Vision: A Survey.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Generative Physical AI in Vision: A Survey

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.304254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.304254Z digest=sha256:4f038365c3da18baaf032cec2a371e58f054971f0ab45eea4c2d904c655067be

Observation 4725a73b-cd00-4335-bd38-b554a4a47c85 · outbound

This paper cites Visual instruction tuning.Advances in neural infor- mation processing systems, 36:34892–34916, 2023.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Visual instruction tuning.Advances in neural infor- mation processing systems, 36:34892–34916, 2023

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.407372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.407372Z digest=sha256:66b133f26fae3887b0f8c6741f2870cf65b71a035235ab203da6be571fb3044c

Observation e9aebeb5-ba66-4fb6-b7fc-a4e731d52021 · outbound

This paper cites Improved baselines with visual instruction tuning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Improved baselines with visual instruction tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.487338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.487338Z digest=sha256:f8ad59e0753fad3623c4dcd7073e0b10374609639c0531cb77287776920771c4

Observation 4b6cbdf4-27dd-42ad-b025-e4b7e614b806 · outbound

This paper cites Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.553933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.553933Z digest=sha256:ab93b524ddf01093ff297628cf569e9f0c732e39dda58f07b1f25ed3e6b2b462

Observation 2c559219-0d9f-435a-a826-b1d462cbcf51 · outbound

This paper cites Mmbench: Is your multi- modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mmbench: Is your multi- modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.605866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.605866Z digest=sha256:2dec03f2d1fa14e14c274630bb15eb20816557b81704d1246f331338f70576db

Observation 75e2c42f-5191-4b9b-bbce-3ebab6794364 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs TempCompass: Do Video LLMs Really Understand Videos?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.644159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.644159Z digest=sha256:ea8c92d4262356847f51c172529208ddf5b0e01c0e3b5cd038d589ea180be103

Observation 0eb26b62-1cfd-4c03-b18b-baa9752989bb · outbound

This paper cites Videoreasonbench: Can mllms perform vision-centric complex video reasoning?arXiv preprint arXiv:2505.23359, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Videoreasonbench: Can mllms perform vision-centric complex video reasoning?arXiv preprint arXiv:2505.23359, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.701446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.701446Z digest=sha256:c80b066daf0f2b99827f71fecdf73de6d6823c48da5be866e172d024e65405bd

Observation 98b90b1a-9759-4ffa-bb92-81e248525fb2 · outbound

This paper cites When thinking drifts: Evidential grounding for robust video reasoning.arXiv preprint arXiv:2510.06077, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs When thinking drifts: Evidential grounding for robust video reasoning.arXiv preprint arXiv:2510.06077, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.749238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.749238Z digest=sha256:d8700f4734a0643a7c8b4b059ab80d9e9a8226c6dc3fa77e831bcab8062a706e

Observation cf187c95-adac-4098-b1bc-88bddfc558b9 · outbound

This paper cites MIT press, 2010.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs MIT press, 2010

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.792172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.792172Z digest=sha256:fcf7ea3c4815575daa45bb66cd0b2a169e33215f4477e369ea37164bdde09aeb

Observation a349a015-7c69-4313-89ee-8323b3926575 · outbound

This paper cites Ba- sic books, 1988.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Ba- sic books, 1988

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.833758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.833758Z digest=sha256:39b7a416dcf63b91a262c820801ce25dfa36d6138eb484d63177304c3f3b2a24

Observation e1d5a2e2-5c32-4bba-a7c1-b626cb290a91 · outbound

This paper cites Clarendon Press, 1978.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Clarendon Press, 1978

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.895137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.895137Z digest=sha256:1edcd5f5442a9fb1e90d36a6697c68604247fac7f9130185aed9636bfb0c85b0

Observation 2e884371-8bcb-4f6f-b24f-2d478ec58b54 · outbound

This paper cites On visual knowledge.Frontiers of Information Technology & Electronic Engineering, 20(8):1021–1025,.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs On visual knowledge.Frontiers of Information Technology & Electronic Engineering, 20(8):1021–1025,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.945973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.945973Z digest=sha256:b3a59f9204236d7ab334fc6e5b78b87ddac7f9d86264634ac33d93c73247a2ed

Observation 8baf9970-6bcf-4ac3-a585-5cac0dc0bd69 · outbound

This paper cites Basic Books, 1954.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Basic Books, 1954

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.986585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.986585Z digest=sha256:74121cd350fee394c419d9e51f2c3ec347daf8b001eeef90787d64f535d844ab

Observation 814d2dfb-5841-4200-b9ad-4dbed6af28e2 · outbound

This paper cites Free Press,.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Free Press,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.041280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.041280Z digest=sha256:b228c9aa21027fb04efa40eff0341f60df2e3f285dda5da68d12f79ca8700e1f

Observation 383f3a06-06dc-41c2-888f-4915c0f5158e · outbound

This paper cites Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1 (4):515–526, 1978.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1 (4):515–526, 1978

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.169383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.169383Z digest=sha256:178216ba4cc153d2e480909561205417571d080ae46552971e4c0d933d122e9f

Observation 8085e3b1-b330-4261-bfe7-ff04d5b2faa8 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Learn- ing transferable visual models from natural language super- vision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.206756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.206756Z digest=sha256:ae27ce893e5ba05841e336415266d9f911002f36219328c7fd03b5556f15685e

Observation 76b23818-ceb9-458e-a77e-a1b1a0fec841 · outbound

This paper cites Robust speech recognition via large-scale weak supervision, 2022.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Robust speech recognition via large-scale weak supervision, 2022

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.241697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.241697Z digest=sha256:e2b617d193af45064293d2fa0de8201895bd1aa67e47646c2cbb72273b298d30

Observation 8e03999a-7e75-490b-8664-0a14bed4c345 · outbound

This paper cites Reka flash 3, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Reka flash 3, 2025

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.292109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.292109Z digest=sha256:6381c3b55ad88d31d3d9297e9469a380e3039873991bfa93e4b9d984516d3c49

Observation 37092672-20ea-423e-a20d-b2de8430b06b · outbound

This paper cites IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.339842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.339842Z digest=sha256:0c1ef21acb1041c91483c94819f853d6765c59d8de89dad0474fd54bdabe0141

Observation 0835e414-529d-4a0f-bb76-38dd27af434c · outbound

This paper cites Wellness Insti- tute, Inc., 2001.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Wellness Insti- tute, Inc., 2001

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.400203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.400203Z digest=sha256:589c95e035c19e94659678f49a7560e3d52c2ad115e057879ac0259fe5f65254

Observation f2da25a9-8ede-48f8-89bd-94ad7b8478ef · outbound

This paper cites Lawrence Erlbaum Associates, 1977.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Lawrence Erlbaum Associates, 1977

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.439459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.439459Z digest=sha256:8808659d23e8b212d2c65582b8f65c245dd137cccae7dbe2159c6f920136daca

Observation 655df557-4eac-465e-882e-bdf5827dc90f · outbound

This paper cites Core dimensions of human material perception.Proceedings of the National Academy of Sci- ences, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Core dimensions of human material perception.Proceedings of the National Academy of Sci- ences, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.482435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.482435Z digest=sha256:67055490a3d5e87a69720bfaaed67b2051fc757e7b3cbfb89310d9f2d9aec9f4

Observation 85483ddc-b72c-445c-a6be-a44e7720ef2f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.535503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.535503Z digest=sha256:f603544a223b461475db9446eb429ec740710394e9bbfb8b35204d2de7f9e199

Observation dd66fe63-c1d4-4f88-802a-62595ad9b606 · outbound

This paper cites Towards vqa models that can read.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Towards vqa models that can read

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.578820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.578820Z digest=sha256:e9f17ba7309c00a1adf2b70fcb5d3fd9da5c85be1028979cd901ad0798e327d7

Observation 4babefde-5f86-4720-bc87-6e121df70ef2 · outbound

This paper cites A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.635646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.635646Z digest=sha256:b07b7a3503abba8b1639b9f0f1a4a6ad2737fb367fd1d00c08ea1428847f874c

Observation bd8084a8-b7ec-405c-a36d-98be1261629c · outbound

This paper cites Origins of knowledge.Psychologi- cal review, 99(4):605, 1992.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Origins of knowledge.Psychologi- cal review, 99(4):605, 1992

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.685266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.685266Z digest=sha256:6c1cb20a822c59c17ad402355051258dffaf6b71037f9717c47aa03cf0944296

Observation c5496cb1-740a-4094-994e-73864e367016 · outbound

This paper cites Mimo-vl technical report, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mimo-vl technical report, 2025

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.722348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.722348Z digest=sha256:d4e774f4bacc6ae8cc6a27e131cc50b87d7143ab4f32f8d51f392de959ff75aa

Observation 595fa52c-9b86-407c-876f-a754121fa28b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.763176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.763176Z digest=sha256:7ead63b62a8ac0ef7cd61f3dd9434fd4928586b28209d0fbbef2d41ee6a10a7d

Observation 19aaa2ea-9a29-43ce-92bb-673ffdf86ce9 · outbound

This paper cites Qwen2.5: A party of foundation models,.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Qwen2.5: A party of foundation models,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.811188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.811188Z digest=sha256:51443c87e970cdc1365db7b2c71f00573567ce84c0b283a0a0b2d640dc0f63b4

Observation 300bd0a9-c482-4590-9ca2-ee6940c64633 · outbound

This paper cites Trl: Trans- former reinforcement learning.https : / / github.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Trl: Trans- former reinforcement learning.https : / / github

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.869311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.869311Z digest=sha256:fc63c1a0ec875a0a012bdf6f0bd7b6b4b18008c250dd7ad198ec4bf102c83b35

Observation a632ed77-af1f-40e5-b70b-9ade9057808b · outbound

This paper cites Make Your Training Flexible: Towards Deployment-Efficient Video Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Make Your Training Flexible: Towards Deployment-Efficient Video Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.938062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.938062Z digest=sha256:f4bfb15ece1b5d19d35cc42b8aa7db6914c696fbe15c03294d66987262833ef3

Observation 7e6f7a3f-13d5-4433-b15e-9a32d1ba58a9 · outbound

This paper cites VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.982542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.982542Z digest=sha256:d8157a83cb4690ff9e9305ef117450fcb513f905ef7449c624064fab654dfd92

Observation 04382e29-c0ac-4578-8504-e3b22ab54b6d · outbound

This paper cites Videorft: Incentivizing video reasoning capabil- ity in mllms via reinforced fine-tuning.arXiv preprint arXiv:2505.12434, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Videorft: Incentivizing video reasoning capabil- ity in mllms via reinforced fine-tuning.arXiv preprint arXiv:2505.12434, 2025

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.036858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.036858Z digest=sha256:bb449732c238d5298852b4ac8bd86172d0854d8bb864792aa6ed84e7e82e8ade

Observation 7976c6c5-d3fa-427d-883a-80ef94acb2d8 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.078012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.078012Z digest=sha256:b9ab47aaa037176f66d7f03c7877c7aa0ba7e5fb0f20283b5a74886b5fcdd24f

Observation 5e3a3b12-e1d4-4821-9fa0-158de14183ab · outbound

This paper cites Visual knowl- edge in the big model era: Retrospect and prospect.Fron- tiers of Information Technology & Electronic Engineering, 26(1):1–19, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Visual knowl- edge in the big model era: Retrospect and prospect.Fron- tiers of Information Technology & Electronic Engineering, 26(1):1–19, 2025

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.123948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.123948Z digest=sha256:069f2f2497719d14d6db450db9267dc5a2aefd4f9962f35e72fe14d3287a7e9d

Observation 446e61eb-466a-431c-b28f-f2d23c2e59e7 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding, 2024.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Internvideo2: Scaling foundation models for multimodal video understanding, 2024

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.176103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.176103Z digest=sha256:dbd3be1cef1a7318a953f830c74430543fd6e3ac20fa1c09b494331dc266dbfc

Observation cf47c66d-3022-43d2-8fab-2778e12eba31 · outbound

This paper cites Social-iq 2.0 challenge: Benchmarking multimodal social understanding.https : / / github.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Social-iq 2.0 challenge: Benchmarking multimodal social understanding.https : / / github

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.201464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.201464Z digest=sha256:503c9aba12f4f22f6579666d096c853281129badbbd5f6a9848146aa9cea1dbf

Observation 4b7087a9-7e98-47b4-92b3-56b13ebd9674 · outbound

This paper cites Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception.Cognition, 13(1):103–128, 1983.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception.Cognition, 13(1):103–128, 1983

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.239523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.239523Z digest=sha256:df5e699552e6a83ebe93b92ea3cd2b0a09bbff52f6ddb0bda1e25d8a07dd980b

Observation 905837be-2b9a-4bfc-bfc4-713a1d332ae9 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.271934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.271934Z digest=sha256:6047176b7cee46cdfc3b6018ce63cc817bc857e84d4563a2761b5d49d6b40087

Observation 7cceb949-8946-4862-ab0c-74e261ec11a1 · outbound

This paper cites Visionary-r1: Mitigating shortcuts in vi- sual reasoning with reinforcement learning.arXiv preprint arXiv:2505.14677, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Visionary-r1: Mitigating shortcuts in vi- sual reasoning with reinforcement learning.arXiv preprint arXiv:2505.14677, 2025

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.319389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.319389Z digest=sha256:c89f902ff0fc731e549f0bff79eb0cccab4a9bd15910ba7a99313c5d91b014d5

Observation 42a0454b-2425-4809-989e-d2fa46412140 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Next-qa: Next phase of question-answering to explaining temporal actions

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.349331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.349331Z digest=sha256:f6e6cd7d890024a4f95985730ffab72e25f7a4f208a5ff030c320d88742063f1

Observation 7330fcfe-4bed-43db-9130-ee96a05e3ba3 · outbound

This paper cites Advancing multi- modal reasoning capabilities of multimodal large language models via visual perception reward.arXiv preprint arXiv:2506.07218, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Advancing multi- modal reasoning capabilities of multimodal large language models via visual perception reward.arXiv preprint arXiv:2506.07218, 2025

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.392482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.392482Z digest=sha256:02dc1a183e1b0b1615533c391e610dbb580218e5bb0db223e3280ee30f07ce83

Observation 36820249-8bf6-4b53-b020-81b06c3cd21c · outbound

This paper cites Expvid: A benchmark for experi- ment video understanding & reasoning.arXiv preprint arXiv:2510.11606, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Expvid: A benchmark for experi- ment video understanding & reasoning.arXiv preprint arXiv:2510.11606, 2025

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.464600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.464600Z digest=sha256:715e9ffc34cc52cde444ed23fc2bef81f35673b3577b4a29fde8c50e85221c77

Observation c280d29b-db4e-4c29-98d3-f92a41a34247 · outbound

This paper cites Qwen3 Technical Report.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Qwen3 Technical Report

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.540368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.540368Z digest=sha256:3171d1ad4a5ae2fd3c1757e734fb2ff33ab02eb88a4fbc2d6de9b538cce2d36d

Observation a88715d9-15f4-484d-b8e2-edd3ccdb777f · outbound

This paper cites Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.705686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.705686Z digest=sha256:42395e09306cddedebbbd2f3bc5330bde3935e65ca0d5d8cab93212f1b031e08

Pith citing papers

Observation 243b57be-37fc-49e0-b526-59420390fa47 · inbound

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction cites this paper.

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:00.259742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T02:46:50.373450Z digest=sha256:61e1f6d831784770f1503da0cafef2e5e1a131d99038e34612cd5ffb51e769a6

Observation 9e3662f1-eca9-4aa7-bc4c-f4783f099257 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

Reference 289

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:00.259742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:8ac3b18c7fd84b1855094f73f6dd6e0641159d96276479ea08bea7a230096acf

Observation d83de7e4-9f52-4e18-acfc-1e5e5fef7c06 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:00.259742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:73ee8b55b309527a09c0f96497ebc0b8352f46b33adffcae635c5d40e8e3f6ff