Pith. sign in

Paper Citation Record · LEDGER

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models

As of 7 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 0 inbound Pith citation observations for arXiv:2509.08538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08538 v2

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:32:57.059594Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

92 of 92 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved87
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e356cbf-7a71-4122-b51a-a651befb776a · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.355191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.355191Z digest=sha256:e8802a1dfc995a2cfc4a0bae5c273bebbdc7bc6cc58055ef00d0aebbc770e052

Observation bc4d416f-196f-496d-ac35-4956244fe320 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.362792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.362792Z digest=sha256:2cb24b0285b34252c029ca87fe4544390d650299055c0add9ff292892795e5f8

Observation 7bbc4ae2-2b48-449f-8b42-69a845ab6de4 · outbound

This paper cites Qwen Technical Report.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.368080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.368080Z digest=sha256:37018f9015f56ecfebf710d426f696bb1e2fff8c4d9ab30674406aa4628e15b7

Observation 506debf6-d33e-4759-ad25-68f24fa4c7d3 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.373710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.373710Z digest=sha256:c5616f0ca51d003a985483d1065f4c3227e4f47cd07bcbf3423d6d5a74927758

Observation e9e3d4af-39a4-47d7-803e-99950a1b1502 · outbound

This paper cites 2020.The visual story: Creating the visual structure of film, TV, and digital media.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2020.The visual story: Creating the visual structure of film, TV, and digital media

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.385920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.385920Z digest=sha256:36d8ef77c881c6366236f8e974c8014e5c3852ac66ab469e0a5229274cc99d2b

Observation 6c802f18-24d5-4edf-b6d1-356ca4c3cd7e · outbound

This paper cites 2005.Figures traced in light: On cinematic staging.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2005.Figures traced in light: On cinematic staging

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.393082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.393082Z digest=sha256:8078ed8ff7dfb33ea287bcd6a85c354be3c21cea1308974f66076c63746bb109

Observation 7c56042b-1401-4670-9243-4994f562161f · outbound

This paper cites Bordwell and K.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Bordwell and K

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.399095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.399095Z digest=sha256:b34742e06ebed9d6b73fa5fd220e06decf06e2805b76e82988461acb41150791

Observation 48aedb5d-29d7-4710-be97-319161b2aa83 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.405345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.405345Z digest=sha256:87ca50a41636dcab0d3ec5797fa0e434d96253483a6cab111b85604ab3280ccd

Observation cd5213d4-144f-4941-a1e6-c02f6c688d28 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.413191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.413191Z digest=sha256:2995cc82f59d2c3ad932fda659068b8757a7cd5eeac8fe59dfe03f3fad78ab36

Observation b5c93435-f166-4389-9de1-3406a681970f · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.424333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.424333Z digest=sha256:5fc658331e55a61af83aedaaf5444dbb4eeb18dee9bfa78c127fb346b54b102f

Observation 163f588a-820f-4e44-96dc-363e3f0a1951 · outbound

This paper cites InternLM2 Technical Report.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models InternLM2 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.418696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.418696Z digest=sha256:0876e8a991d2865eccac6219b79f95aefd1ca7dfd5e4c396d49957be139041f6

Observation a75f2100-c4ac-4857-8313-27be9a211d3a · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.437988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.437988Z digest=sha256:76621a0ab70ac00b1671dfb2d35921ac2a7a79f81bd6916712e5b41c4f931918

Observation 89a7aa44-fea8-48ec-afbb-113bec72c9a3 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.430841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.430841Z digest=sha256:f11e04b91d8e5e625faa6f7ca4fdb789806eba0259712aef89cfcf118dbd6171

Observation c689760a-0b09-4ad3-88ed-e4cb4f398809 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Gonzalez, Ion Stoica, and Eric P

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.449752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.449752Z digest=sha256:2bc719b5befa235eeada8e206ca044d49f21387c23a147f27bc5c5b5d33ed552

Observation 129077e5-ac55-4b2d-8b39-f1a3bbc40f0a · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.443650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.443650Z digest=sha256:579ab1a8dc89d19f2666c241467b2e8f220e6301d54d300912c903ac532e3aa7

Observation 55d36d7d-628d-4b31-8ccd-196bb6d8b1d3 · outbound

This paper cites 2013.Human information processing: Vision, memory, and attention.American Psychological Association.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2013.Human information processing: Vision, memory, and attention.American Psychological Association

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.461933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.461933Z digest=sha256:5e7755999ad74f8971f7186dce166adae43965a1ecf4d936efb9fb2f5be6c367

Observation 1ff7a2c9-2aae-46fc-a09c-759ecf7c37ab · outbound

This paper cites VidHal: Benchmarking Temporal Hallucinations in Vision LLMs.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VidHal: Benchmarking Temporal Hallucinations in Vision LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.455277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.455277Z digest=sha256:684b4e080dde48821f089a734d579a2c10832159d834ff0d8529b5a564898d2e

Observation 01e836b6-5343-49bf-8ddc-8e925eb88abb · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.471710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.471710Z digest=sha256:411117acd6c2e8ea7aff31fa5e0045eadb6f6949b671bce0f39216325f8fe28f

Observation c3ba4ab5-7b46-4cf8-a28f-7f50626fb6bf · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.482556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.482556Z digest=sha256:4150023b2752889decc325bffcd81604da5e3060ba791c190f667e9f4b6aa0a4

Observation 0b62f425-e764-4385-a3c0-7c7e35137384 · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.477678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.477678Z digest=sha256:6351e1a28e89f1fd0d220da9c08f9c532af2aa1e089a011d846969fbc0a6514b

Observation 5a94453e-fb10-4388-a852-2c8bff38d7bd · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.493996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.493996Z digest=sha256:768bc833b6cca811c63ca7225fe06779f3a33deaee703d51bc0e391686e2d485

Observation 9b19fdc3-d003-42a6-b2f2-2260e3534d61 · outbound

This paper cites DeepSeek-V3 Technical Report.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models DeepSeek-V3 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.488472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.488472Z digest=sha256:b57fcbe5d7f7c23b6ae2b80a39e83eeb89ac2eece82fbad02eafe30e3eb9ab41

Observation 01668f78-84b6-4797-9512-062dda1e2696 · outbound

This paper cites 2002.Mise-en-scène: Film style and interpre- tation.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2002.Mise-en-scène: Film style and interpre- tation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.505747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.505747Z digest=sha256:c03260293da26add2ba1c7701f0a70307f6f8fbc93e30c451a97b0059c2c91a1

Observation bac15c89-d678-4e26-aca4-202fc2396257 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.499149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.499149Z digest=sha256:055129a6286b3d4d96874c258d4d9ca6a1f713f4d9cdb09e6b12c9b025fc5db8

Observation bdf5e9e9-13c3-476d-bb22-a90c3b2bf7e8 · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models ImageBind-LLM: Multi-modality Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.516191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.516191Z digest=sha256:c18bc954e61598e1ba44de7faa1df4fe2cad6a97fa09a8f19836b0daa48fc134

Observation fde065e8-3b2e-42e2-93f2-e98e9cf51af7 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.523017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.523017Z digest=sha256:bd85631dcd444fe08f0c480bc5737987037ecdce9e6d8d0ef36b2b217b3de803

Observation 985e3870-8e8e-4c16-9e1b-8bf5a67d78e6 · outbound

This paper cites GPT-4o System Card.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models GPT-4o System Card

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.537462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.537462Z digest=sha256:66dacefa3bcb51c98006d1bdfa09aa42a86655f051e00e02100587dddc12c5d9

Observation 379ccc35-1a65-4357-98b5-e12bd82b8bac · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.532290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.532290Z digest=sha256:4a1e62a87fd926f9235e2f4a8ca9bba135331106055d97b24afc38eb8222e2d4

Observation 87382e3c-6127-4ec1-a8b8-a85f99667977 · outbound

This paper cites A Comprehensive Survey on Visual Question Answering Datasets and Algorithms.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models A Comprehensive Survey on Visual Question Answering Datasets and Algorithms

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T20:33:54.768847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:32:56.548575Z digest=sha256:d154d3c241deb5c83336b41564aac824cda101e1d35c945f11e6cae2292b04c2

Observation b7e168e3-710a-4620-8786-2f19419ce369 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.543391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.543391Z digest=sha256:24c8970bdc79dbf8254e90b74f38b8783fac3c51184ed59d763fd71c7fe3406c

Observation 26f2bd1a-148b-4313-acf8-1c91b77cd277 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.563400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.563400Z digest=sha256:fab916d7813de5cd796ec879f94a2348de42258ff8104d8f2d0c4b59fc18d7df

Observation a6efcdce-56d0-46e0-ba0b-dbc53ae73aa1 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.554860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.554860Z digest=sha256:9b5f6e34192ec6b1071d3f334605ecac1737754d00c57ad97974262e120d7c24

Observation d8d35c71-7537-4e24-ad74-7058134d5267 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.574322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.574322Z digest=sha256:3379b101252415a664a5f75eab2f0cbdcb9d93dab118c516dbee8b0d36ad92ca

Observation 0c100652-0ca4-4266-9173-dd453fb46011 · outbound

This paper cites Berg, and Mohit Bansal.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Berg, and Mohit Bansal

Reference 37

Resolution
malformed identifier
no resolver link, observed 2026-08-04T20:32:56.567981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.567981Z digest=sha256:aa53b8040a59fadf4e4229fab89ccece47a7fb8257b7a2b24f31b86fa7edeb7d

Observation a133ce72-f350-4e46-b0ab-30c9f0976478 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.586319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.586319Z digest=sha256:d5352b5ee5b74ee1c21ea25562467e269d16ac6bfd779cc3208a9818daabbc5d

Observation 6536a079-2501-4bae-aa6f-40347e8612a4 · outbound

This paper cites VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.580184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.580184Z digest=sha256:f41df8e4d808f9cdfaa4db4c1250ddcb4ba18382c19dc000923734c30fb584fb

Observation e5f377ff-4841-4fa9-a423-5bb7e32e9fdc · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.597946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.597946Z digest=sha256:a1661e5b51f6e4a6665283324d5c04a7abd29ad0c49fc20bad6da56da7646c32

Observation 46b695c2-4d8a-437a-b8c4-2b14d639210a · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.609106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.609106Z digest=sha256:ac2cea0d1ca2d55f362901b98be90fe4a752671c13a7757df8c5a8bc97639703

Observation ad4666fc-3b11-4d3a-9e02-7d8192d1a559 · outbound

This paper cites Making Long-Context Language Models Better Multi-Hop Reasoners.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Making Long-Context Language Models Better Multi-Hop Reasoners

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.603444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.603444Z digest=sha256:736ea658a0e12b01141ecb4890000ed611526b09125984a75615dd7bb49fa12f

Observation 1354d313-04a2-4e86-a48a-d63ce58b7a6e · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VILA: On Pre-training for Visual Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.622377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.622377Z digest=sha256:df84fe036eee2daf629ee6749e9fe6b66674905d325699073336e0c8af5202b9

Observation bb973a2c-59d6-40c4-ace3-c32c41566e5f · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.615854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.615854Z digest=sha256:11268b9206caa670f28bd31bc25c53ddffcc24ebce4240452551e0ae0c53b3c0

Observation d376f94d-7f43-4ee1-ae00-3a6a48c7a26b · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.638414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.638414Z digest=sha256:12cb1ba1e45172be55681c3e965f5387231d75f7439b7d9898ee359a1b40ddd7

Observation 495627fe-3d36-4200-a04e-5b3679e4326d · outbound

This paper cites MM-VID: Advancing Video Understanding with GPT-4V(ision).

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.628654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.628654Z digest=sha256:78bcbce986f63cdedff24f58756b4d278c51663b82c6dcc964fe7dda1ce0cf03

Observation 06da0f1b-aa4f-42fb-8eea-179b37a4ad6f · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.656822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.656822Z digest=sha256:2c9c9de4869039eeec6ddb0fa6e39702de44a744796046beb7fc003ad286e46b

Observation f6b3f7e5-12bb-4488-a772-6471d83c52a9 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.649625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.649625Z digest=sha256:f010a2c79c47d585b1a061782ae14b9ad03bac5ccfa7df64024c9c08013063c5

Observation 97bd2deb-23c0-4f14-9651-fec031ddc604 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.670143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.670143Z digest=sha256:f85ccf55ec9f0001d77fbc4f0fd6c7ea35304f6f182a8b2a9b72945ad53fd8ce

Observation 8bf8dc53-b87c-4285-bdcf-5ad10395a3ea · outbound

This paper cites PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.662679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.662679Z digest=sha256:c616cf171cdb8c3b8baaa5a2435f9aae4e614095083e3e13ed47c6fe0262096d

Observation 58f373fb-7795-410a-b6db-af1f1b950951 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.689614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.689614Z digest=sha256:0eb1d9b31325b74f6c2468e9db4968d635bc7a981b94314daf4fa7f5cd75ad94

Observation fd07ad05-d3d8-4c9c-8428-c674b264f753 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.676255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.676255Z digest=sha256:d8050219b20fa7587a05fff1fa3058c2db7c1d828549e8b4ed55247aacd8bd85

Observation 8291684d-8999-4361-96c2-dda099eae3b5 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.684255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.684255Z digest=sha256:4731709beabf6389acf2218ee81c87cbf9af16e69b1e840be995336bbfa08590

Observation 25003f71-9556-4097-860a-2ba84db3ad0d · outbound

This paper cites O’Connor.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models O’Connor

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.708218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.708218Z digest=sha256:7bad92f2cb4f2d616d0e93294ff869248cb53caae147e0f5ceb27cebfea28262

Observation 0ee60542-d7c1-4097-af91-440d060e6b87 · outbound

This paper cites Foundation Models for Video Understanding: A Survey.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Foundation Models for Video Understanding: A Survey

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.695590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.695590Z digest=sha256:bc459f8dabf93b8fa3dddfb0d89000be91d5c63cdab48f0b34c4c9b32cbbd768

Observation 2f6e6c7d-485e-4426-acb8-9bdb093bfc71 · outbound

This paper cites 2016.From Human Attention to Computational Attention.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2016.From Human Attention to Computational Attention

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.701856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.701856Z digest=sha256:14fb0eda2bcbe6b3d2eefd3f0c80b2e432db9cc48206609b59a39cd2fce346eb

Observation 18ffc1c8-dede-49d3-9eb5-78e45b38b78a · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 58

Resolution
verified exact
doi, observed 2026-08-04T20:33:54.420702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:32:56.726333Z digest=sha256:0ca4b06760c380e47b4617b71294b374a40fb6c3e36eaa7335a12274407cb250

Observation a9428a4c-7806-414d-af13-5fe53049cac8 · outbound

This paper cites 1999.Foundations of statistical natural language processing.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 1999.Foundations of statistical natural language processing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.714924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.714924Z digest=sha256:6decdedd644efcb0451986ff10737bf88315d20bdd0c5bb393067d5a415218dd

Observation 0caae250-134e-44c9-b60d-100de4137922 · outbound

This paper cites Large Language Models: A Survey.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Large Language Models: A Survey

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.719979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.719979Z digest=sha256:a29cfc4bc30599bea2dc777cf186d487c4c5406a86648e8cd06fd78e9722da05

Observation 3e437a63-3153-4989-8783-835832e93c72 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.753734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.753734Z digest=sha256:ee52a4207d50a5c094966ff4ee2fd24c7ca997cd880abb98cb3d1a73738aca29

Observation 21a2c4cd-3046-4363-94bf-87e42334fa85 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.736581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.736581Z digest=sha256:9eb0b8504133ff8720fd3755069e2fbdad4a69eb46a972d709c78ff98430308b

Observation cf239a5f-b4c5-488b-84dd-7fca0def9139 · outbound

This paper cites GPT-4 Technical Report.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models GPT-4 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.744472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.744472Z digest=sha256:f410d844f7319593d1d1e4a99e6a1547c93a1b96a5e28b9bfcfca4295ba05b14

Observation 4fad3c1e-f926-404a-96d2-75968684131d · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.787102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.787102Z digest=sha256:6f1ad4431b4c1932e291042848f6863226a0b6e58b698bee6b5edc67115a4c76

Observation d2e64a15-bf18-4644-86c8-68f0a76d935e · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.761348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.761348Z digest=sha256:cb527c4e4320faed4e1838d5cefd46e0328bd2e00b20f24cb3f171b11c364a3b

Observation 52cab061-e964-47fb-af3e-dff4cd902496 · outbound

This paper cites Girshick, and Ali Farhadi.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Girshick, and Ali Farhadi

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.771366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.771366Z digest=sha256:5a8b294315e5dc5c1d905bf762a2fa4977e2cc4a174af04ad31e4a31cf879770

Observation 48ab59c9-9d8a-4f19-843d-204ab322525c · outbound

This paper cites VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-04T20:33:54.254429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:32:56.811561Z digest=sha256:c5ac8b30d626af9d72cf6fce5fef1c2474e57f7cb4258927be31c9dfa3ed007b

Observation b7efc35b-e33d-4610-9fea-8e2297ab618c · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.818052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.818052Z digest=sha256:a98593f5ec7269f0f55f8fa46e1852ecf7ed4ca14985d7dd303ff18e8095c665

Observation 8502143b-8ff9-4926-ac1c-d057501e47a2 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.791245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.791245Z digest=sha256:c955ae1b1e34f6655dd239a2c2eb6241b88e74f6bb7f5b081c4804222fd06f98

Observation cd7a7a2f-5127-4713-9ec0-14b0dc18d603 · outbound

This paper cites A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.800975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.800975Z digest=sha256:10262d57a90f92f191fe3dd8f71156facb5ca7c3f0ca85ee7f543e64b067e5ce

Observation ba898d62-0a29-4e56-8bf4-2fa5ddf24e8e · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.860424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.860424Z digest=sha256:c68403394c256eae5bd6fdbd118513794ed14396c5ddf31ddb69303cbb01fe9e

Observation 2a774b62-1def-48a4-8bf3-6a5b37fdb8cb · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.881841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.881841Z digest=sha256:81af66c9caae16bf2f0cc420319f44126cbe0ed9c99775309b30b7aaa5899662

Observation fd36a8e4-b0fc-46e9-98bf-99b284787fc0 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.838540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.838540Z digest=sha256:2bbaff1a44635484994858e0624f9091dcef6708405a622f3da65a999b5fdaa2

Observation 18729337-7b79-43b7-ae5b-c8c879c8914c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models LLaMA: Open and Efficient Foundation Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.919983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.919983Z digest=sha256:487cdd9c4f5ee108cbee86d8d8c4fbf9ec34450787100b07d1c195c68c20313b

Observation a0306755-04af-44b5-b69f-1995e547ec4a · outbound

This paper cites 2013.The Oxford handbook of sound and image in digital media.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2013.The Oxford handbook of sound and image in digital media

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.929596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.929596Z digest=sha256:f96c8d200127b522ba5720dbfdba08712eb239f619481586fbf0d69c73c062cd

Observation a0173072-d145-4abc-903e-ca8a286a69e4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.938002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.938002Z digest=sha256:5c6bec7a91fc9667a22bbc429fecf535ea909a1ee6fc038294671b36d00ff42e

Observation 0c613ce5-c7c3-4050-9a71-e8656ada9263 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.907053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.907053Z digest=sha256:56d702b074c6aff270d8e50338ea6e935499560628c1076ff617d3781c90557d

Observation fd98d0cf-6b39-4998-a054-ef3e5fcfe619 · outbound

This paper cites 2009.Multimodal Sig- nal Processing: Theory and applications for human-computer interaction.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models 2009.Multimodal Sig- nal Processing: Theory and applications for human-computer interaction

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.914794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.914794Z digest=sha256:804577022bb1ad3e6a7f67f139aac92266a8dee033d3f5a2d5577dc6228d7e81

Observation a55a2557-5730-4d44-8d43-a7dfa88ea0b1 · outbound

This paper cites PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.967932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.967932Z digest=sha256:16634bf883ad7f2f955020c2510b904867f3f4bd8fba36bc499793520b9d7df9

Observation a1cf3dde-b5ac-47fb-af1c-7cd6dcf65ebe · outbound

This paper cites Vript: A Video Is Worth Thousands of Words.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Vript: A Video Is Worth Thousands of Words

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.975944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.975944Z digest=sha256:d37a5a45f003854f40c3c5c2ac0cfe2de106ad11d06f338775fb7653f2d21901

Observation b4a199f5-70ff-4355-8060-0d6ce4e43a0d · outbound

This paper cites A Survey on Multimodal Large Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models A Survey on Multimodal Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.981480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.981480Z digest=sha256:5bcb326a2b77faceee60b95ff588d9a961f992f6061ae6ecadb901d12d66fb0b

Observation fbe2bbd1-7f04-4423-aaf0-1df9d433dd98 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.945883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.945883Z digest=sha256:5147fb7ea27ee3ae63651495ca61a8c5af546750035fe1e0b32afe876bbebc87

Observation 14760259-ed39-4b68-b20a-974d045ae3cb · outbound

This paper cites VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.952259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.952259Z digest=sha256:efda25a78f293428eaed406122051d4edd62a0fbb9a7b11e8bcb2a673b9def24

Observation 13186908-81bf-4add-9d15-bf2d0dd2677d · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.960144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.960144Z digest=sha256:be2060fe6fb96aec6ebf27a7c662840c0b11b1f82ea42916d9f7a91e20136b22

Observation cc89babc-b975-4d58-9f7b-b9b38196ef01 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.016518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.016518Z digest=sha256:6f324835dcdba1661291c78df07b22c6a47fc4cd58711c03b34e564dc606b187

Observation 96c2936d-b272-44d5-a83f-ba6c9623791d · outbound

This paper cites DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.027292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.027292Z digest=sha256:f3110e56051f2e736b49750c4cfdf8380b82086b7574d8fe2a781f83753f9245

Observation 9f0fd19d-6798-437b-a629-a7d532ad0a5c · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.038704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.038704Z digest=sha256:d8ae81a6791acf9a10c747a91e0301ea497d4a8fc0850ac58dc15ed8b9d24ea0

Observation 288f0ac9-7ebe-4473-9e8d-6d6f16129ce4 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.987009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.987009Z digest=sha256:08c39f0c6846ab1be9a56a0bb730206ee1baf9be7d4cce6a75e18b26173c7b58

Observation 25887e44-7304-4c90-8082-a4d8b91cef5f · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.994039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.994039Z digest=sha256:cd39d0b12f02696bb9640bbf0526c63ea0f78d83efc999e873922e23fb8d0eaa

Observation 3b0cb272-5819-462f-a3d7-d55ee8d43571 · outbound

This paper cites arXiv:2409.16597 doi:10.48550/ARXIV.2409.16597.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models arXiv:2409.16597 doi:10.48550/ARXIV.2409.16597

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.000075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.000075Z digest=sha256:4e5fb082a1196a479fd060a64b595c0a7b8188b3c4f4f099af7c6bb82b67b414

Observation 78f725de-5138-4e5d-8427-d6f0628d5ddb · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.007950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.007950Z digest=sha256:a5972ddd8ca6f547c7bdc58eb6d7a63efe1b74f61b0651ba334669e9ba708cb6

Observation 2999ef6d-299e-4698-8d5e-77d45e0f5f99 · outbound

This paper cites an unresolved cited work.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.050146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.050146Z digest=sha256:eccddc236ec093e8b9bd627a6ccca28479efcc360e4830ac505ccad51bc0f6e3

Observation fea97aaa-4d6d-47aa-b7ce-35c4e290a4f6 · outbound

This paper cites (00:38 – 00:46) (CIV) Q: Which individual was the one who spoke in the video? Options: A: A man in a white vest was speaking in the video.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models (00:38 – 00:46) (CIV) Q: Which individual was the one who spoke in the video? Options: A: A man in a white vest was speaking in the video

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:57.059594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:57.059594Z digest=sha256:68d4505ac8306d4fa23126188b1b2cf3e36f9aaba3a3710d74414be83a1122f2

Observation 7303eb7a-f8ea-4545-a508-6cda5b5daacb · outbound

This paper cites In2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models In2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016

Reference 2016

Resolution
malformed identifier
no resolver link, observed 2026-08-04T20:32:56.778550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.778550Z digest=sha256:8c9b51aeb9584b5c9d8cd017b563e26dc48cb23c44c01ee80c3bb082f5926be4

Observation 7c1dc52b-a6f6-401a-90ea-ad035364cf8f · outbound

This paper cites arXiv:2312.17432 doi:10.48550/ARXIV.2312.17432.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models arXiv:2312.17432 doi:10.48550/ARXIV.2312.17432

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.895347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.895347Z digest=sha256:d723d4224dc8b6f3fde59e7faf24cd543996edd7b24c9bba77fa0641f9abc614

Observation 1f7800d9-7adc-415d-b298-9ecdd297d6bf · outbound

This paper cites In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.379260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.379260Z digest=sha256:9a8d9cd94d2fb12147d66298c2edcbf797e9ada9fc3dabd2745adced07bb4f63

Pith citing papers

No inbound Pith citation observations are available.