Pith. sign in

Paper Citation Record · LEDGER

Valley: Video Assistant with Large Language model Enhanced abilitY

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 70 inbound Pith citation observations for arXiv:2306.07207.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.07207 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 70 of 70 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:01.147624Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.409899Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 897e4d1c-875a-441b-b2f8-26c72af5d579 · inbound

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension cites this paper.

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:59:50.650439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T16:59:50.495335Z digest=sha256:5740707f9cf94b101730caa4832ff3fb03e407a6d6478ce4ae76a0a81e01f2eb

Observation c342be09-1bd7-449d-b81d-5e098411cd7c · inbound

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection cites this paper.

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:08:01.242531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-14T18:08:01.166072Z digest=sha256:ea1e97bf5bad53389d7c4d719d3067c45370f52183e471a4b653bf94360b5137

Observation f560f83b-f72e-4b21-9b3f-27f69a86c5a5 · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.014045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ffd31ed43ca0bab23d797f2afe7f624f5c6ce70125c3ca216bc7d57de958ba4b

Observation da2205a2-3c68-494f-bd58-63204b3bf502 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.769669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:b87d52c0758307e84a0ec10ac5b72090de2a78d5ebbe40a2ef24a12cfafdebcc

Observation d8272357-eb90-40ef-b937-06bdfc05800d · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.884272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:cbbc83e1ab651a0256aca023fff145ca210290ad369435a4d2c2990272ad3c35

Observation 634a0eee-6bd6-4860-99f0-19f0893c45f8 · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:53:33.675561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:fc41722899b56c37554701d2a45450204449f27770d27405a63050161b65e21c

Observation dc8993eb-df65-4539-8b6b-4bd24028e0e2 · inbound

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance cites this paper.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.706077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:f84d10d33c33ef88dcf88ed33ddcbf915351c8da1a54c0e7d1d555baa3715584

Observation 6af2b771-a107-4a7a-a8d2-f466382770b7 · inbound

On the Consistency of Video Large Language Models in Temporal Comprehension cites this paper.

On the Consistency of Video Large Language Models in Temporal Comprehension Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.783631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.783631Z digest=sha256:94fdda3a703676e4621776624d2abec1037bae1f8d878ac5a21f40490828d9f8

Observation 6810e1e3-ec47-45c6-8423-7108dfa86601 · inbound

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation cites this paper.

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T16:42:55.853138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:42:55.853138Z digest=sha256:422f25a5162d4be9787233403e0830230c27f6afdc815ea926c601577d9248fe

Observation c073fbc5-43d2-4f94-954e-6da67bfaee97 · inbound

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models cites this paper.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.701759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.701759Z digest=sha256:ee3a7c2805d11c28bbb11ef8b91075f7df589674d30ce5f319037c89f31d0e12

Observation b880f91c-ec04-4505-b538-480c237f60b2 · inbound

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy cites this paper.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.817645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.817645Z digest=sha256:1632b167ff3e9294f25ae4903cf0e55939540546fe3608df0618d553b1836b2f

Observation 3c47b35c-78e1-4398-81a5-ad030fec65db · inbound

VideoOrion: Tokenizing Object Dynamics in Videos cites this paper.

VideoOrion: Tokenizing Object Dynamics in Videos Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.486398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.486398Z digest=sha256:45367d949e73ac8141370a9b045e902891675e85e6d6504bc7c767c12c7a4fe4

Observation 63f7993f-0778-486b-8eb8-85cf83c68f70 · inbound

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding cites this paper.

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:47:21.256599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:47:21.256599Z digest=sha256:24cb130684617f6b5d36ab926a13266d66ead70a1a2926b46367452b66ef1d3e

Observation 33a71314-8054-4c1c-821c-3ddd7218c9d4 · inbound

TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability cites this paper.

TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:56.156668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:56.156668Z digest=sha256:42ed5c366504070fe98ce421a87a374f86ebdedd4ff0dfd308d38cd1ea3a0cd4

Observation d5acb9b1-de30-4f44-b5e1-d187ce1b4212 · inbound

VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation cites this paper.

VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T04:56:01.335609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:56:01.335609Z digest=sha256:236a3e90162ebf8e724a171ff2e9ded2ca3a0b0d1719e173a34b29a8db87040a

Observation 8f5e0e6f-f569-4ead-809a-2a3b4e6d6d3e · inbound

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos cites this paper.

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:12:43.870984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T08:08:01.889675Z digest=sha256:055054b5e6e134c4d06061f33607ac8e2c63b99cba83fddb57df70509b28cdc4

Observation 3bdb0bea-db6f-4b54-983f-14f23fce4394 · inbound

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation cites this paper.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.390529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.390529Z digest=sha256:460cc592ff11f39c945d10112a15dcf877688c900c498aeb0e694c009525cc06

Observation db3d5011-caec-4608-bf12-0e187d993267 · inbound

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios cites this paper.

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:00.147486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:00.147486Z digest=sha256:a1eed70224e23b7c4eb05832e842db52e658262bc419dfb4f6cd94858799fdf7

Observation b26d8b05-3d70-4934-bfb3-f67df0002500 · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.305623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.305623Z digest=sha256:5b7951651be9e16578cef66ebda4bbea572736bf7891bd6a0d980f93280c1421

Observation f11a2fd3-218d-4c88-92cc-c18cb430c74f · inbound

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM cites this paper.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.861339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.861339Z digest=sha256:05b821a60e815445325f590c14f34094d69a0cf0ad8e82960e5c5cf2827198a0

Observation ddb41289-f379-443c-a354-ad0be15de7cb · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.298328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.298328Z digest=sha256:b7ca44bc50a235fa196f7c25e4c4de98f174fd3070a8e9ef624a5f393cad2910

Observation e4ff90b6-9572-4e9e-bb04-3ebbee2a1dcc · inbound

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens cites this paper.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.262723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.262723Z digest=sha256:e0a8bd3ab958e3290ab6bb0f0c15939f7a7d7c3a46a48925a2588239be6feee4

Observation dc3a7763-ffdf-4f00-81bd-f23db2bba192 · inbound

PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension cites this paper.

PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T14:32:24.171691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:32:24.171691Z digest=sha256:8b7b42ecd24e60ba56b7072d6a8c76b6df889410f1163462d89aa74ae61b80a0

Observation d41b05a8-27f4-4d7e-9a71-fa220b004333 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.594849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.594849Z digest=sha256:e022eaccbe7c668f4169542e6f8fc18c28be02a33e948ec6672270366c0ead49

Observation a070a401-af41-4cf3-bf4b-5f1b4a630075 · inbound

PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask cites this paper.

PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:07.404751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T06:00:07.404751Z digest=sha256:fe490abe76a9ffa461d81e16b50333dcc081b5482a19baaf2471c5d963c29ef1

Observation 1e0484b3-77f8-466c-92d9-979364e2f649 · inbound

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models cites this paper.

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:55.448710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:55.448710Z digest=sha256:f7b88e9c48d0e1e8f7ab4f3847f4febd3dec5f11271043fb07f51d08b7ce4c7c

Observation fad0de72-75f8-4fb7-8187-ea511a871abb · inbound

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs cites this paper.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.420232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.420232Z digest=sha256:87e8b90222a58f8bba3306c13b86379afc0603f5ce27b2be3b43f13c8c4b3d5d

Observation 3e0d8370-5271-4774-b198-8a87f395500b · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 149

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.468244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.468244Z digest=sha256:55af615030308bfb77c47af25ab8633c2d5a51511a6cdbea28245681f65371a6

Observation 0d7c2300-c773-404f-b17e-e1d945aebbbc · inbound

H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving cites this paper.

H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:11.591661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:11.591661Z digest=sha256:5199b6e6e9616e4f8583239364b4fbc85496afb974b491a3a4ad1b4d42123b7f

Observation 8893fbe4-de43-400e-b3e2-bd7511f87085 · inbound

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs cites this paper.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.647673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.647673Z digest=sha256:664c669045f9bea38cad2e17ec3807873f2b2175f3967240d33a217aba5bddc1

Observation 2c3bb458-c345-4ce0-8ed2-259c8fcd3909 · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.097996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.097996Z digest=sha256:370a828c623b09c0778cb9e91b0cf5e735ffb09ffd249b07c471e68584cee279

Observation f9abaec4-93d2-42c8-8110-ea82c247bd49 · inbound

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis cites this paper.

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:03:12.928970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:03:12.928970Z digest=sha256:26286399b8ea70a9653a577906f33b2c1e20cd63c3357b43e182307755745af5

Observation 38287e0a-7220-4daf-9cbd-e6fcfdf0004d · inbound

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler cites this paper.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.857839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.857839Z digest=sha256:f08e6bfceca4ab9a2f8f8b8b2f738c8674f821cf52b10689b7a39ff20ca63886

Observation 0e1d8574-032b-4117-ab75-eb7050bd2c38 · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.158259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.158259Z digest=sha256:ec30a2aa57d7e72e523c2427965b53b44d0583fff11a575903035793b07c9ea9

Observation 08271f24-be62-4741-8711-197e91098711 · inbound

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO cites this paper.

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:22:18.746828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T00:21:51.621582Z digest=sha256:b1919094dc081022ae324564a60a2851427f7988b1b562dda7b153f0750985de

Observation 561ecb5f-7874-444c-8124-e84ec3bc5839 · inbound

ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations cites this paper.

ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:52:16.810074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:52:16.810074Z digest=sha256:6b3cece351787f4761d27d7b3547a8a7dd30ef64e02529773a0a223d60cad989

Observation 53acd390-7574-4ba7-b50f-d3af457f750e · inbound

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task cites this paper.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.147624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.147624Z digest=sha256:6de5e762c3ce43321d649031f8f08dbfa18717526295ed53a13a2aafdd4bda60

Observation 2e37c068-59d9-4c0f-9b1d-cfd7baf896a2 · inbound

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection cites this paper.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.077492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.077492Z digest=sha256:87ed7a002e9f571d986124ff779bd8747c6a7eac278a9474ff17d8c3f1295f3d

Observation 9d462060-56a2-4417-bf57-3e403f24cf89 · inbound

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding cites this paper.

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:21.375731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:21.375731Z digest=sha256:744a1d5a866890a841775c12d348506f5075c4183ee00e1b673a420f80cd6f31

Observation 8cc7f7e2-19f1-46a9-960c-85daf3909906 · inbound

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models cites this paper.

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T05:12:18.391581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:12:18.391581Z digest=sha256:656763eedd46219a100852f8738e1c8e7b142a3552e3fdd206720650383da863

Observation e4a58acc-aa56-4f3d-8ace-d31fb1ac8d54 · inbound

Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos cites this paper.

Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:15:24.032863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:15:24.032863Z digest=sha256:1f5404164426ac906b7dbd867c222698372236a61c09817f339163a1cb23d03b

Observation 6a34b274-097a-452d-b1ab-c0d48b1f9d54 · inbound

Sage Deer: A Super-Aligned Driving Generalist Is Your Copilot cites this paper.

Sage Deer: A Super-Aligned Driving Generalist Is Your Copilot Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:17:05.355117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:17:05.355117Z digest=sha256:2b38731d77ef54b2c7329a5c6acbbdf485288520595c5cfbd3dde04b27b49758

Observation 4cc09498-1b67-4618-8e8a-9608a54336a7 · inbound

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding cites this paper.

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:46.826689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:46.826689Z digest=sha256:fe7403ddc8fb0cdedb7f59dac46a1dfbf6b05456ece8b0d541f65f48e07fc0ca

Observation 1b3c9493-c1dc-4320-aa56-c06c69c66d39 · inbound

Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding cites this paper.

Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:09.970206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:09.970206Z digest=sha256:489e0ec4a8d4f633d68f7fcdfbc33721e983ac09a6ea8df575ea2e719ee4177a

Observation b589f826-7aa3-4d32-9b16-895eee05959f · inbound

Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling cites this paper.

Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:22:51.754421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:22:51.754421Z digest=sha256:82d8cadd6b840fa8d9e3b147090b6a1bc00027e87d69bf81d27febd8689cb8d2

Observation 13fc8891-7b8a-4efc-b6b5-fc1689b8a69c · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:13.167183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:13.167183Z digest=sha256:bee1e5a50370359a0939a80c45bfbea9634e5efe7089b668941e5f6b16061769

Observation c1b43079-4c42-428b-a47a-aff1f18e0f89 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.929275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.929275Z digest=sha256:e73a66a261b1f6668b3deed60ee992653c53825040de764c0c3bcf5cf2fc9365

Observation 1646c9f3-a123-4005-9125-201509576e94 · inbound

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model cites this paper.

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:26:38.671225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:26:38.671225Z digest=sha256:a5a11b9e8d199cafb754acf329263b31a3c5ab1257d96f28eba7abb7b63f5cb6

Observation cf23a219-89bd-4daf-9ae1-0fcb6e0aa809 · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T19:00:53.834507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:00:53.834507Z digest=sha256:6ee8e1bcdb6fadb6ee96c9f9308836b1d478c1772efc37ccd2baf6ab38dceecb

Observation c4229669-f0ae-4e90-a50f-ac08f48c9807 · inbound

UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding cites this paper.

UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:47:10.369499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T07:43:26.736409Z digest=sha256:3d53b2599de28e907229ca46070c173e040545bd8f04d79e63c5eaf3802c0565

Observation d25c8878-8950-4f13-a1ec-3ad356218ae2 · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:37.503819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:37.503819Z digest=sha256:9ca67c999755f4198985ca5fbb0a1ddd1686c65a8e379bce0dcd508928145631

Observation a50540c1-0041-4ace-8dd5-f98bb031f217 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.176899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.176899Z digest=sha256:b079953189091dc7a6414ab5282c8329804c8254899e50c6a20608232c32242d

Observation f21d3b11-8dd8-4f2a-8183-1df5bf4165df · inbound

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding cites this paper.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:35.652829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:35.652829Z digest=sha256:acb7a4b9a498b952eaab63f91481d14e823b2df85f79c7279021c9e2fa1ea153

Observation a6630497-dc85-42bc-83c5-efaa1c28942c · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:54.875262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:54.875262Z digest=sha256:d463fbd1f31042b5e299d08fc0ee66b5f1ee2e0aaa7d6288e6f26cc089762f20

Observation 34b2886f-ccd6-447f-b1c7-55ef0e380930 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 248

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.248728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.248728Z digest=sha256:d7ff6a42aac17f6961b6302aa10461644b201ac622b84e4a6dc998092137bcc6

Observation d5431206-c977-4ddb-9a17-869ded7abf2b · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:47.855446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:47.855446Z digest=sha256:1924f0e5683c29869f48b331e2a3d0de52fa4bb1575c9ff557e150f42ac8033f

Observation 0536ee5a-5b72-4bb6-b5f3-d8f5ecdfcc87 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:30.906964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:30.906964Z digest=sha256:cd121a4504e37e081166e9f9321752dd7d0335e9947e800ae82283ca75106722

Observation 3780dad6-ab2a-414b-8f9d-9e81fe75ccdc · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.556812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:d84877737c91443759d443f1ee589d6cd082e92e360a95b2b5370c7bb4da59ec

Observation 89f4997e-db72-495e-a46d-391c9c729afc · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:49.044127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:5f2e4e3f5edb64d1a72aa87c277c99d80e1cf40b2334fe600dc902ad53709505

Observation 053cc833-2807-415b-b9d3-2620cdc741e8 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.924293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:bffff038362db033e42c8d97b33b36a3be6f6ae90ae5350aceff24795012a7c8

Observation 6d2e3842-bbb3-4e92-84ad-800e932d1c89 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:e2eacea39478471fb5bf47251b31f3668641a2bfacf4fc57d0ce8b8c3bfbea65

Observation 6ef3b4ba-347d-4592-8f41-d38510568d4b · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.566195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:c64282df8ec0d7cb394a629cc9d76d67adef0cb2e1f8ef452607118069d2050e

Observation 9a223249-f3da-42a3-af03-839e912aeace · inbound

ClimateVID -- Social Media Videos Analysis and Challenges Involved cites this paper.

ClimateVID -- Social Media Videos Analysis and Challenges Involved Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.448614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T05:25:05.760276Z digest=sha256:6810c7753383550cf555d3cc6ab6e1a1abc836eec3924b73d37a0c6e3e06951c

Observation 696d60da-b03f-4d11-a459-80b296cf68c3 · inbound

Dynamic Model Merging Made Slim cites this paper.

Dynamic Model Merging Made Slim Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:18:25.408077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T15:16:24.651868Z digest=sha256:79acbc7ed0897c2fefc55f2cd96d46d0a7368d41fdc1bbcae570ee62c347e371

Observation 54e5d202-bf23-46d1-98fb-d3bd45514b1f · inbound

Closed-Form Spectral Regularization for Multi-Task Model Merging cites this paper.

Closed-Form Spectral Regularization for Multi-Task Model Merging Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.379995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T22:40:00.510742Z digest=sha256:b9cb481588f83be91b68862ac76d1a245085e1ba890d83380f965280f5be66e1

Observation 3b7dfae4-d4b9-41b7-8a07-ff1bc762f762 · inbound

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning cites this paper.

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.879950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T13:53:53.520545Z digest=sha256:612ead548d60f6022ae419f621b6796eb517ec0c0c7343d247b25f9efc092391

Observation c2a7f144-34cd-4dad-9773-a3100e589982 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:02.979640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:5e84e2fc7550ef1c3ebda5b5807b334861d92bfd864e0f088f0c4c7898df843b

Observation 0b1ceda6-c948-4963-b7a4-e89a65659738 · inbound

On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning cites this paper.

On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.125759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T11:13:49.266859Z digest=sha256:7baea2bfc6af4bcefe5f30db5ba5e2cd01f6aaf79a6840337ec98001180250b0

Observation 2bf24e0f-0851-454d-97e7-f1d6c87a3583 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:06.411854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:707d39420069ad701579277006c24edf897b93e17c1dd76bf52db652e288f225

Observation c1801018-9997-4953-8a27-fcc10b024224 · inbound

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models cites this paper.

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.685914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T13:44:48.250838Z digest=sha256:a2ab6b3a4e7cc494c34bf4de8c18092c1af41e681c147e171c55d1e4c9284406