Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 70 inbound Pith citation observations for arXiv:2306.07207.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:01.147624Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:20:06.409899Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 897e4d1c-875a-441b-b2f8-26c72af5d579 · inbound
SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c342be09-1bd7-449d-b81d-5e098411cd7c · inbound
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f560f83b-f72e-4b21-9b3f-27f69a86c5a5 · inbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation da2205a2-3c68-494f-bd58-63204b3bf502 · inbound
TempCompass: Do Video LLMs Really Understand Videos? Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d8272357-eb90-40ef-b937-06bdfc05800d · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 634a0eee-6bd6-4860-99f0-19f0893c45f8 · inbound
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dc8993eb-df65-4539-8b6b-4bd24028e0e2 · inbound
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6af2b771-a107-4a7a-a8d2-f466382770b7 · inbound
On the Consistency of Video Large Language Models in Temporal Comprehension Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6810e1e3-ec47-45c6-8423-7108dfa86601 · inbound
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c073fbc5-43d2-4f94-954e-6da67bfaee97 · inbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b880f91c-ec04-4505-b538-480c237f60b2 · inbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c47b35c-78e1-4398-81a5-ad030fec65db · inbound
VideoOrion: Tokenizing Object Dynamics in Videos Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63f7993f-0778-486b-8eb8-85cf83c68f70 · inbound
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33a71314-8054-4c1c-821c-3ddd7218c9d4 · inbound
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5acb9b1-de30-4f44-b5e1-d187ce1b4212 · inbound
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f5e0e6f-f569-4ead-809a-2a3b4e6d6d3e · inbound
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3bdb0bea-db6f-4b54-983f-14f23fce4394 · inbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db3d5011-caec-4608-bf12-0e187d993267 · inbound
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26d8b05-3d70-4934-bfb3-f67df0002500 · inbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f11a2fd3-218d-4c88-92cc-c18cb430c74f · inbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddb41289-f379-443c-a354-ad0be15de7cb · inbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4ff90b6-9572-4e9e-bb04-3ebbee2a1dcc · inbound
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc3a7763-ffdf-4f00-81bd-f23db2bba192 · inbound
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d41b05a8-27f4-4d7e-9a71-fa220b004333 · inbound
Do Language Models Understand Time? Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a070a401-af41-4cf3-bf4b-5f1b4a630075 · inbound
PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e0484b3-77f8-466c-92d9-979364e2f649 · inbound
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fad0de72-75f8-4fb7-8187-ea511a871abb · inbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e0d8370-5271-4774-b198-8a87f395500b · inbound
Visual Large Language Models for Generalized and Specialized Applications Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 149
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d7c2300-c773-404f-b17e-e1d945aebbbc · inbound
H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8893fbe4-de43-400e-b3e2-bd7511f87085 · inbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c3bb458-c345-4ce0-8ed2-259c8fcd3909 · inbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9abaec4-93d2-42c8-8110-ea82c247bd49 · inbound
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38287e0a-7220-4daf-9cbd-e6fcfdf0004d · inbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e1d8574-032b-4117-ab75-eb7050bd2c38 · inbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08271f24-be62-4741-8711-197e91098711 · inbound
FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 561ecb5f-7874-444c-8124-e84ec3bc5839 · inbound
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53acd390-7574-4ba7-b50f-d3af457f750e · inbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e37c068-59d9-4c0f-9b1d-cfd7baf896a2 · inbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d462060-56a2-4417-bf57-3e403f24cf89 · inbound
MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cc7f7e2-19f1-46a9-960c-85daf3909906 · inbound
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4a58acc-aa56-4f3d-8ace-d31fb1ac8d54 · inbound
Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a34b274-097a-452d-b1ab-c0d48b1f9d54 · inbound
Sage Deer: A Super-Aligned Driving Generalist Is Your Copilot Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cc09498-1b67-4618-8e8a-9608a54336a7 · inbound
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b3c9493-c1dc-4320-aa56-c06c69c66d39 · inbound
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b589f826-7aa3-4d32-9b16-895eee05959f · inbound
Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13fc8891-7b8a-4efc-b6b5-fc1689b8a69c · inbound
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b43079-4c42-428b-a47a-aff1f18e0f89 · inbound
DisTime: Distribution-based Time Representation for Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1646c9f3-a123-4005-9125-201509576e94 · inbound
Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf23a219-89bd-4daf-9ae1-0fcb6e0aa809 · inbound
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4229669-f0ae-4e90-a50f-ac08f48c9807 · inbound
UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d25c8878-8950-4f13-a1ec-3ad356218ae2 · inbound
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a50540c1-0041-4ace-8dd5-f98bb031f217 · inbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f21d3b11-8dd8-4f2a-8183-1df5bf4165df · inbound
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6630497-dc85-42bc-83c5-efaa1c28942c · inbound
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34b2886f-ccd6-447f-b1c7-55ef0e380930 · inbound
Video Understanding by Design: How Datasets Shape Video Models Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 248
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5431206-c977-4ddb-9a17-869ded7abf2b · inbound
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0536ee5a-5b72-4bb6-b5f3-d8f5ecdfcc87 · inbound
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3780dad6-ab2a-414b-8f9d-9e81fe75ccdc · inbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 89f4997e-db72-495e-a46d-391c9c729afc · inbound
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 053cc833-2807-415b-b9d3-2620cdc741e8 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6d2e3842-bbb3-4e92-84ad-800e932d1c89 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ef3b4ba-347d-4592-8f41-d38510568d4b · inbound
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9a223249-f3da-42a3-af03-839e912aeace · inbound
ClimateVID -- Social Media Videos Analysis and Challenges Involved Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 696d60da-b03f-4d11-a459-80b296cf68c3 · inbound
Dynamic Model Merging Made Slim Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 54e5d202-bf23-46d1-98fb-d3bd45514b1f · inbound
Closed-Form Spectral Regularization for Multi-Task Model Merging Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3b7dfae4-d4b9-41b7-8a07-ff1bc762f762 · inbound
Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c2a7f144-34cd-4dad-9773-a3100e589982 · inbound
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 142
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0b1ceda6-c948-4963-b7a4-e89a65659738 · inbound
On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2bf24e0f-0851-454d-97e7-f1d6c87a3583 · inbound
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 135
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c1801018-9997-4953-8a27-fcc10b024224 · inbound
MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.