Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T22:26:25.786061Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.15778.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T22:26:25.786061Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 71770873-c96e-41b7-8ea5-5f4f414eb406 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Visual instruction tuning,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b60c17eb-72d5-4278-9532-11abf5a921e2 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Vtimellm: Empower llm to grasp video moments,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca80b31-fcc4-4c19-a038-2e67946393ab · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Moviechat: From dense token to sparse memory for long video understanding,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 516af506-3490-4cae-8328-df5f228a802f · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Longvu: Spatiotemporal adaptive compression for long video-language understanding,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc751885-29c3-46e8-a23b-9341c4188ed2 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Adaptive keyframe sampling for long video understanding,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26a6ad1-2c4c-4c83-a035-d270a02de89e · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Videotree: Adaptive tree-based video representation for llm reasoning on long videos,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4514555-9fbd-427b-b890-8d02cb6bd112 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding LLaMA: Open and Efficient Foundation Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 619415b4-fd29-417c-a927-e3f6caaa5805 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Multi-modal generative ai: Multi-modal llms, diffusions and the unification,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99a99d8b-abf6-454d-9c75-d4977f9faf5f · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-llama: An instruction-tuned audio- visual language model for video understanding,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d9003b-0352-48f2-97e9-ae08892790e7 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Fuyu-8b: A multimodal architecture for ai agents,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 186f6d8e-07d2-4b51-bde7-d135f99109ad · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 397a33f9-1e21-4be9-8587-1e01bb531ae1 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Multi-sentence video grounding for long video generation,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06667faa-010d-4fcb-98b3-ba8c654cf06f · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Modularagent: A task-aware modular framework for joint optimization of multimodal large language models and world models,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27f6580e-b61b-4b21-8470-779f462246bf · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-rag: Visually-aligned retrieval-augmented long video comprehension,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9a81adf-f114-4fdd-b6ea-507aa91a415b · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-xl: Extra-long vision language model for hour-scale video understanding,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e370153-360e-4d06-9b8a-c84ad2d9ff37 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc3e9312-03be-4c05-9bc7-dd2aa0f38df6 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e94b226-070e-4c91-8f2f-c5abcb729756 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Qwen2.5-VL Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffdb04a3-c6cc-41a7-8f42-bc62895b3193 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Llava-video: Video instruction tuning with synthetic data,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d97a38c-90ec-48ce-8061-3c336499d9a1 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Deep reinforcement learning from human preferences,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ff3a8e3-9878-42e5-9d63-566d7077ad1c · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Direct preference optimization: Your language model is secretly a reward model,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cab6b849-c83f-4440-976c-279ef2738788 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Modularized self-reflected video reasoner for multimodal llm with application to video question answering,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2f565c9-4d53-4318-b85c-26a6049b825d · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Learning transferable visual models from natural language supervision,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 349e2906-80de-4009-ae36-b0068cf5dddf · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ee821ea-8269-4eae-b24c-152814391dbc · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Lvbench: An extreme long video understanding benchmark,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf72d852-b689-4c96-af23-6c98535cefeb · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Mlvu: Benchmarking multi-task long video understanding,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 157dd742-d96c-4a01-82fe-16c9341215c6 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e56562a6-61e0-413a-9acd-5a1a5c1133d2 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e6a8668-c044-4dbd-8956-2902276bf543 · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Cg-bench: Clue-grounded question answering benchmark for long video understanding,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b7562bd-dc6d-4bb7-b34c-a1c11e98e26e · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Videoitg: Multimodal video understanding with instructed temporal grounding,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d76e0c01-b8d1-4867-8c1c-aa60e84153fb · outbound
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding ordering,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.