Pith. sign in

Paper Citation Record · LEDGER

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding

As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.15778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15778 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:26:25.786061Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71770873-c96e-41b7-8ea5-5f4f414eb406 · outbound

This paper cites Visual instruction tuning,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Visual instruction tuning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.595392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.595392Z digest=sha256:2bfc100732be91d17e149eb0c7dfdbf3f354d8515a73e2393f96aa215d73972a

Observation b60c17eb-72d5-4278-9532-11abf5a921e2 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Vtimellm: Empower llm to grasp video moments,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.675835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.675835Z digest=sha256:9daa00c2e6d8c87fe11234c6972d9d8160695644f1ac840ef1091c3c899a4ce0

Observation bca80b31-fcc4-4c19-a038-2e67946393ab · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Moviechat: From dense token to sparse memory for long video understanding,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.762387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.762387Z digest=sha256:b6fe534f339e8c067c4a0e4f2914f700881e66d8b125596db22b643b76acf60b

Observation 516af506-3490-4cae-8328-df5f228a802f · outbound

This paper cites Longvu: Spatiotemporal adaptive compression for long video-language understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Longvu: Spatiotemporal adaptive compression for long video-language understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.851331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.851331Z digest=sha256:0ee05892c4c60808b53e815496252877e7ec8040682c99e6dda1a683d953988e

Observation fc751885-29c3-46e8-a23b-9341c4188ed2 · outbound

This paper cites Adaptive keyframe sampling for long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Adaptive keyframe sampling for long video understanding,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.929607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.929607Z digest=sha256:a2fbc31651bc2531d95fca5f14879da50207ed54009f4bef1893a95d93f56827

Observation b26a6ad1-2c4c-4c83-a035-d270a02de89e · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Videotree: Adaptive tree-based video representation for llm reasoning on long videos,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.014552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.014552Z digest=sha256:d8beebaeabc3c6acc5dd1a6332cc5f0f0eb8f80ccc1d22f197514b573a8198df

Observation b4514555-9fbd-427b-b890-8d02cb6bd112 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.098278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.098278Z digest=sha256:755e0e15dfc6b7cb565f96e51876242f6b021faeaa1dfd5ed50757eb8a58d277

Observation 619415b4-fd29-417c-a927-e3f6caaa5805 · outbound

This paper cites Multi-modal generative ai: Multi-modal llms, diffusions and the unification,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Multi-modal generative ai: Multi-modal llms, diffusions and the unification,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.184455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.184455Z digest=sha256:117c331789dbffa61ebcf01356acccd4994c6d4273c2c564bb19a875bf19d48e

Observation 99a99d8b-abf6-454d-9c75-d4977f9faf5f · outbound

This paper cites Video-llama: An instruction-tuned audio- visual language model for video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-llama: An instruction-tuned audio- visual language model for video understanding,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.262670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.262670Z digest=sha256:cc63597cca571f980e43492bd45983037f3ec6201e223c3e7d687544f7d1fbc9

Observation 52d9003b-0352-48f2-97e9-ae08892790e7 · outbound

This paper cites Fuyu-8b: A multimodal architecture for ai agents,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Fuyu-8b: A multimodal architecture for ai agents,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.319973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.319973Z digest=sha256:730018ee3970b3d66d118d4b41c767fa1859d9909be35f7602bf0b4d3226cd37

Observation 186f6d8e-07d2-4b51-bde7-d135f99109ad · outbound

This paper cites an unresolved cited work.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.409429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.409429Z digest=sha256:4fb927a46abc096cc6d79ad9033b6dd22856430d0de54552ce5b312a37b7e248

Observation 397a33f9-1e21-4be9-8587-1e01bb531ae1 · outbound

This paper cites Multi-sentence video grounding for long video generation,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Multi-sentence video grounding for long video generation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.490559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.490559Z digest=sha256:76c93f9a7a8f389a9c1faf4f6cf28f95f582ad9853590e572dd0a5dfca0d1b7d

Observation 06667faa-010d-4fcb-98b3-ba8c654cf06f · outbound

This paper cites Modularagent: A task-aware modular framework for joint optimization of multimodal large language models and world models,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Modularagent: A task-aware modular framework for joint optimization of multimodal large language models and world models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.568442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.568442Z digest=sha256:7a576a5e267f8cbaa5b625acf986879a3f358a7096134d5d6e49ed7c8877ea65

Observation 27f6580e-b61b-4b21-8470-779f462246bf · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-rag: Visually-aligned retrieval-augmented long video comprehension,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.659409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.659409Z digest=sha256:6dc317404560893c1213fb31c0dee77eb3cd74dd45ee627019fa553d6e516c0a

Observation b9a81adf-f114-4fdd-b6ea-507aa91a415b · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-xl: Extra-long vision language model for hour-scale video understanding,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.726194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.726194Z digest=sha256:95d3836e4707c7c22a052553429f4b8a78e6d45348e7197bf73888570b1afb87

Observation 3e370153-360e-4d06-9b8a-c84ad2d9ff37 · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.782850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.782850Z digest=sha256:067cbf6a1b1e9ddef7b766e852d4a3979ae5be578248c2fa2466cab20ced8bad

Observation bc3e9312-03be-4c05-9bc7-dd2aa0f38df6 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.831412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.831412Z digest=sha256:1a34d11316e5c5ca23e289b3e953d975d94426017f7bedfee8cce748b1e6845e

Observation 5e94b226-070e-4c91-8f2f-c5abcb729756 · outbound

This paper cites Qwen2.5-VL Technical Report.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Qwen2.5-VL Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.886162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.886162Z digest=sha256:eab48c10b9ea2e16090bc929b39b287c54d2a5f14e5f49737d7fc95b5c565f4a

Observation ffdb04a3-c6cc-41a7-8f42-bc62895b3193 · outbound

This paper cites Llava-video: Video instruction tuning with synthetic data,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Llava-video: Video instruction tuning with synthetic data,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.942904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.942904Z digest=sha256:9cfbbc12c64b0ca751cd286777e73127fae0acd6adbffe5a38492790c4103e7d

Observation 1d97a38c-90ec-48ce-8061-3c336499d9a1 · outbound

This paper cites Deep reinforcement learning from human preferences,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Deep reinforcement learning from human preferences,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.981811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.981811Z digest=sha256:b222f31d9dd3a3de8dd6da836c69b406d416b936373e97a7ee8d222542bcd449

Observation 6ff3a8e3-9878-42e5-9d63-566d7077ad1c · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Direct preference optimization: Your language model is secretly a reward model,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.055475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.055475Z digest=sha256:504482ea9e8c89b689b936fc5236dbf068167ed96d1df2cbaace468e0848663e

Observation cab6b849-c83f-4440-976c-279ef2738788 · outbound

This paper cites Modularized self-reflected video reasoner for multimodal llm with application to video question answering,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Modularized self-reflected video reasoner for multimodal llm with application to video question answering,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.141517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.141517Z digest=sha256:c86196fd82c950664066893ca13bded1242f728b0bd9f4dc750178b9d600c402

Observation c2f565c9-4d53-4318-b85c-26a6049b825d · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Learning transferable visual models from natural language supervision,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.216497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.216497Z digest=sha256:408ce38c352c899544ede61a565998675adeb9cb7b10913d085bbb9e706bc830

Observation 349e2906-80de-4009-ae36-b0068cf5dddf · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.266600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.266600Z digest=sha256:aef41d009f78b864da3c6ab50dd353be730302cd97439475ad6da19fab577963

Observation 7ee821ea-8269-4eae-b24c-152814391dbc · outbound

This paper cites Lvbench: An extreme long video understanding benchmark,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Lvbench: An extreme long video understanding benchmark,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.321124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.321124Z digest=sha256:ee28b167e8ac128da8d421a76c9c3669e4c7f014274b69e8bf23e7f3d754ca16

Observation bf72d852-b689-4c96-af23-6c98535cefeb · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Mlvu: Benchmarking multi-task long video understanding,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.405821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.405821Z digest=sha256:e0a2935a9733a1dde85c31319043c0b00ccec80f6dd3539bae3784118f0e8a42

Observation 157dd742-d96c-4a01-82fe-16c9341215c6 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.477899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.477899Z digest=sha256:876b2e02fbfd2668123633051fd9d3487ae6990b8a4055133be834f4ed6cbfc6

Observation e56562a6-61e0-413a-9acd-5a1a5c1133d2 · outbound

This paper cites Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.559526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.559526Z digest=sha256:506ca1d402e7733d96a6b80b8af4aeb010373447b75ec1f24b853a804ba1a952

Observation 8e6a8668-c044-4dbd-8956-2902276bf543 · outbound

This paper cites Cg-bench: Clue-grounded question answering benchmark for long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Cg-bench: Clue-grounded question answering benchmark for long video understanding,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.637198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.637198Z digest=sha256:8c235cb83a35dbf2efa2a160aaac875e666a4d440459e2aca0da6d549ea3c984

Observation 6b7562bd-dc6d-4bb7-b34c-a1c11e98e26e · outbound

This paper cites Videoitg: Multimodal video understanding with instructed temporal grounding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Videoitg: Multimodal video understanding with instructed temporal grounding,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.723913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.723913Z digest=sha256:c05fcab4a1be77d79fa40e837d873f9f806f7ca66efd89fb3dd09487f621d10d

Observation d76e0c01-b8d1-4867-8c1c-aa60e84153fb · outbound

This paper cites ordering,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding ordering,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.786061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.786061Z digest=sha256:f7839871a79e1fc87dcd8ab64433afecd22a89a8ec9677eed8850af0cc0b9511

Pith citing papers

No inbound Pith citation observations are available.