Pith. sign in

Paper Citation Record · LEDGER

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding

As of 10 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.15778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15778 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:26:25.786061Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71770873-c96e-41b7-8ea5-5f4f414eb406 · outbound

This paper cites Visual instruction tuning,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Visual instruction tuning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.595392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.595392Z digest=sha256:6abf87fc1d90ec779ec4eae7c16659890f28476b7befbeda517e798f54b4ed3b

Observation b60c17eb-72d5-4278-9532-11abf5a921e2 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Vtimellm: Empower llm to grasp video moments,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.675835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.675835Z digest=sha256:e6c5b511a7a7f86fa786792b2e0b4287d10fa5dbdc822947c07a8846dfec94a9

Observation bca80b31-fcc4-4c19-a038-2e67946393ab · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Moviechat: From dense token to sparse memory for long video understanding,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.762387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.762387Z digest=sha256:7ef99bb19629caaf514f61ccda4a7e6bcbb04e64361503f51e9a16b6b3429a7b

Observation 516af506-3490-4cae-8328-df5f228a802f · outbound

This paper cites Longvu: Spatiotemporal adaptive compression for long video-language understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Longvu: Spatiotemporal adaptive compression for long video-language understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.851331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.851331Z digest=sha256:eabbb1593d194db885bddd0dad3d187f6748cf45b960109dc2297b5ab461ab48

Observation fc751885-29c3-46e8-a23b-9341c4188ed2 · outbound

This paper cites Adaptive keyframe sampling for long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Adaptive keyframe sampling for long video understanding,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:23.929607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:23.929607Z digest=sha256:f1576c195500cb409bdf6d1abaf35294238f96f073660920044c72d9cdc346a9

Observation b26a6ad1-2c4c-4c83-a035-d270a02de89e · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Videotree: Adaptive tree-based video representation for llm reasoning on long videos,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.014552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.014552Z digest=sha256:4b7498c29cbd08d04b1252beb8a2129caab82ab909a62e863f376a31b42f232c

Observation b4514555-9fbd-427b-b890-8d02cb6bd112 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.098278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.098278Z digest=sha256:effd73dd2285c8185d9f1c477092716f3ff64f0f70f6b55aed07a036899b8c43

Observation 619415b4-fd29-417c-a927-e3f6caaa5805 · outbound

This paper cites Multi-modal generative ai: Multi-modal llms, diffusions and the unification,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Multi-modal generative ai: Multi-modal llms, diffusions and the unification,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.184455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.184455Z digest=sha256:77cc964c8232799f4da824fa8c722ab1644c7c335b050ada7518d68b0d259072

Observation 99a99d8b-abf6-454d-9c75-d4977f9faf5f · outbound

This paper cites Video-llama: An instruction-tuned audio- visual language model for video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-llama: An instruction-tuned audio- visual language model for video understanding,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.262670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.262670Z digest=sha256:a796a52dd5e2538b95d32cb4f3741ffc04a782d3a0f53012dee65337f36f2a98

Observation 52d9003b-0352-48f2-97e9-ae08892790e7 · outbound

This paper cites Fuyu-8b: A multimodal architecture for ai agents,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Fuyu-8b: A multimodal architecture for ai agents,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.319973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.319973Z digest=sha256:2a695f8f1b587d1aca9f79c8e04969c1e5f054ec4b0534f3f32cffcf78d00e02

Observation 186f6d8e-07d2-4b51-bde7-d135f99109ad · outbound

This paper cites an unresolved cited work.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.409429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.409429Z digest=sha256:5c0b2f4e85da96d89ad8dd7eb273f430f0643435140fbc6cf7f2b3466ae99abf

Observation 397a33f9-1e21-4be9-8587-1e01bb531ae1 · outbound

This paper cites Multi-sentence video grounding for long video generation,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Multi-sentence video grounding for long video generation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.490559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.490559Z digest=sha256:c4870b8e54d17df8966eafbef8573cb258eb62a88c9381d6d86cfd908e18b16c

Observation 06667faa-010d-4fcb-98b3-ba8c654cf06f · outbound

This paper cites Modularagent: A task-aware modular framework for joint optimization of multimodal large language models and world models,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Modularagent: A task-aware modular framework for joint optimization of multimodal large language models and world models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.568442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.568442Z digest=sha256:47facd30891095bd78b5bd07abd3230a88b551f054544ba4e530bd46342d99e9

Observation 27f6580e-b61b-4b21-8470-779f462246bf · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-rag: Visually-aligned retrieval-augmented long video comprehension,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.659409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.659409Z digest=sha256:ac2d0203340ec6ef3982997b39f5f9b4e9b938e05a08317160678eba8571752a

Observation b9a81adf-f114-4fdd-b6ea-507aa91a415b · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-xl: Extra-long vision language model for hour-scale video understanding,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.726194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.726194Z digest=sha256:56d63693ee5a2c553b1f5dda89a368de0306109d20e5d05437e93a4e83e8656b

Observation 3e370153-360e-4d06-9b8a-c84ad2d9ff37 · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.782850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.782850Z digest=sha256:aaa8ac08475f675fdd4997fedac982b5d91febdb6302ee8697d4b343456c7dd3

Observation bc3e9312-03be-4c05-9bc7-dd2aa0f38df6 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.831412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.831412Z digest=sha256:f3345b59800a835f3e693914c2282b4fa71fa544b62e4f55d0f968fbf4c93f56

Observation 5e94b226-070e-4c91-8f2f-c5abcb729756 · outbound

This paper cites Qwen2.5-VL Technical Report.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Qwen2.5-VL Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.886162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.886162Z digest=sha256:aaea55ec956a7645b610b4eb32bcacdd61c0f1ae04dee3de6a80881740841e58

Observation ffdb04a3-c6cc-41a7-8f42-bc62895b3193 · outbound

This paper cites Llava-video: Video instruction tuning with synthetic data,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Llava-video: Video instruction tuning with synthetic data,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.942904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.942904Z digest=sha256:7faca743538e1b695a115fb3d5a0155b1f44f7985948cff453de02f95821ac89

Observation 1d97a38c-90ec-48ce-8061-3c336499d9a1 · outbound

This paper cites Deep reinforcement learning from human preferences,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Deep reinforcement learning from human preferences,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.981811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.981811Z digest=sha256:e3fad5ebc80f04aaddbfcab9d9dfe484ee6f6ee6831c6076c96acd2a65c6b956

Observation 6ff3a8e3-9878-42e5-9d63-566d7077ad1c · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Direct preference optimization: Your language model is secretly a reward model,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.055475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.055475Z digest=sha256:153ce544bf4f877952a2bfec685a79e91b7e7f04fe0d609d2a3e8deb69f5a563

Observation cab6b849-c83f-4440-976c-279ef2738788 · outbound

This paper cites Modularized self-reflected video reasoner for multimodal llm with application to video question answering,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Modularized self-reflected video reasoner for multimodal llm with application to video question answering,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.141517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.141517Z digest=sha256:9cc18f75355bd2b734d9875aced96f9e4c18e78bd00c648c76afc5a21cbb9e1a

Observation c2f565c9-4d53-4318-b85c-26a6049b825d · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Learning transferable visual models from natural language supervision,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.216497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.216497Z digest=sha256:31aaf8b3d1efd65d9fa3c78c2716c20423dbca5045ee5b9b2826b7d70a390f10

Observation 349e2906-80de-4009-ae36-b0068cf5dddf · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.266600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.266600Z digest=sha256:55afa9eb63a69c1e057ac0decb4cc2eee1c5cfaee41c7e95db22645062028c4f

Observation 7ee821ea-8269-4eae-b24c-152814391dbc · outbound

This paper cites Lvbench: An extreme long video understanding benchmark,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Lvbench: An extreme long video understanding benchmark,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.321124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.321124Z digest=sha256:ad43b8cefb782029a6b45d0b54e1492d787b73f6d2bcf10355b8bc0b3b7557fd

Observation bf72d852-b689-4c96-af23-6c98535cefeb · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Mlvu: Benchmarking multi-task long video understanding,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.405821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.405821Z digest=sha256:ea39b53e9e7c761b9b99d24f3d447084b4eba5d5f2b41295ed419483fe4aed41

Observation 157dd742-d96c-4a01-82fe-16c9341215c6 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.477899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.477899Z digest=sha256:9173fece8551d2f0f59d3d4d2697a94ec755e7141b297f436b0fea8d60c96a22

Observation e56562a6-61e0-413a-9acd-5a1a5c1133d2 · outbound

This paper cites Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.559526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.559526Z digest=sha256:752c8d7621da8c78aba881b6fd33db2703e71a749122d848003474069cae22f7

Observation 8e6a8668-c044-4dbd-8956-2902276bf543 · outbound

This paper cites Cg-bench: Clue-grounded question answering benchmark for long video understanding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Cg-bench: Clue-grounded question answering benchmark for long video understanding,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.637198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.637198Z digest=sha256:b4f99a4ef3021e2e5594b8cc41878ffed76357219280ec68d70297f78faac002

Observation 6b7562bd-dc6d-4bb7-b34c-a1c11e98e26e · outbound

This paper cites Videoitg: Multimodal video understanding with instructed temporal grounding,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Videoitg: Multimodal video understanding with instructed temporal grounding,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.723913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.723913Z digest=sha256:6b6889275d75495a6555811addd1b974ea7434af291bde541edd4d007b47c8d4

Observation d76e0c01-b8d1-4867-8c1c-aa60e84153fb · outbound

This paper cites ordering,.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding ordering,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:25.786061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:25.786061Z digest=sha256:7167f8a9b6643037ee3a68f68b009ae3dee2ce831376ddbb030c4dd9e29ca29e

Pith citing papers

No inbound Pith citation observations are available.