Pith. sign in

Paper Citation Record · LEDGER

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding

As of 13 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2411.12355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12355 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:43:03.400998Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2192909b-a3ec-40ac-aa7f-5a5b0ac5a6cd · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.840167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:02.999634Z digest=sha256:19e77d76bd00f4ce6f050d4ad68d13259c269fe9bfba744cdcaade4687f38cb7

Observation ccf63729-4ef7-4c31-9099-15657401fcf3 · outbound

This paper cites Textvqa: Towards understanding of visible and invisible text in images.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Textvqa: Towards understanding of visible and invisible text in images

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.821791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.005333Z digest=sha256:38bc6c144bf99d0d0bd6811e7bc489fc5a582238ab621769237a5b1d62e5de25

Observation a7df40f1-6a21-465e-8c84-95e39da5afaa · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.010674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.010674Z digest=sha256:f0f9dc703c8f2241b22c3c0a85c2708ef34408d86035c736f05247a3b1246f70

Observation ae255e3c-586c-4dc1-aa72-d127c3ac2d32 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.802082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.016186Z digest=sha256:cc077d8dade0a704acfb3cf97212cf5548843c17eeb506ee658f5ab67eb79b6f

Observation 036feb4b-a3b7-4b69-93c4-d10d7eba5f3c · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.782344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.021259Z digest=sha256:a0e2fe6b043650e9935e46ecdf83da669318a7614d055dddf2b3eeea4109c8ed

Observation 47c3774e-a343-4e4e-9a7a-3209d40cd88d · outbound

This paper cites Learning with Differentiable Perturbed Optimizers.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learning with Differentiable Perturbed Optimizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.026435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.026435Z digest=sha256:40448882e330137520e8ed9e93c91a245874ff8bf705fe0bb270d3cc13c554d0

Observation 169e7c30-2c2d-47b9-b16c-25b842bf533a · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Activitynet: A large-scale video bench- mark for human activity understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.762216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.032611Z digest=sha256:b6e9a9c280cfabe0909e17b2c3f0fe86e1fceda6d8d732be97d7108c2cb57d83

Observation e2174da7-e12b-41df-aeb8-06e665057820 · outbound

This paper cites Chen and William B.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Chen and William B

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.741759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.037293Z digest=sha256:a52a9efe6487ccb651756fee639ffc78e019e9ee2d1d7292f1bc198f1a86838f

Observation 8c0bab9f-0d08-4aef-b355-b32aadb75f55 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.047959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.047959Z digest=sha256:b17f44ce260e55578db5db51801e5a59c7cf92901469dda418f9e4177282aeed

Observation 7c0a61f2-f1a3-46bd-9800-de413e0a100f · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Gonzalez, Ion Stoica, and Eric P

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.053816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.053816Z digest=sha256:842250277bda918df4048015c791bea4771ea285e335cb3de06ab075aac1db31

Observation 53588708-e7e8-44db-86f2-a24f74c2d2c5 · outbound

This paper cites Palm: Scaling language modeling with pathways.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Palm: Scaling language modeling with pathways

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.058756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.058756Z digest=sha256:771ddce11742d9acb8bd90122aa603881fd1e01183330156dcd9d7d67280af4b

Observation a94e7e06-e704-43d2-8ba4-b05789559123 · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.677150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.063662Z digest=sha256:1dafe708804e956ab5d5b8f93731e92163d0ad90b7eec11fbd1faef0c81570e2

Observation 8923c3c4-b5a4-4d60-b4ec-53ed7bae0b2e · outbound

This paper cites EgoQA: Egocentric question an- swering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding EgoQA: Egocentric question an- swering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.659167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.068378Z digest=sha256:63ad1ef3c7438602c8c9ab23afbb6fadb2acac573927efbb72295fea9b98b6f8

Observation 5178ef37-024c-4463-8b7d-feb5f9a4adf1 · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Study on density peaks clustering based on k-nearest neighbors and principal component analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.640474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.073060Z digest=sha256:8eeebe8dd62a969f62662390b1da5e4c5006b60e868d1ef387a29e57a09a263e

Observation 114aef3b-3e26-4775-92a2-779f5607bec3 · outbound

This paper cites EV A: exploring the limits of masked visual representation learning at scale.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding EV A: exploring the limits of masked visual representation learning at scale

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.622796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.077586Z digest=sha256:a400525a1e1c7f01248a41883c00d64b280743c00fcd647a936182429046a890

Observation 7a073534-89e0-4cd5-8319-8f235bfaaaef · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Making the V in VQA matter: Elevating the role of image understanding in visual question answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.604589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.082185Z digest=sha256:4a641cc98345a6e8ddc1d3e335855a322c5cb68f21798d3e80f6285cf58b2d44

Observation 2b2efb6a-fb1c-44e9-ad25-53a1c8e3cfd0 · outbound

This paper cites Semantic-aware modular capsule routing for visual question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Semantic-aware modular capsule routing for visual question answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.587504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.086860Z digest=sha256:8861113c8ab763b95e82bcc24c08022765a95de79aa01c184d60ff4e869ce8e2

Observation 967e23a5-a606-4c25-8945-23b63da30962 · outbound

This paper cites MA-LMM: memory-augmented large multimodal model for long-term video understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MA-LMM: memory-augmented large multimodal model for long-term video understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.568683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.091661Z digest=sha256:ef363de575360806f584e5b66c96bbaed62a44badef431528896adf90ae3ea60

Observation 5de07298-2e8b-4d31-94ad-f6ed51167c0a · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.551934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.096166Z digest=sha256:b53d85f50acdcf75654ca6246eab46bd8df85e8d1e242678f83fffd2dc927aac

Observation 6d4b9996-60c8-438c-b156-f73174c9c75e · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Activitynet: A large-scale video bench- mark for human activity understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.535708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.100797Z digest=sha256:2ce21c284c00b3d33d623369b0ae41bf455ebeb8920c1a79874c4f89edc7c873

Observation 3d217d76-55f0-4919-988e-a7158c81f2f2 · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.518297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.105442Z digest=sha256:12563fc9152a7de5c2c3cd42e0b6ab84106def49f8efec74fa9604bbc6e13147

Observation 02a8aba9-ed37-42e7-a542-a2bc61b21d20 · outbound

This paper cites VTimeLLM: Empower LLM to Grasp Video Moments.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VTimeLLM: Empower LLM to Grasp Video Moments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.110131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.110131Z digest=sha256:93bc78f235ed25aeb5ba92c5a87f151e3b715ddd2d87992619148868e8e4133c

Observation a4a3e019-fd4d-42d8-8cec-21765358ad1e · outbound

This paper cites Ng, Hongqiang Rong, and Zichen Li.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ng, Hongqiang Rong, and Zichen Li

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.501170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.115496Z digest=sha256:15944db649c4a58f9486e6d7045c7bc7abe59b566298d2471db973a233c44426

Observation 1c862dea-09fb-40c2-8280-c60f3386f8ad · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional questions.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Gqa: A new dataset for real-world visual reasoning and compositional questions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.484544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.120120Z digest=sha256:f353bba1c783635855d1771ca8e25e998e186a5bb41bcb90072699687c1704c4

Observation a63abfb8-567c-49a0-aabe-73c348fdd3a4 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.124790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.124790Z digest=sha256:29bd629076fd4b9c2e4dc711e985caea3792c832f4dc5a9fdebd2bc79fdeee9d

Observation 2611f948-041c-4fcb-8df2-73f54e44b3ed · outbound

This paper cites Scaling Laws for Neural Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Scaling Laws for Neural Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.130132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.130132Z digest=sha256:101f0ca19642808dfe29f32947c211c349a0f142d0297c31b34f58ddf147170a

Observation 7b32a324-e24f-4f7b-ac0d-ebe87d4d61ba · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.463976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.135145Z digest=sha256:42bc04af7a1b1cf4db4f0ff8eab24e6967be1c9929da2ceb90d88d238a1bc675

Observation d1aedadd-e8e9-4b7c-bc4b-b87f8f061e44 · outbound

This paper cites Kingma and Jimmy Ba.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Kingma and Jimmy Ba

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.444757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.140583Z digest=sha256:30b2f08da846ff73df71bdeca034e7ed237c476e29803cd5d83f810bc61726aa

Observation 1358aabb-24c5-4dfd-84b4-abf6534493aa · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.428366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.145305Z digest=sha256:6f970865066f40886a48b7d1e16ca43002c2ccb18b7a05a7ccc007f98d88d245

Observation fe1aeb9e-49b2-4b88-8ada-5996889d545e · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.411158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.150322Z digest=sha256:cbbb4bff18d2f0c99e59003903644dd6faf8128f203095a505d663f5663ec4d3

Observation 8dd26375-64c0-4b7c-87be-1d82ecc6d9e3 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.155428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.155428Z digest=sha256:00bd4c9f5f69fe90fd19a9f5f3eed17b577294123d7f904f204ed58762ca1dd2

Observation 4b36e757-7f2d-47c3-b746-91d1986490f7 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.160467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.160467Z digest=sha256:a99c957f393a939ce85e01be1641c7458770abd69246743411515c636ca76831

Observation d816dd29-779c-4244-9cf9-18ca9a8829ea · outbound

This paper cites Scienceqa: A new dataset for science question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Scienceqa: A new dataset for science question answering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.392817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.165529Z digest=sha256:79f557cd4b92de53bda7c82a15cf24f77932017c205150e9b6602ee17f32f233

Observation afc0463e-dede-48a9-9cd3-afd740fe1de1 · outbound

This paper cites Learning dynamic routing for semantic segmentation.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learning dynamic routing for semantic segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.373424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.170423Z digest=sha256:96f929fa1d8dcdfcfa7d2913a72c7a3c75c4bff2ea7e35a2c0470cda08f83415

Observation eb7e9455-6eaf-461f-8f19-563a7e6d1ab0 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.175123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.175123Z digest=sha256:7877ced20e530779f8efcc79f4baee4879d4626ecac39812cb562f85f0df32f4

Observation 66d3c174-85f9-447b-a60a-1f540dc777cb · outbound

This paper cites Video-llava: Learning united visual representa- tion by alignment before projection.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-llava: Learning united visual representa- tion by alignment before projection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.351160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.180928Z digest=sha256:7df11e012a7231e4fb336e990775a1aa11b74d0db35a1a59bc8ec30e7bfae20c

Observation c1e97409-d45b-42bf-98a6-7a84d74224b8 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.186135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.186135Z digest=sha256:8c90836fd8f6718597bcbc46a5ee1ae2c9a002c5fa4612dd5e50e4ae30d9b1c7

Observation 90fcb0c1-438f-463f-8bb4-6fd700a3c9d0 · outbound

This paper cites Vila: On pre-training for visual language models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Vila: On pre-training for visual language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.333933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.191688Z digest=sha256:71648baa94bdd4f803da02b027366c93872f5780cc230550fd4c624991656e64

Observation 79bff67f-cb71-4e27-9c1b-2d75451d3a95 · outbound

This paper cites Lawrence Zitnick.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Lawrence Zitnick

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.315970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.196680Z digest=sha256:0fd80fc7f2958da719f3ecd9b25de6bfbca42a3569c921cf9ca321a4669f41c1

Observation 25d71ced-da19-4f2f-a516-afa49035b62f · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Improved baselines with visual instruction tuning, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.203557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.203557Z digest=sha256:6b5c5ab3305c70097bc074c722b7c23a47fbeb6dd7fc44f870745ee198e307ed

Observation de0d25fb-e774-4205-9732-0b5a640fbf72 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.208853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.208853Z digest=sha256:005b4fcea51895c14b253312baead715434b4be1bdb6c9ff30598ff00b0aae64

Observation 60a8e0cb-1022-4047-a7ed-825647eab5e4 · outbound

This paper cites BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.214641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.214641Z digest=sha256:549dcbe93542c949734cb65b5291c736c56c773859a84524440b537038b0474f

Observation 1593324f-320e-4af4-9cb9-cfc8b1ed7bc7 · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.220788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.220788Z digest=sha256:b5667bf3f242530461d3cf5109dab8ce8bf9cf135a56ae8ea7eef6fa31a77265

Observation 2af014f9-11db-4cdd-aa50-d2a8f3d2c860 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.226223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.226223Z digest=sha256:49bb1bba243983c46cffed73718577eb64e15c8190eeac08e0e14fcb34fdcce9

Observation 2ffbb597-19fb-4961-93fd-8e4f31a16a59 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.231069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.231069Z digest=sha256:ab91385dad431dd064bc26af992197c7f2635f85cd21600e25615ba9a0730ae0

Observation b8a8055a-73b3-44f0-91dd-8e70d97aa0ee · outbound

This paper cites Some methods for classification and anal- ysis of multivariate observations.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Some methods for classification and anal- ysis of multivariate observations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.272680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.236284Z digest=sha256:5b3857217208f53b4ef64d2d29745a86f6669ebda380571c0a3f3cb8906c095a

Observation 1f1405bd-bdfa-4a08-b6e1-179d714a455f · outbound

This paper cites Yuille, and Kevin Murphy.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Yuille, and Kevin Murphy

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.253109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.241058Z digest=sha256:148126606607bf53646c7f385621fa4596d5562b82883893f9aa4b3cac015cee

Observation fbd14a40-e224-4093-8831-e1162e9cf647 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ocr-vqa: Visual question answering by reading text in images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.245850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.245850Z digest=sha256:09b0770cce37da5d2cd82c5ef1c18556e3628655146f18f4f5a5222ca2113814

Observation ba16ab19-d20a-4e06-9572-00797ee7c6ab · outbound

This paper cites Webvidqa: A large-scale dataset for video question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Webvidqa: A large-scale dataset for video question answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.224590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.251180Z digest=sha256:3d453c87b7851d591969a57dd1369592e98b95b3399eb5711c1d562285d2551f

Observation 6ea3099d-e5d3-48e4-be36-b99bf90505b1 · outbound

This paper cites Introducing chatgpt.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Introducing chatgpt

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.208955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.256692Z digest=sha256:13e58af40860ba42b2e077ece5ce10377bc5816a1bb4c9411074ef2d8169204a

Observation b9e6b080-9048-43f3-8b18-406637cedd42 · outbound

This paper cites GPT-4o system card, 2024.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding GPT-4o system card, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.192873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.261844Z digest=sha256:f01be1b4675c1596549c1064a887caed256ca364527166b4d0be0675c86db7d5

Observation 6ecae718-bdbf-4511-a75c-0ff1b842447e · outbound

This paper cites Learning transferable visual models from natural language supervision.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learning transferable visual models from natural language supervision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.177914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.267235Z digest=sha256:39f643ce88142f4a6b82bd36843d5a31426bacee780aa8fcde14966fcde31bca

Observation b2a3c4cd-944a-46b1-a257-bab491b451bf · outbound

This paper cites Improving language understanding by gener- ative pre-training.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Improving language understanding by gener- ative pre-training

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.162143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.272111Z digest=sha256:ac8c4d5e148434bf7bb0ae7394da9527eedc92527b62661b308e7fdc01f5e195

Observation 80ab209e-d25d-468f-b4f5-f2c99e816a88 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.146210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.277089Z digest=sha256:e10c3db37a9fb6795063e7f1c55f5fdccf8cbc933d1a07930e6d638d4f52dec0

Observation 849ce177-89ab-4c57-84b2-263947569873 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.129764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.281962Z digest=sha256:3bec4f5f4055438ba35330c25a94dab8f9547b298424b06298e4e9030b8f5c01

Observation c8db052f-139d-4721-92db-375d350333a0 · outbound

This paper cites Massof Sarah L.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Massof Sarah L

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.112583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.288087Z digest=sha256:5e5de92cfe500a8f445102a16c068202ffa1aa26d5a7430366a0c4df40ec1f15

Observation e74eb7fb-e81a-4a1a-aedb-f5c5f9d1a06d · outbound

This paper cites A-okvqa: A bench- mark for visual question answering using world knowledge.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding A-okvqa: A bench- mark for visual question answering using world knowledge

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.095257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.292697Z digest=sha256:ba414865ba4e33596e414d388d4a6ebab8aa957c0fb0e09dde440420e0a2808a

Observation 35c00115-a116-4de1-9640-eabf185e24a1 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.078763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.297505Z digest=sha256:7312660e14ed195120e82529f254e61ee0cb82d0da0529bfefdf5fbb2f95f112

Observation 04cbea46-a537-4ccb-a7c1-a20ce0943b2b · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Textcaps: a dataset for image captioning with reading comprehension

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.059597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.302047Z digest=sha256:0ba9925501bb37d98cba26a8d4ce3129b4b6ee89bf6411fc5dacb391690ae246

Observation 07fcd0e9-4a9d-48f4-913e-16a82ba77ff9 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.306611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.306611Z digest=sha256:c14e3e6db54cab5f2b818709722466c4ed0b79aa08182f06c03fce58c3df5063

Observation 244acf83-61ea-4a6b-9dd9-53248aa8ebca · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Moviechat: From dense token to sparse memory for long video understanding

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.038794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.311747Z digest=sha256:aa17b0d951b8dc05a6055c64ea0d2bbc43ef0fcad3ae55bf27733669b9f371ea

Observation adc1f000-41c0-4e15-beea-8be646f352d0 · outbound

This paper cites DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.316445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.316445Z digest=sha256:fe5c9fe76fb6120600da86abce516585be4690ff853f88453bb3cd03ad33b92b

Observation f27cfd0f-7d32-4381-9795-91a3214f24e6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.321491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.321491Z digest=sha256:f63fe3260e7303519490cb385a55b520e0ba67c7bab7bf7ac0531c21f630c790

Observation 43a75b67-c0a8-435b-9aac-d5c43325d460 · outbound

This paper cites Ocrvqa: A new dataset for optical character recognition in visual question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ocrvqa: A new dataset for optical character recognition in visual question answering

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.020499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.326429Z digest=sha256:aea40b7284bd167d5ce9e536298137cdf63b74162906baf5c6f438469a6c00b6

Observation 7e892c22-933f-4318-965b-dc19ed6cf8a9 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.331072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.331072Z digest=sha256:0d6f29f75cb6366f7b869edbad2fb108e96dc25ae01d55f600bcb657206a7350

Observation 97f740c3-73c6-4c9e-8958-67f8bdfdf36e · outbound

This paper cites VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.336459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.336459Z digest=sha256:7a8dedb7bb96f9bcea3e38e8a9d545f1f75511c01188d9a8ab02b68062edba93

Observation bd29c8d1-fd60-4a56-b0ad-e2dc3962211b · outbound

This paper cites Videollamb: Long video understanding with recurrent mem- ory bridges.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Videollamb: Long video understanding with recurrent mem- ory bridges

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.003342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.341593Z digest=sha256:444aabd8e40996fd437d3ae0fa129ef4d8c7b45948a202b12a141c6ff4d988e5

Observation cc9702e0-bbdb-40bd-a4b7-33a4e5b83220 · outbound

This paper cites FreeVA: Offline MLLM as Training-Free Video Assistant.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding FreeVA: Offline MLLM as Training-Free Video Assistant

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.346619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.346619Z digest=sha256:31518b86c0012798a51a79a25d3b69538539f9b90ea2a1ae9c86eda4773755ba

Observation 0d789eab-03a8-4918-a2c7-8a4e8d87ea3a · outbound

This paper cites Davis, Kristen Grauman, and Rog´erio Schmidt Feris.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Davis, Kristen Grauman, and Rog´erio Schmidt Feris

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.986333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.352245Z digest=sha256:905732973b56dcc374885078f50b7db6ee616c961db4dfae82d5dad582d5b505

Observation c824db95-0761-4a49-8aef-c606a9170226 · outbound

This paper cites Deep learning for video classification and captioning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Deep learning for video classification and captioning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.969334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.356916Z digest=sha256:b6bae916d725b47c60dae7553804a160d1034b9fa110b2bcb5ebe7cef8a9c7ee

Observation b0b61000-7204-46e3-9c89-70cbe5eb90f7 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Msr-vtt: A large video description dataset for bridging video and language

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.361819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.361819Z digest=sha256:e52945ab66a3b42a52930f723eed6ec3d49f40058f1df3eb7a4723c1324a1aca

Observation b417a12e-c613-4623-b461-5dcfb2fc48c0 · outbound

This paper cites MSR-VTT: A large video description dataset for bridging video and language.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MSR-VTT: A large video description dataset for bridging video and language

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.940819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.366481Z digest=sha256:27d55eb4f704f78ba5d52abe670c15b3090cc6a8d733247c4bad4cb27a1c1a13

Observation 580d49b8-86fd-4237-b6e6-0aaf0f50a8b1 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.371221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.371221Z digest=sha256:3150f80787a2fa543e14d823943819a97dfac9c1a4ec9590c1a7e1367f4f6f72

Observation 0886b48f-09a5-4d1f-a63a-4601f2da034d · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.376477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.376477Z digest=sha256:606a8ebfd4faf877182dc93ae83b7c3bf082d63d389b01a133f23ae1a644474e

Observation 86ed1b74-7c25-43e6-9888-5a20b69ec836 · outbound

This paper cites Tenenbaum.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Tenenbaum

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.924319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.381552Z digest=sha256:d53c540cc0d380cf9f0279a2285c20c934fccad563e5077fdb849467cecb43c4

Observation ac5110db-7965-434d-a3c8-be105974b135 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.386720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.386720Z digest=sha256:9cf6e25095a6cdce9f08c9c329e721368f8ef0593a1e0f2372e7feae562cc617

Observation 5cbfdf1e-ed6a-4579-8036-f4c7b9d9ea87 · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.391576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.391576Z digest=sha256:31456ed7eac4261b59a6f544801dff8d280ff545a2d587ffbe26e4795a88bceb

Observation f67af649-7e95-4e51-9057-9c8392cd8a03 · outbound

This paper cites Llama- adapter: Efficient fine-tuning of language models with zero- init attention.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Llama- adapter: Efficient fine-tuning of language models with zero- init attention

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.906623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.396276Z digest=sha256:a91276183c59dbfe66333400b415500465b1f244e6ebd4d8150c3ef81c2a3501

Observation f7d1270e-acba-4a7c-b2a9-061431d94cc9 · outbound

This paper cites Please Carefully Think.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Please Carefully Think

Reference 79

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T17:43:03.887845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.400998Z digest=sha256:2db02ea8074916eb95ef628e8a207645e49671692151e83cfea193b9ce426068

Observation 831ab334-8259-405f-a2cb-5c78bfdc1331 · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 200

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.719668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:43:03.042499Z digest=sha256:da30c7fc8683fcd9f307a57c6a6b27d42e602968682b712cbe77fc3c27127b30

Pith citing papers

No inbound Pith citation observations are available.