Pith. sign in

Paper Citation Record · LEDGER

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

As of 8 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2608.05780.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05780 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.360123Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved83
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88193e5c-508a-4353-a368-0423a0a7db08 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.967551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.967551Z digest=sha256:6c6b2a357f2e43d3f6e6b7b596673fe1e33e3351b8c3bda7f2f23f83ec466f17

Observation 322731e2-edaa-4f28-8a97-e0626bd31fe9 · outbound

This paper cites 2025 IEEE/CVF International Conference on Computer Vision (ICCV) , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding 2025 IEEE/CVF International Conference on Computer Vision (ICCV) , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.972401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.972401Z digest=sha256:0ea6327b5868751649e10a1f521171417f8d27543a84f67451edb7e205dd02c2

Observation 95c97b24-1dcf-4dab-98bd-8c7ddd47d332 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.976757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.976757Z digest=sha256:fd6c684aa6de2995149c1909f4da26b9bb2a13c462884ffcdef6d70358b74dd8

Observation fe9af7a0-9666-432a-bea8-55eb0c103604 · outbound

This paper cites PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T23:42:38.263471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:34.980982Z digest=sha256:1e0f1dd10566d955bb6323b1dfd074b61cb52f282cc32c3c8001bd17e77a1cf2

Observation 95e3be72-963b-4e16-8dc1-e63d4aaedbf0 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.986046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.986046Z digest=sha256:1cec4a99cb6b4996f29c258afa17d9a79762de26eac98a375e433815cc7f53d2

Observation 91da7fbf-cf4c-42e8-859b-57cdbc5c4565 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.990368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.990368Z digest=sha256:0f7ef96651c8e78b5e488d4706f4931f64e6b228b8d15665c6bf0b3441851416

Observation 1468d980-9e42-4d3b-9c1b-9749f003a400 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.994738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.994738Z digest=sha256:097379e60e71905347bca2d2d538f90bf3e441857915c56f6df577b79c0b92ec

Observation da832fe9-bd47-4155-84ab-af0c83ea644c · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.999140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.999140Z digest=sha256:8ae031b7f1f6c75407c18650746fadfc15d45fbc819ee5b8057bc7c19e51c3e8

Observation 791ff586-6a7d-43bc-80a3-1bcf324214ed · outbound

This paper cites arXiv preprint arXiv:2507.07966 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2507.07966 , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.003049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.003049Z digest=sha256:d17f556dabb14015dbca600178be552e2aeea7279dadc48b535b1456b4266d8e

Observation 1632a862-58f5-49aa-adfd-28d2d8552ee2 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.006944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.006944Z digest=sha256:238ac3e37e34e67938430220961603978ad63bad6b8e10d56ff84c3a32e0e0a6

Observation aece37f1-91a7-455a-af84-5edb8d9f13be · outbound

This paper cites Qwen2.5-VL Technical Report.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Qwen2.5-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.011144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.011144Z digest=sha256:8b3dc9c52ad4c60513a99a776baed9f760a3ad24f94eb32823bea86903129a16

Observation 5e7d020c-7a32-4010-9062-a055c58b0abf · outbound

This paper cites Journal of Visual Communication and Image Representation , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Journal of Visual Communication and Image Representation , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.019850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.019850Z digest=sha256:d67ea595fce4962223f2e9ce3a92c7ca6ca622fb119ebb7324215348439a94c6

Observation 93f062a9-c40d-4f58-b11d-11d09aa7a2b2 · outbound

This paper cites Pairwise is Not Enough: Hypergraph Neural Networks for Multi-Agent Pathfinding.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Pairwise is Not Enough: Hypergraph Neural Networks for Multi-Agent Pathfinding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.024660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.024660Z digest=sha256:c9580f7b18bb66ec795e26edd9178db60c5efa10e2c1ea76b1d81742dab7cfb3

Observation 076c26d2-d334-4f85-ace0-5774c15b4a2e · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.029043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.029043Z digest=sha256:4dd67c64ad26a917f83a2b34b94334473bead422061b61015b6882a9ea9290bb

Observation e6977576-efe3-4033-b901-1ae2ef49a408 · outbound

This paper cites CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.033889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.033889Z digest=sha256:b15c3e5c4daa41512db110baef6b238c1c49411271a5ae41cc7e7e50523e1f03

Observation 49e25a13-f186-45c8-85a8-c1ed82e9cc6e · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Describe Anything: Detailed Localized Image and Video Captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.038915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.038915Z digest=sha256:769979ce66088e56fc74f536608ba416e5e296dda6992df2f75543c4720d144b

Observation 40435e09-dda2-458e-b446-000310e5d923 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.044026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.044026Z digest=sha256:dbda1a755bfeab7b391bcdf0dcd05ad6e41901392c2e892b451d6730c9f74b11

Observation b18ccfcd-fb5a-49d3-a109-90f2e479835f · outbound

This paper cites Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.047867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.047867Z digest=sha256:f298e607602e13142877c0f1d96d68d085746020aa0c76aeb0fe33779ba2c57d

Observation a50bf22f-14fb-4ea0-bc91-1d05444a9348 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Advances in Neural Information Processing Systems , volume=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.051735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.051735Z digest=sha256:7a57dbe079fe60dec06557f0caf4b45db0529c736f7c1fbb5c66000259a422fb

Observation 67558f6b-def8-4ec3-9a80-11676bf87a67 · outbound

This paper cites arXiv e-prints , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv e-prints , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.055629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.055629Z digest=sha256:0cea0a40ed79382bd894779834f3360c61cd49e2211f87a4eb63fa07d3e9ed36

Observation 532e0f22-be28-400e-b515-63fca5abadf7 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.060058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.060058Z digest=sha256:72dd72e73734b2d894213246ea9eef5247bb35d4c1202d857029641a2ac19b59

Observation bc3329b6-0a19-48fe-9578-f8319e3a2843 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.063900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.063900Z digest=sha256:15a0ddc280899253263cf71e50c513a8a046687e6ec906c5a0739be96867d180

Observation 9b1d950f-353f-427b-b794-37e0b9e0cecf · outbound

This paper cites arXiv preprint arXiv:2508.04369 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2508.04369 , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.067606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.067606Z digest=sha256:b3b39f69e484b46e1a61f95ec657a7f35b2f50c5d6b669bd3cc8c2b1b9f12aa7

Observation 577608bd-94de-4b36-a3c6-6450c71c66ed · outbound

This paper cites an unresolved cited work.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.071742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.071742Z digest=sha256:f3b12fdb403c60fbb93d7cff9374ea828464163e98dc3e0f4a5b693558a43298

Observation 9fc0d62d-9555-459d-94bf-0351f54c83b0 · outbound

This paper cites an unresolved cited work.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.075368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.075368Z digest=sha256:6220d0032fcb624104ddc531767a86eaf252dd6de91f20809306ca0c19748360

Observation 1df357a6-03e0-4789-a052-6628592dcd04 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.079279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.079279Z digest=sha256:72b6a8152854379e94e086fe0d9d5f2df2f80bd8e04ab52f4296cc3c48c89c0e

Observation a5af1714-1c85-4d3d-8b9e-8d47f4503caf · outbound

This paper cites Advances in neural information processing systems , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Advances in neural information processing systems , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.083243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.083243Z digest=sha256:5a732ae17e3856059b6ccae7af9b0922278e00e621766b6a8d0589d112cad2ff

Observation 2faaf74f-9daf-450d-a617-3a4466d330b6 · outbound

This paper cites International conference on machine learning , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding International conference on machine learning , pages=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.086792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.086792Z digest=sha256:af9be77806de41dc2c7913320c07f92fac09aeb266ad603bd08678df85f16510

Observation 5a542cce-ab43-4480-b6da-baf3ac9d6403 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.091644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.091644Z digest=sha256:6bc3e89a5a2470ba5b9eb1578e11e7669122ccaba4f75b5c631f11a0ab4a36c6

Observation 308c53c7-5f8d-4242-8dec-85a3af4acfb3 · outbound

This paper cites Long Context Transfer from Language to Vision.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Long Context Transfer from Language to Vision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.096042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.096042Z digest=sha256:4baa7a0ac9444ee1e591e09ce4486a0d691a362111cc5c43720ce438807cbc8a

Observation 1a34726a-c37f-4148-b236-4e7398ac6ccc · outbound

This paper cites arXiv preprint arXiv:2502.05177 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2502.05177 , year=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.100870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.100870Z digest=sha256:33ce9216c2630962bcc630788eb366ee7ab75e0dcab18910e48a575d07c7d9b9

Observation 62eb75a8-496f-4c5b-9297-5a32e09166bc · outbound

This paper cites European Conference on Computer Vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding European Conference on Computer Vision , pages=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.104559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.104559Z digest=sha256:550fea4fffd876a8e7971aa810f3c57b30cba363b802b87fb7526e6519c17a1a

Observation 9d0d5924-46e1-4636-bc0c-4ef0ccd7060d · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.108718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.108718Z digest=sha256:70942011d239da9213076e91074e4a6e1702e46dba9086813d39f3b866620b69

Observation 23615e5f-63f5-473a-ae1e-29277313f056 · outbound

This paper cites LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.112820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.112820Z digest=sha256:312fdf15cc16e935bee5cb4210682a86715aa650065304ced0fe49431dc6d1fb

Observation 3a363538-f653-4df9-86c7-043fdc5295c8 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding DINOv2: Learning Robust Visual Features without Supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.116786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.116786Z digest=sha256:443aea04b250e78dce682c812228178366c18d65bc6ba3a9e231d973e47e8da2

Observation 2528b329-713a-44be-acd8-f0ffde781769 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.120837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.120837Z digest=sha256:57234513a36df8a753be4aa674ca9dc25f97849a2fe781318ee3e613b27cbf9a

Observation 31af810c-b90a-453a-8506-cbaa23064c22 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.438990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.125004Z digest=sha256:3cd212031e5e02dbb2dfcbdb9f6e0b26ce687d3de843d397fa83d2218c53ebb0

Observation fc745383-719e-47c8-9c37-c19f8d4f6247 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.128973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.128973Z digest=sha256:2aedd8afe40eb921ad7b0c202cf56b7e9f3c20b3426daef0c308c28d0d31361e

Observation dd591fcd-b595-4d6b-8978-6099f1586b65 · outbound

This paper cites arXiv preprint arXiv:2510.27280 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2510.27280 , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.133666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.133666Z digest=sha256:4891edc8155e55f9a84d777f0b84557a8dcf1fdd9f0337f0e576a2acde5eaead

Observation 83ce1718-bef9-4d79-a24a-13456167b10f · outbound

This paper cites Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.137666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.137666Z digest=sha256:b12413bfa5859ad0e53c947c16a9645af4d5c2adc9b2d20171e5bbaeb18c3f2d

Observation 1b80a305-06b8-4a90-b461-e9b9d104db82 · outbound

This paper cites FlexSelect: Flexible Token Selection for Efficient Long Video Understanding.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.141746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.141746Z digest=sha256:772081daa66d3ce022d1234ad3200177f1f2b7ed38c3ddf0c1bbee243f998abd

Observation d680571f-867a-4a75-881e-03ebc212f6d2 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.426552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.146226Z digest=sha256:9fe79a98c5f5d1813f8e7e24f828bf3849295e050120f841c813c6a1bbe30a69

Observation 104a7dd4-af46-47c8-989c-76bb7ba2552c · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.413824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.149849Z digest=sha256:696be3e2dd99decf7056632150da34a015f8966ca320696e1172447d22ffb502

Observation 9a674896-953e-4cc1-aae0-29c2ccddcae7 · outbound

This paper cites arXiv preprint arXiv:2510.13891 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2510.13891 , year=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.154446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.154446Z digest=sha256:04be34dcab6759113a5d79ecbe13839eec2add52ac88cd161608494b54d92772

Observation 949b5c8a-2cf7-4ce3-a17a-63fc99fbed09 · outbound

This paper cites Frame-Voyager: Learning to Query Frames for Video Large Language Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Frame-Voyager: Learning to Query Frames for Video Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.158387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.158387Z digest=sha256:6053585e44c4437ca38ab47892903880912c1cd8cf0dd8e4206adee4eb3e0d72

Observation b9026e97-5b11-4dc2-88ff-3fe443f527e5 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.162434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.162434Z digest=sha256:1dcf760e65c8735aa3cb0b5be6cdf14311a12bc6b5f399d19510232f6bf1b064

Observation 4ac55cbc-22b6-42bd-9fff-474e2edde272 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.166581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.166581Z digest=sha256:6792e55003e18f1971a9558d19185e4ca09d00bdb4896bc20ff6cbdc4e4e1507

Observation a341499c-793e-48c2-9e41-3b2136d50295 · outbound

This paper cites arXiv preprint arXiv:2512.11534 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2512.11534 , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.170212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.170212Z digest=sha256:1a53f0f4020db35089f624806437df1f757cec656a36485a6c4e68f96dc59067

Observation f24d6f3c-e584-445e-93d1-99d3b7bafd5b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proximal Policy Optimization Algorithms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.173910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.173910Z digest=sha256:0f2110db1e089430dc2f48c4cd4a63838e10cec20b112e798992aabc292c6a8d

Observation 454c92c5-a9de-4550-9945-7fb52ba7f392 · outbound

This paper cites Advances in neural information processing systems , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Advances in neural information processing systems , volume=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.177728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.177728Z digest=sha256:b703d20c64cd91bccc3943a11a864e10edb112f7846aac6a9a9cb4ee0c9fb3a1

Observation ca11859a-759c-41e5-a5f9-85d0663ace74 · outbound

This paper cites Advances in neural information processing systems , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Advances in neural information processing systems , volume=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.181274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.181274Z digest=sha256:2ae4b0af2c39ce46f13839e837920f0fed08a706c9ef4d747a594f9381f9feaa

Observation ef3320d5-b85e-4927-9168-7fd73b355b84 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.185360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.185360Z digest=sha256:0e53ebffaa984f47fa91a130ea9c3fe7aab72a48d8727c9ca9a14d4fe6608668

Observation 9acbfab0-4321-4bc9-b2a6-ab619266c7e4 · outbound

This paper cites Proceedings of the 2024 conference on empirical methods in natural language processing , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the 2024 conference on empirical methods in natural language processing , pages=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.190392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.190392Z digest=sha256:30bf23568c261a0bd5c3a691ce68620724dd7455022a99becbcd40bbb996caf6

Observation 0dfa2b5b-5896-496c-ba1e-69fde5afd8e0 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.195518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.195518Z digest=sha256:a4f880a0c39dd927c710a05d918ee59bf41fcae2407cbc3c8f2f196d8a031205

Observation a3fbdcb7-97d9-4efe-8d32-f232731bec8c · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.376340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.199407Z digest=sha256:45cb089b4bbf65594cf3c0dc0a9f55b1655e9bae48fa91df536a176bbe74192d

Observation b44f048c-53c6-4815-8e91-8912f439fc1a · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.203009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.203009Z digest=sha256:aa3f6a0c9fc38f9bbd1d3e892a004f159d53bf501de220ee5ed10f3d321b305c

Observation dce91058-7b8c-47d8-99cd-fdaad28ec38e · outbound

This paper cites an unresolved cited work.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.207163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.207163Z digest=sha256:9db1c37cbb94ec32391275c60c7855a06ff55364ad6b20457c3fbe314a4ad9b8

Observation 108275d8-c8cc-4622-9a7c-4a52af0a8537 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.211189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.211189Z digest=sha256:78c0b41d4983377073bb6835f8b284b40a3818bba2a6d802b1e0b9d04dcd5911

Observation 2adaf088-b632-4f60-a978-61943e5b8a5f · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.215228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.215228Z digest=sha256:df59d01f9a313f20e11e955cce7c82db842da8fdf479631a0205ce65303d6e8e

Observation 735d3b90-9c25-467b-b563-83d0b51517eb · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.219360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.219360Z digest=sha256:6a908774a70ec96a4432ecc717a56d5c21dbc0f41652d8a3089414cd52b62f27

Observation 6ff40cdd-5706-4348-972a-deb5b3b8976f · outbound

This paper cites GPT-4V(ision) system card , url=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding GPT-4V(ision) system card , url=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.353966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.224392Z digest=sha256:941ee093d509fc441995c2671d655384a81428307518a85ff65f62275d064014

Observation e7ee0b09-26c2-494a-9af5-f554f2efc499 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.228769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.228769Z digest=sha256:d95b9530f2fe4c3bb007a148c4a5c1052f2ca89111361ebdb3d4c1d130b5c190

Observation acfb088c-807c-40ce-a2ea-8892e3ff6b96 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.233288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.233288Z digest=sha256:f57288f61e1fd4d593afc3426d0e90c1e8c82e75556b27a96ff38141482ce924

Observation 29663a2e-d21d-4b52-8965-a6238af709bf · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.343764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.237827Z digest=sha256:dec94e4c994c1520c305eb9a5d7e5d5452239768631a687f032a68c60b8f07bb

Observation d3ac00bb-b2ca-4bee-bfe6-89ec6ab877b7 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.333721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.241909Z digest=sha256:ecd6a53a6207cef74046042b78948211d32cf1959d6a7b1c87db35ddbe0affcd

Observation 1b31cd46-dfdf-4aac-94d0-baf39675fb3c · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.246208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.246208Z digest=sha256:05944d7d05338a480820075f07cba1d444465ac43c145416d8217d8cd363f1e8

Observation 4b778ed7-8140-41ab-a1f1-b6e80088b44b · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.250414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.250414Z digest=sha256:82dba989c487670562292415af41c61797165e4eeecdafb07ef5eabaa982cb16

Observation 8be2849b-f821-4698-84c7-b874683bd49a · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.254789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.254789Z digest=sha256:887e4c995c03253b70ddd4d91a7764e4fd956b0ce0c23c41a8d5645804a8ca2f

Observation 2adc3bb7-81f2-47ca-942d-a0b06f11b82b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.259354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.259354Z digest=sha256:97c390f5b7da930f5df99280582c3d367c97c9c7f901dedc263fe835e0866fcb

Observation a3caf4fb-be58-49e2-ac26-315351eaeb75 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.263951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.263951Z digest=sha256:0917c524fa9aa314f491e50c9340befa49f00d19cc0eafbcfe80273381fea7d2

Observation ce3af3d7-e684-4f93-85a3-4ea4f6cb84df · outbound

This paper cites arXiv preprint arXiv:2510.20622 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2510.20622 , year=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.267788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.267788Z digest=sha256:78315d0fb605dbde6b66bf057b80e5f93f383b3c7ef33aba1c039c76f0781da9

Observation d5d5aada-913e-4dd2-af7c-e63665cd10d5 · outbound

This paper cites KeyVideoLLM: Towards Large-scale Video Keyframe Selection.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.271317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.271317Z digest=sha256:25264aa3807a7f5ec3bd48f5516945a1116aa15d3029c1c1aef0297172304d1a

Observation 3edf887e-ae98-487f-8ca2-118820cb80a6 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.275164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.275164Z digest=sha256:ad581427ad5b26be2f18d4d8c5c237ae8dba110c55c01fb394cd26086bc10665

Observation ea224450-abc1-42c2-b9ae-bca17a4f097c · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.279225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.279225Z digest=sha256:cf4793ca5b560a7d9bf14ca4ea21d413301c026fd0426ef78ec6c3c2d0467485

Observation 64354bbb-8bc1-43d1-8297-6a99e744addb · outbound

This paper cites European Conference on Computer Vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding European Conference on Computer Vision , pages=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.282895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.282895Z digest=sha256:584ace1a02d04d6cbb2c7b810b40361d9e3385bc535da27f5bb40caed1c65d60

Observation 70505639-07af-4d2d-ac65-935bf4ad9567 · outbound

This paper cites arXiv preprint arXiv:2512.04000 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2512.04000 , year=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.286834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.286834Z digest=sha256:ddec47db23703d57d5f8468bb908490977a446f4df133c9c96e8a76854eb12bb

Observation 7231cc6d-7072-4e4c-bb8f-8dc3ddd4dd39 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.290721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.290721Z digest=sha256:cb900804e065c659c59f53789757018af694167904b4ed9e4c0a5cc1fd2b79ad

Observation 0f75b74b-7551-4bf3-a952-b6523142cad7 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.302472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.294520Z digest=sha256:172343c4949ec4b096845a68afdf0177bb5e7b94bf0a7f8ca90c0aa3205e3a13

Observation 3f747f0b-d0cf-4dcc-9a4d-28bae416a214 · outbound

This paper cites arXiv preprint arXiv:2512.06866 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2512.06866 , year=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.298403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.298403Z digest=sha256:365be44f29fa5f43f0bca50f03c6c85be9a2f57fc426f7deb260a1fad027fadb

Observation e82103f8-687e-4358-be36-fd9b36733719 · outbound

This paper cites International conference on machine learning , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding International conference on machine learning , pages=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.302345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.302345Z digest=sha256:21a1acc32b5419da2a1aff05fde4b36839c34eaed055c4981553b561e2c239c2

Observation 386423a4-cb6a-4ec6-9647-fab143900a02 · outbound

This paper cites DeepSeek-OCR: Contexts Optical Compression.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding DeepSeek-OCR: Contexts Optical Compression

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.306068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.306068Z digest=sha256:8e6b516d217d7fe1fec4a325579f971e4248c29d13c2e0173561f74aa3062d06

Observation a8e0168f-e7cf-496b-be61-0b99e380d554 · outbound

This paper cites arXiv preprint arXiv:2502.02770 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2502.02770 , year=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.310474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.310474Z digest=sha256:a466641500b844ab7a5236422ac26439e4fefc917912c415bbce120580b795a0

Observation 8651f988-a6eb-478d-9397-9933409b53f1 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.314074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.314074Z digest=sha256:be8c1c6e0b47a368db31755141ae525724fb9df68f94fd6f1ed90f0d29f32e61

Observation 0c39dbf2-088b-4d6a-9299-72b5745b961d · outbound

This paper cites arXiv preprint arXiv:2601.22582 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2601.22582 , year=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.318166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.318166Z digest=sha256:e0c952c3b05bd114982feb3fbb02fedaa0d74c8b27f25788b419742cbe633cf3

Observation 89644a61-823b-477b-9050-d893594e292f · outbound

This paper cites CoS: Chain-of-Shot Prompting for Long Video Understanding.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.322337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.322337Z digest=sha256:8df3c2a62196d9df7c3cb59f0827101f6a36a84c86bd3b5c9f849b33eb874753

Observation 0f00d2cd-5d32-4612-8501-582a08c55afb · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.326390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.326390Z digest=sha256:96606be459a3ff95a909d16fb5cbbbb035a7359eb45c8dd7601f167e469fa974

Observation 6e9bdbae-b87b-4365-9025-a400c1fbc98a · outbound

This paper cites Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.330473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.330473Z digest=sha256:f9188980bbbc6231e329f944a9bba049f996db89791bf1f6cfe8eff73b26b7be

Observation ef3c325f-9b54-45f3-b907-81c8326c3dbb · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Group-in-Group Policy Optimization for LLM Agent Training

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.334661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.334661Z digest=sha256:34db42029f00c09ad333fce2b3caa5ade1dfac95f8880064785fa8587cddadfe

Observation a46cfd20-1e86-4467-b0c4-b3d875e8158c · outbound

This paper cites On the Emergence of Position Bias in Transformers.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding On the Emergence of Position Bias in Transformers

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.338843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.338843Z digest=sha256:cbfdbbab89937f05b5deef63a5b2ea0544545fb81c1bb01268735b510b332fb3

Observation df6bbcc7-9e88-4fbf-96b0-9b851f16bda6 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.284992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.343208Z digest=sha256:ae71a2540d7f0ef5cfb6957e5cd34080f2449ddce0d28652e8ca103004f1f4c2

Observation 8aace232-8f19-4705-b6a3-084255d68645 · outbound

This paper cites Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.274866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.347966Z digest=sha256:27024d01c60bcea62f6d688ec6f82330b66153f0eb4bd8d68608690ab51212da

Observation 8a677408-6190-4db3-9c70-bd63b83027e4 · outbound

This paper cites arXiv preprint arXiv:2504.18579 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2504.18579 , year=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.351819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.351819Z digest=sha256:15d223785a8070b01236af3dba45001a45fa0795af3b9711cebe06753eadc9ec

Observation f91a1cdf-23b2-4ce0-b2ec-332e7591b29c · outbound

This paper cites arXiv preprint arXiv:2603.06199 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2603.06199 , year=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.355809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.355809Z digest=sha256:19456bf0bf7a026e10f19b45f21ae44e12bb968ad5a0c083bde22e25b2fe1113

Observation 6083fe37-9c5f-40e6-97e3-d61f110a605a · outbound

This paper cites FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.360123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.360123Z digest=sha256:3576b992bc676c11c406168585a36135181b9b947173cf7d3563f97522acf7b5

Pith citing papers

No inbound Pith citation observations are available.