Pith. sign in

Paper Citation Record · LEDGER

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

As of 8 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2608.05780.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05780 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.360123Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved83
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88193e5c-508a-4353-a368-0423a0a7db08 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.967551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.967551Z digest=sha256:2d58eb421738143673915c11493fd4bc6ed2e15096bfa4291c2fb683b62d847d

Observation 322731e2-edaa-4f28-8a97-e0626bd31fe9 · outbound

This paper cites 2025 IEEE/CVF International Conference on Computer Vision (ICCV) , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding 2025 IEEE/CVF International Conference on Computer Vision (ICCV) , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.972401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.972401Z digest=sha256:0ad2a341043c9bd061a20d70703ac0f12df0d427153ba438c28f6edf4777e425

Observation 95c97b24-1dcf-4dab-98bd-8c7ddd47d332 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.976757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.976757Z digest=sha256:4925775e31919272e13a7470b03a1c6702d850c4eb8911d48014cdd2b11ec035

Observation fe9af7a0-9666-432a-bea8-55eb0c103604 · outbound

This paper cites PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T23:42:38.263471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:34.980982Z digest=sha256:3c6fdf830437d98512aa3cc9ca98ea327e1808bcba9ea3550024f10e08c49bb9

Observation 95e3be72-963b-4e16-8dc1-e63d4aaedbf0 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.986046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.986046Z digest=sha256:a82424d6d26f3d66fc01474740bd741b8f60b8763b4c8c682643f66a6f9fe4fc

Observation 91da7fbf-cf4c-42e8-859b-57cdbc5c4565 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.990368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.990368Z digest=sha256:daa91e7d49943911624e769193d0a258b32cbec7d443678fbbf6b7dc4c475462

Observation 1468d980-9e42-4d3b-9c1b-9749f003a400 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.994738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.994738Z digest=sha256:c0b85df84e2d1bfb30275395731eacd553aeb7b5fc1e6555f95884901ba3585d

Observation da832fe9-bd47-4155-84ab-af0c83ea644c · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.999140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.999140Z digest=sha256:3f98e0c71b03e0825ed2191a658dd6a3d4505908b962656b2a8e0671bb76fd2b

Observation 791ff586-6a7d-43bc-80a3-1bcf324214ed · outbound

This paper cites arXiv preprint arXiv:2507.07966 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2507.07966 , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.003049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.003049Z digest=sha256:ae9e5533784fca01c508f484a598b2ae590df4968e136141fdf2b8494500a390

Observation 1632a862-58f5-49aa-adfd-28d2d8552ee2 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.006944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.006944Z digest=sha256:b3c5366209ffd5bee0753c13f2285c5168ace0c638efd31b2dff4959a82d16c2

Observation aece37f1-91a7-455a-af84-5edb8d9f13be · outbound

This paper cites Qwen2.5-VL Technical Report.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Qwen2.5-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.011144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.011144Z digest=sha256:69f983befab917853e2704679958ecdcbd7febe9b12ca27ac7560bae653b1eb6

Observation 5e7d020c-7a32-4010-9062-a055c58b0abf · outbound

This paper cites Journal of Visual Communication and Image Representation , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Journal of Visual Communication and Image Representation , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.019850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.019850Z digest=sha256:272ad680755283adee7be524a88d7ed99500a3291c61656d423527effe422ded

Observation 93f062a9-c40d-4f58-b11d-11d09aa7a2b2 · outbound

This paper cites Pairwise is Not Enough: Hypergraph Neural Networks for Multi-Agent Pathfinding.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Pairwise is Not Enough: Hypergraph Neural Networks for Multi-Agent Pathfinding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.024660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.024660Z digest=sha256:a235a639030d94aa13f17965e2386bdc94062820dd0d67958198f0af797f6f7f

Observation 076c26d2-d334-4f85-ace0-5774c15b4a2e · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.029043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.029043Z digest=sha256:0c10d1e2343bb5f4d606bfd1dd0e2cc2783e79b853266507d5d916652113f498

Observation e6977576-efe3-4033-b901-1ae2ef49a408 · outbound

This paper cites CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.033889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.033889Z digest=sha256:4309824b613e8195afb313cf7613966d7b4626dce850cb15770a3f5ebc615d0f

Observation 49e25a13-f186-45c8-85a8-c1ed82e9cc6e · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Describe Anything: Detailed Localized Image and Video Captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.038915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.038915Z digest=sha256:2bf0641208b042f5f2abaa781ba7805c1876e98476dd7eca6148a8ae9c14f388

Observation 40435e09-dda2-458e-b446-000310e5d923 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.044026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.044026Z digest=sha256:f93d9bde68f4a2eedbc45e2abb8ebdad1d068ce27e14162f48fa4ae77f977e7e

Observation b18ccfcd-fb5a-49d3-a109-90f2e479835f · outbound

This paper cites Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.047867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.047867Z digest=sha256:010909fdd8a5dc9903ebb38122e06dc98f7e19f0ce50a88c0ad25bc6969cfb72

Observation a50bf22f-14fb-4ea0-bc91-1d05444a9348 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Advances in Neural Information Processing Systems , volume=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.051735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.051735Z digest=sha256:48d9000bc3a0a247e172269453922174ce3e0eea88716f8782213f053bf0cdc9

Observation 67558f6b-def8-4ec3-9a80-11676bf87a67 · outbound

This paper cites arXiv e-prints , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv e-prints , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.055629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.055629Z digest=sha256:04c93ca3cc15481cd2334f288ca8044dae02aaa3e5eb0f598425bab61c528ce0

Observation 532e0f22-be28-400e-b515-63fca5abadf7 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.060058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.060058Z digest=sha256:223e60df53617d02b345a74a3de3664a4f5a3b69b2ad5dd762a874f2851550a4

Observation bc3329b6-0a19-48fe-9578-f8319e3a2843 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.063900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.063900Z digest=sha256:be7ad1670756d79629ff2dc60baf9663f3ac8893bf239aebf9799aadad4ce6a4

Observation 9b1d950f-353f-427b-b794-37e0b9e0cecf · outbound

This paper cites arXiv preprint arXiv:2508.04369 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2508.04369 , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.067606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.067606Z digest=sha256:7df7a3247acf04a4716b121bd34e1a9a5176a0c3448cd740960f0c7a138ba67c

Observation 577608bd-94de-4b36-a3c6-6450c71c66ed · outbound

This paper cites an unresolved cited work.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.071742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.071742Z digest=sha256:58d567f6ece8a7dd43caa78f95e22d9d79be0276c2030fe22d82c5f9ef3a4153

Observation 9fc0d62d-9555-459d-94bf-0351f54c83b0 · outbound

This paper cites an unresolved cited work.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.075368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.075368Z digest=sha256:c1b4f6536baeb694a7205f5e23a2dbe1524d99d4b8c291325d66bef8f522201e

Observation 1df357a6-03e0-4789-a052-6628592dcd04 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.079279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.079279Z digest=sha256:6f9b6a40ee135bff432bde9bb09eabce6f8a70450610e94f663b30e4b0bb14ea

Observation a5af1714-1c85-4d3d-8b9e-8d47f4503caf · outbound

This paper cites Advances in neural information processing systems , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Advances in neural information processing systems , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.083243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.083243Z digest=sha256:b41bab701fa3eef6a2142e3871b151b051349f8084562453bf72694a42ffff08

Observation 2faaf74f-9daf-450d-a617-3a4466d330b6 · outbound

This paper cites International conference on machine learning , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding International conference on machine learning , pages=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.086792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.086792Z digest=sha256:e37fb56f23e42977559a36f1244491d5b63190d5eee54384f9290ccd4b58e42c

Observation 5a542cce-ab43-4480-b6da-baf3ac9d6403 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.091644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.091644Z digest=sha256:ee0a5bde43d8e4a985332ea0a4ebc5ede2329ca8fbf98fb043fde8654d073ebd

Observation 308c53c7-5f8d-4242-8dec-85a3af4acfb3 · outbound

This paper cites Long Context Transfer from Language to Vision.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Long Context Transfer from Language to Vision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.096042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.096042Z digest=sha256:ed9e874c17b5acb378366d0f0fd910f3d77f9b6d1361a149968329d960599b23

Observation 1a34726a-c37f-4148-b236-4e7398ac6ccc · outbound

This paper cites arXiv preprint arXiv:2502.05177 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2502.05177 , year=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.100870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.100870Z digest=sha256:afefa2103131dbe3106db7c015a4beca87d763881419d3d04469a272db3650bc

Observation 62eb75a8-496f-4c5b-9297-5a32e09166bc · outbound

This paper cites European Conference on Computer Vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding European Conference on Computer Vision , pages=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.104559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.104559Z digest=sha256:f9b921f6faf51494e60ef0394a8a22f05d002b6b6753b3c06b1d7fe335cd56f2

Observation 9d0d5924-46e1-4636-bc0c-4ef0ccd7060d · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.108718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.108718Z digest=sha256:73204384a55b7d56d4ea22392e5a2d582104d01daf4006f615f83425625df44a

Observation 23615e5f-63f5-473a-ae1e-29277313f056 · outbound

This paper cites LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.112820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.112820Z digest=sha256:fd9481f8fd1f9c1946e36f416280fca99cedcdf69eec877a9b8ec777ee151677

Observation 3a363538-f653-4df9-86c7-043fdc5295c8 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding DINOv2: Learning Robust Visual Features without Supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.116786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.116786Z digest=sha256:0a971d139b5f9a0368a919881d98c4a30487f8d32e87c3105df0d6f87a6e2eca

Observation 2528b329-713a-44be-acd8-f0ffde781769 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.120837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.120837Z digest=sha256:b721a9dd118a7ecaed7824f6d79c62b8bf0f07f5cb132d9693d1748e315bfad6

Observation 31af810c-b90a-453a-8506-cbaa23064c22 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.438990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.125004Z digest=sha256:070ca1f8e59ec61752ff08b834b86da679c08d03895ff214b5caba40c9097de6

Observation fc745383-719e-47c8-9c37-c19f8d4f6247 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.128973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.128973Z digest=sha256:2649d13e7daed2f20fa88496387b8128a9947b1320c581dcecad462e1cf701f1

Observation dd591fcd-b595-4d6b-8978-6099f1586b65 · outbound

This paper cites arXiv preprint arXiv:2510.27280 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2510.27280 , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.133666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.133666Z digest=sha256:1bc5385eeec75cc7e0a81d7e15b2b35dc768e46d778bc9bbcc5526626deef798

Observation 83ce1718-bef9-4d79-a24a-13456167b10f · outbound

This paper cites Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.137666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.137666Z digest=sha256:88ce29ba3a515b691bffccbefc0519e92503cc85b12ada480e56520c18e98e79

Observation 1b80a305-06b8-4a90-b461-e9b9d104db82 · outbound

This paper cites FlexSelect: Flexible Token Selection for Efficient Long Video Understanding.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.141746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.141746Z digest=sha256:0f1770322fd8e167da37738f1dd0459dece031f1e04472ec8fe47ae4c04b3583

Observation d680571f-867a-4a75-881e-03ebc212f6d2 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.426552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.146226Z digest=sha256:323e226b2cec963e1f2c6ec7d60d08de053e8523a9fe8129662f78bf7453c361

Observation 104a7dd4-af46-47c8-989c-76bb7ba2552c · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.413824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.149849Z digest=sha256:5892307ce0d0c909f54edfbb98d5fc37253b950dda7adfb0673ceb93e9a86f2a

Observation 9a674896-953e-4cc1-aae0-29c2ccddcae7 · outbound

This paper cites arXiv preprint arXiv:2510.13891 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2510.13891 , year=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.154446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.154446Z digest=sha256:6e6e9805966388f774c60588d3bc26009433c696ab8a4fd71d4ce6f7c41da829

Observation 949b5c8a-2cf7-4ce3-a17a-63fc99fbed09 · outbound

This paper cites Frame-Voyager: Learning to Query Frames for Video Large Language Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Frame-Voyager: Learning to Query Frames for Video Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.158387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.158387Z digest=sha256:a08479cbc3a372991256908f0591c2e6e653ea668b2ba2046350d9af678d0c87

Observation b9026e97-5b11-4dc2-88ff-3fe443f527e5 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.162434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.162434Z digest=sha256:d7f3d9f4814c4976c39594cbc915ec386de5c9afbfac921a9d084238c9da2eda

Observation 4ac55cbc-22b6-42bd-9fff-474e2edde272 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.166581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.166581Z digest=sha256:7740c4e6ee9e4fb6a72e5202ba2f9c781cccb2cc3ce8e21ddd34dea3eba13efb

Observation a341499c-793e-48c2-9e41-3b2136d50295 · outbound

This paper cites arXiv preprint arXiv:2512.11534 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2512.11534 , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.170212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.170212Z digest=sha256:4a6590fa718a20676689483f02a684444672a83f5983060d755d87993ef015b7

Observation f24d6f3c-e584-445e-93d1-99d3b7bafd5b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proximal Policy Optimization Algorithms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.173910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.173910Z digest=sha256:333feb33ee420a81ab31a11e98276e3a6e78f6d120247ea277acdf730b314592

Observation 454c92c5-a9de-4550-9945-7fb52ba7f392 · outbound

This paper cites Advances in neural information processing systems , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Advances in neural information processing systems , volume=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.177728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.177728Z digest=sha256:2c31f16029a8319616ee86ce852d59007cbdf86547ebbdd754d6d58dfb98d4d4

Observation ca11859a-759c-41e5-a5f9-85d0663ace74 · outbound

This paper cites Advances in neural information processing systems , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Advances in neural information processing systems , volume=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.181274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.181274Z digest=sha256:5e5ea9a71ff3a192a395310465758becde6829d2657868704636a462d7095733

Observation ef3320d5-b85e-4927-9168-7fd73b355b84 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.185360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.185360Z digest=sha256:6490f05f04fad43823e30635593b0c382d86a05b47422dd00384b99e0ffc2192

Observation 9acbfab0-4321-4bc9-b2a6-ab619266c7e4 · outbound

This paper cites Proceedings of the 2024 conference on empirical methods in natural language processing , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the 2024 conference on empirical methods in natural language processing , pages=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.190392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.190392Z digest=sha256:3fdd386d34e1c783d934b5adefee0f9c19b632e83ec1b2e998133907c4ba4017

Observation 0dfa2b5b-5896-496c-ba1e-69fde5afd8e0 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.195518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.195518Z digest=sha256:c63dab9407e5c4505b55832b8e06e151852fa1878398fe8023396d296afa4618

Observation a3fbdcb7-97d9-4efe-8d32-f232731bec8c · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.376340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.199407Z digest=sha256:85b085deb57b45b0381bb0dab05955d25da22d5c1d69a549dce6c6f97e6eff01

Observation b44f048c-53c6-4815-8e91-8912f439fc1a · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.203009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.203009Z digest=sha256:fb2a380748baac9d159e50e633448414424a0f251354f98cf9c5b7e3cf32d405

Observation dce91058-7b8c-47d8-99cd-fdaad28ec38e · outbound

This paper cites an unresolved cited work.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.207163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.207163Z digest=sha256:c02ac24e6eaaf0a7c63762cd09ea4b761d9277629138bdda348e890975080907

Observation 108275d8-c8cc-4622-9a7c-4a52af0a8537 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.211189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.211189Z digest=sha256:28a65f6bfb353e53ee6adc205ee4392157859b7d89f5e676833fd1073c3ad1bf

Observation 2adaf088-b632-4f60-a978-61943e5b8a5f · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.215228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.215228Z digest=sha256:68afceb8783c14e6e983aa6f3b742d6cde231440fafc8e873c509272f3ad0fd2

Observation 735d3b90-9c25-467b-b563-83d0b51517eb · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.219360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.219360Z digest=sha256:0b7b857c540741b19842f0116565dad14eea6d781301e868f45af155938e91da

Observation 6ff40cdd-5706-4348-972a-deb5b3b8976f · outbound

This paper cites GPT-4V(ision) system card , url=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding GPT-4V(ision) system card , url=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.353966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.224392Z digest=sha256:d9ef1d996d7f43e2f6e81f466116c39b61d3961df5d4694e6bcc092c9ea24812

Observation e7ee0b09-26c2-494a-9af5-f554f2efc499 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.228769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.228769Z digest=sha256:94259050e252dec4490797a84536acd5be2d08062a29de6c7b4d8c91f1b11267

Observation acfb088c-807c-40ce-a2ea-8892e3ff6b96 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.233288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.233288Z digest=sha256:bcb850ea921852d1e0360424fba1ceaa504397c78527e40b8bb0a95b68064930

Observation 29663a2e-d21d-4b52-8965-a6238af709bf · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.343764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.237827Z digest=sha256:bbad976b6277611ae370e61a062e77c02c6d94334e6f3d5be0bef8bbbdb0a784

Observation d3ac00bb-b2ca-4bee-bfe6-89ec6ab877b7 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.333721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.241909Z digest=sha256:69f043d6dc7cc554ca90a7a7ea29cb92cf840c8e9f74fbcdf22e9f74d7b0598d

Observation 1b31cd46-dfdf-4aac-94d0-baf39675fb3c · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.246208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.246208Z digest=sha256:0ebbbb8f7f123b15df878eed551d27be7382ced9f688cb96d48c39812c5cd655

Observation 4b778ed7-8140-41ab-a1f1-b6e80088b44b · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.250414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.250414Z digest=sha256:b3957d402c976577aa8ecdc8439cc31cd47a20f228bca13cee13d15ccff3c272

Observation 8be2849b-f821-4698-84c7-b874683bd49a · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.254789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.254789Z digest=sha256:d3aef2852f47ae45d2c3f06b11288dca2c38bc81654579f84c80c9a745228592

Observation 2adc3bb7-81f2-47ca-942d-a0b06f11b82b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.259354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.259354Z digest=sha256:d5ddbe556ae50492f83e8b819ffbcf7ecfd26909c1257c32617d7c1c44baea5f

Observation a3caf4fb-be58-49e2-ac26-315351eaeb75 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.263951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.263951Z digest=sha256:1cbf38d5648e921e869b65a040a1867f5fe4331f9e9ec7b18cfcc46258c87223

Observation ce3af3d7-e684-4f93-85a3-4ea4f6cb84df · outbound

This paper cites arXiv preprint arXiv:2510.20622 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2510.20622 , year=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.267788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.267788Z digest=sha256:07a37b5642f9e6ab9634bf1d4ccdc1a297b73bc50ef425ab06f4b881884aca41

Observation d5d5aada-913e-4dd2-af7c-e63665cd10d5 · outbound

This paper cites KeyVideoLLM: Towards Large-scale Video Keyframe Selection.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.271317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.271317Z digest=sha256:81befa7175b91bcc10880339a6b429919bf694b4e6184113d007d1d981c162cb

Observation 3edf887e-ae98-487f-8ca2-118820cb80a6 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.275164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.275164Z digest=sha256:47fcbd125331d0f274642236935a626ec1e633c4c12820292ba0547e4e6bd40c

Observation ea224450-abc1-42c2-b9ae-bca17a4f097c · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.279225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.279225Z digest=sha256:384e2b623957dca275d99ddb7752d96cb656e0653b9b5b0da3d9d2f66451430e

Observation 64354bbb-8bc1-43d1-8297-6a99e744addb · outbound

This paper cites European Conference on Computer Vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding European Conference on Computer Vision , pages=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.282895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.282895Z digest=sha256:1e91cecdcf1d495debc7f7b7f012ea388fc560e0d2b015b2a4d21ce1cbf0a888

Observation 70505639-07af-4d2d-ac65-935bf4ad9567 · outbound

This paper cites arXiv preprint arXiv:2512.04000 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2512.04000 , year=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.286834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.286834Z digest=sha256:6cb0f1dbfcd81146d0c3df7d8490449acb6d6c918f42af8a0ec2ea71ff3c71e3

Observation 7231cc6d-7072-4e4c-bb8f-8dc3ddd4dd39 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.290721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.290721Z digest=sha256:36240063f544598701b2b92d2d83ce0f3c70761703e868b3473b71d82e39440a

Observation 0f75b74b-7551-4bf3-a952-b6523142cad7 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.302472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.294520Z digest=sha256:d97617c39cbd2afe2295931a8cfc0d3e70d8cc5dc4831573c6543455587bfb35

Observation 3f747f0b-d0cf-4dcc-9a4d-28bae416a214 · outbound

This paper cites arXiv preprint arXiv:2512.06866 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2512.06866 , year=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.298403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.298403Z digest=sha256:a20cd9ab5304f0c1a3739b1cf7dd337f6b2233a6f5c5dcb39a0139f6817a02d2

Observation e82103f8-687e-4358-be36-fd9b36733719 · outbound

This paper cites International conference on machine learning , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding International conference on machine learning , pages=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.302345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.302345Z digest=sha256:9e9c56d38d15e8a4a8a76f47e9554b3297231c8d6b2f69bcf5637aed0eb017da

Observation 386423a4-cb6a-4ec6-9647-fab143900a02 · outbound

This paper cites DeepSeek-OCR: Contexts Optical Compression.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding DeepSeek-OCR: Contexts Optical Compression

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.306068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.306068Z digest=sha256:77c0f3c3e139427ba8e4dd5e6876f7302cf1763ce8ddeea9d314201c54213e24

Observation a8e0168f-e7cf-496b-be61-0b99e380d554 · outbound

This paper cites arXiv preprint arXiv:2502.02770 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2502.02770 , year=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.310474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.310474Z digest=sha256:28d4aa20a4b8549dfa1ff89273f871ac93ea3d5610b9eab70c2df3538ae6a12e

Observation 8651f988-a6eb-478d-9397-9933409b53f1 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.314074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.314074Z digest=sha256:5e0b54a97a2e90e15ef968a1a7f9fa4165e9dce0d1354dd732b048a47fd499d7

Observation 0c39dbf2-088b-4d6a-9299-72b5745b961d · outbound

This paper cites arXiv preprint arXiv:2601.22582 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2601.22582 , year=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.318166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.318166Z digest=sha256:96b1f23ab50e8306376c769067c793987ea57a1e7722346c810565f300fa46f0

Observation 89644a61-823b-477b-9050-d893594e292f · outbound

This paper cites CoS: Chain-of-Shot Prompting for Long Video Understanding.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.322337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.322337Z digest=sha256:8f34eceb14a803759290c37a3a434a7b6d65772f8f8bdf25ae54302f25443bd2

Observation 0f00d2cd-5d32-4612-8501-582a08c55afb · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.326390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.326390Z digest=sha256:6ba3872a59e4b6e200a60a8c4c3971af9e7079f8a86005618e52dd984aec7351

Observation 6e9bdbae-b87b-4365-9025-a400c1fbc98a · outbound

This paper cites Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.330473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.330473Z digest=sha256:e55e3040f51c5f6c91cfa0c01ec6a095f42e2e1e8fe39fc5e74f9f3b3af4814c

Observation ef3c325f-9b54-45f3-b907-81c8326c3dbb · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Group-in-Group Policy Optimization for LLM Agent Training

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.334661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.334661Z digest=sha256:4eb03e1d038487bc7af5af492eae29880c9ffcf47934fab3d3ad32dfed2371ba

Observation a46cfd20-1e86-4467-b0c4-b3d875e8158c · outbound

This paper cites On the Emergence of Position Bias in Transformers.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding On the Emergence of Position Bias in Transformers

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.338843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.338843Z digest=sha256:ccdffa194bfbc88ea675b1d860dfa9ed955612b793cec1a5df7b279e91958d47

Observation df6bbcc7-9e88-4fbf-96b0-9b851f16bda6 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.284992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.343208Z digest=sha256:8190686781dcbde86883a27783a9a756eca505560a3d71db5fb2cf08d5241ff9

Observation 8aace232-8f19-4705-b6a3-084255d68645 · outbound

This paper cites Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.274866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.347966Z digest=sha256:d492ad84ddf2e02cd16c9a012486f5c3a5264e75812b53ffd78dc03bc0560239

Observation 8a677408-6190-4db3-9c70-bd63b83027e4 · outbound

This paper cites arXiv preprint arXiv:2504.18579 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2504.18579 , year=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.351819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.351819Z digest=sha256:81c3cb2b013184fadf1b1b99cf4dd6ab7d39c8b1ef32732516a44500bdd6b6ca

Observation f91a1cdf-23b2-4ce0-b2ec-332e7591b29c · outbound

This paper cites arXiv preprint arXiv:2603.06199 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2603.06199 , year=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.355809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.355809Z digest=sha256:63c4f9315ce59f75c51f84d2498f274f5f9e60af03566797ae9822b5392d5d61

Observation 6083fe37-9c5f-40e6-97e3-d61f110a605a · outbound

This paper cites FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.360123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.360123Z digest=sha256:44a6cf5f8ceab984ea1fb02915bb1c3ee9e4e82eeed871927e649947f7131cf9

Pith citing papers

No inbound Pith citation observations are available.