Pith. sign in

Paper Citation Record · LEDGER

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 5 inbound Pith citation observations for arXiv:2506.00993.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00993 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:59:11.354129Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.141746Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T10:11:28.172880Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved42
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2b7e372-c127-4f3a-90d4-efd27c3a854d · outbound

This paper cites Qwen2.5-vl technical report,.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Qwen2.5-vl technical report,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.319800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.319800Z digest=sha256:fed3f44bf0c302dc5b4d4de03f1ec419e677638bdac8e8b402e9bbee4b02d9f2

Observation 142e73d1-bebc-49bf-aba7-c98074de1e36 · outbound

This paper cites Fast Differentiable Sorting and Ranking.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Fast Differentiable Sorting and Ranking

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.529141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.529141Z digest=sha256:d279e63ce79dc22c6dd8cf79d8cdb8c66bcc6d837616507dd1a7f7dfefa858f6

Observation 0c9fb958-105e-4065-a867-de5b55fd6aa6 · outbound

This paper cites Principal component analysis.Analytical methods, 6(9): 2812–2831, 2014.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Principal component analysis.Analytical methods, 6(9): 2812–2831, 2014

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:13.275919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:59:05.644985Z digest=sha256:d561d6b5478273a3ac4f23ab06e5b83b22dbb7dd2fc34279681e4548d24ad631

Observation ec529799-0e10-4035-b845-b9cc040f1b06 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.780757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.780757Z digest=sha256:afe7e7df5c0cd4b5429ccd31c58cba431aca4034e986c89465611163c34cb45d

Observation 6fbfa54c-6909-484a-814c-ee2ecabe4087 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.898445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.898445Z digest=sha256:c52cacd0ffa160ec8304eb73197a70b09e1d54e5d945affdebe7ea93c1efda01

Observation 765a8170-44c3-4b5f-8f69-cb30a2eebca8 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.001743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.001743Z digest=sha256:5be98283a5112179059b25506e881c84f0f3283e3cbed785b980d101ebe4fb56

Observation d536abaf-3295-4fbb-a432-666bc4a20e34 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.098310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.098310Z digest=sha256:a7cf66ebc8a35d3c6d398e4385ae7a22148fed990814968a274edc5564bde1a0

Observation 4a6a7819-c052-4f17-a2a0-c9e20b22d499 · outbound

This paper cites ReWind: Understanding Long Videos with Instructed Learnable Memory.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding ReWind: Understanding Long Videos with Instructed Learnable Memory

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.204066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.204066Z digest=sha256:82e708766cb697057e279801e80df7e3ea54020a7aeda4e9d30026ff7f424af7

Observation 0cf1cc51-d3b1-4ccd-95ef-a55c6c0a8e66 · outbound

This paper cites Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.327785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.327785Z digest=sha256:825fede9ae7f022103e54a18d80267da224d76d7a83bd061b0cc0df36fd861ca

Observation 889779c6-2937-4cab-bb15-295b9ff8aeea · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.421786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.421786Z digest=sha256:71d01f4eeeb0f9beff8d61b0a677fcff7cc273d0b3d1bf1041cf154ea54db73e

Observation edb3e71d-43c9-4f30-aeab-8340890c5497 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.561620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.561620Z digest=sha256:35ffe4bea579efc5e6fd3eec165aef948c85ecf73924232ed0fdd36771f30c1e

Observation b1499470-2857-43b4-9b4a-d9b606b4758b · outbound

This paper cites Framefusion: Combining similarity and importance for video token reduction on large visual language models, 2024.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Framefusion: Combining similarity and importance for video token reduction on large visual language models, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:13.133611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:59:06.662818Z digest=sha256:80af42a165e622505255087e4c33a7120267f54b2b082de49ff1e05ec170957a

Observation 7a28d731-d390-4e5c-9c8d-1465ae651bdd · outbound

This paper cites LinVT: Empower Your Image-level Large Language Model to Understand Videos.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.759496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.759496Z digest=sha256:1e5bdb54b50e589b5199ddadfa16eb799c846eead69ee8778a2daca37feb26cf

Observation e2fd70d1-368d-4c2e-a2d5-30a39a559de3 · outbound

This paper cites Llava-onevision: Easy visual task transfer,.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Llava-onevision: Easy visual task transfer,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.898702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.898702Z digest=sha256:e64c7a712b6fa1fafe3264d54f99455915ced73e49f7dfd0380e9d44b6c7b4e3

Observation 21cc2dfd-dfd4-44a1-b1c5-5e5e52635d85 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.178112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.178112Z digest=sha256:fc9798eec84068ea44fb715c9669c927f00aa1be2ca6ab1756365a303d5b5147

Observation cade06a3-70a4-4910-80ba-31e38fac5da2 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.303609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.303609Z digest=sha256:49df83f7acf6c18b5c7b37fd01117a5f404238ec54d00fb7b0c3a9e9a4485fc0

Observation c0f4db96-2152-4a32-8749-5ef93fdf9da5 · outbound

This paper cites Temporal preference optimization for long-form video understanding, 2025.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Temporal preference optimization for long-form video understanding, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:12.981096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:59:07.398143Z digest=sha256:eb8387878f039417659398d244daf97743fb0fef8a8789886ad52adca659740e

Observation 02ef2d55-2f23-47e2-b7f1-85fa1b1795f0 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.537007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.537007Z digest=sha256:78dbf571cf99df965e77debe8eaa9a75d874384dfab3a30930ba3ec8d02b3c67

Observation 45e15b40-6cc5-4825-8b1a-1dea06352fcd · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.645844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.645844Z digest=sha256:c18de2e7b156be5dd23a618effdefcb5bd88f87adee9bfa080b2b4206e7ed3c1

Observation 8b3bc549-41c2-46f4-95bc-16ad15f0b52f · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.742693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.742693Z digest=sha256:1bcda20d26b75c0102230baefa1362a3830cd9f013b6d85cc2caa83456f5a02b

Observation 53815c49-612f-4fdc-ace2-72e165abdbdc · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding NVILA: Efficient Frontier Visual Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.844726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.844726Z digest=sha256:a12ad881fbff92984c07619e280ac5dee77a259ec938bfa89bf25524ee0b5719

Observation 76f94cf2-f665-450c-bbba-a82b44d337a1 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.957257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.957257Z digest=sha256:8638560a8da217a6debc7db0de84bb2fd25b2bab7cd1ffd3cb14feb02c1a0779

Observation b223a9c7-6b65-4bea-8703-b7752342912c · outbound

This paper cites QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:59:11.720174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:59:08.080043Z digest=sha256:6f1d1a5ea0cc6d0d7d5233e140614df9687082fad79d889dad85310a3f9aa247

Observation 8c88a420-dec8-4220-ae5d-15760d672cf5 · outbound

This paper cites Hello GPT-4o, 5 2024.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Hello GPT-4o, 5 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:12.829537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:59:08.188771Z digest=sha256:db5eed778ded93d24734e8df3c342abdf379b8bd91b79300acd5f41f04efafe9

Observation b5051d15-0865-43d9-a9a2-a195f343f37b · outbound

This paper cites Yarn: Efficient context window extension of large language models, 2023.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Yarn: Efficient context window extension of large language models, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:12.565984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:59:08.296154Z digest=sha256:fba63a09a47f30fb81197aeef25fd43b8c617ac778f472560b784a83c3516f25

Observation 44eec4db-7d7a-43fd-86e8-4d3f6f8c8d70 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:08.402346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:08.402346Z digest=sha256:ea9eb4acd339bcc840a5fdac7bf0ff23c1c60755a2ec53b0aff82c73ed387b66

Observation 1367ce42-ec36-4505-b7ec-4aae10fc5a3b · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding,.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-xl: Extra-long vision language model for hour-scale video understanding,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:12.397566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:59:08.534974Z digest=sha256:726f80b6922ee6c9bd800953b051d47cf8f36aa35b010e9baa960a08bfe6dd89

Observation ac8e8494-8784-4eee-8003-c3e14569b47c · outbound

This paper cites The proof and measurement of association between two things.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding The proof and measurement of association between two things

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:08.752562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:08.752562Z digest=sha256:ce002886a0d2c13f6104b1866074b2d00baf0732d9598f447e61b202cc405eef

Observation 44971cd9-b215-45e0-bfe9-0d8d6dc5fc3c · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast video large language models, 2025.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Dycoke: Dynamic compression of tokens for fast video large language models, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:12.268873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:59:08.868437Z digest=sha256:3bca172d46efe596f023eaed06cd6a00261d15f23f05a1fc1ba7eb4d1be7c470

Observation 358355fd-9998-444c-aae7-b87c1418de3a · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:08.637236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:08.637236Z digest=sha256:4f983d0736a6affaff337a0a468c323f3cc72a0abdeccfc3abb0c13c05faec7c

Observation f1b035eb-1fbc-4349-8fb3-6a2e48470e27 · outbound

This paper cites Qwen2.5 Technical Report.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.068469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.068469Z digest=sha256:589338a1b74f5bbb920bac779d257a12c43385cb5442d79ae059e1d8962d5a11

Observation 5ee70352-c5af-403b-9cb6-f1beddcc5450 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.173823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.173823Z digest=sha256:a50cff47f03fb0b8992c659b32180c7d174c9e2216c7cb63ac1ff1aa9ec58ed3

Observation 8e3b39db-a13d-4434-9faa-b797beceb8ab · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:08.992098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:08.992098Z digest=sha256:95895bf8b9bdac00754eb46b6403d0bc8cb989de65cee7ce55ebb16935a6cb1d

Observation 59d56e0e-84dd-4318-9897-74785c7466c4 · outbound

This paper cites AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.407693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.407693Z digest=sha256:e9db893e982df63b7a258ddc577ec7ecb1599105d4a0b77889f4a6040714f9d3

Observation f35f5d8d-1efb-41e6-96bd-6b0121220b77 · outbound

This paper cites Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.565561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.565561Z digest=sha256:265a1ec07d3c0a3b490bdd52983b0a0b4df004eb23b2a0e97962e9b4afc42454

Observation 85466586-ec1c-4645-8ef9-f0792b8a866d · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.320069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.320069Z digest=sha256:33bf1c6c36fbffbf4388de376d73634ac648928bc9dbf2836947c57279b95687

Observation 03995c8e-38cc-42d9-8556-b7994c2b04e9 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.820042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.820042Z digest=sha256:51214c89c5b1789b0d1764f4527170a34aba69dde328b67bcf06ea3b8b03dd1a

Observation fdd0506c-935c-43ec-b472-bfb3ddaf2094 · outbound

This paper cites SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.935958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.935958Z digest=sha256:03bba0d001524802b1315cccb87e46bb758e55950da28047954cf48679e12eee

Observation 07b2a234-bd0f-4ea4-832d-8429a0bf96d2 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.727032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.727032Z digest=sha256:3dd7e9327673ec56c239b62e58de1b722186e5e58ece44e9294be49ce3d0b8b2

Observation 98b0f7b4-8bb5-4331-ba7a-d661b7b8ed23 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.291998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.291998Z digest=sha256:9c7d82dbfb41a8df58cc7c08b22549b6102563f4fafc2175eaab6132f5483c71

Observation 317dbfad-ce0d-4ebf-b807-df11483ea597 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.075062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.075062Z digest=sha256:5491412f321f94e738c2d91648a4f017ce40bbf2d1901ef03868ec9a85efee8b

Observation b56cbe40-4a7d-4461-89f2-913ff3f5f7bc · outbound

This paper cites Long context transfer from language to vision,.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Long context transfer from language to vision,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:12.081596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:59:10.503442Z digest=sha256:0508688a64b5d954d748465348995567893997be8c6de965650bc4baa2fdcff6

Observation dfa4e81f-0ea0-49af-aad9-41a331e93e78 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.721952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.721952Z digest=sha256:a3b4ae76b3d51e6c33f2926d4f5c634bdfed0f231a840881e9b662a739dc0875

Observation 8826dfff-bcbe-4a51-859c-962a6a42fd2b · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.390042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.390042Z digest=sha256:620a951aed0a301fa2945061953a49d2a4cd08f2103183fa0c1c9628d32b906d

Observation 947f0de2-438a-4d18-be4e-eb00dfad343f · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.962445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.962445Z digest=sha256:de7ba51c0b2e90ebb39b50733a2bdd4cc41d05d59eb7ce5b8764210dd1fd020d

Observation 31892af6-6997-4b2c-9b72-0e9cc0709858 · outbound

This paper cites Long Context Transfer from Language to Vision.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Long Context Transfer from Language to Vision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.606703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.606703Z digest=sha256:6e982981f7d1629064905fce2236b1fbf75212d9139563bd0aae6ea6a7555597

Observation 28156de9-2eef-4bfc-95fd-494707660083 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:11.232538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:11.232538Z digest=sha256:980a7feae28b1f8c9ed8525328a6bc7f13e52a387881f816c7658bad433bd934

Observation cb4cd5d7-3937-4748-ae83-c7fe93378ead · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.869569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.869569Z digest=sha256:b5a402ff37dcf5c09ec368c4003443c868329a206aa5843b77bb7fcdff113e8b

Observation 7f6d8b62-bb23-49b7-b3fc-4b9f9fd99e74 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:11.136490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:11.136490Z digest=sha256:83918bd9cf6471819345da65e8a0e3722128f195e492aced109252541ec99592

Observation 0ff19277-4d81-42ea-a6ca-e0b1fdb91644 · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 53

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:59:11.354129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:11.354129Z digest=sha256:7a68c1a8a899bf43f448b574cd75e791afaae5719265ec9cdb7fb44cb5450b59

Observation 6e10be65-1ea8-4a7a-b4ef-a848e5b51ac4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.050092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.050092Z digest=sha256:831fa853da4e76fc7caebf95be033e98350af61551c7f4305fdbdc07c88d0a63

Observation 0add14b5-d3c4-4cbf-becc-c9a872c9f851 · outbound

This paper cites Qwen2.5-VL Technical Report.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.425676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.425676Z digest=sha256:654e4269ab71d9ded85b680ff5c922903525a24d03c6ae0830e69462707b75c6

Pith citing papers

Observation aab852e4-f310-444f-81f7-d11913bbc228 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:08.270295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:08.270295Z digest=sha256:a93b58a3ddbd405c7d0690b67fc258a2dc1a6407dfe6f5b61413d6aba4e83899

Observation ad69ed62-b0a9-4128-bfef-26435c8257e2 · inbound

Stateful Token Reduction for Long-Video Hybrid VLMs cites this paper.

Stateful Token Reduction for Long-Video Hybrid VLMs FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:05.045461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:05.045461Z digest=sha256:0af5a13c237133c53ff4311460dc4317e33c575162f61f0b4130237eea5ff1e5

Observation 8232cef7-438a-4279-b437-41f59df8bfab · inbound

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG cites this paper.

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:11:28.175133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T07:19:44.125479Z digest=sha256:6408da65c7c2fb0072308efdd2dc976428b9401f5608bc8f2f9c8e7d742ca350

Observation fd526ce5-1a3b-4b14-91b0-56a7c9883c5e · inbound

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding cites this paper.

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:00.719844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:00.719844Z digest=sha256:b11a0dfaf30fd02459ec3e44f9768ce2f643f811f2b5e038078b3102e4b3ee30

Observation 1b80a305-06b8-4a90-b461-e9b9d104db82 · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.141746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.141746Z digest=sha256:e8400e75c2548692079f5a551e608680614260da7da7bed55a978863fc27c940