Pith. sign in

Paper Citation Record · LEDGER

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding

As of 21 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.05707.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05707 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:39:08.840327Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact4
  • verified fuzzy12
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fee22884-ee9d-4572-a6e7-eb3430ad59c6 · outbound

This paper cites LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.655128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.655128Z digest=sha256:0fb04675aeea01baeae83c9d0ef5719da15487bea72f088dc427716b366b0cca

Observation 7fdf519b-8ff5-46c1-a4aa-950f2871d53d · outbound

This paper cites Qwen3-VL Technical Report.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.661233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.661233Z digest=sha256:ddb5ced15bd559daa51a73ccd919dcd0ae0d3296862e64fa922272477d40eee4

Observation a255419b-f8da-42cb-bc6a-271fcbbc047a · outbound

This paper cites Matryoshka Multimodal Models.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Matryoshka Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.665758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.665758Z digest=sha256:79bfa7a4c1609aaa0b27a5d598375abdd9b9661c219bb5b8afa5dea0cb7fe9a5

Observation 7693e0ee-425f-4544-9324-faf97550a288 · outbound

This paper cites Event-anchored frame selection for effective long-video understanding, 2026.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Event-anchored frame selection for effective long-video understanding, 2026

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-15T14:39:09.580633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.670128Z digest=sha256:df1d194f03237c4ac4a77a73f6cc7ce6065d0c70af3bc9c0f0f2ddbaa128e05b

Observation e95530e5-1449-46c2-8a37-58a5d5b8faa6 · outbound

This paper cites Wavelet-based frame selection by detecting semantic boundary for long video understanding, 2026.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Wavelet-based frame selection by detecting semantic boundary for long video understanding, 2026

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.674102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.674102Z digest=sha256:75e7286d2429fcdd9e764e009ed1da9667ee49a1034f00ebed0941e925f7e63f

Observation 830fb610-1057-4df5-bf95-459f62dc55f8 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.679341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.679341Z digest=sha256:63cc63713a8c3845a78712a07a3bc8e75000ed129187319b62cc574abbfaaf5d

Observation 3ff31317-a781-4cec-9271-dfed0adf97a0 · outbound

This paper cites The blur effect: Perception and estimation with a new no-reference perceptual blur metric.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding The blur effect: Perception and estimation with a new no-reference perceptual blur metric

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.684598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.684598Z digest=sha256:2efad230ce2d9fdc077f3a8fafffca11aba76bf1ccf4f7531df18498866084e1

Observation 4b2f008c-e1cf-41ed-b984-0d536141ad48 · outbound

This paper cites MatFormer: Nested Transformer for Elastic Inference.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding MatFormer: Nested Transformer for Elastic Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.689169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.689169Z digest=sha256:7c601e8acdb687eeba0e6c185dd4368b6608555360abfc97944ac52b9f0adbe9

Observation 912ec55e-287c-46f0-bba3-64e5e3f65f79 · outbound

This paper cites Agentic Keyframe Search for Video Question Answering.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Agentic Keyframe Search for Video Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.694284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.694284Z digest=sha256:ea77d9ff46c039d8a11cb01f4bdb6c7cac4233ecf4026ed314ed7061e870a0a2

Observation df5b6b56-6f9f-4886-9414-484068385707 · outbound

This paper cites Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.699220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.699220Z digest=sha256:8642bc9575340d94e81fcbcc334234ad646612d84b7ca68532602c80edbee46a

Observation 3f9de5e5-869f-4562-bf73-36b4badbadfa · outbound

This paper cites CaptionFormer: Unified segmentation, tracking, and captioning for spatio-temporal objects.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding CaptionFormer: Unified segmentation, tracking, and captioning for spatio-temporal objects

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.869309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.705289Z digest=sha256:ae25edbf025da9b9cd124b4bb16efb1619045e3e5b76380272d6355a2c5ecbf3

Observation 4fe15752-7c47-4557-b6f0-b4ee1a3f0d06 · outbound

This paper cites Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.855302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.710011Z digest=sha256:5bee31cec0e1e9f486f2f1976f9bdc2d020191ec2924ef92ed17449d797558a0

Observation d84f121f-d0be-4c49-aee4-7260e2bf2880 · outbound

This paper cites Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.714567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.714567Z digest=sha256:438ceb22590fe1094ca8014586c90a915bb1145cd9e08968be1a80cf6a79cfd1

Observation d71be289-6e55-4f99-9554-61f1a5ab511d · outbound

This paper cites Matryoshka Query Transformer for Large Vision-Language Models.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Matryoshka Query Transformer for Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.724129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.724129Z digest=sha256:b001a18bb2614f9bdccfd20d3368771c06b26575c81b8cfc9a8919ab4d7ebb5d

Observation b3808f1b-af17-42b9-bdfc-b39c3a2ae6e9 · outbound

This paper cites Matryoshka representation learning.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Matryoshka representation learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.840792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.729135Z digest=sha256:39a28d58d1f4b4961719533083219b690da06d0aa2e6b599e9ff91c035306219

Observation ce44ada9-19d2-4f02-8643-505729c67a4b · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.734046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.734046Z digest=sha256:4cd3d4f8282d49bcaf2ab5b24fa4fd34eeb364cc7fd7e209d0c8ae36adec5445

Observation 78ebd74b-d4bd-435f-bf7f-a145dff6af3b · outbound

This paper cites VideoChat-Flash: Hierarchical compression for long-context video modeling,.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding VideoChat-Flash: Hierarchical compression for long-context video modeling,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.825865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.739356Z digest=sha256:c6a6f901942379067817c7ff88c6fcebcdbe43f28a011c8ff4dbb7c7e08179b9

Observation a5132a82-e81d-4dba-8861-3d12af0cfa9e · outbound

This paper cites BOLT: Boost large vision-language model without training for long-form video understanding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding BOLT: Boost large vision-language model without training for long-form video understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.812709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.750547Z digest=sha256:20932449999ff374625297ba1210512858f0ac56ac1de45dd140121df0ed6df3

Observation 93203fe6-26bd-4347-bfce-bdd996d0f35a · outbound

This paper cites Keyframe-oriented vision token pruning: Enhancing efficiency of large vision language models on long-form video pro- cessing.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Keyframe-oriented vision token pruning: Enhancing efficiency of large vision language models on long-form video pro- cessing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.797691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.754446Z digest=sha256:1910436dde536a4f50c9f27fe14e8b54874a5b8fe3c2762bb3208d0d6e10054f

Observation 12be401c-c4d7-4588-bd69-867102effd16 · outbound

This paper cites QuoTA: Query-oriented token assignment via CoT query decouple for long video comprehension,.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding QuoTA: Query-oriented token assignment via CoT query decouple for long video comprehension,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.782019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.758366Z digest=sha256:09c1001dbedcfd1e603c76165eeeb49ff1baae8bf498d1a18eebec893c88dfab

Observation 2f53d855-b540-453b-aad5-1640138216df · outbound

This paper cites Exposure fusion.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Exposure fusion

Reference 23

Resolution
verified exact
doi, observed 2026-08-15T14:39:08.878619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.767588Z digest=sha256:d0ddfa3c2f5871d7010da6053ba96559d5e55e3465aac07c1b6c6aa5b2e44284

Observation 6cd299df-29fd-4924-8246-2bf884d2ad26 · outbound

This paper cites QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.762405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.762405Z digest=sha256:66a2ed31f3ac8dbed6c2e1dc97c75745ca8014bbd386b816de42deb31f5f458b

Observation 9e51d890-69e9-485e-8a0b-904661c36be9 · outbound

This paper cites MovieRecapsQA: A multimodal open-ended video question-answering benchmark.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding MovieRecapsQA: A multimodal open-ended video question-answering benchmark

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.751933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.778334Z digest=sha256:80a2d4e783de9100bf6e4ce7fe775581ef5e00a09b0ae221485da8c48c34c557

Observation 8db751ac-510a-4904-8c1f-e94bffb5ed87 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Qwen3.5: Towards native multimodal agents, February 2026

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.766190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.773546Z digest=sha256:0f40b0dbe864ee59bebff7f91af9ee112348b5e6ac0ce0b0485ba5637f073563

Observation b6d8aced-93b9-43b1-a17b-33134db38047 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.791550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.791550Z digest=sha256:f56db22fcbd2d5bd0fdddae6e2f4259a881f4c67fd15ef8fd5420fa0aa341c40

Observation e0d74a05-62c8-4b5f-80c0-a4853cca8f5b · outbound

This paper cites Adaptive keyframe sampling for long video un- derstanding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Adaptive keyframe sampling for long video un- derstanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.738796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.782828Z digest=sha256:46e188cb0864d116c1748725fa0ff69082ad2ead8ef5febc83863a2d98b03a82

Observation a04f6f4e-fc86-46c7-b29b-9ce45b09b6e2 · outbound

This paper cites URL https://arxiv.org/abs/2502.21271.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding URL https://arxiv.org/abs/2502.21271

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.787076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.787076Z digest=sha256:2897b632071c089e707a0e7dbf2d4cd4323ed3144e325e08a30f8d04eee6fb6b

Observation 9ee14c71-720a-4f8f-bdc4-67e1c9e184ff · outbound

This paper cites TimeLens: Rethinking video temporal grounding with multimodal LLMs.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding TimeLens: Rethinking video temporal grounding with multimodal LLMs

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.724560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.806292Z digest=sha256:5516dbd48f73ae8939fa23b4367c6247d30b46fde7540b77dac2ba76dae630e1

Observation 6b05253d-2d9d-4dcf-ace6-0d079ddfc991 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.796357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.796357Z digest=sha256:c8813b85c06e04c789c97b9ea69af3ba49e0d267845dd0e781d2d500205a3851

Observation 7d69fa49-2367-4afc-8969-2a60642d6871 · outbound

This paper cites WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:39:09.210487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.801658Z digest=sha256:b78b1388f7397a94a8c9248c90f9273e2f07531e02fd0cd990a067a0efa1d0f0

Observation 72b1b799-b471-4040-9058-8a6dde17144b · outbound

This paper cites Deep video discovery: Agentic search with tool use for long-form video understanding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Deep video discovery: Agentic search with tool use for long-form video understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.820556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.820556Z digest=sha256:58c2f12cab6598219d408a555854e7d6f7f3335db18cc8e3cd79e8a395aa7d98

Observation a32642e9-2422-489c-9344-74ce05aa9e8c · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.811166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.811166Z digest=sha256:3c99e6b3086bbfedeace50d199f845c56507e3d4615001bf1a6bf839cf9fe5b8

Observation d01cf865-37df-4d79-b516-0c466b9fe484 · outbound

This paper cites Q-Frame: Query-aware frame selection and multi- resolution adaptation for video-llms.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Q-Frame: Query-aware frame selection and multi- resolution adaptation for video-llms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.816251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.816251Z digest=sha256:b3f8f37b01aef5402b1e13aab397a72e2193bcebfca02138d690931ab62911c9

Observation d238c998-8baf-485e-8b81-b41f25a13895 · outbound

This paper cites A.I.R.: Enabling adaptive, iterative, and reasoning-based frame selection for video question answering.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding A.I.R.: Enabling adaptive, iterative, and reasoning-based frame selection for video question answering

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.709128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.833888Z digest=sha256:a21aa16a67bab934fe4633ed51363fef0b7eb8c8b1075e36893d0c588213fb9a

Observation 94845f54-84f0-41dd-be1d-7ff3699dc832 · outbound

This paper cites Shot-aware frame sampling for video understanding, 2026.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Shot-aware frame sampling for video understanding, 2026

Reference 37

Resolution
verified exact
raw_fallback, observed 2026-08-15T14:39:09.110701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.825006Z digest=sha256:e436659ef07ae8507980c7b3b88c550ae8af20c116c5705fd4e64cfec45023e4

Observation 220788f3-2194-4ba0-8810-794b0e15eafe · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.829183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.829183Z digest=sha256:2592148557f5a8fd0278af694e578607e34e35749a5c43049f0c44d2e0399abd

Observation ff6a968b-88c5-4d75-9640-05dcc13fcd9d · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.744823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.744823Z digest=sha256:f8d4c3e7e9bc14bfe18556290f397f5fc662d154ec94fdbc7636e3aaa9455d34

Observation 10affe58-b0fb-4a0d-ae6e-c6f4bddb9e3e · outbound

This paper cites 13 MAC-AutoML Appendix overview.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding 13 MAC-AutoML Appendix overview

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.840327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.840327Z digest=sha256:711185e70c4485e4f3cabeafe3f6347c14693b85991851c6e2a4d02f5b4c97bc

Pith citing papers

No inbound Pith citation observations are available.