Pith. sign in

Paper Citation Record · LEDGER

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding

As of 20 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.05707.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05707 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:39:08.840327Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact4
  • verified fuzzy12
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fee22884-ee9d-4572-a6e7-eb3430ad59c6 · outbound

This paper cites LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.655128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.655128Z digest=sha256:751e9af12bc213e351ab44ae410414caf73442affc8101d9751e1d818c035daf

Observation 7fdf519b-8ff5-46c1-a4aa-950f2871d53d · outbound

This paper cites Qwen3-VL Technical Report.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.661233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.661233Z digest=sha256:074aeef2906c002547249abdc9dda2e7fba976b12b7b048f9fd564c1e004df63

Observation a255419b-f8da-42cb-bc6a-271fcbbc047a · outbound

This paper cites Matryoshka Multimodal Models.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Matryoshka Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.665758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.665758Z digest=sha256:62d92fff87f496f1a1e7d52cfeaf488b3ec55e87ed4ac39a5e56f7704c3672cf

Observation 7693e0ee-425f-4544-9324-faf97550a288 · outbound

This paper cites Event-anchored frame selection for effective long-video understanding, 2026.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Event-anchored frame selection for effective long-video understanding, 2026

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-15T14:39:09.580633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.670128Z digest=sha256:499e1643fc1e8c269cd4f546cbbce47885c6a562c1bc1ec3117a2d683a59bfbc

Observation e95530e5-1449-46c2-8a37-58a5d5b8faa6 · outbound

This paper cites Wavelet-based frame selection by detecting semantic boundary for long video understanding, 2026.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Wavelet-based frame selection by detecting semantic boundary for long video understanding, 2026

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.674102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.674102Z digest=sha256:8e7cbaa2d53c652ac358c36efc6d59b54b8f786647dce749b1bc28eb608a74d8

Observation 830fb610-1057-4df5-bf95-459f62dc55f8 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.679341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.679341Z digest=sha256:cb25aae18e5826fbbc027e58cb0b2e3e4ddae242c10116b0885ea0170fd62ab3

Observation 3ff31317-a781-4cec-9271-dfed0adf97a0 · outbound

This paper cites The blur effect: Perception and estimation with a new no-reference perceptual blur metric.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding The blur effect: Perception and estimation with a new no-reference perceptual blur metric

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.684598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.684598Z digest=sha256:4ad05ba01234e20f36c3da545eb75eea44f4904fef455a86524265b00b29c1a0

Observation 4b2f008c-e1cf-41ed-b984-0d536141ad48 · outbound

This paper cites MatFormer: Nested Transformer for Elastic Inference.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding MatFormer: Nested Transformer for Elastic Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.689169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.689169Z digest=sha256:f7ad477ed1056ce2e4d1e0e5a76ec05a89b5447291fc317bdcb2adc74398547d

Observation 912ec55e-287c-46f0-bba3-64e5e3f65f79 · outbound

This paper cites Agentic Keyframe Search for Video Question Answering.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Agentic Keyframe Search for Video Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.694284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.694284Z digest=sha256:3962256708d43d635fccc291fe69f5f391a56d61125659b9bc5033a4b41759ce

Observation df5b6b56-6f9f-4886-9414-484068385707 · outbound

This paper cites Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.699220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.699220Z digest=sha256:7fe18c226ec8be99aab02e07e448d306698343b7f2b05d6f1b621d63c6bb3098

Observation 3f9de5e5-869f-4562-bf73-36b4badbadfa · outbound

This paper cites CaptionFormer: Unified segmentation, tracking, and captioning for spatio-temporal objects.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding CaptionFormer: Unified segmentation, tracking, and captioning for spatio-temporal objects

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.869309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.705289Z digest=sha256:a7b0f26e85d983411b4c9745d8fe476f8ef7afe3cff36459386b21d9add97bdf

Observation 4fe15752-7c47-4557-b6f0-b4ee1a3f0d06 · outbound

This paper cites Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.855302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.710011Z digest=sha256:4bb9a04a4d0eb1efbb775dda33ae7d9cec98abdae139ef9c760f9e17ffe555a0

Observation d84f121f-d0be-4c49-aee4-7260e2bf2880 · outbound

This paper cites Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.714567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.714567Z digest=sha256:6069ce2bec8e2cb3323358b3f2b01c7f55cdae643a67abbc9f02e94b2b03536a

Observation d71be289-6e55-4f99-9554-61f1a5ab511d · outbound

This paper cites Matryoshka Query Transformer for Large Vision-Language Models.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Matryoshka Query Transformer for Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.724129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.724129Z digest=sha256:4ea1705c73e59a669608d3a08af6bd35023eee949c97c88bf9af291fdc19d4b4

Observation b3808f1b-af17-42b9-bdfc-b39c3a2ae6e9 · outbound

This paper cites Matryoshka representation learning.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Matryoshka representation learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.840792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.729135Z digest=sha256:0377bd6c2f7d7c7e326d43d6a0fe921b877dcb8bb0c24664411d0351e0fc5472

Observation ce44ada9-19d2-4f02-8643-505729c67a4b · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.734046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.734046Z digest=sha256:ea8b2883a000093f70807edefd5e7af16897dd222ff2764b6c288abdf0dc9bc3

Observation 78ebd74b-d4bd-435f-bf7f-a145dff6af3b · outbound

This paper cites VideoChat-Flash: Hierarchical compression for long-context video modeling,.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding VideoChat-Flash: Hierarchical compression for long-context video modeling,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.825865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.739356Z digest=sha256:c327322febf94f37287af2ef1c6512a0a3a81a75a9b6686568e3d18e9493cb16

Observation a5132a82-e81d-4dba-8861-3d12af0cfa9e · outbound

This paper cites BOLT: Boost large vision-language model without training for long-form video understanding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding BOLT: Boost large vision-language model without training for long-form video understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.812709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.750547Z digest=sha256:f57a4b94b76b2fa744fedfebc417c65bead581142893568155af573d06437bbd

Observation 93203fe6-26bd-4347-bfce-bdd996d0f35a · outbound

This paper cites Keyframe-oriented vision token pruning: Enhancing efficiency of large vision language models on long-form video pro- cessing.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Keyframe-oriented vision token pruning: Enhancing efficiency of large vision language models on long-form video pro- cessing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.797691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.754446Z digest=sha256:82bd61a92f25df576280fc18535e384d2da672235dac652aa88f3def59a72c9d

Observation 12be401c-c4d7-4588-bd69-867102effd16 · outbound

This paper cites QuoTA: Query-oriented token assignment via CoT query decouple for long video comprehension,.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding QuoTA: Query-oriented token assignment via CoT query decouple for long video comprehension,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.782019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.758366Z digest=sha256:5a3b15b24c6ba9495b3c9fa9a1871099eb47e9c0861d90a79d0d440bf75b160d

Observation 2f53d855-b540-453b-aad5-1640138216df · outbound

This paper cites Exposure fusion.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Exposure fusion

Reference 23

Resolution
verified exact
doi, observed 2026-08-15T14:39:08.878619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.767588Z digest=sha256:dc187adce0507265e8aaa2a0cca4df37c70d7004163a59d67079d501634e2c67

Observation 6cd299df-29fd-4924-8246-2bf884d2ad26 · outbound

This paper cites QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.762405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.762405Z digest=sha256:ce948d72c34e0ed2b02a28b84b8366d224691e743dc75420e079a06a72d4537e

Observation 9e51d890-69e9-485e-8a0b-904661c36be9 · outbound

This paper cites MovieRecapsQA: A multimodal open-ended video question-answering benchmark.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding MovieRecapsQA: A multimodal open-ended video question-answering benchmark

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.751933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.778334Z digest=sha256:7fd8a9ddad17cb025dd38f0a4691f022771bc6f2c6efd0c1985f6fd32f229291

Observation 8db751ac-510a-4904-8c1f-e94bffb5ed87 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Qwen3.5: Towards native multimodal agents, February 2026

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.766190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.773546Z digest=sha256:be2b823701e7fcdf3226e19f3cd5d30f40f6d5f2c58e6d99cb6254b6635f29e8

Observation b6d8aced-93b9-43b1-a17b-33134db38047 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.791550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.791550Z digest=sha256:cea17ce9be836d88b2affeb3f3e990cdab28ae4b19335133caf0259ba52d3cff

Observation e0d74a05-62c8-4b5f-80c0-a4853cca8f5b · outbound

This paper cites Adaptive keyframe sampling for long video un- derstanding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Adaptive keyframe sampling for long video un- derstanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.738796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.782828Z digest=sha256:0f2b70346e86276e2f0d445e4d23824c7cb1fb0ad7cce1d7e76952866a36a9c9

Observation a04f6f4e-fc86-46c7-b29b-9ce45b09b6e2 · outbound

This paper cites URL https://arxiv.org/abs/2502.21271.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding URL https://arxiv.org/abs/2502.21271

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.787076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.787076Z digest=sha256:f94b9486a72767ec12875178537615c0c21fc3d0e9df3bdfffed3a5dec77ef7d

Observation 9ee14c71-720a-4f8f-bdc4-67e1c9e184ff · outbound

This paper cites TimeLens: Rethinking video temporal grounding with multimodal LLMs.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding TimeLens: Rethinking video temporal grounding with multimodal LLMs

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.724560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.806292Z digest=sha256:aa2781dd726f38d7b79b5361821a5dcfad32a8ca943db3de6eb0a59358b926eb

Observation 6b05253d-2d9d-4dcf-ace6-0d079ddfc991 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.796357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.796357Z digest=sha256:8d3128fe497883e6b8911479149405596056730862f0d9554a560f38aa7a36cd

Observation 7d69fa49-2367-4afc-8969-2a60642d6871 · outbound

This paper cites WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:39:09.210487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.801658Z digest=sha256:8a61731f2ec14258e652792c861c8b78c985db9688bcb478245f1ffc02f27ac0

Observation 72b1b799-b471-4040-9058-8a6dde17144b · outbound

This paper cites Deep video discovery: Agentic search with tool use for long-form video understanding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Deep video discovery: Agentic search with tool use for long-form video understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.820556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.820556Z digest=sha256:4952a04afa71d7043492a58f2b6cdcc4ced9f550526a6cb4edf57c01c3571202

Observation a32642e9-2422-489c-9344-74ce05aa9e8c · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.811166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.811166Z digest=sha256:3688a8076924015f1ba8f7a985b425e2e762dfdb49e2ca52888ae0a688be65e3

Observation d01cf865-37df-4d79-b516-0c466b9fe484 · outbound

This paper cites Q-Frame: Query-aware frame selection and multi- resolution adaptation for video-llms.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Q-Frame: Query-aware frame selection and multi- resolution adaptation for video-llms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.816251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.816251Z digest=sha256:8dd93d85e9c35edaebfacc830db57df20fcd16a97a68ebde3043b5b397220889

Observation d238c998-8baf-485e-8b81-b41f25a13895 · outbound

This paper cites A.I.R.: Enabling adaptive, iterative, and reasoning-based frame selection for video question answering.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding A.I.R.: Enabling adaptive, iterative, and reasoning-based frame selection for video question answering

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:09.709128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.833888Z digest=sha256:6bf522c345b32e2ff5f9c3d5d5d9dd6b87bf0d8821ccfa11188e925f4f367ff1

Observation 94845f54-84f0-41dd-be1d-7ff3699dc832 · outbound

This paper cites Shot-aware frame sampling for video understanding, 2026.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Shot-aware frame sampling for video understanding, 2026

Reference 37

Resolution
verified exact
raw_fallback, observed 2026-08-15T14:39:09.110701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:39:08.825006Z digest=sha256:5f6316f938d6bdf6ea332020141e5232a06ee8cfebb0ea7e42edeb56052421e9

Observation 220788f3-2194-4ba0-8810-794b0e15eafe · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.829183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.829183Z digest=sha256:17803692ec75b2b251f71ed04a44dc60f3822d2036f69f595a6819de58311056

Observation ff6a968b-88c5-4d75-9640-05dcc13fcd9d · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.744823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.744823Z digest=sha256:dc082ae6dc38a076342f2b8ecf6dde9497abb5470876f69c15efc0b1a2904756

Observation 10affe58-b0fb-4a0d-ae6e-c6f4bddb9e3e · outbound

This paper cites 13 MAC-AutoML Appendix overview.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding 13 MAC-AutoML Appendix overview

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.840327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.840327Z digest=sha256:7e6bb53e8b20029b051b09e75d2b1e65607a1b211fd97ee5c3cde9a31e692ee2

Pith citing papers

No inbound Pith citation observations are available.