Pith. sign in

Paper Citation Record · LEDGER

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 8 inbound Pith citation observations for arXiv:2506.22139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22139 v3

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:15:05.248695Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T18:39:24.915547Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:09:57.149192Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2c548b6-25d1-44d2-ac91-7ec88012bc3a · outbound

This paper cites Qwen Technical Report.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:01.915030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:01.915030Z digest=sha256:2e3731f20d755b462257846eae318137133419ef036d29068bc25feaa4ec3343

Observation ea63bb34-9624-424a-a837-d4b9cafa0642 · outbound

This paper cites Robust motion-guided frame sam- pler with interpretive evaluation for video action recognition.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Robust motion-guided frame sam- pler with interpretive evaluation for video action recognition

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:09.053397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:01.967902Z digest=sha256:65118394a7525e3e44b352496279327f2199155d1337c6b1207d5d9a8a1fb162

Observation 86249126-8260-4bac-8d4e-d5b7523ba862 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Sharegpt4video: Improving video understanding and generation with better captions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.884299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:02.033532Z digest=sha256:540e83ffd22706b1ee9cedb6be1fb52752748954e6cc99f59f9e8e52c3b9ec44

Observation 5ca6dea2-4ce0-4342-9bda-680746d4e6d6 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.128356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.128356Z digest=sha256:e37f3194842c5b665c9d7b856ff2cc5b77256add8f4093f2de2157ceb0f09847

Observation 636f60ca-4311-4df5-bf28-12d428f28cbd · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.241212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.241212Z digest=sha256:43cabe4324596bae024cb65c3dd315a5eaf32cdcff3062c7dfe6f7386d3b9685

Observation 55a5ff53-44d3-4273-8475-82950400326c · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.340057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.340057Z digest=sha256:9966026075b5f1499d6ea5e6f69471aca0f8d26fee7881b9b2b2cf2705c8cb81

Observation 260f516a-eb60-4a70-aa94-5a542647b211 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.450017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.450017Z digest=sha256:2c931f2af6bf08c0ed1afac3408345415549143e8aaf3e657a57aa5c90532ba5

Observation d7ed91ed-796f-4940-b039-802d573b9f48 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Categorical Reparameterization with Gumbel-Softmax

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.527447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.527447Z digest=sha256:0858aae0940d711b318deec6e9a024779804096c152cb386e291976879869b85

Observation e1403065-aac9-4fcd-bfdc-403074e446ef · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.677569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:02.615899Z digest=sha256:30b2207fde704028ee001c830259b23fb10afb81e68461049d25e8c1c73b89bb

Observation 49eacf54-21f3-4a01-b6c3-b164f5b11144 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.714570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.714570Z digest=sha256:6b870ed8ffb02d8e93dfd6a996a0deac5f2f873791b3675cb1e3f5dede957ad4

Observation f53e097e-80e5-4101-b645-9a7a49a929b7 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.537709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:02.778441Z digest=sha256:e34e35419810c0ac75df3ef1197c0c79406bc9729629f7ebfc6155dbae0deb78

Observation a797c863-8993-4d59-b303-c2213ace2331 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.389761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:02.859142Z digest=sha256:255f4714d4cc4ba2041d796ce513d1549ac9df2c6faf3849c0c03df2ac9ee787

Observation 76dc95d8-f7e4-4b88-9a02-94e10a7c9b77 · outbound

This paper cites KeyVideoLLM: Towards Large-scale Video Keyframe Selection.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.922173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.922173Z digest=sha256:ae3d0c6a2d69b0f3e9eb2750587cfd3ad1e72c53af1883b5408b6ecb59099581

Observation b2dd8b82-c552-47d0-885a-c4f771f400d5 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-llava: Learning united visual representation by alignment before projection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.243146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:02.991356Z digest=sha256:65c0f007dbda27ecac8f456661ed46ebc75090951ba9a5150a5f951d145556f4

Observation 2475e21a-6cd2-4261-8e08-9faca5ee815c · outbound

This paper cites Vila: On pre-training for vi- sual language models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Vila: On pre-training for vi- sual language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.048504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:03.062699Z digest=sha256:27bfe55beda1bada9cadeb689c119e152e721b2333679f3a5692b087c5a6f433

Observation d5d2932f-6139-47ee-8d33-e1b0201e24e6 · outbound

This paper cites Visual instruction tuning.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.838909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:03.144052Z digest=sha256:13fc685127a762c7a0865455d202d820b109092a454987770e205a70c177a966

Observation 73badf97-5289-4ab2-9a12-ae2e30d2d29d · outbound

This paper cites Improved baselines with visual instruction tuning.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Improved baselines with visual instruction tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.202524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.202524Z digest=sha256:069d3751c41c43add2dde297d96351f538c5f051dae369a1272c018ffa409968

Observation b8379fdf-f320-447c-9623-29220e3249d2 · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge, 2024.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Llavanext: Improved reasoning, ocr, and world knowledge, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.653968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:03.286726Z digest=sha256:75b71262da29c03c889bfc06dd3050d17466e08e654e7d72cdc392acda61f2a1

Observation 0f33d806-78f9-4ccf-9bd3-34caa259491c · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.351183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.351183Z digest=sha256:4a437cf645a20083cb68b3c4243890da0f4409b546a1972aad01d0b23f9a5b2a

Observation 666db4a5-e386-4eac-b8a7-5681a16d0392 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.490542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:03.417538Z digest=sha256:6abfaff513c7bd710047215e9965997ed5ad72006fe80cbd4544cfa6e97d3fc7

Observation ad62d436-89cf-4000-b98e-3e65e6ed29ca · outbound

This paper cites Hello gpt-4o, 2024.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Hello gpt-4o, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.287773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:03.481220Z digest=sha256:d7f8c55c2a1ffef93bdfc71cde9538b6ebbe320ca722939b120b81638435cce5

Observation 9fbe48c4-ba7c-415b-a164-a6f7af190310 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Learn- ing transferable visual models from natural language super- vision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.110497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:03.561078Z digest=sha256:7b8224b21da19555314cfe07438556c18372c79df43e2e0c530c2c0413c28757

Observation 20018ddc-e8da-4d12-8dd6-b41804ca24cd · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.906491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:03.678869Z digest=sha256:e293a0086b43e80272ef54cf69990414a5f378716e65acb72540e7eea21bac23

Observation 51b1b8f9-be18-4d35-bc41-a649916768d0 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.755141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.755141Z digest=sha256:168038c62d30b8936bfd1cf0eacf87d7546cc8d3e649bd795a65cf69591bfb89

Observation 5f3b6777-e92f-4325-91a5-bf8b25fc6660 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Moviechat: From dense token to sparse memory for long video understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.824422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.824422Z digest=sha256:f7c8333e5997e4b787590a7fa22f076dd48b30c82997b2097384cc5b7235e51a

Observation b47ab2da-1955-449d-998b-88df156c4469 · outbound

This paper cites Video understanding with large language models: A survey.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video understanding with large language models: A survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.875000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.875000Z digest=sha256:76b15141a170e78d04940e46dbb67fd3a5760e3e94429b315c069a156c63f381

Observation e4359f59-bc8b-46fe-a718-628a619f961b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.971684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.971684Z digest=sha256:a7624a9f4be9d3cd9c3f4c2875458d2205659c61ba8376dec479990847bdb8c9

Observation 1f4c3484-0531-472f-ad6f-3aa4f348d22f · outbound

This paper cites Internvideo2: Scaling foundation models for multi- modal video understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Internvideo2: Scaling foundation models for multi- modal video understanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.755357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:04.029772Z digest=sha256:a202275dc863fc55a0074d9a5b1cbd4a91caaac7d06df657922f23575bdc4458

Observation e39c4189-926f-46bb-b906-7f1542fec27a · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.629724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:04.102916Z digest=sha256:322d0d45bc485a1ff6c2464f8ae90f0f5e22d464a08375ce0e080b4e32873b08

Observation e73ebea1-c1cf-4f09-8858-be514963c3a3 · outbound

This paper cites Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.502690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:04.241569Z digest=sha256:e5a1b8e260f2bb37f43f13b738507fd5cb19d379fccaffc26a5e124687f44bab

Observation d916726f-0327-4842-ae41-4fcc2571c33c · outbound

This paper cites Adaframe: Adaptive frame selection for fast video recognition.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Adaframe: Adaptive frame selection for fast video recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.387469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:04.299593Z digest=sha256:b59e7e82bb50b8be65817d0877d27f61ab5168bad0dedaa908a9317e314100a5

Observation a4b2f519-101d-44cc-af46-640da2b1c2df · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.382683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.382683Z digest=sha256:0557c26a21079597a7a0c7655a13c543e3a867774e7ee5cee0dd2f8755e8f81f

Observation 665bd746-852e-4a8b-a786-607ecfbe37a9 · outbound

This paper cites Frame-voyager: Learning to query frames for video large language models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Frame-voyager: Learning to query frames for video large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.261410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:04.480606Z digest=sha256:2deb3a4ec397c4d27732e645a26ba1b8667fae14afae83f704870c470edf9659

Observation eb4d00fa-9cc6-44eb-bf77-de94c0c3682b · outbound

This paper cites Sigmoid loss for language image pre-training.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Sigmoid loss for language image pre-training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.143689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:04.561512Z digest=sha256:631719e7f107e7f8e60bbf7c8d7dcc3028a1a3234a30cf31599a66d001ed55c0

Observation e0e680ee-9a99-4961-857a-fd21ac5307dd · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.623326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.623326Z digest=sha256:ec5e05b611f73c4433959e9008680ab083e9abac8da31debfe3e2871a8710f27

Observation 5d985fbf-4b73-481f-b255-2b737e68829e · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.686714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.686714Z digest=sha256:094d5f56c25bcd5d60eb399c7f305a633507b89b5738b1e9fd8a69d39993bae9

Observation ef8de3d9-0280-4974-bb2f-47853121def6 · outbound

This paper cites Long Context Transfer from Language to Vision.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Long Context Transfer from Language to Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.749344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.749344Z digest=sha256:407db9b63b22867b1b1a53c90e58258f6d5ae99d59eae26f2af70efa051dc8d3

Observation 5a1f0d10-f715-47c1-b335-f47ffddde9c3 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.864133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.864133Z digest=sha256:c9abedf12a65065cb8514307ad78959b60ff3b27f447ba9581229f7f89a8a618

Observation 0c307f2e-7cdd-4c08-b992-c4fc22cea03a · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:05.019965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:05.019965Z digest=sha256:8c06ea99e2c85f4fb031aa0cd68130175430465da06d829b197423e7032d626a

Observation 9fc18216-6abf-4dc7-915d-f115ba96b59f · outbound

This paper cites Mgsampler: An explainable sampling strategy for video ac- tion recognition.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Mgsampler: An explainable sampling strategy for video ac- tion recognition

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:05.983536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:05.086874Z digest=sha256:d75bcee117eaf1391e4f87d99951ecd17f9b930f033b5fe1e5fbe58ff1d59006

Observation f71645bf-2cfe-46ad-8924-4ae73dd3d962 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:05.172214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:05.172214Z digest=sha256:1dbe2bf5827481b85b9c24401b52563a12561a716086529a5ced1d4803b23605

Observation 2eef322c-e13d-4329-90c3-8916f65c57a9 · outbound

This paper cites Limitations Q-Frame enhances query-aware video understanding, but it depends on pre-trained models, lacks explicit temporal modeling, and operates within a fixed token budget.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Limitations Q-Frame enhances query-aware video understanding, but it depends on pre-trained models, lacks explicit temporal modeling, and operates within a fixed token budget

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:05.867099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:15:05.248695Z digest=sha256:edee3989fa8d349f5a235f37546ad7143ae122f7450357f39c0e5891b9f08cc1

Pith citing papers

Observation 0b6ad907-50e4-4dc3-bfa1-ecd47740152c · inbound

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding cites this paper.

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:56:04.269758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:21:47.439019Z digest=sha256:ee12286d97b5cedfa487affa733204713387a3934043289216cde2adb56dc607

Observation 33844b79-8431-4eda-bdce-8ca1e28b38dc · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T15:34:37.002001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:34:37.002001Z digest=sha256:a247887827269cb66f3c4abe11bc0ed7c72012cef2b4c81f6788e0dae8de9d6a

Observation eacdb4a2-0578-4ee1-87d6-3cc48c57148e · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T18:39:24.915547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:39:24.915547Z digest=sha256:b5d1489c8da725b0a542f801549d39db1694f70ef97b3c8733c58a77eaed0ee2

Observation 079f6c34-65d1-42bd-aab7-82261c6a427e · inbound

PEEK: Picking Essential frames via Efficient Knowledge distillation cites this paper.

PEEK: Picking Essential frames via Efficient Knowledge distillation Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:01.318490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:46:53.704892Z digest=sha256:82147940ff64ef6353bdc46dd925f381cac926de3c2461da712ffa00a22a69c0

Observation 9498d452-63c8-44f2-8693-0f8a72695bb3 · inbound

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation cites this paper.

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.027565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T10:54:02.188634Z digest=sha256:202f00a22361da053d7fd46049e1b1166dce9c361e1aa66ff938e728d5400da4

Observation de4ae604-e237-4b7a-9565-2c18234cc729 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.057718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:542fd1db11556139fd32d39c319b6f420fced450b493a8823f27e900d8e4db3d

Observation 9824f3dc-f80b-4c09-b06c-904f35ade8c6 · inbound

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling cites this paper.

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:09:57.150830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T00:53:43.629684Z digest=sha256:1333d50b0238f84d621043630c37a8a387ddda8a2ff61aa2c87e8481c6ca4ee8

Observation 5051ba87-dc57-4d7a-946d-cbf469c36426 · inbound

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding cites this paper.

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:17:02.668832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T14:09:54.549499Z digest=sha256:498ae716eb227c833cd4a5a010adc49275bc36d8d41fe2c2180efa40b6850df0