Pith. sign in

Paper Citation Record · LEDGER

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

As of 14 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 7 inbound Pith citation observations for arXiv:2411.11066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11066 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:03:32.584154Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:37:38.250264Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

25
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 49d949da-4101-48f5-88b1-b60044caf604 · outbound

This paper cites Qwen Technical Report.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.904212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.904212Z digest=sha256:ea4e8bbee5a251270c403b4f99c422c7e7c739a91a1a7bda8dfa64ff33484a19

Observation 8cf08a90-6df7-4ff9-8eb8-214db30a03f9 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.930824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.930824Z digest=sha256:d7cbb1834e7a63cab387e4882378d33c5cb8804af0e8a5489831f42b11cec915

Observation 25aafd0b-38a8-4169-9d1f-998625a32d2d · outbound

This paper cites Matryoshka Multimodal Models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Matryoshka Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.940342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.940342Z digest=sha256:00e1fdd983d2013301d74925ae9f7282b7fb45e968a79934516bf0c57927e925

Observation 063c4c28-daa6-45c7-b103-b0989029fd18 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Collecting highly parallel data for paraphrase evaluation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.620276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:31.943621Z digest=sha256:7847e54c232194bb66ce2a26d99be28925a3569c579513dba965f90cace5cb1a

Observation b467affe-0ee3-4eef-963e-26673668d44d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.946661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.946661Z digest=sha256:2253146f98bc3cfb8089a577efdbca8d1d98dcb2893f992c3b5265d98a0fd375

Observation a3fc6f6b-dba2-4b7c-a953-ed9cda08e7cd · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Gonzalez, Ion Stoica, and Eric P

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.550546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:31.950260Z digest=sha256:bbc8867fb8735e60df7e706908822fc855d6e3af52498970e3092a54bfb05e8a

Observation 86dfe8fb-1ca0-44fa-87f5-658e4ade1cf3 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.954093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.954093Z digest=sha256:61c49c3c6d121bc8185130d8197def03b1768376984ac98bc6a5366f02b37c36

Observation d835c0fa-309c-44fd-b5cd-7cf8ca2383b5 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.957310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.957310Z digest=sha256:0f4c7c0760c391fe4f8eceb661e9df9b5e6a9f9738a0d7e8e6fd09314cc32c6b

Observation a6da4ce0-5e6d-4df6-b0ae-959caac389f5 · outbound

This paper cites Slowfast networks for video recognition.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Slowfast networks for video recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.531487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:31.960775Z digest=sha256:eb143e376164aecca203b585f8f0114b3e7051ca77be584f6b29a65aebca1c0f

Observation 8c4402a2-4819-4d01-9714-55fafd592f34 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.964048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.964048Z digest=sha256:cf9c906864376eb45792e5e516f68e749970e0df59292e4f594855d5eaf44c93

Observation 026bce24-78e3-4d60-abe2-19d7abe416a6 · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models LITA: Language Instructed Temporal-Localization Assistant

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:31.982517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:31.982517Z digest=sha256:d0a74997c4ec0b6ae70b86b1392b252dc2d5e66a648e5468c2302eb2012a4e4d

Observation bf677fb0-5780-4252-8d81-fc61ba1acf29 · outbound

This paper cites Mixtral of Experts.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Mixtral of Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.019710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.019710Z digest=sha256:b677217c29d0b7094d41aa4335b6f63c9f80a6c2dc2500d76763a7c6ae6a3727

Observation 629b42d8-c99d-4fbf-9987-35bcecbf71c9 · outbound

This paper cites An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.050286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.050286Z digest=sha256:3475597e0e5441075f209e63ab5f3d408e3c61a195742a94fa3903a7c54ecb8e

Observation 6fe6cce8-ec58-4c31-b4d5-83e2d9e174d4 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.523341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.089359Z digest=sha256:aab898fd4bfa29865e4ffa3edf4c07ef468dc974b53eb8102a83e6eda1826d85

Observation 99278d1d-5ca3-4f63-9b70-514c2e126e18 · outbound

This paper cites Inten- tqa: Context-aware video intent reasoning.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Inten- tqa: Context-aware video intent reasoning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.515890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.110166Z digest=sha256:a53baf6be2d28de8f2b4dac11e67f3551664024297117a8cfaca573414b5f964

Observation 8a26ddf1-7c2f-4a28-b60b-0d06a77b490e · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models VideoChat: Chat-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.113190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.113190Z digest=sha256:ec62d7c5a0d1454ea0eb845ad8c613f3432f87fdee0a2d033ddf2a313eb80f29

Observation 0eba0374-34d4-4b17-ab94-035a947aa78a · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Mvbench: A comprehensive multi- modal video understanding benchmark

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.508446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.126048Z digest=sha256:25a40944dc507827012a00b45aa261e0e6ab9df7da113d45fef8d738e405ebb3

Observation 3e925e5b-2366-479e-93be-f68f161e4f0a · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Tgif: A new dataset and benchmark on animated gif description

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.487482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.156621Z digest=sha256:3fe01dc4b5a341f752adcd9c511c6570e0ea06385ac500a2ee371210c5e993c4

Observation 48d78477-2644-40f0-8843-779fb6822449 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Llama-vid: An image is worth 2 tokens in large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.455390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.162477Z digest=sha256:3043f980ad30241240b0c4928a2947b967af4e7778449224776c7763c7ef68de

Observation 966c248f-6e73-4062-bcb7-fd9abb640bb9 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.165398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.165398Z digest=sha256:dbd8e5576ed5ca1a05d3c99548dbbf2c835346fbb48ae7100cc7cba93edca67a

Observation 065f7bc0-e07a-4975-ab56-355775f5a668 · outbound

This paper cites Visual instruction tuning.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.376748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.168388Z digest=sha256:3d8da3082d5db478327ea71d719dfe4b33c70b36f2d53b9fe0d254095aa41ffa

Observation ea3865ee-5d69-4452-8071-078583c44807 · outbound

This paper cites Improved baselines with visual instruction tuning.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Improved baselines with visual instruction tuning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.298728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.170871Z digest=sha256:834456e968cec17cba097b41f751be5b12d73c02d451140cafcdefa308aac0f5

Observation 6a8855a7-698c-492d-b184-4795c63dc146 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.225453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.173551Z digest=sha256:6fe4b50725431241740dc62c766691086efc9bc8f9a57821deec4c3b54ccacdf

Observation 6fdb8bc3-08c8-4df8-8837-be660e658518 · outbound

This paper cites St-llm: Large language models are effective tem- poral learners.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models St-llm: Large language models are effective tem- poral learners

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.171618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.176107Z digest=sha256:0123d3f9e5c64eac1ab3826af7a4271c03fb749772cb29c79927670b58f416cf

Observation 8b501599-1bf7-4069-90dc-9a65dee36862 · outbound

This paper cites Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.163744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.178588Z digest=sha256:c2fda906038553b85d5e115efccf6dd7dd7a3da0090134779e553be5ba8db64e

Observation 80b922fb-32be-4c83-9ae1-26c6f76d7b00 · outbound

This paper cites Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.155251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.181013Z digest=sha256:f5d03b2d1bc84c53b46928b724cdafddb54b060bb6181e57a296bf59ea0a4977

Observation dfa5c33a-bcac-4d8b-ac98-b59275345f64 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.146272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.184057Z digest=sha256:fb78fa0ce0b6f4065c52b93f38175f8b61ec4298575e7627da70f45a311c4201

Observation 4d65746e-6007-420e-885f-691aa1ddb11d · outbound

This paper cites DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.186910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.186910Z digest=sha256:252dc720f63381e7bfedff916cc0b6506c5fe30b436ffedc934c6e23df0187c5

Observation 64b4855d-4bd5-437b-a49b-481ba80a0497 · outbound

This paper cites Nous-hermes-2-yi-34b model card, 2023.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Nous-hermes-2-yi-34b model card, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.136959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.190184Z digest=sha256:3b08a3f14074430ee7d70122e2bc75761deff3e34341dc05ea0e19ba9f970165

Observation f44c9f46-5076-4aaa-b059-05d27b8992af · outbound

This paper cites Gpt-4v(ision) system card, 2023.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Gpt-4v(ision) system card, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.193233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.193233Z digest=sha256:8d12fe1d8aeea9482c0d9c3ae65576620e883f6e4b414a5612b8ab96c94d9341

Observation 52ae98c6-b9ae-429d-85f0-971bf09715f3 · outbound

This paper cites GPT-4 Technical Report.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models GPT-4 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.195669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.195669Z digest=sha256:1679848ce531da17cc43a48e392d90a76271968a2f41178c42038c7d60d8dc1e

Observation 062d0e32-bcfd-44f0-b41a-f06eeae5b588 · outbound

This paper cites Introducing routing functions to vision-language parameter- efficient fine-tuning with low-rank bottlenecks.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Introducing routing functions to vision-language parameter- efficient fine-tuning with low-rank bottlenecks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:33.089594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.198582Z digest=sha256:e2a3d228f7f7d34f42fd1862a42117e6fb5dc35d86cdd244ba53e3048ab7b010

Observation 86c645ed-e928-4f3b-b381-1124b575ed83 · outbound

This paper cites Learning transferable visual models from natural language supervision.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Learning transferable visual models from natural language supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.201141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.201141Z digest=sha256:b0f07fb795bbf64e1fb80e521b9def2153abbc4b34889b6ccf98e1c1c3990fae

Observation 8dadc60d-6921-45fb-945e-4233ee06fd1a · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.204134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.204134Z digest=sha256:9f6f0e430b5aefebd7fefb0662cec84177de79dc5d7e75e78ec71e6fb911de0d

Observation 07b575ed-eb43-4530-9b83-7eb9ad6cbc9f · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Moviechat: From dense token to sparse memory for long video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.997350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.207504Z digest=sha256:d5b7572ea44195e3077b7ec907adf2b4511f76b8c4e51fe1789576266d20cd54

Observation c59ad5fa-7152-4191-9dd5-9fa4c2acdc0d · outbound

This paper cites MovieChat+: Question-aware Sparse Memory for Long Video Question Answering.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models MovieChat+: Question-aware Sparse Memory for Long Video Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.209929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.209929Z digest=sha256:913a72cac7cda1dbb5053cb8d079b44e2a0ef6332cd8ce7692e89b97fc96acc8

Observation 357115af-0b2a-40e2-b44d-f4237438b1d1 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.245623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.245623Z digest=sha256:3d34264b15550cb554f1210f53eb5219a2eaefd189356ad4d9d94a7d50679bfe

Observation 30c069cb-d8a1-4ee6-b41d-681485490811 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Videoagent: Long-form video understanding with large language model as agent

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.966106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.290029Z digest=sha256:61b3877bf7ed24e10956200a1b90961b9ac75fe78a0656849611b8569c6f38c9

Observation bd605700-9e6c-4c29-8b8d-1dabd561fca6 · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Videotree: Adaptive tree-based video representation for llm reasoning on long videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.940983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.393563Z digest=sha256:16769a15dee8bc6f92267532433c86a7da7f593674ff8a0f10ab149d255782a9

Observation df1c2007-d344-4dd7-9213-d609fb4ed462 · outbound

This paper cites FreeVA: Offline MLLM as Training-Free Video Assistant.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models FreeVA: Offline MLLM as Training-Free Video Assistant

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.413331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.413331Z digest=sha256:0135f98be865104f609524cbdadb9f423a5e626eea6de29ab4649c8dbcdbdf1f

Observation d8aa5db7-4889-491b-9f80-ac9b2b609948 · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Audiovisual SlowFast Networks for Video Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.416966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.416966Z digest=sha256:76360a45feffcde4a17137497e8159b16aadb877da1555207b03e152fcc82bbc

Observation 9e9814f2-3e7a-495e-96e7-e2c7504d0be3 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Next-qa: Next phase of question-answering to explaining temporal actions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.420208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.420208Z digest=sha256:c56b9ec8838c6c46b672b16882dc5a34a63c1b96eace2d67baecfe9272bad8fc

Observation 454efdcd-bade-450b-947b-92cdef74de9b · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Msr-vtt: A large video description dataset for bridging video and language

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.897591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.424070Z digest=sha256:28241b92cbdbf93bd8a4b2e24b948ed73013bd47f8958067b13b3a672bbeb9a2

Observation 960ececd-5f18-4a8c-8b78-dfb1f4466133 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.426845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.426845Z digest=sha256:58791d1da9544264295f424fd8ba0bb5728f88e4edd570ba00e02289291855ad

Observation 90f4958e-b8eb-4ca9-80d1-ca51d38a00e1 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.430071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.430071Z digest=sha256:8b0c74c411b30ae9a1c4728276d7b3a0a3f77c7436d9440b03f7bbd5d0c3c7c5

Observation 6bf81f57-f298-49cd-a8ba-365e17bce8b4 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.433441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.433441Z digest=sha256:3f9ab9ddc787b3f7e4c98f7dfeed21702ba27d6d8b580e50ca4754e510ff860b

Observation f7504c50-82e8-4672-a0e5-d1691be65961 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.888047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.436405Z digest=sha256:b1a8a692fee0c6ea8dc097d7b0fd0d9f8cc6bbc23b6708318bd0263337e43e40

Observation 0c2cd52a-0aeb-411f-8a50-df560546a403 · outbound

This paper cites A Simple LLM Framework for Long-Range Video Question-Answering.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models A Simple LLM Framework for Long-Range Video Question-Answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.439057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.439057Z digest=sha256:cd9b3b36080be8d62f2742456938b4f1850f9797f95cd4401fa2925cacc9ce54

Observation 5251b9e7-9c04-4f8d-80a2-baca23deffc2 · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.878571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.442459Z digest=sha256:0b04b8ff73dac5584a7b5c5453c6c37fa775fd5ac6b91751b23a75290e322407

Observation d5f31825-3ca4-46ec-9b5c-2aa88069d694 · outbound

This paper cites Long Context Transfer from Language to Vision.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Long Context Transfer from Language to Vision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.445716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.445716Z digest=sha256:4681e9a8d6ac2c2d4e9157c93214037f9a79a393acddeb9aad839f52ab28351e

Observation ff5e4446-e2bd-4034-9277-7defca510274 · outbound

This paper cites LLaMA-adapter: Efficient fine-tuning of large language models with zero- initialized attention.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models LLaMA-adapter: Efficient fine-tuning of large language models with zero- initialized attention

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.449162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.449162Z digest=sha256:51281e64019f482b4af3355310b592c8807711afaeb1ec2a6261f3270a1840e0

Observation e7f5868a-7c34-4aeb-9f05-91b1058f5499 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Llava- next: A strong zero-shot video understanding model, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.861900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.452314Z digest=sha256:963335eabbc4a95f100fda9cf10724f4ebba3c0ad1a254edce8697dc83b46ced

Observation 9290f609-327a-4826-870a-9b568ad1d897 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.455173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.455173Z digest=sha256:2f610c0f62bd72753b966bb80676648e87d8c31299a4f610e7f4247880cdb022

Observation 9bd6bf4a-077d-435c-897b-011540ce64f1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.482619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.482619Z digest=sha256:46706121420af9b35e77fd702f653a159317bfcee5a58f1642ecea5e93258bad

Observation 5fb455bb-ce26-4a11-aed0-35a49164e72f · outbound

This paper cites Each frame is resized accordingly to fit within the resulting 336 ×336 thumbnail image.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models Each frame is resized accordingly to fit within the resulting 336 ×336 thumbnail image

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.850844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.513984Z digest=sha256:1a73e224dab8113966499f981a6b6db32c0e369a8bad8fe7428a041b0ebd3dd8

Observation 6f72a5fd-f31a-4f6d-8dfb-366a28852cda · outbound

This paper cites We start with additional experiments conducted for the study on compression strategies.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models We start with additional experiments conducted for the study on compression strategies

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:03:32.839951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:03:32.584154Z digest=sha256:04ac5e8deca53621b9c82a8c27b3b03c6d83ff09bfb69830470ebba8a006386f

Pith citing papers

Observation f1827bc4-9c67-4cb6-ace2-6ab52c299c2a · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:38.250264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:38.250264Z digest=sha256:c1e3d6e6297f0116743a906243b87d2d2edf501b08934e810abfbe9ca109a812

Observation 3bb5536f-a601-4a5b-a460-1e52aa4c1314 · inbound

Direct RNA sequence design under codon constraints using expressive tensor-based secondary structure models cites this paper.

Direct RNA sequence design under codon constraints using expressive tensor-based secondary structure models TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:29:46.584977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T00:28:35.611378Z digest=sha256:633ab5762530ddd49fd686782fb0f419acb5b6cc566876f638360070ae91c880

Observation bbac7741-6968-45af-9c5e-35fc5092ab46 · inbound

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization cites this paper.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:07.914366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:e87b7edb5abbeb557ba2ff71be3f71141ab1c51e42acf7daf261a04588036caf

Observation c5335c59-93c7-493c-a6f2-bcc8ab59db26 · inbound

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding cites this paper.

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.852517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T06:16:32.972099Z digest=sha256:493559c7a35849ca909fe4e6426173c4fd9d2555ca6a3229f390df7f395ec9a8

Observation 5348825f-e32a-4020-9bf9-db996aa84538 · inbound

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation cites this paper.

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.047994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T10:54:02.188634Z digest=sha256:71c44b2ef6a2ee181cf3eaffe2e35faa7be19fed3ffd486ae7e6a483be9bd253

Observation 9e73c270-c79c-40ff-b7cc-3c2a95686929 · inbound

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering cites this paper.

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:46:56.879237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T01:52:33.768494Z digest=sha256:0d31932cf6a1a45704601d558e7dc391e14ffad90a3fb2ea706f6dddd0bdcadb

Observation 6464f441-c331-49fd-ad4d-a67d7c43f29b · inbound

TreeSoc: Tree-Structured Dynamic Reasoning and Tool Synergy for Soccer Video Understanding cites this paper.

TreeSoc: Tree-Structured Dynamic Reasoning and Tool Synergy for Soccer Video Understanding TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T07:49:34.639380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:49:34.639380Z digest=sha256:17da580fb3f694c0236edec1a57c0324857e93f8066048864e68ba2134b7443d