Pith. sign in

Paper Citation Record · LEDGER

LinVT: Empower Your Image-level Large Language Model to Understand Videos

As of 23 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 7 inbound Pith citation observations for arXiv:2412.05185.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05185 v2

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:54:16.716594Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:51.011685Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:56:47.335501Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5dc14061-e0f3-481c-9e21-ea4fd459a802 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:15.910033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:15.910033Z digest=sha256:a5d8b39fee4382abdce61e1357dd65898625759e76525616af41126c53e4987f

Observation 852db616-abf6-47ed-9907-869d2d9bdfaa · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

LinVT: Empower Your Image-level Large Language Model to Understand Videos MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:15.917336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:15.917336Z digest=sha256:acbd3e32aa1ee3b8d7f39c324e6ba76d91539186801f952652d0d4e6ceca15bf

Observation f635496e-235a-49ef-82c9-965cfabbf208 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:15.923774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:15.923774Z digest=sha256:f4da8242937b6e18be3c3cbbe1dbbacddcf2b73dddd30c5adefdcfd293a5b6f2

Observation 9b740c24-d2c2-4211-a045-07097d4f31b6 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:15.935643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:15.935643Z digest=sha256:dd4d6b4f17d856d065519eab4cac2b5f4c36a485fb690e263fd9817ad74a6cb1

Observation 53c06836-8ede-4c36-b014-1544b13184ed · outbound

This paper cites Revisiting the” video” in video-language understanding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Revisiting the” video” in video-language understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:15.943141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:15.943141Z digest=sha256:b8fba4bc78c5c0a32b73d6c516d732f953d946234e102dd72cd7621cbe10c8f7

Observation e7f41d36-55e7-4f44-b779-125fc154ff97 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Activitynet: A large-scale video benchmark for human activity understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:15.957431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:15.957431Z digest=sha256:9792687d91354891179ecb73f97b24dd7aace6f38c3951a6ebfb0154a3b9ae49

Observation 885d6d3d-6697-418d-8e0f-df75883c3c6d · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

LinVT: Empower Your Image-level Large Language Model to Understand Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:15.963977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:15.963977Z digest=sha256:18f4ed82f94b192db2b846f860e122d2cc3940a568ae3581ea98fc6e34b294fd

Observation 7d70b0b9-8b00-40b6-8b24-8c9847ce786d · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

LinVT: Empower Your Image-level Large Language Model to Understand Videos How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:15.970790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:15.970790Z digest=sha256:b7599ed196f1834b1d94b3f336ced9690229b8c0ea28708f35435aaa4068fb9b

Observation 291a1813-5a22-4529-b9ad-c1c194930586 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

LinVT: Empower Your Image-level Large Language Model to Understand Videos VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:15.984621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:15.984621Z digest=sha256:f7d3a8863df6d29ff5e2e3131f7970aa041296f9309b456832ff9e995aaebc65

Observation e238d3ae-6c7b-4833-a16e-4dc544dcecc8 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.002888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.002888Z digest=sha256:21e8b2b505d50157e93b4119dc09fdb8ff3b3f60d78c2c75669c401cd037f38f

Observation 0d85e74d-a4bb-4a77-8962-5a4837f587ae · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.014239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.014239Z digest=sha256:a364c0efd5480fb4e838625732d9096f1ce6ccf3363119d5a5f59a3c3ab61bed

Observation 850977aa-d673-47e4-8b7d-43a92542ba1e · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.020195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.020195Z digest=sha256:0e0c53b89be188310c75d36e47b05d4ee571bef7f0c77e02473156514ba98aaf

Observation 4550844a-d677-4d83-8c5f-d04588f1f544 · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.027352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.027352Z digest=sha256:d8069c8eb18d7a59b8f42f20eca62fe474a32a2300b7f83b7ac6374cda4f276f

Observation 396c5639-611d-424f-99e4-d82669d17b38 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.034272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.034272Z digest=sha256:1324aaf17b91d582b6b3814c98adc25efe222c76abf019f7b4634993a844991e

Observation 13d072b5-f133-4f36-8e09-58c51f3b58c8 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.040767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.040767Z digest=sha256:5eed8da10c2c782497c619dd16f2da16e2e9755498d05156b49cc0191fe4e333

Observation 27bdf96c-db8e-434a-8f83-eb19c5148028 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.046710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.046710Z digest=sha256:6720bfd3371b85aef99842b8e85d754a960accb60d03630df738008eebf4035b

Observation 88ca61d6-2467-463d-9ec8-119d1d6d4566 · outbound

This paper cites LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference.

LinVT: Empower Your Image-level Large Language Model to Understand Videos LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.056197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.056197Z digest=sha256:ea4b667489bd2a925d8bbffcb9722856bfcda1cfdbbf0b9c54dfa52207fe9b49

Observation 7a41d7d7-a942-4c84-ad28-26de6c7e37a8 · outbound

This paper cites Saliency-guided detr for mo- ment retrieval and highlight detection.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Saliency-guided detr for mo- ment retrieval and highlight detection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.064473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.064473Z digest=sha256:db7970b62f79b256e63ef93532fe086860b4e28c523055d8ced6e7ba54041333

Observation 31ecd2a9-c40c-4263-8fbd-2d667c2faed0 · outbound

This paper cites Infinity-mm: Scaling multimodal perfor- mance with large-scale and high-quality instruction data,.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Infinity-mm: Scaling multimodal perfor- mance with large-scale and high-quality instruction data,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.071284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.071284Z digest=sha256:e38bbeb8d46dccd0ac42dceee300dd0774b3a492c936a339858ef5f1a769e949

Observation 859c2de1-bb90-4cca-85f8-c8c77b60910a · outbound

This paper cites Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.078424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.078424Z digest=sha256:db6a974019694bd4f96e0c65dd9ee6ee49ebe797e029f3ab20c8317bc9aadbda

Observation 792263f1-7a0e-467e-8146-5c6ad9088974 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos LoRA: Low-Rank Adaptation of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.086722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.086722Z digest=sha256:79c7e2660d0fb520fe49fc7607222feaefba059320f67c4256546edd25854034

Observation b61746de-c447-4709-aca1-098094180db4 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.100787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.100787Z digest=sha256:35cca7eb31418366499a3017d57db9d51392ebb1b9cbaf68103ae86e71cbed67

Observation fc7c82e3-ca10-46fb-b449-26e0a73ce3d6 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.109596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.109596Z digest=sha256:ec4e39445582e1bb5126ed06a61d2a24ab068f6e3402d1ee0c41e8052e5b018d

Observation 9428fb93-49c5-4164-a31c-a4649442ffa1 · outbound

This paper cites A diagram is worth a dozen images.

LinVT: Empower Your Image-level Large Language Model to Understand Videos A diagram is worth a dozen images

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.115802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.115802Z digest=sha256:8c2f1a86bed9196f17f45a95ba7c997ad546ed89ab771138193e38bcc1db47e0

Observation f1585358-d089-4f2f-bd40-0056b058c3e5 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Detecting mo- ments and highlights in videos via natural language queries

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.974776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.126201Z digest=sha256:670a4fec80e6ae540a62de7e59354f58326a198f634d94852f621e6237d9c53f

Observation 67c25b43-a5c3-4a1b-9c40-aa1947da8c33 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

LinVT: Empower Your Image-level Large Language Model to Understand Videos MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.135074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.135074Z digest=sha256:9626c46230bc46f1fbbcfbb1e651d86963fd0b246ab44b6fd15563dae8bee6e3

Observation f2c9feb3-c9e9-4ad5-a7f0-ea21234d9dba · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Seed-bench: Bench- marking multimodal large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.943022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.142045Z digest=sha256:e8ab60f089808830318ea007eab0aa2c71d5233c87f59c45c9c8fe0f0d723062

Observation d88b740c-8e4e-49a1-b8cc-57b91c137d99 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LinVT: Empower Your Image-level Large Language Model to Understand Videos LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.147964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.147964Z digest=sha256:ec4b4f3551123b8225ff6aeb1a6901681ce8f15056b365ef385b9cb0b8b448ac

Observation 09172287-8ce2-4372-8ec6-512f66059795 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.916248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.156877Z digest=sha256:2949065f8db6b2f067a5f3e9908fc2e8407f93d46407806467538da260a84141

Observation a302d60a-ea03-4908-82c1-477ba1b94e62 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos VideoChat: Chat-Centric Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.164884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.164884Z digest=sha256:4c3fe95e444e367bbcdbc1d70297b8379f3f30e6aefd4bd6b4623e7a355ecf39

Observation bd69e5f0-b7b5-4691-a043-157a5b50eeae · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.887668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.171614Z digest=sha256:c60deb80727ddc25e3738e85d30124208a4e772f7569de34895c56d12dda7ce6

Observation 1aed4ae9-c534-45fd-88bb-1c390699e2ef · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

LinVT: Empower Your Image-level Large Language Model to Understand Videos M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.179148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.179148Z digest=sha256:6be836a9df59058dd75d579d32d6bf871d281059342097b05efe1339920e37ac

Observation 5d7277bd-e588-4874-9fe8-6f783d5fbe5e · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Evaluating Object Hallucination in Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.188330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.188330Z digest=sha256:a3cfc1232d9a94effc65031ae24de62b22366dbeb861a7e7a4ee71540bf80b8f

Observation 83bce1cc-4f62-40e8-9a35-40d97de11af2 · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

LinVT: Empower Your Image-level Large Language Model to Understand Videos VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.195188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.195188Z digest=sha256:804edfdaab461add810875a20e7ee69951d36a7f64d983e9fca6e24330912834

Observation 8b20b214-48ec-4c71-8588-01e00e3ed1c6 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Llama-vid: An image is worth 2 tokens in large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.842858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.201551Z digest=sha256:27db4ff966570dc8ebb9c14a294f13b48011f77f7fb890b81768902d8e1527dd

Observation 72bd0409-4cbf-4573-a813-95a36e2cd732 · outbound

This paper cites Detal: Open-vocabulary temporal action 10 localization with decoupled networks.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Detal: Open-vocabulary temporal action 10 localization with decoupled networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.808513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.208070Z digest=sha256:bd036af2bae46569bd82a28d70280548cb3b011a4d5fa0f50989d6926aa5724d

Observation 0b962e72-1c3c-46db-979c-c39df16e618c · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.215020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.215020Z digest=sha256:b6d80d4b79c2c6acef7d260a2081e8bf9f8ad858183f087efb721504a9a7a917

Observation c2037de3-473e-427e-b2fb-46634319ed74 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Vila: On pre-training for vi- sual language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.772516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.221010Z digest=sha256:e82352cd48a850d5b5ca59a70b6914a6bb85763275f2c5564032294456086a4e

Observation c4152585-fa01-4d5c-a33a-10787db2a4a6 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Univtg: Towards unified video- language temporal grounding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.748183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.229463Z digest=sha256:6b0e0ee4cde28c19f7ba39386fbeed29f4593a2a66e8d2801faedd9f85f38765

Observation ad53c99c-4772-4714-8c9a-63b4970cbbe5 · outbound

This paper cites Visual instruction tuning.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Visual instruction tuning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.712901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.240693Z digest=sha256:08ba87b947b53571724b46df12a8bd405e4bc87f5c1d3b1d826607e705fc6720

Observation 3b77870b-956d-45b3-9b29-202d1f92d7df · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.256483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.256483Z digest=sha256:b0c18809f12501ab48e3e361b9b20b4b6b5c4ae40254f05b92ca8676adf4d8ae

Observation 3f584a53-6ec4-4382-9eaf-7fdea08665ee · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

LinVT: Empower Your Image-level Large Language Model to Understand Videos TempCompass: Do Video LLMs Really Understand Videos?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.270074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.270074Z digest=sha256:17e7df77a929c5ce19cd25ca8da047899ab2219d38363c83e139713dbb78b1d4

Observation 86425931-5824-470a-b991-7e2b0337dc5a · outbound

This paper cites E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.286499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.286499Z digest=sha256:987c4c0faebec80f8260da77d0f824b9e7ab052416d89db8b33e93a4506f2b3f

Observation 945ebb97-a10c-4d9b-97af-c503e35b8d95 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.295108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.295108Z digest=sha256:a00df24107ef042f0a42a15896b6c0bbabc1b311c7d6b0692728309aa949844c

Observation b26d8b05-3d70-4934-bfb3-f67df0002500 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.305623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.305623Z digest=sha256:5b7951651be9e16578cef66ebda4bbea572736bf7891bd6a0d980f93280c1421

Observation e4bfa45b-ed4e-4206-ac5e-1d5e9fd8cc3c · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.316927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.316927Z digest=sha256:b17331f9f5a9503f8e44eae75f107a116d014d7b872726900fc41d47ed9a62dd

Observation 825b849e-405d-44f1-80cb-08e4ccba5722 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.330529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.330529Z digest=sha256:36fefe4d3c1ebc8d1660a7d402495985e49fc4c330212728be169a48153c5142

Observation d118f865-05bb-447f-9d43-656536eacb13 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Docvqa: A dataset for vqa on document images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.338453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.338453Z digest=sha256:8b60d978f7129a09e71afe27848d98de5f9e8b9f10a622029601d2ba2c9a1d62

Observation f27365f7-de19-411f-968b-05a9ec37efcf · outbound

This paper cites Snag: Scalable and accurate video grounding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Snag: Scalable and accurate video grounding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.610559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.347424Z digest=sha256:4ce0596949cae4c1e44fc7b3385a7ca064cf1b54d9e034f5dca2b65e92ca8ab0

Observation b58015fd-920f-49c7-a41d-c964c24d5baf · outbound

This paper cites 4v (ision) system card.

LinVT: Empower Your Image-level Large Language Model to Understand Videos 4v (ision) system card

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.560201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.355885Z digest=sha256:d1dc674df17e7728be2b6a16718cc7ba73b3cd6159d466caadf96e699c8f5864

Observation feb3f638-c942-464b-a23c-dace1f571fb9 · outbound

This paper cites GPT-4 Technical Report.

LinVT: Empower Your Image-level Large Language Model to Understand Videos GPT-4 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.363366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.363366Z digest=sha256:c84766943a930e45787a413dcf27845e6ae4c2d245937b71e89a1ff6b04731cf

Observation 4664d066-03e3-49fa-aaff-b713413da57b · outbound

This paper cites Scanning only once: An end-to-end framework for fast temporal grounding in long videos.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Scanning only once: An end-to-end framework for fast temporal grounding in long videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.520947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.370469Z digest=sha256:a2b9d34618c52812a4d382e561f0c10c2710470c5c2b3c4884bef88952c334e6

Observation 13edb34d-4d65-4d0a-b6ae-f772c38069b5 · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Per- ception test: A diagnostic benchmark for multimodal video models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.487513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.377465Z digest=sha256:c130739ca4a6e50ba8b37eec32f206e95cb2f090c9c188d86984de553efe11ff

Observation 5410bbec-2005-4be8-ba32-49007412a2e4 · outbound

This paper cites TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.384673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.384673Z digest=sha256:66570c96516888f80dbdad64b27939cb0036d59a4198c99c3b1f0f9bdc8be193

Observation 1ca72c69-b34e-4a3b-a4f6-dcb8b7090296 · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.458276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.392422Z digest=sha256:c15f537fed3dc391be341a7514c0756792089aa2b80b292203c8e51efa9dd6cd

Observation 7bfd6d50-e9e7-4de4-b680-50eacabd53a1 · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

LinVT: Empower Your Image-level Large Language Model to Understand Videos xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.399375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.399375Z digest=sha256:74c0a6127958516409605a2576487b93584b5c6078eab2191430ff05a1b466d5

Observation fd02c596-ccb0-4908-a753-aa0254423ca1 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.410121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.410121Z digest=sha256:1fa8114ae10edd8a2aa7b9a9196b511091b3f10b869ad6d8cbdbc52a8dca2b3d

Observation c22e7847-d52e-41c7-b0c7-009d22ecfc0f · outbound

This paper cites TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval.

LinVT: Empower Your Image-level Large Language Model to Understand Videos TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.416990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.416990Z digest=sha256:b868c92cbfa6f3018f5802028ead48e0fbab49c6847cc75c7423af15c10d7334

Observation 6b5b8fe8-7797-4d7e-b20d-118d64623924 · outbound

This paper cites React: Temporal action detection with relational queries.

LinVT: Empower Your Image-level Large Language Model to Understand Videos React: Temporal action detection with relational queries

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.424431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.434362Z digest=sha256:9ed4c640d88b32c66320c2d354040872047f83e514207076729eddae7a4f44fb

Observation dbaddcac-98b0-468b-ae6b-844078d3d9a1 · outbound

This paper cites Tridet: Temporal action detection with relative boundary modeling.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Tridet: Temporal action detection with relative boundary modeling

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.382023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.444441Z digest=sha256:d1cab1fe8bdf99d590fe15b7c74ca95fc3ee2d53b95e7c509496ed2eecd97ed8

Observation 9d2c0c10-baa2-4edf-9738-cda68b6d61e0 · outbound

This paper cites Towards vqa models that can read.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Towards vqa models that can read

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.452275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.452275Z digest=sha256:279023fa6332eadd9278caa3a1916300ecf2ce0b326c93f5cecdf829a11ad0f6

Observation cee085c5-9f9f-46ee-a889-3bcf5c3143d9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.459122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.459122Z digest=sha256:d2fcc5069154bc7dddc815c4cec891c1a4e934c482b40c9ceaf939941ead0452

Observation 732da97b-b6d9-41b6-891a-c783e8bfae8f · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

LinVT: Empower Your Image-level Large Language Model to Understand Videos LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.464792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.464792Z digest=sha256:e2c7fffe1bdd24222924c686f3714636ac4e70e8905a35ab64629048978c6b6e

Observation 1f9c87a1-a224-42c3-87d2-2b5769e407f0 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.477303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.477303Z digest=sha256:f7e7a50a89bcb9938024ef13b25c4d246da2c1da56ece6cb15de2cf667327b41

Observation 4aebf4ec-6ebc-4667-8d08-baf1b7f33c44 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.489581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.489581Z digest=sha256:76453d19eb840fea63d505b65e14a3cf4e8e5df896db94784bfeca60f0d04b1a

Observation e48cf4af-68ce-4e96-8819-9634f4d81204 · outbound

This paper cites Non-local neural networks.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Non-local neural networks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.499101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.499101Z digest=sha256:7833dda79fd0a81a1d7119c69a2c4bf55ba4c3aa9ca3adc62b6616f0b50b4095

Observation 63dfb193-06fd-487b-8e67-027fdcb92d8e · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.504426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.504426Z digest=sha256:4b3937824d27ea8235e64150be4ab0e5db7c046079beb305fbccd72f1f8a8c4c

Observation 99ce7e4d-717d-4991-8488-4779a4090651 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Next-qa: Next phase of question-answering to explaining temporal actions

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.514092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.514092Z digest=sha256:df56d9308617fcfe01caf7013e96f77b630dfae2e365de62442de0f9b748db10

Observation 45694ab0-fe91-4a2e-a2f1-7e89eb16a041 · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Can i trust your answer? visually grounded video question answering

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.296269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.520494Z digest=sha256:4f1aa05479681422e49c08d3e2cdd0b4063912f6a8b5eca38f92ac2c9117ff30

Observation 243ad04b-20a6-4c1e-b9fe-c5d21c0ec617 · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.526505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.526505Z digest=sha256:baab6b28cc1f8f9ebe62ced1c24953fd143e21b2c69d363412601b8a5674a514

Observation 5e89a1e8-347b-4b8c-b5a7-6f886ffd551d · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

LinVT: Empower Your Image-level Large Language Model to Understand Videos PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.535561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.535561Z digest=sha256:481eb8c7fcecde0912aa1e393812dfe031ba1a25a17e26d15f2999be18433ea5

Observation 47d7345f-2cd6-49c2-ae8a-22677989dc35 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.543027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.543027Z digest=sha256:a6c1cd7cd8608890636dc3e104476225cfbfe17885434a6c7d9062867bf87bc3

Observation 508953d0-0975-4cf6-9d4d-d9398d4d629c · outbound

This paper cites MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning.

LinVT: Empower Your Image-level Large Language Model to Understand Videos MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.549792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.549792Z digest=sha256:ff0eab9d964a0faf6b4030dbb8bbb370cdd2be7948d8ffd7b683bfac79eefb3f

Observation 7d32a76d-932e-4a78-aa7a-52758bd63697 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

LinVT: Empower Your Image-level Large Language Model to Understand Videos LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.556012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.556012Z digest=sha256:ce7664de4ac3a02b1d4687ddfa72070154f279b41fefa7363d586cb2e6dd3741

Observation 29606d90-03a6-458a-aeba-b39c76bbd570 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos xgen-mm (blip-3): A family of open large multimodal models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.573195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.573195Z digest=sha256:edcd3bca9421ae493b0e444cc0e3413bba5f1b48559641398ea699e6f2d018a8

Observation 1dafeab4-7543-4f4f-a368-9ac9526c3059 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.237279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.583346Z digest=sha256:d91d3f27dbf0270c110f33835bbe14027c7990ac83696a87db06ced7b5543675

Observation f2dc37d0-bbf1-453a-a883-dff3de098cc0 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.204222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.591817Z digest=sha256:65b777c885122e7ab5954a82edcf21f005348002b592a45164ba4b47a3ed7e3d

Observation 195b3517-6d8c-43fa-a4d8-7b7f33a06d35 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.597463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.597463Z digest=sha256:911b4d4df9febadc9e7f4f773dd14199448e00335eb6be9fada0abdc62525507

Observation 73dcccf3-f323-4ea8-8967-1c371ef89db9 · outbound

This paper cites Unimd: Towards unifying moment retrieval and temporal ac- tion detection.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Unimd: Towards unifying moment retrieval and temporal ac- tion detection

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.132984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.607062Z digest=sha256:bdae674dfc864ea76303da87f1ab80ae45e6db980e434566fb02c3eeaf1b44b2

Observation 97f0f6dc-81ee-4c09-880a-e2bd5ca4651a · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.615199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.615199Z digest=sha256:83a8366bcd3f03b892fcbda7c168e542cc4eeb0179a3aadd4fcf64d6455a2a16

Observation c533a558-e491-4a95-aebe-462d1261e8fe · outbound

This paper cites Long Context Transfer from Language to Vision.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Long Context Transfer from Language to Vision

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.630647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.630647Z digest=sha256:22a2cbfb4fd6aa1988f795ac78168787fd888d8c2513803b089d92924b260a34

Observation eee166cf-f9c9-4556-9cd2-ee9f2bafac97 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.639694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.639694Z digest=sha256:aaa1434e5143e28cfac0909a8b4063a21d15ec8f1774441ba9d8754657a993b8

Observation 90fdb764-d941-40a2-864d-b49ed48f944e · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Llava- next: A strong zero-shot video understanding model, 2024

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.646904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.646904Z digest=sha256:89a6abeea65c435050bb700e4f37333bf31c57b5090ae937ef8f3142489e16ec

Observation 0c835781-9d0b-4789-8b74-4ad4b54b8ca8 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

LinVT: Empower Your Image-level Large Language Model to Understand Videos MLVU: Benchmarking Multi-task Long Video Understanding

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.660595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.660595Z digest=sha256:228a7fb43297e28eb50a6a804cd52792474fce9f1a3caade03a1baa93539f51d

Observation b35c9e04-9cfb-4cc8-b41a-83591684ea07 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.668052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.668052Z digest=sha256:8fa956744e7bd1fa8d2e1b9e96438161a172ae237de739bec26610fc0c468d18

Observation 0c855a13-d7ce-447e-969e-9181ab7ba09d · outbound

This paper cites Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.675248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.675248Z digest=sha256:dbf65d38e02f300d7cff0e77b741490868ce4c012626f523b4ad0943da924597

Observation bb804248-7373-4f6c-8270-6efe8fa73bf4 · outbound

This paper cites Openai’s gpt-4o in surgical oncology: revolu- tionary advances in generative artificial intelligence.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Openai’s gpt-4o in surgical oncology: revolu- tionary advances in generative artificial intelligence

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:19.012512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.682870Z digest=sha256:c1a5c8ee9fceb09fc779251bdabd20b767e95d9e2af61831b88492f92ce1acf7

Observation cd6393ad-0425-4d5a-93f9-30e577ff0eff · outbound

This paper cites 12 Autoshot: A short video dataset and state-of-the-art shot boundary detection.

LinVT: Empower Your Image-level Large Language Model to Understand Videos 12 Autoshot: A short video dataset and state-of-the-art shot boundary detection

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:18.978265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.691146Z digest=sha256:8417a0fa6103030aa04426e4a1cb5dd4dd02fe9f0d2cf40c1bdf6fd36e5a8436

Observation e1e9fc91-9055-4050-b8b3-2c69f4484726 · outbound

This paper cites 64, 32, and 16).

LinVT: Empower Your Image-level Large Language Model to Understand Videos 64, 32, and 16)

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:18.942446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.700943Z digest=sha256:96ddf75e5858d86682e482150c80d10e98bd2c6722a741965fee6f0b5442a6c0

Observation df64a703-d495-4241-8493-f3a2eba7fc94 · outbound

This paper cites More validation results are presented in Tab.

LinVT: Empower Your Image-level Large Language Model to Understand Videos More validation results are presented in Tab

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:18.910891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.710807Z digest=sha256:8af0ea7d93a6a7d5928df121374efd1615248a2b9c445a88c6d0f26129ad6726

Observation 4978add1-5682-48f8-ba9c-97a87048fd99 · outbound

This paper cites Image patches of the selected tokens.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Image patches of the selected tokens

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:54:18.880649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:54:16.716594Z digest=sha256:b90ab1da635d1074693005740ffb53a4c55b3cc30192009e50c4a0f0399a57d4

Pith citing papers

Observation e2bdee07-8166-41f5-a9a7-78f04a2aaf8d · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:51.011685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:51.011685Z digest=sha256:45c36f9662101373738b3ed1380df1a251528420e8fc41ed97c7eb9c06793d50

Observation 7a28d731-d390-4e5c-9c8d-1465ae651bdd · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.759496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.759496Z digest=sha256:316ede6138566fa4fb919bed3488f07c6af9b322e8c9d56d8bf97643da99754c

Observation 47912e91-9c6a-4594-8839-40ba677a6da4 · inbound

${\mu}^2$Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation cites this paper.

${\mu}^2$Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:24:45.561143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:24:45.561143Z digest=sha256:bafdadc11d0f59e2710505ace877f8497280752555caafa353f26720440746d9

Observation e555be81-c022-47df-b305-a1faed9b8205 · inbound

AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multi-modal Information Between Reality and Videos cites this paper.

AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multi-modal Information Between Reality and Videos LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:25:02.720820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:25:02.720820Z digest=sha256:b0b23f148251b228424e2ba7d197ba30fb19863efaa704ed4b546eb0c58ab00a

Observation 9b54c22b-409b-4082-828b-5cac0d08ff71 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.511427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:931a3103760a329db65be0fb2c5c56a1a66b40f721a34ada5cf3091f17c55458

Observation 944fc3e0-f602-4b05-be63-7a75b5014222 · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.337270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:675c7af8cf9cdeb87eeb7baacf4aac67270cc4430f0f7e0932d816611e68b4e0

Observation c3ac952d-456c-4709-bf4b-b1a34cc5cead · inbound

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction cites this paper.

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:22.394593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T07:48:01.719339Z digest=sha256:45162b1c30c3cabf7a354b8b9550272360dba0824b7d1090ccee3b897d79f6ef