Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:54:16.716594Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 7 inbound Pith citation observations for arXiv:2412.05185.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:54:16.716594Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:51.011685Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T07:56:47.335501Z
91 of 91 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5dc14061-e0f3-481c-9e21-ea4fd459a802 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 852db616-abf6-47ed-9907-869d2d9bdfaa · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f635496e-235a-49ef-82c9-965cfabbf208 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b740c24-d2c2-4211-a045-07097d4f31b6 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c06836-8ede-4c36-b014-1544b13184ed · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Revisiting the” video” in video-language understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f41d36-55e7-4f44-b779-125fc154ff97 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Activitynet: A large-scale video benchmark for human activity understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 885d6d3d-6697-418d-8e0f-df75883c3c6d · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d70b0b9-8b00-40b6-8b24-8c9847ce786d · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 291a1813-5a22-4529-b9ad-c1c194930586 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e238d3ae-6c7b-4833-a16e-4dc544dcecc8 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d85e74d-a4bb-4a77-8962-5a4837f587ae · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Instructblip: Towards general- purpose vision-language models with instruction tuning,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 850977aa-d673-47e4-8b7d-43a92542ba1e · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4550844a-d677-4d83-8c5f-d04588f1f544 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396c5639-611d-424f-99e4-d82669d17b38 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d072b5-f133-4f36-8e09-58c51f3b58c8 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27bdf96c-db8e-434a-8f83-eb19c5148028 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ca61d6-2467-463d-9ec8-119d1d6d4566 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a41d7d7-a942-4c84-ad28-26de6c7e37a8 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Saliency-guided detr for mo- ment retrieval and highlight detection
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31ecd2a9-c40c-4263-8fbd-2d667c2faed0 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Infinity-mm: Scaling multimodal perfor- mance with large-scale and high-quality instruction data,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 859c2de1-bb90-4cca-85f8-c8c77b60910a · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 792263f1-7a0e-467e-8146-5c6ad9088974 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos LoRA: Low-Rank Adaptation of Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b61746de-c447-4709-aca1-098094180db4 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc7c82e3-ca10-46fb-b449-26e0a73ce3d6 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9428fb93-49c5-4164-a31c-a4649442ffa1 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos A diagram is worth a dozen images
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1585358-d089-4f2f-bd40-0056b058c3e5 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Detecting mo- ments and highlights in videos via natural language queries
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 67c25b43-a5c3-4a1b-9c40-aa1947da8c33 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2c9feb3-c9e9-4ad5-a7f0-ea21234d9dba · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Seed-bench: Bench- marking multimodal large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d88b740c-8e4e-49a1-b8cc-57b91c137d99 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos LLaVA-OneVision: Easy Visual Task Transfer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09172287-8ce2-4372-8ec6-512f66059795 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a302d60a-ea03-4908-82c1-477ba1b94e62 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos VideoChat: Chat-Centric Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd69e5f0-b7b5-4691-a043-157a5b50eeae · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1aed4ae9-c534-45fd-88bb-1c390699e2ef · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d7277bd-e588-4874-9fe8-6f783d5fbe5e · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Evaluating Object Hallucination in Large Vision-Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83bce1cc-4f62-40e8-9a35-40d97de11af2 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b20b214-48ec-4c71-8588-01e00e3ed1c6 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Llama-vid: An image is worth 2 tokens in large language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 72bd0409-4cbf-4573-a813-95a36e2cd732 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Detal: Open-vocabulary temporal action 10 localization with decoupled networks
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0b962e72-1c3c-46db-979c-c39df16e618c · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2037de3-473e-427e-b2fb-46634319ed74 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Vila: On pre-training for vi- sual language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c4152585-fa01-4d5c-a33a-10787db2a4a6 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Univtg: Towards unified video- language temporal grounding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ad53c99c-4772-4714-8c9a-63b4970cbbe5 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Visual instruction tuning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3b77870b-956d-45b3-9b29-202d1f92d7df · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f584a53-6ec4-4382-9eaf-7fdea08665ee · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos TempCompass: Do Video LLMs Really Understand Videos?
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86425931-5824-470a-b991-7e2b0337dc5a · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 945ebb97-a10c-4d9b-97af-c503e35b8d95 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26d8b05-3d70-4934-bfb3-f67df0002500 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4bfa45b-ed4e-4206-ac5e-1d5e9fd8cc3c · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 825b849e-405d-44f1-80cb-08e4ccba5722 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d118f865-05bb-447f-9d43-656536eacb13 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Docvqa: A dataset for vqa on document images
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f27365f7-de19-411f-968b-05a9ec37efcf · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Snag: Scalable and accurate video grounding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b58015fd-920f-49c7-a41d-c964c24d5baf · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos 4v (ision) system card
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation feb3f638-c942-464b-a23c-dace1f571fb9 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos GPT-4 Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4664d066-03e3-49fa-aaff-b713413da57b · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Scanning only once: An end-to-end framework for fast temporal grounding in long videos
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 13edb34d-4d65-4d0a-b6ae-f772c38069b5 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Per- ception test: A diagnostic benchmark for multimodal video models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5410bbec-2005-4be8-ba32-49007412a2e4 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca72c69-b34e-4a3b-a4f6-dcb8b7090296 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7bfd6d50-e9e7-4de4-b680-50eacabd53a1 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd02c596-ccb0-4908-a753-aa0254423ca1 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c22e7847-d52e-41c7-b0c7-009d22ecfc0f · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b5b8fe8-7797-4d7e-b20d-118d64623924 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos React: Temporal action detection with relational queries
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dbaddcac-98b0-468b-ae6b-844078d3d9a1 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Tridet: Temporal action detection with relative boundary modeling
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9d2c0c10-baa2-4edf-9738-cda68b6d61e0 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Towards vqa models that can read
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee085c5-9f9f-46ee-a889-3bcf5c3143d9 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 732da97b-b6d9-41b6-891a-c783e8bfae8f · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f9c87a1-a224-42c3-87d2-2b5769e407f0 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aebf4ec-6ebc-4667-8d08-baf1b7f33c44 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e48cf4af-68ce-4e96-8819-9634f4d81204 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Non-local neural networks
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63dfb193-06fd-487b-8e67-027fdcb92d8e · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ce7e4d-717d-4991-8488-4779a4090651 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Next-qa: Next phase of question-answering to explaining temporal actions
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45694ab0-fe91-4a2e-a2f1-7e89eb16a041 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Can i trust your answer? visually grounded video question answering
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 243ad04b-20a6-4c1e-b9fe-c5d21c0ec617 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Video question answer- ing via gradually refined attention over appearance and mo- tion
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e89a1e8-347b-4b8c-b5a7-6f886ffd551d · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47d7345f-2cd6-49c2-ae8a-22677989dc35 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 508953d0-0975-4cf6-9d4d-d9398d4d629c · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d32a76d-932e-4a78-aa7a-52758bd63697 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29606d90-03a6-458a-aeba-b39c76bbd570 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos xgen-mm (blip-3): A family of open large multimodal models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dafeab4-7543-4f4f-a368-9ac9526c3059 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Mm-vet: Evaluating large multimodal models for integrated capabilities
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f2dc37d0-bbf1-453a-a883-dff3de098cc0 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 195b3517-6d8c-43fa-a4d8-7b7f33a06d35 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73dcccf3-f323-4ea8-8967-1c371ef89db9 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Unimd: Towards unifying moment retrieval and temporal ac- tion detection
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 97f0f6dc-81ee-4c09-880a-e2bd5ca4651a · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c533a558-e491-4a95-aebe-462d1261e8fe · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Long Context Transfer from Language to Vision
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee166cf-f9c9-4556-9cd2-ee9f2bafac97 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90fdb764-d941-40a2-864d-b49ed48f944e · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Llava- next: A strong zero-shot video understanding model, 2024
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c835781-9d0b-4789-8b74-4ad4b54b8ca8 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos MLVU: Benchmarking Multi-task Long Video Understanding
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b35c9e04-9cfb-4cc8-b41a-83591684ea07 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c855a13-d7ce-447e-969e-9181ab7ba09d · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb804248-7373-4f6c-8270-6efe8fa73bf4 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Openai’s gpt-4o in surgical oncology: revolu- tionary advances in generative artificial intelligence
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cd6393ad-0425-4d5a-93f9-30e577ff0eff · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos 12 Autoshot: A short video dataset and state-of-the-art shot boundary detection
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e1e9fc91-9055-4050-b8b3-2c69f4484726 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos 64, 32, and 16)
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation df64a703-d495-4241-8493-f3a2eba7fc94 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos More validation results are presented in Tab
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4978add1-5682-48f8-ba9c-97a87048fd99 · outbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos Image patches of the selected tokens
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e2bdee07-8166-41f5-a9a7-78f04a2aaf8d · inbound
DisTime: Distribution-based Time Representation for Video Large Language Models LinVT: Empower Your Image-level Large Language Model to Understand Videos
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a28d731-d390-4e5c-9c8d-1465ae651bdd · inbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LinVT: Empower Your Image-level Large Language Model to Understand Videos
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47912e91-9c6a-4594-8839-40ba677a6da4 · inbound
${\mu}^2$Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation LinVT: Empower Your Image-level Large Language Model to Understand Videos
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e555be81-c022-47df-b305-a1faed9b8205 · inbound
AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multi-modal Information Between Reality and Videos LinVT: Empower Your Image-level Large Language Model to Understand Videos
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b54c22b-409b-4082-828b-5cac0d08ff71 · inbound
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding LinVT: Empower Your Image-level Large Language Model to Understand Videos
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 944fc3e0-f602-4b05-be63-7a75b5014222 · inbound
GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning LinVT: Empower Your Image-level Large Language Model to Understand Videos
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c3ac952d-456c-4709-bf4b-b1a34cc5cead · inbound
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction LinVT: Empower Your Image-level Large Language Model to Understand Videos
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.