Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:39:51.010552Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2412.09907.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:39:51.010552Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e02ffcd6-58d6-4a15-b389-789cab693e83 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Language models are few-shot learners
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b8a991d5-6b0b-4def-930f-1a8017c5852d · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60d4aa57-f0f0-4aae-bcdc-fdc7841a407e · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LLaMA: Open and Efficient Foundation Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba488df6-c1d5-49e5-93d3-e8c434ff5840 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d914eadc-9670-4b8a-b592-a95e649f1f77 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47baf3dd-106c-437c-ab2b-4f3267a0dd2a · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Large Language Models: A Survey
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93db207e-1bd1-4efe-864d-574316170933 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs A Survey of Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 198297c3-56fb-40a5-af1a-930268614b8c · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7051ad6-9582-4943-a991-5f1b3f762708 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Retrieval-Augmented Generation for Large Language Models: A Survey
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff95a4aa-43fb-4058-b81f-a1875abb4bd2 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Large language models in finance: A survey
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d3bd051e-56f3-4363-ae45-38c748120ad9 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs ChatGPT for good? on opportunities and challenges of large language models for education
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cfa1b187-2365-48ca-bbaa-c171a10ae1a8 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6405444b-f2c4-420c-adc6-54d42b3e6a9b · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 62ecb2e1-2ac1-4ecc-80ae-ab648109dc97 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1fd1fcf-879b-4f4c-9811-38fcf5d07b5c · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Transformers in vision: A survey
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e4b91528-cec1-48c3-ba14-24d1623d5f83 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Multimodal few-shot learning with frozen language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4f3b3a45-e132-4b39-a8db-1ada099ed34d · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Video understanding with large language models: A survey
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a75d1a-adf9-44d3-b2a1-a4922ef02ac5 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f0615bc-fe54-49ab-b707-6a890db5d7b0 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs A Survey on Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3348bc08-d6d7-4a01-98a1-80a6b32f561f · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Vision-language models for vision tasks: A survey
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 478e8dfe-10fd-48c0-9f8b-c1fa39e5f388 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs MovieChat: From dense token to sparse memory for long video understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 628d42a5-194f-4540-a0b6-571c0dd4267c · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs MA-LMM: Memory-augmented large multimodal model for long-term video understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7ca2d4c2-f622-4168-b330-c8e425bf9916 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb5cceb3-9cb7-4f50-919a-1636faa6c94d · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Evolving conceptions of memory storage, selective attention, and their mutual constraints within t he human information-processing system
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7df3325e-06e8-4e3c-94f4-970f5c8fe60d · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Gorillas in o ur midst: Sustained inattentional blindness for dynamic even ts
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 986f31a7-4b68-4462-a0fe-8ea12a53f39c · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6f7290c-b168-4ce8-90ef-4e4180549252 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Towards Reasoning in Large Language Models: A Survey
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b80c3c7-1390-4ca3-a541-0d9e7a71cb79 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs BLIP: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 459e376f-6ea2-4216-ae39-b4cca2348859 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Improved baselines with visual instruction tuning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cb17de36-5de2-44a5-a6a6-36a309a7d109 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Visual instruction tuning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0a196353-6970-48be-a561-6475df71fb37 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Learn- ing transferable visual models from natural language super - vision
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a730e3ca-6ed9-4f89-ba7c-313898b53c0f · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Gonzalez, Ion Stoica, and Eric P
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4ef69fd8-c8cf-428a-886e-69c097dd6643 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LLaV A-NeXT: Improved reasoning, OCR, and world knowledge, January 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 26db7f9b-4f97-4657-96ae-b30f1f594096 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Chat-UniVi: Unified visual representation em- powers large language models with image and video un- derstanding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f88579e9-fa6d-4200-b263-d4b8af2eccfa · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2987bfbe-8160-4cc4-82e7-04dd86388976 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LLaMA-VID: An image is worth 2 tokens in large language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 036ff4c6-2a4c-4f33-843f-ad7a63e79d0d · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Video-ChatGPT: Towards detailed video understanding via large vision and language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bff15451-a104-4096-8d85-99a4a3608245 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs VideoChat: Chat-Centric Video Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ed335c6-8428-477c-bc0e-df5493f02079 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Video-LLaMA: An instruction-tuned audio-visual language model for video u n- derstanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation be4ece27-1405-4655-a5b2-eb40b3fd5212 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Long Context Transfer from Language to Vision
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e125a25d-9c3e-4c63-9e97-ce526d0e041f · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs MM-VID: Advancing Video Understanding with GPT-4V(ision)
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11af3d05-e190-49d4-a95f-e93605ad857d · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Artemis: Towards referential understanding in com- plex videos
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bc2997e8-fe62-4226-ad6f-73f01a0bf2e9 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Language Modeling Is Compression
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1e2a817-ff0d-4bcc-9ac4-056c70b034ba · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Learning to compress prompts with gist tokens
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8b0d3a57-a938-4473-8bb8-46d60f090cc9 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Adapting Language Models to Compress Contexts
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b7ba47-f645-4d15-b8d6-c635ace6a03e · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs In-context autoencoder for context compression in a large language model
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 44a69e30-3231-4489-b71d-c1e357755cd6 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LoRA: Low-Rank Adaptation of Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36121cd2-b067-4a17-8c4e-862db36d924b · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ ar, and C
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c8122340-37a4-493e-8b54-e86355b4bda8 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee9f54a9-fe60-4ab9-8322-2a6126eb3595 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Towards VQA models that can read
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b91841aa-e94c-484e-b18d-581ad71db4a4 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs OCR-VQA: Visual question answer- ing by reading text in images
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a44ff3c8-5ee0-4ec5-9ba8-30e9cc4b7686 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Shamma, Michael S
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5df1571a-4fb6-460c-8846-abbf364b02a8 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs ActivityNet: A large-scale video benchmark for human activity understanding
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 22b0ec88-4294-43cb-830d-f43a972b8534 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Decoupled weight de- cay regularization
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d500fb90-886e-4633-8b52-c1f4a61a9da6 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs In- finiBench: A comprehensive benchmark for large multi- modal models in very long video understanding
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 830fac7b-03d5-4a5c-93ae-870ead615c7a · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs MLVU: Benchmarking Multi-task Long Video Understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c017d31-1616-4b52-9b68-3eff82284fd9 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LVBench: An Extreme Long Video Understanding Benchmark
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34ca34d3-ca2a-4c55-a4e3-523c1cb8ff76 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs NExT-QA: Next phase of question-answering to explaining temporal actions
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation daf2b5a1-113b-455b-aa4c-32eaa5c58fdf · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Video question answer- ing via gradually refined attention over appearance and mo- tion
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6ccf13c5-6b7b-4808-8755-2956f5e49417 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18562d97-8e0a-4fd6-ba9e-ff9b5af20b43 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb3bfd6e-af12-462e-8e2e-3dc2546552d4 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Sheldon,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 45e7e380-d35e-46e5-b4aa-40a29260f057 · outbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Invisible Gorilla
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
No inbound Pith citation observations are available.