Pith. sign in

Paper Citation Record · LEDGER

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs

As of 14 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2412.09907.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09907 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:39:51.010552Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e02ffcd6-58d6-4a15-b389-789cab693e83 · outbound

This paper cites Language models are few-shot learners.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Language models are few-shot learners

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.780473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.777896Z digest=sha256:404bcb4c9409e6f93cd49a607491f087bf5f3f49ed46c7a1190038aa1c99d9a0

Observation b8a991d5-6b0b-4def-930f-1a8017c5852d · outbound

This paper cites GPT-4 Technical Report.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.783088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.783088Z digest=sha256:2b179c23d6f8cb1079a02ca92c837c9fc35806db79bf22e1c7be216593e626f1

Observation 60d4aa57-f0f0-4aae-bcdc-fdc7841a407e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.787925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.787925Z digest=sha256:666726e70857b293df13129ab54bda202dd72d6761918f78d9e9bccca307f245

Observation ba488df6-c1d5-49e5-93d3-e8c434ff5840 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.792033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.792033Z digest=sha256:2082c98122ce91aa50807ea2ded095ae4e8719c90fa22c2c535dc97d6f0fbde7

Observation d914eadc-9670-4b8a-b592-a95e649f1f77 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.796435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.796435Z digest=sha256:a1b4d62223bdb899201c293c3f7de096eb4218bac227a144dfd17e477a462dd7

Observation 47baf3dd-106c-437c-ab2b-4f3267a0dd2a · outbound

This paper cites Large Language Models: A Survey.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Large Language Models: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.801165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.801165Z digest=sha256:99597b080088767b135608d8bf0cd93f16c0359bb930a22da981501bac983bb2

Observation 93db207e-1bd1-4efe-864d-574316170933 · outbound

This paper cites A Survey of Large Language Models.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs A Survey of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.805536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.805536Z digest=sha256:a1c367749788de97964c9fad27e5c6f3c7bfa6937ea5f45b8623f7d99f8c3168

Observation 198297c3-56fb-40a5-af1a-930268614b8c · outbound

This paper cites Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.809349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.809349Z digest=sha256:6beebd859b51986ded14c712095c3e9649e336f348e8d289a0499f78723ab378

Observation c7051ad6-9582-4943-a991-5f1b3f762708 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.812759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.812759Z digest=sha256:a8ce7e4d33ffd5a9ab3ff96392adf97d6b47249800d14032c158cd1a7cfa32d6

Observation ff95a4aa-43fb-4058-b81f-a1875abb4bd2 · outbound

This paper cites Large language models in finance: A survey.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Large language models in finance: A survey

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.761879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.816569Z digest=sha256:76b7d969b1908e3b145002c50e1ddfea9efeb545f14eba0049d8c6859dfad4c2

Observation d3bd051e-56f3-4363-ae45-38c748120ad9 · outbound

This paper cites ChatGPT for good? on opportunities and challenges of large language models for education.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs ChatGPT for good? on opportunities and challenges of large language models for education

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.751309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.820085Z digest=sha256:bfed322f534e560215d7f419ac99a901adb23f1906bd161f21821f4eabf4989c

Observation cfa1b187-2365-48ca-bbaa-c171a10ae1a8 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.823522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.823522Z digest=sha256:73e225de364b05e8e535be520d27ab8f63487782397195a29c46f62cc2944369

Observation 6405444b-f2c4-420c-adc6-54d42b3e6a9b · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.739802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.827195Z digest=sha256:895205a08efa73a2372ff46fd607b6661d787680aa6ab309a4294363a3e53acc

Observation 62ecb2e1-2ac1-4ecc-80ae-ab648109dc97 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.830665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.830665Z digest=sha256:ad23b7f88a70c25860093e65c1c1035d6859b37c5afebcf6c02dc77853febf2b

Observation b1fd1fcf-879b-4f4c-9811-38fcf5d07b5c · outbound

This paper cites Transformers in vision: A survey.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Transformers in vision: A survey

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.728417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.835225Z digest=sha256:b90b0f8a33573d337d912780c14b7a6d5645e621dc7d07dafde26da2fbe19681

Observation e4b91528-cec1-48c3-ba14-24d1623d5f83 · outbound

This paper cites Multimodal few-shot learning with frozen language models.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Multimodal few-shot learning with frozen language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.716906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.838764Z digest=sha256:1ce45cdbc323b6c79ba59267115cccde24806a6f9e4ddd8a87c4d27a7a69d73f

Observation 4f3b3a45-e132-4b39-a8db-1ada099ed34d · outbound

This paper cites Video understanding with large language models: A survey.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Video understanding with large language models: A survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.842227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.842227Z digest=sha256:ba2d441755bd1088092ca8d900202bcab533423d5ca5f0d09c4f8333823bb27a

Observation e6a75d1a-adf9-44d3-b2a1-a4922ef02ac5 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.845769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.845769Z digest=sha256:7e4ca1fa6e62521e0a8c600b42ce74104d779aedfb42cd5808ae800552564d50

Observation 6f0615bc-fe54-49ab-b707-6a890db5d7b0 · outbound

This paper cites A Survey on Multimodal Large Language Models.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs A Survey on Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.849625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.849625Z digest=sha256:2a820854fa4e1c3eeb210b4ec521957bd7b18fa7e1d817d0af22c58cd28a7f84

Observation 3348bc08-d6d7-4a01-98a1-80a6b32f561f · outbound

This paper cites Vision-language models for vision tasks: A survey.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Vision-language models for vision tasks: A survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.853080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.853080Z digest=sha256:7be569d25c5fdf614a0a7e9fab8f031dbfc118038513302b961b827599df2328

Observation 478e8dfe-10fd-48c0-9f8b-c1fa39e5f388 · outbound

This paper cites MovieChat: From dense token to sparse memory for long video understanding.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs MovieChat: From dense token to sparse memory for long video understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.699025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.856551Z digest=sha256:9b5e97dd0aaf6aa221034c1df0932aa3306204403169db6055c17a9e2a561a74

Observation 628d42a5-194f-4540-a0b6-571c0dd4267c · outbound

This paper cites MA-LMM: Memory-augmented large multimodal model for long-term video understanding.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs MA-LMM: Memory-augmented large multimodal model for long-term video understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.688677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.859735Z digest=sha256:45c292a2522741472a6cf7de0d05d86d3a0faf251ffdca3d4c62dd5fc44e8569

Observation 7ca2d4c2-f622-4168-b330-c8e425bf9916 · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.863339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.863339Z digest=sha256:f7790dd60eea0bac600c35f4b6d9503e89202f3055a8500b3d24f26c9b053dcd

Observation cb5cceb3-9cb7-4f50-919a-1636faa6c94d · outbound

This paper cites Evolving conceptions of memory storage, selective attention, and their mutual constraints within t he human information-processing system.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Evolving conceptions of memory storage, selective attention, and their mutual constraints within t he human information-processing system

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.678038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.867004Z digest=sha256:39666fb43998c40d52834add5138652bd1508cd5fb9bf50066ed7823b8a82025

Observation 7df3325e-06e8-4e3c-94f4-970f5c8fe60d · outbound

This paper cites Gorillas in o ur midst: Sustained inattentional blindness for dynamic even ts.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Gorillas in o ur midst: Sustained inattentional blindness for dynamic even ts

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.667871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.870229Z digest=sha256:94c6626dbcddedc6a5517eed67ad9ca0f36fc2ac3e0cd69a68b1a04116af27a6

Observation 986f31a7-4b68-4462-a0fe-8ea12a53f39c · outbound

This paper cites A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.873451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.873451Z digest=sha256:732d8f6b681a92d7ac28baadb8b91578f66826dca06db73a42b93ccf08c3e83e

Observation a6f7290c-b168-4ce8-90ef-4e4180549252 · outbound

This paper cites Towards Reasoning in Large Language Models: A Survey.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Towards Reasoning in Large Language Models: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.878552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.878552Z digest=sha256:f2cebf4ffdc3f202fadcadd78928ed106cd91778b2a565ad847fe0d888830ec9

Observation 0b80c3c7-1390-4ca3-a541-0d9e7a71cb79 · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs BLIP: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.657305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.881988Z digest=sha256:35fa8e40cc9e8012371fb72725dd32dd1a29cff75c9ccffc98ad88eef5843031

Observation 459e376f-6ea2-4216-ae39-b4cca2348859 · outbound

This paper cites Improved baselines with visual instruction tuning.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Improved baselines with visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.646969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.885452Z digest=sha256:50fb15ac025b7d4dda5da27bc618421d9aab7fcba7cd0a2789612035c253ac17

Observation cb17de36-5de2-44a5-a6a6-36a309a7d109 · outbound

This paper cites Visual instruction tuning.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Visual instruction tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.636455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.889051Z digest=sha256:397a0ffaab1a68c2da6addd9109a79e30b43291c6ec81121bfeea545dbc83946

Observation 0a196353-6970-48be-a561-6475df71fb37 · outbound

This paper cites Learn- ing transferable visual models from natural language super - vision.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Learn- ing transferable visual models from natural language super - vision

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.626114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.892415Z digest=sha256:b18d09f72230f94b7fefeb2e65bf90aa888e56ae6a4f10e96cb222a0bc0546c0

Observation a730e3ca-6ed9-4f89-ba7c-313898b53c0f · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Gonzalez, Ion Stoica, and Eric P

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.615563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.896011Z digest=sha256:a05b8b99592b0f13988a4d69c7a7c096b4db2d25527e6cdc8c30f14568459f23

Observation 4ef69fd8-c8cf-428a-886e-69c097dd6643 · outbound

This paper cites LLaV A-NeXT: Improved reasoning, OCR, and world knowledge, January 2024.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LLaV A-NeXT: Improved reasoning, OCR, and world knowledge, January 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.605002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.899721Z digest=sha256:1d016e6e3a7b0ab8b2486997ea61d677261bf970955049c00a6f47024c0200ca

Observation 26db7f9b-4f97-4657-96ae-b30f1f594096 · outbound

This paper cites Chat-UniVi: Unified visual representation em- powers large language models with image and video un- derstanding.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Chat-UniVi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.593509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.903238Z digest=sha256:f7c6ecda16622a53f8bd742786f40c2ee343973423cc7659ca9a48d8104517a0

Observation f88579e9-fa6d-4200-b263-d4b8af2eccfa · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.906942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.906942Z digest=sha256:5958d82a86469619520208b1c5dcca83e5cf027f3b30dfe02d7117213446fe14

Observation 2987bfbe-8160-4cc4-82e7-04dd86388976 · outbound

This paper cites LLaMA-VID: An image is worth 2 tokens in large language models.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LLaMA-VID: An image is worth 2 tokens in large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.582514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.910622Z digest=sha256:7b4120be232218680ee1b4b0e5bbcd6cc402271961fe0474f69b0d3f02c01fb2

Observation 036ff4c6-2a4c-4f33-843f-ad7a63e79d0d · outbound

This paper cites Video-ChatGPT: Towards detailed video understanding via large vision and language models.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Video-ChatGPT: Towards detailed video understanding via large vision and language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.571257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.914242Z digest=sha256:5d5da3182c943ea7122b39348489d098492217c9eadca10904662a0fdda2a0a9

Observation bff15451-a104-4096-8d85-99a4a3608245 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs VideoChat: Chat-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.918069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.918069Z digest=sha256:d5324d1e83756a9bf7b3b600cda05bf0b7b1abfc6488d9e0cce3f265e45dca41

Observation 9ed335c6-8428-477c-bc0e-df5493f02079 · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video u n- derstanding.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Video-LLaMA: An instruction-tuned audio-visual language model for video u n- derstanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.560231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.921787Z digest=sha256:5d9b4d0fd26ff78995fae79422ff374418bc7cc76316bbd0f9dcec6e3cd4a477

Observation be4ece27-1405-4655-a5b2-eb40b3fd5212 · outbound

This paper cites Long Context Transfer from Language to Vision.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Long Context Transfer from Language to Vision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.926246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.926246Z digest=sha256:5895e58818b0358fc389b198da8dce7fb6414080d6bf86431ac2bfc9016cff73

Observation e125a25d-9c3e-4c63-9e97-ce526d0e041f · outbound

This paper cites MM-VID: Advancing Video Understanding with GPT-4V(ision).

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.930220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.930220Z digest=sha256:02c5f050243a605468738c5ca36463b2d525d4bcda1767cde572c602eaa47a9a

Observation 11af3d05-e190-49d4-a95f-e93605ad857d · outbound

This paper cites Artemis: Towards referential understanding in com- plex videos.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Artemis: Towards referential understanding in com- plex videos

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.549077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.934205Z digest=sha256:85d2bb38c7e49e33dd4f00b1858d41c9cbca4f5033b958de8bacf50c4038083c

Observation bc2997e8-fe62-4226-ad6f-73f01a0bf2e9 · outbound

This paper cites Language Modeling Is Compression.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Language Modeling Is Compression

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.938041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.938041Z digest=sha256:846a3c9ca9ddd5de7478abb032258e56346daf4a5a9d22add6973eb47357bfd0

Observation d1e2a817-ff0d-4bcc-9ac4-056c70b034ba · outbound

This paper cites Learning to compress prompts with gist tokens.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Learning to compress prompts with gist tokens

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.537249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.941664Z digest=sha256:bcfaeeecd7fbba9c08565de5d869cd124ade8826bcda5c6eda833ed873c2842d

Observation 8b0d3a57-a938-4473-8bb8-46d60f090cc9 · outbound

This paper cites Adapting Language Models to Compress Contexts.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Adapting Language Models to Compress Contexts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.944973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.944973Z digest=sha256:0cb90cb0fa53176bd0d694531e50d7d0fe72fa4a4c372ff35302eb9f93b1b0b8

Observation c1b7ba47-f645-4d15-b8d6-c635ace6a03e · outbound

This paper cites In-context autoencoder for context compression in a large language model.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs In-context autoencoder for context compression in a large language model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.524994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.948606Z digest=sha256:9d33fe05e5cc607b61ba1346fca76e3eac267dbc02b256ca5834d729ebff159c

Observation 44a69e30-3231-4489-b71d-c1e357755cd6 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.952139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.952139Z digest=sha256:729fc5a7d99f06e41d9453899468c4fc5e156dc7763eaf99e90530f38e0bd5d4

Observation 36121cd2-b067-4a17-8c4e-862db36d924b · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ ar, and C.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ ar, and C

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.512853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.955683Z digest=sha256:03beaf27729c54ca453ae2001cc20bb98fdf8c975f1d4815ba414ab3cad9c691

Observation c8122340-37a4-493e-8b54-e86355b4bda8 · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.959092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.959092Z digest=sha256:77a1c3362ef7baa4c42ce26439de68b2f0fc3b9dfc83a8d5cfd82b5f4296671a

Observation ee9f54a9-fe60-4ab9-8322-2a6126eb3595 · outbound

This paper cites Towards VQA models that can read.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Towards VQA models that can read

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.501402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.962779Z digest=sha256:f6ef08be32ab8309952a1cf443e1ec7744dc55db263149cffb2b50a7e05a4dc6

Observation b91841aa-e94c-484e-b18d-581ad71db4a4 · outbound

This paper cites OCR-VQA: Visual question answer- ing by reading text in images.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs OCR-VQA: Visual question answer- ing by reading text in images

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.490420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.966933Z digest=sha256:3ab35feee038b6c61b3f0e1de83f0977766b6f60ba81a54fe2b4dbbfc4fd2f8f

Observation a44ff3c8-5ee0-4ec5-9ba8-30e9cc4b7686 · outbound

This paper cites Shamma, Michael S.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Shamma, Michael S

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.479193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.971037Z digest=sha256:b64ae692b28601725715cffef3042c7f5c565a2b4c1ae08019635df778ddb695

Observation 5df1571a-4fb6-460c-8846-abbf364b02a8 · outbound

This paper cites ActivityNet: A large-scale video benchmark for human activity understanding.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs ActivityNet: A large-scale video benchmark for human activity understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.467827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.974555Z digest=sha256:30d404d8da12ab9f50de6996e0a1a4d286e83d26ec39a71ad3f05a72e909d11c

Observation 22b0ec88-4294-43cb-830d-f43a972b8534 · outbound

This paper cites Decoupled weight de- cay regularization.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Decoupled weight de- cay regularization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.977785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.977785Z digest=sha256:3d65a2889986e1d7264c23e35e159dbe4533274556cd0fec278a02e5e3468012

Observation d500fb90-886e-4633-8b52-c1f4a61a9da6 · outbound

This paper cites In- finiBench: A comprehensive benchmark for large multi- modal models in very long video understanding.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs In- finiBench: A comprehensive benchmark for large multi- modal models in very long video understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.981107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.981107Z digest=sha256:beaecd25e2fa08cdd91dc6164d47f46f569bd3332c2fba273de56f46681678e7

Observation 830fac7b-03d5-4a5c-93ae-870ead615c7a · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.984462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.984462Z digest=sha256:ef3d0e8e4a19be2248f2604f1db84bc3b7685cccc6d199a4eaa0055b6aed38b1

Observation 1c017d31-1616-4b52-9b68-3eff82284fd9 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LVBench: An Extreme Long Video Understanding Benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.988258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.988258Z digest=sha256:141c6bfbe96ba8d8c7d492e4bcd70a5214f4867e582b3a12b5f8db2050ce1100

Observation 34ca34d3-ca2a-4c55-a4e3-523c1cb8ff76 · outbound

This paper cites NExT-QA: Next phase of question-answering to explaining temporal actions.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs NExT-QA: Next phase of question-answering to explaining temporal actions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.449957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.991774Z digest=sha256:a24839fe8080fcf520fec3dc4eeaf7d575acf53a29618b44700b35af84ca6152

Observation daf2b5a1-113b-455b-aa4c-32eaa5c58fdf · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.437540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:50.995521Z digest=sha256:e4be810c50b1d7448e76f053793263751be4f591a52f8baa34aeecab8da413bb

Observation 6ccf13c5-6b7b-4808-8755-2956f5e49417 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:50.998994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:50.998994Z digest=sha256:7fef0a7544f52f4bc32ccfabeb41bd03028b891e31d581a12217a3f7c97d52e4

Observation 18562d97-8e0a-4fd6-ba9e-ff9b5af20b43 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:51.002585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:51.002585Z digest=sha256:f1e7768ca91ebdc60ed6f10c5568c8662e1ea8b2b4c4605a427cdfda2b6352d1

Observation fb3bfd6e-af12-462e-8e2e-3dc2546552d4 · outbound

This paper cites Sheldon,.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Sheldon,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.425862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:51.006668Z digest=sha256:b6f29154cc6929f6a17d8c364582fe318b148d91bdfe2853b2c8fc188f3ad0f2

Observation 45e7e380-d35e-46e5-b4aa-40a29260f057 · outbound

This paper cites Invisible Gorilla.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs Invisible Gorilla

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:39:51.414569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:39:51.010552Z digest=sha256:febb501de6231c138b0f8ecd5105428036cd87ead017736e7275faf3deeed073

Pith citing papers

No inbound Pith citation observations are available.