Pith. sign in

Paper Citation Record · LEDGER

Long Context Transfer from Language to Vision

As of 11 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 100 inbound Pith citation observations for arXiv:2406.16852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.16852 v2

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T07:08:35.946669Z

measured 194 of 194 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 173 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:57:40.232042Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

94 of 94 outbound references displayed

  • verified exact6
  • verified fuzzy84
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c366485e-fc67-4d54-95ac-b43e3ba6b31c · outbound

This paper cites Llm testneedleinahaystack.

Long Context Transfer from Language to Vision Llm testneedleinahaystack

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.691536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:c9dfb03f4f8d0e1d9dcbe0ee44786c577c51559c46139f0b0e9f5609761a4b61

Observation 50743afc-919a-4f25-b93d-a8717ed4bf0c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Long Context Transfer from Language to Vision Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.699662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:887b0ede034d98c7d5db045030f868fb194aa477c115047d4aac5371d6bfc49a

Observation 659fe65a-0d19-44bc-80f1-ee50f65e057d · outbound

This paper cites Vqa: Visual question answering.

Long Context Transfer from Language to Vision Vqa: Visual question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.703685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:2bffd52161eaf8ce7295e6467b475f105cc153b0b7995f5c32d713d08e572e08

Observation 4623201c-56f0-44f2-86ae-8bb725a7a67c · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Long Context Transfer from Language to Vision OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:52:01.553388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:26ba7e6c4b1c83dc7c2fbbcd710c0c59b92d20423f3ba6fadab5a0f33db40da0

Observation f2501790-375e-412d-8855-0935e1a217f7 · outbound

This paper cites Longalign: A recipe for long context alignment of large language models.

Long Context Transfer from Language to Vision Longalign: A recipe for long context alignment of large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.707886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:6ff00efada24253aaea90361058907eb2c2f74208efd031b7635430b06dce7ba

Observation ecac27f0-511d-42f0-8085-7ce36732e94f · outbound

This paper cites ntkaware scaled rope allows llama models to have.

Long Context Transfer from Language to Vision ntkaware scaled rope allows llama models to have

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.712440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:41b74d5000ef1580a4d5a25de2dbb45a8cdcdd84f6f71139006cabd7805e0be1

Observation 6d758d12-f987-4289-bbec-c07c07d1faff · outbound

This paper cites Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation.

Long Context Transfer from Language to Vision Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.716803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:a9f70f020ed5320cb225b4e88c7978e12d3a8189d970c40cb75353d3b182be4d

Observation 905e55d7-98cd-4c23-9dc4-6e7238239d85 · outbound

This paper cites Language models are few-shot learners.

Long Context Transfer from Language to Vision Language models are few-shot learners

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.721167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:dacba6d11406e7019f6797e0bd9729d7b0e98c5e1c99f622836ab1616bcfa6b0

Observation 5ca33b95-7cad-422d-a50d-64fa184583d7 · outbound

This paper cites Matryoshka multimodal models.

Long Context Transfer from Language to Vision Matryoshka multimodal models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.725016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:bba05f4529e1e8720a391f551b679eba1844655c2c958117f573afa52822d46a

Observation bf1bfdc5-d2bc-4d63-a79f-b38aa19b3a16 · outbound

This paper cites cerebras slimpajama-627b.

Long Context Transfer from Language to Vision cerebras slimpajama-627b

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.180077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:9f0a0f32f1546d5d27759fdcbb187073e40d035d7c557d64fa305a5d36187029

Observation a489bee3-2c34-48ed-acc4-12e190345998 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Long Context Transfer from Language to Vision An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.186534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:eb9a4f2095785f60502ff4607f1dfadb35f946c99b046a95030d80650475f894

Observation 220cbefe-a31b-48ed-af1c-e2c36a307709 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Long Context Transfer from Language to Vision Sharegpt4video: Improving video understanding and generation with better captions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.194617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:912ed0ebc4be7a30a8944966601917b370ac34d0320db2e3e4ab8866e2c9d050

Observation ed4463ef-97d7-4ee2-ad65-4b7cea21d963 · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

Long Context Transfer from Language to Vision Extending Context Window of Large Language Models via Positional Interpolation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:20:58.016961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:33dbc000948881da376d72ecb482eb2f69253b020ce711c121f0d015c9113f30

Observation c84878b0-bf24-47dc-b82a-86533c7295bf · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Long Context Transfer from Language to Vision InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.650334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:48fa17def79533b0917ee31ea6bad2b8aea7bcb1322157f03bab5004ea81e19c

Observation 93a0653b-d59c-4bce-a1fe-4d7fbac8707f · outbound

This paper cites Videollama 2: Advancing spatial- temporal modeling and audio understanding in video-llms.

Long Context Transfer from Language to Vision Videollama 2: Advancing spatial- temporal modeling and audio understanding in video-llms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.203855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:bbe9b5e9d8e97150cd04ec4a17f3d1372290a92b24f3631f70f6d7725cd8d026

Observation c7426fe1-5d31-4fce-b4e6-9a468340e314 · outbound

This paper cites Generating long sequences with sparse transformers.

Long Context Transfer from Language to Vision Generating long sequences with sparse transformers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.210061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:8c1851aa641e0dd8a2e14195f14ac29a9e6ba426b30f2a379854232430ec4109

Observation 56417ca4-2c47-41c7-9a90-51ecd373eaab · outbound

This paper cites Introducing command r+: A scalable llm built for business.

Long Context Transfer from Language to Vision Introducing command r+: A scalable llm built for business

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.224542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:b21d0a030fb5eff0efb760dbc5d7a7ee2432581c0ee329032c9906d5450bcf8a

Observation 9d2e526e-1e09-453c-9c98-eebfe1bd5f4a · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Long Context Transfer from Language to Vision Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.233354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:edadb76b6cf3dc74e5d9072ada516d768d1450ef5256f327fac2a77b31588f0b

Observation ee4847fd-1b8b-4220-a6be-7b7a231fbc9b · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Long Context Transfer from Language to Vision Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.241659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:46057b2600a81281aaea17d5585804e4753f79b0ebdf2fbb0843aba44aea7c2a

Observation 4b2960b7-7bc7-4481-9f00-ce19b2ce278e · outbound

This paper cites LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens.

Long Context Transfer from Language to Vision LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:31:04.851790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:e43127366794dda24ac05958419d455bbb1bd695a7fbc6be25b7ca614714fb5c

Observation bb0033e1-04b0-4d9c-a744-711677d53bb6 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Long Context Transfer from Language to Vision Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.250356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:7e1bc169a5cb314512e10a628c4706f46669de579963523aeada323abf8b08d7

Observation ccb39f39-a40e-484c-9d15-d677930e0011 · outbound

This paper cites Data engineering for scaling language models to 128k context.

Long Context Transfer from Language to Vision Data engineering for scaling language models to 128k context

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.260561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:d9aa337144428609df4fc72d6d12504b4ee5ec46fc5aa7351827762bf89e83e9

Observation 797d774f-cc90-4cb2-8c8c-b08ef8b7f0e1 · outbound

This paper cites Llmtest needleinahaystack.

Long Context Transfer from Language to Vision Llmtest needleinahaystack

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.280355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:694480c31dac4ac2df8a0571da8f33f2896f439b60260783ccae54a9a003c876

Observation 08ee7fe6-3659-4de2-8c9a-c843b8d993a9 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

Long Context Transfer from Language to Vision Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.286095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:07337f01728493cb468760ce061928a4890d341d61f2017066b19ee239d4daae

Observation b59f8ffe-9c0e-4f6f-8456-14319553b247 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

Long Context Transfer from Language to Vision Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.302349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:05d585752033216f0ddf9e425a8226792a7aea5102ae4494f218f47b27a01a0a

Observation c59119fb-e4b1-46b1-bd9f-e2a18e02848a · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Long Context Transfer from Language to Vision Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.307938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:98d04762ebc278a81e8ff5d03dbe5692df7c07e87c361e6cc1a7dcf71b00ba64

Observation 875ff106-2979-44c8-b1aa-406718bb2297 · outbound

This paper cites Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models.

Long Context Transfer from Language to Vision Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.317510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:97385053eccb44e9e18110ea828f8e414e12308d9f3b0ee096ad060af9fa74c5

Observation 431282a5-8059-404a-9b69-e2a2f43a0b22 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

Long Context Transfer from Language to Vision Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.323053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:2cc27830390f1723698d0bb630cd55d43252cee66fb66ab355a73f5170100917

Observation 8ccc644b-33b7-4d06-8459-36c890c5d0b3 · outbound

This paper cites an unresolved cited work.

Long Context Transfer from Language to Vision Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-12T07:08:36.332674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:4a72d8511088571ef9cd000f02102a1a8374194925f00576b16f742b73b4cbe7

Observation 5871f125-0798-467b-bab6-ae427a9a2ca4 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

Long Context Transfer from Language to Vision Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.163331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:cd4c15998998bc497ea31e13c536ffa195c7bab2b1f45f5e1aa30df6d2f2d2f5

Observation 0877590c-f5dd-4a73-9700-8486f7cc4447 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

Long Context Transfer from Language to Vision Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.344358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:08f6cb125797c7c57bc9faac954ac955ce32597e31f52df2473ef0e423dd4efb

Observation aa696503-b6b5-40f9-b092-e2abb5276edb · outbound

This paper cites A diagram is worth a dozen images.

Long Context Transfer from Language to Vision A diagram is worth a dozen images

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.352361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:7826e902a1c51757cfdfa96cacf0ea84e7b901ca8be971b08cee5a3e6b5873de

Observation ad2ea75d-fdb9-4749-8943-b4db2459357b · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Long Context Transfer from Language to Vision Gonzalez, Hao Zhang, and Ion Stoica

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.357791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:1068476afc24a9c8d6bfd48162b076832ed5797fa33e309767fea6d7cb279663

Observation 6f1e1d1a-ad11-4aa1-9ad8-d9e3780d8638 · outbound

This paper cites Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh.

Long Context Transfer from Language to Vision Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.364272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:b0f9cb7cb45391f1b99e2ee6a17c7971021dc2b24a90077577e4770ba6dfb11d

Observation 4148427a-4e00-4bf7-9dc7-4495c96d2938 · outbound

This paper cites Llava-next: What else influences visual instruction tuning beyond data?, May 2024.

Long Context Transfer from Language to Vision Llava-next: What else influences visual instruction tuning beyond data?, May 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.368998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:da345d723289523be91447d95edae2194e6bc920c98480e725d95d3726b16f3f

Observation 9d162a63-2d51-4649-84b6-4574efb6e34d · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024.

Long Context Transfer from Language to Vision Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.373769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:b2663253878e895e3413f6d8c38f25a39512565c1bfc2a4fa2377b64da4c620d

Observation 31f68640-8092-47cc-919c-748e02322fe8 · outbound

This paper cites Otter: A multi-modal model with in-context instruction tuning.

Long Context Transfer from Language to Vision Otter: A multi-modal model with in-context instruction tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.378469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:63bc37eb1fd7c4e16f6b28d306e9986ae3f5c71bc5f651bb2296bc06e6f083ec

Observation 9ccd396b-2634-4d0b-92b0-0547650d9813 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Long Context Transfer from Language to Vision Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.391851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:e79f059f5270ea370299478aa576857775c1a3fa3ab58eefc2b2d54f3d148482

Observation 3e0bd944-f4ae-4368-b5de-b7ab0728c3cc · outbound

This paper cites Videochat: Chat-centric video understanding.

Long Context Transfer from Language to Vision Videochat: Chat-centric video understanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.396864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:d6862260eabb8b3e73d1097afa4a521b2ebfc9f8bd44ce2756156d07145506ce

Observation 1068bac3-20bc-41c1-90e9-06c95207df4d · outbound

This paper cites Sequence paral- lelism: Long sequence training from system perspective.

Long Context Transfer from Language to Vision Sequence paral- lelism: Long sequence training from system perspective

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.401539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:aef192b35fc63ccc39197b6044a77ae0fba152f5421ac9a9d6cf5ff99889562f

Observation 6187fefe-a1dd-4be2-9db1-77b4da39c36f · outbound

This paper cites an unresolved cited work.

Long Context Transfer from Language to Vision Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-12T07:08:36.406051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:db011e9962415619bd7bad5374c2a202a4af46ad9de344cfe443d2636c9a954d

Observation dd08ca2c-2755-405e-bd68-230d44525070 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Long Context Transfer from Language to Vision Llama-vid: An image is worth 2 tokens in large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.411011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:ac68612eefb256793dfa734d27dd687c5c41ef79dfa9099faf8f184f41343560

Observation 05ff61ff-ab17-4d8c-b015-49d4738e6899 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

Long Context Transfer from Language to Vision Video-llava: Learning united visual representation by alignment before projection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.415948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:c42f672aaf70615d93bea26c5e755c8a96a2bcd6c9c91d4564d603d86c4cdf6c

Observation 5313ed3a-bd0c-45ad-ae20-d32052be9d23 · outbound

This paper cites Vila: On pre-training for visual language models.

Long Context Transfer from Language to Vision Vila: On pre-training for visual language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.421099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:5c714c7ea8287e0c2e18eb1ad8bfaf854dfdebfcc956b0fb7a018bedab51fcde

Observation bcf40407-7b27-431e-a3a0-41c56003b6cc · outbound

This paper cites World model on million-length video and language with ringattention.

Long Context Transfer from Language to Vision World model on million-length video and language with ringattention

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.426166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:fe36fa7c2cac5d911a6666c67f266072721214fdadb11f771612684b92ead641

Observation 6ba7ff8e-32a8-44f6-a7a6-327a4b115342 · outbound

This paper cites Ring attention with blockwise transformers for near-infinite context.

Long Context Transfer from Language to Vision Ring attention with blockwise transformers for near-infinite context

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.434632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:45e62be48619ca900bf9a5427b24f5f1a3bd6afdd0d3b9c16d9f8305ad8cf5e6

Observation b1937ddf-b590-47ab-a549-e1ec801731e6 · outbound

This paper cites Improved baselines with visual instruction tuning.

Long Context Transfer from Language to Vision Improved baselines with visual instruction tuning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.440592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:ca69ac9704b848df3485f2dcd1b82c20a4099a434e9162a4f503ff1b4f822607

Observation 3cd987c5-616c-420a-8aed-1b523a984114 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Long Context Transfer from Language to Vision Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.445356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:a8426c1c7fcfb24cfba4ceaeb76a68802b332f7a2e203cde4fef98872761778d

Observation eacd7ac8-72f0-4f01-b779-7ca5b4c29200 · outbound

This paper cites Visual instruction tuning.

Long Context Transfer from Language to Vision Visual instruction tuning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.450601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:248a36dbdb7e13ef4a47e0cac4cc1eac8370411ebcdaacc16eab23aefdaf9afa

Observation edd700d0-bac1-4967-8998-027020572c2c · outbound

This paper cites St-llm: Large language models are effective temporal learners.

Long Context Transfer from Language to Vision St-llm: Large language models are effective temporal learners

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.455090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:51695b2d1611506de0733857267bfe063be288ccc8bda3a738ae221c156480ca

Observation d4c7d364-6ab5-4c3d-a24f-00d27f75c165 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Long Context Transfer from Language to Vision Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.459819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:d17e99beb33c8b5eb4deb6eee438b6964d7ba9743c2c3f540597860418bc2af9

Observation a079384e-8f2a-4e1e-95ed-6bebe30694ea · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Long Context Transfer from Language to Vision Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.468784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:121e0cb2e304d7f9e42888fe8650278be7598dfeb4641710b882000fbc22e26d

Observation 94a7c3d4-7e31-4ac8-a5cc-9bf6284d3177 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Long Context Transfer from Language to Vision Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.473576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:cc4f8ab18b57764bf48a9266a11c1d52b844524dd7b36e266e0530714b2d1b7d

Observation 9fc75294-2741-4978-a25d-8c0ab778eea9 · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

Long Context Transfer from Language to Vision DocVQA: A Dataset for VQA on Document Images

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:08:36.126572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:aa0975ec9cbb0e7e088eded5e13bf00d35ec85c33dc765770720a9284dc54368

Observation a4947d69-546e-4380-bd39-e5541fb74f0a · outbound

This paper cites Mixtral 8x22b: Cheaper, better, faster, stronger.

Long Context Transfer from Language to Vision Mixtral 8x22b: Cheaper, better, faster, stronger

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.480410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:df85a6415df093dde073b891e8b55da2445e7c10f2e21d3b0a1485edf2f0f275

Observation c1bc88f2-a0fe-4a4b-b17b-341b0e9ff8bf · outbound

This paper cites Hello gpt-4o.

Long Context Transfer from Language to Vision Hello gpt-4o

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.484594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:20968d9def1c38a830f0bffddd7fc19b0c19aaf5fec1905390cd0597137aad5b

Observation f429b08d-65b1-41f1-9c39-98f38a7d72e6 · outbound

This paper cites Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models.

Long Context Transfer from Language to Vision Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:08:36.171513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:6d457d139438e69ad044b086c304f61ce926347394ee2d1f9ec47d9575ff71b1

Observation 3a0800b8-10ef-4189-b47e-8c78998a67ec · outbound

This paper cites Yarn: Efficient context window extension of large language models.

Long Context Transfer from Language to Vision Yarn: Efficient context window extension of large language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.493601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:70843493179452cbd9382486e7766693581c89e37a8bb9c3412425d19d516ca8

Observation 362b40f3-e9d5-4e64-b3a3-0416a3148c13 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Long Context Transfer from Language to Vision Learning transferable visual models from natural language supervision

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.499135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:467657cb7f805d4d2f2727bfbca85d2a24628350c9ce2895bc4ef980d67342c8

Observation 6c4ee3c3-bd57-4ac5-a831-05d94d02d326 · outbound

This paper cites Zero: Memory optimiza- tions toward training trillion parameter models.

Long Context Transfer from Language to Vision Zero: Memory optimiza- tions toward training trillion parameter models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.504551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:adf387131269f73e2054bfa13b9f4b5e8c64daa479529802e31b7a67f9f0d68a

Observation 22c56be9-862e-450a-9f82-e8c2896aeafa · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Long Context Transfer from Language to Vision Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.512176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:98b368d6e6cbb9cfa996958ab4df66f6c51d020bc2498f1d119af9c404ff01ed

Observation eca5eac4-d395-4a51-b8da-559c783345c0 · outbound

This paper cites Code llama: Open foundation models for code.

Long Context Transfer from Language to Vision Code llama: Open foundation models for code

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.518359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:d09bccde7d533262382a40e54042b6acc5b75e62f31381fef6d103e4587b5252

Observation 0a8a4f38-3623-4c03-885a-5fb996eda2b5 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Long Context Transfer from Language to Vision Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.522924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:68a9b43037fe588967b95aa8b6c8f9a58d23ca8bf4ddf3a4be3b03a77a595327

Observation 1674f7d2-ad0d-4d10-b2e2-2f60d6fddc22 · outbound

This paper cites Milebench: Benchmarking mllms in long context.

Long Context Transfer from Language to Vision Milebench: Benchmarking mllms in long context

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.527502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:6c78ee5eca533ef1d7939f4ab0b86561bd7e5692ff79700bc6de3da623aa988b

Observation 583f157c-35fc-4d17-bc05-3f8668164a16 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Long Context Transfer from Language to Vision Moviechat: From dense token to sparse memory for long video understanding

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.536612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:9b010bdf9b24cfe8677ad2723bee0da3729d52212ac30e292a09e9d0f8cd92bf

Observation b7848974-2c5c-4136-a03c-02b33deed058 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Long Context Transfer from Language to Vision Roformer: Enhanced transformer with rotary position embedding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.542118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:44b4306aefb88dcd00f1a4e5e3659cd2c288b8b7caebae7bdaa27f329dffe9ad

Observation c029ae3e-6fdb-44fc-bb6b-fc9c8dcd80a8 · outbound

This paper cites Gemini: A family of highly capable multimodal models.

Long Context Transfer from Language to Vision Gemini: A family of highly capable multimodal models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.546290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:74015c4e0929a8104fd773b6cd79cbac29fe5dbc3a48f73ee418d5ccf6cdab9b

Observation bac81060-eebc-4633-b315-0f6d9ccc293b · outbound

This paper cites Palm 2 technical report.

Long Context Transfer from Language to Vision Palm 2 technical report

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.550347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:f89b5fd99dd946145574cf79573ad10804e388db8b58f7cc9042f0cca86c15cc

Observation 890e38de-801d-4718-939d-7bebca8c8dfd · outbound

This paper cites Introducing qwen-vl.

Long Context Transfer from Language to Vision Introducing qwen-vl

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.554914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:a22f50121e565d401b752709f4bdfa01633838f1142766d4b8da33ca19edf32c

Observation 1d8b253b-215e-4eff-ad1b-67c920b7621e · outbound

This paper cites Qwen2 technical report.

Long Context Transfer from Language to Vision Qwen2 technical report

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.563452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:5275e4dea0bd2512fdd9208f8da121841c2ccc0988dbb50956d9e02d0c6954ed

Observation cd07319c-223f-42b7-bcd9-21871d48c9dd · outbound

This paper cites Llama: Open and efficient foundation language models.

Long Context Transfer from Language to Vision Llama: Open and efficient foundation language models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.568777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:31e9ef7211807fadb2f24908ee3952a1960fa720b73162d24bb2461d519f1309

Observation 60c8d302-c1c5-4a5f-90a3-d4a61b5f0a8b · outbound

This paper cites Multimodal needle in a haystack: Benchmarking long- context capability of multimodal large language models.

Long Context Transfer from Language to Vision Multimodal needle in a haystack: Benchmarking long- context capability of multimodal large language models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.573521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:72f4b912d04a1176157949dfa8a12eb3754f03939b3bcf7610fb7e822388eae2

Observation c89d548b-72b7-40db-a0ec-f0fee689e4f6 · outbound

This paper cites Lvbench: An extreme long video understanding benchmark.

Long Context Transfer from Language to Vision Lvbench: An extreme long video understanding benchmark

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.577663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:bd3958462a6302f4b213a88f2b31e15384b78628017026c65c7ea029a6bc99e9

Observation 7de1d54a-32a6-4c68-beba-338643ade4f5 · outbound

This paper cites Needle in a multimodal haystack.

Long Context Transfer from Language to Vision Needle in a multimodal haystack

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.583061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:c82640bed37b2012dc32f24013bf2735b359dda50dde84c4f73caebcc5a10c1d

Observation 815cf966-3643-4aa2-b0c2-96a36dd569ef · outbound

This paper cites Star: A benchmark for situated reasoning in real-world videos.

Long Context Transfer from Language to Vision Star: A benchmark for situated reasoning in real-world videos

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.595155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:19738f423058a8929addd31b6bd9f37d4ac5b1bb2eba7c0736791f7876278682

Observation d27cf5b0-d787-4d81-901f-26b4cdd0c1fc · outbound

This paper cites Grok-1.5 vision preview, apr 2024.

Long Context Transfer from Language to Vision Grok-1.5 vision preview, apr 2024

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.600394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:0a58a18dcaee379587414842e8dfebb5aa02338488cd2d3b2aad8f2856c7b8db

Observation c9bf17de-a013-41a3-881a-48e082babbd4 · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

Long Context Transfer from Language to Vision Next-qa: Next phase of question- answering to explaining temporal actions

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.606025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:0f8b99a41078be0d9b3db14fc2b3b94530949ff38a64b7539d62cddb9939b5c6

Observation 50eeabdb-7b33-4489-92a2-6b9a02b69970 · outbound

This paper cites Next-qa:next phase of question- answering to explaining temporal actions.

Long Context Transfer from Language to Vision Next-qa:next phase of question- answering to explaining temporal actions

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.613458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:831043f1ecdc47135415bb8df0fbfa81caab04bec4927276fe5c95f5b81d9808

Observation 46ec1724-4ac3-48eb-8de5-ec81857ad3aa · outbound

This paper cites Funqa: Towards surprising video comprehension.

Long Context Transfer from Language to Vision Funqa: Towards surprising video comprehension

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.621512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:e9673bad05879446b5597551b058ab151b9addc5f5571974e883d1d552eedeb8

Observation d3f810b7-8414-48ae-893f-9cbc8e4c28e1 · outbound

This paper cites Effective long-context scaling of foundation models.

Long Context Transfer from Language to Vision Effective long-context scaling of foundation models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.626200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:c744e2629149518c07fda2fad7a27097f2aae00b5d15eaf95251f0d265565c0a

Observation c2a5815c-bbda-47d8-a9c0-54cd7c814fc0 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

Long Context Transfer from Language to Vision Video question answering via gradually refined attention over appearance and motion

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.631082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:31878d0c7ad4647c914fc5918e9ae54b2f0fcede1aa32d5a29590c7562211985

Observation c44cc93b-e53a-4149-8024-8f5617e657aa · outbound

This paper cites Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events.

Long Context Transfer from Language to Vision Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.635163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:944386271ad821b875a75d9d0a15a605625a196ad2eb5f9eb246076a8fbff02f

Observation 2f87c9ed-ce08-4778-816f-5ac8eca9957e · outbound

This paper cites mplug-owl: Modularization empowers large language models with multimodality.

Long Context Transfer from Language to Vision mplug-owl: Modularization empowers large language models with multimodality

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.639613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:9f6f2e06707e6e8c300ce8cee75560c3b41d17c801686e4ce5625fa921606548

Observation 1217a881-608e-4e7d-b897-e1f6cfcf5548 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Long Context Transfer from Language to Vision CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:01:50.618167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:b2fee02d24d837306494de2d3acc5bee83b4b86ef42fd53cdccbe60935e48faa

Observation 44614c79-4a71-45db-923e-2a8c9200fa7e · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Long Context Transfer from Language to Vision Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.644209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:916eb5b7eff94831ae538e155d037cacc3db0fbcacf8ef7563a0ede982c82b26

Observation d2a1a64a-f7d7-4198-8486-7ecd377754d9 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Long Context Transfer from Language to Vision Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.651693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:d2e41b417d23ec9aa85f036d29978bf891c72a8f426174a8fcd0841971eb18de

Observation 1010d5e6-5683-4497-b6c1-516232eac650 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Long Context Transfer from Language to Vision Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.659926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:777d1c4eb56d8dd40aeb9a91141890d5f0f6811aed14462207822649c40fb347

Observation 16d9716a-4914-4a9c-9aa2-e5951806d5ff · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

Long Context Transfer from Language to Vision Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.664644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:f8f6dcde5510edeb32c4eb768d693593994cad9bdbaac745ecabc65a550e6bc9

Observation 6d7b7e90-727c-4797-8520-f6de88b6d6bb · outbound

This paper cites Direct preference optimization of video large multimodal models from language model reward.

Long Context Transfer from Language to Vision Direct preference optimization of video large multimodal models from language model reward

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.669125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:403296bca5a4c4c91976f99a768b655fce0f2d765f3fa7e2cce61bcb22ca752d

Observation f209fbee-f0be-4c60-b5a1-79fbef8c552d · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024.

Long Context Transfer from Language to Vision Llava-next: A strong zero-shot video understanding model, April 2024

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.673760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:064b16866999c7153cc75d2b9a98a5ff60245d29db261faadcecaf0a3fb89dfc

Observation 67836fc9-a424-498c-b75b-ecf15dfc5b77 · outbound

This paper cites Mlvu: A comprehensive benchmark for multi-task long video understanding.

Long Context Transfer from Language to Vision Mlvu: A comprehensive benchmark for multi-task long video understanding

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.678185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:779a60beefa9fe1d9356192586321e1cfa88ef514f10587f447997ed1b5184a2

Observation 01160dc2-3cda-4ca2-8568-cf65190e437e · outbound

This paper cites Streaming dense video captioning.

Long Context Transfer from Language to Vision Streaming dense video captioning

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.682616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:a9b2f916d5d691a62b91beae916651ba9c4e271aebd524647787227245181c02

Observation bd86627d-912b-4fed-94ca-7b72dffa1ecf · outbound

This paper cites Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models.

Long Context Transfer from Language to Vision Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.687674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:d6bd59e263a0de6e2c0560bacbc418bfd8b62815f4412261579761492f5bb0e8

Observation 4ff52b4b-5133-43f6-8835-03c7cf80f707 · outbound

This paper cites Ring flash attention.

Long Context Transfer from Language to Vision Ring flash attention

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:08:36.695375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:08:35.946669Z digest=sha256:3face2444aaecf3d5e7dcc63d102b8a7c76a266f5fecec762a1ede0dc4532cb4

Pith citing papers

Observation 1251f619-4ad9-4d84-8eba-de96e809bfcb · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding Long Context Transfer from Language to Vision

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:55:26.496537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:07bea3ccfa1069ffae521da6403b434b967b2645bf4a54c373b48ed560ef34b0

Observation 3929cb71-7217-45bb-9bbc-5cb5361efda5 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Long Context Transfer from Language to Vision

Reference 175

Resolution
verified exact
local_arxiv, observed 2026-05-17T10:46:28.873209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:618b4b529a70495e25e0bae9be7cb6283a8cbf57b81fda7635697d06e21b4d30

Observation 2a30ea10-0ddd-4fa5-8371-e8240ea9ff57 · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer Long Context Transfer from Language to Vision

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:84116ebcc63b49586a5591336c4616ac184f519b53d995956bc732bd7035f36e

Observation 1194b3af-a1bb-45c0-a5a6-b9f2c4766120 · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Long Context Transfer from Language to Vision

Reference 156

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:db0e85bb4e1cb3fb15b8038c13b5dfd5938644bbdac1079957eccdddc9723512

Observation 69c4ae92-980c-4438-bcce-fa92c87cbe02 · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding Long Context Transfer from Language to Vision

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:53:33.635949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:5961cff495612982dd1822bacf36384923404a1c2723b5804c63503617a72ee2

Observation e9fb2502-0eb6-4e15-a7ff-d8cb773cf645 · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces Long Context Transfer from Language to Vision

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:27:44.064590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:f5e27b1fee38cf496f6d0f7c2d6f0423004a6fe1c9960f9f1d4e273ad78045b6

Observation e272bae0-47ed-44e8-9638-54b01db80967 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Long Context Transfer from Language to Vision

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.616047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:692d27e7bc51f8d026649f049074288d62e9a58053d7d69f67c099b8e6e07a5f

Observation 1d0859dd-9a8e-4160-865c-1f35042ff9c3 · inbound

Online Video Understanding: OVBench and VideoChat-Online cites this paper.

Online Video Understanding: OVBench and VideoChat-Online Long Context Transfer from Language to Vision

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:40.232042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:40.232042Z digest=sha256:971814735031f17c8af41ad2f47f48138fcf849f706e10ca2b9f3e86621126dd

Observation fa6abe33-b726-44dc-a3f6-008cf740de43 · inbound

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM cites this paper.

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM Long Context Transfer from Language to Vision

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:30.347198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:30.347198Z digest=sha256:69a5e18d0db98ea338770bfac433b8fbe60fdffcf1d74c9b155121f043554260

Observation c0790486-e797-4a9b-a1d8-d0c314216d93 · inbound

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding cites this paper.

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding Long Context Transfer from Language to Vision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:56.460436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:56.460436Z digest=sha256:0c6388dfb785b5593f324eb4c11839c56f6c015744c262a4500c1dded9c560a3

Observation 85cd279f-e32d-49ea-bb59-543be1c4b444 · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Long Context Transfer from Language to Vision

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:08:19.641181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:790781ef7cdf484706787228699f47b3a744232bcbc4df9e90349957bdaaea54

Observation 5e4720b0-756f-4074-a3ec-614e11c3be19 · inbound

LongViTU: Instruction Tuning for Long-Form Video Understanding cites this paper.

LongViTU: Instruction Tuning for Long-Form Video Understanding Long Context Transfer from Language to Vision

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T21:23:58.059208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:23:58.059208Z digest=sha256:162c19ff47a5e8111b4cada643f0690b8e81000415274742c6344ac36afa3abd

Observation cdfd847a-ef76-48f1-a66e-f25636e27b09 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding Long Context Transfer from Language to Vision

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:02:37.639944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:8bb8fec451d107014292796038c3dfc3ab1b2c1f6352fec4158eea2a5ec8bb1b

Observation c62dc1ae-7bce-4b19-8ef8-a581782e4001 · inbound

NExtLong: Toward Effective Long-Context Training without Long Documents cites this paper.

NExtLong: Toward Effective Long-Context Training without Long Documents Long Context Transfer from Language to Vision

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.916097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.916097Z digest=sha256:1fc49a3f1e549462aee85999bb1f9a3473215d1ae53f3f7c7f4b8c142189725d

Observation 987e8f26-1a31-4758-85dc-5e9458ed81d4 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Long Context Transfer from Language to Vision

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:133919c6184061dd06bb85074b555d5175e24cbe4ed10fe43f9acd26715184db

Observation f5040c62-020c-4353-bab4-32171b018be4 · inbound

Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge cites this paper.

Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge Long Context Transfer from Language to Vision

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T15:58:24.473539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:58:24.473539Z digest=sha256:baa9c09fde9c5ee3ecb3dd7aa7af4adba63153e1fc51dffa354206c3e4eb7f74

Observation c28d9e8e-ea83-486e-a407-c8508b62acbc · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Long Context Transfer from Language to Vision

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.190008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:40a3c61c2bf7b80b9687566bad50d6810d18e213c9b21c6e74bd8496c1a8edca

Observation cfb261d3-c08a-4034-8835-35aad62cee14 · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding Long Context Transfer from Language to Vision

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.305277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.305277Z digest=sha256:a13554ee3083686860327f3d7d13c4e30f6dec4bc62878af26ba30470a68b40a

Observation ccba3e18-c063-432f-a6f1-204bd7e412ce · inbound

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation cites this paper.

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation Long Context Transfer from Language to Vision

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T21:21:45.820352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:21:45.820352Z digest=sha256:4735b43a887d04554f1da46039999c674ff0daef52006b91e76ec31960e80e3e

Observation 078991f5-1663-4b24-ac01-f52c2f54a09a · inbound

VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos cites this paper.

VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos Long Context Transfer from Language to Vision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:36.998608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:36.998608Z digest=sha256:785be8e7a3e807c177a47aebc94dd5485b93ab409c3d2be584c513c6efe8d753

Observation e79ee3c6-d068-4c32-9b59-2ac9a81cf4ef · inbound

HD-EPIC: A Highly-Detailed Egocentric Video Dataset cites this paper.

HD-EPIC: A Highly-Detailed Egocentric Video Dataset Long Context Transfer from Language to Vision

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T23:26:20.592106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:26:20.592106Z digest=sha256:81da208f1191b26e29d18d941e340d6ed84c4f2b6687b35b923ffffe854cfeef

Observation d38a661a-07d0-4337-94fe-949196180adb · inbound

CoS: Chain-of-Shot Prompting for Long Video Understanding cites this paper.

CoS: Chain-of-Shot Prompting for Long Video Understanding Long Context Transfer from Language to Vision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.994358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.994358Z digest=sha256:dbd92d67011708b9b0567e6aacca73f9b58f3b4c9f1527997be80137f9196aeb

Observation 5facfe1c-fbc9-40fe-938b-35110551efa7 · inbound

LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation cites this paper.

LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation Long Context Transfer from Language to Vision

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T13:02:23.640079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:02:23.640079Z digest=sha256:2e4262657d5765a35e69ec458b2328b3ffae359cd6f74edda09ed7d1d9d02082

Observation e99c838d-a56e-425e-8168-ddef02a2589a · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Long Context Transfer from Language to Vision

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:43:00.307194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:95b8bbf8aba5aa8d1c33981a425dcb4a863cba148f5e56762d68d18ffe973ebb

Observation ada76126-f8e5-4621-8c0b-fd110515a6dd · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Long Context Transfer from Language to Vision

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:ac466627bab026bebd2b501ecac66d6fc73a90221087defb984cf5b2147fc0c5

Observation fbb17e75-ffea-475a-947c-10ab364bd87b · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Long Context Transfer from Language to Vision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.769437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.769437Z digest=sha256:57cfae8ad7a9949a58faf874ffd5ec14b6d2a11e31eea9e81519ad3fc72626c4

Observation 686d382c-f69d-4325-9780-7c48d5248f4d · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Long Context Transfer from Language to Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.552615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.552615Z digest=sha256:db22f33783ff247c6c9d458edd53f097cd2a4592947255173ff9fe576c2047ff

Observation 896f522a-1fd3-4c1f-9435-e1462eb735e9 · inbound

Clapper: Compact Learning and Video Representation in VLMs cites this paper.

Clapper: Compact Learning and Video Representation in VLMs Long Context Transfer from Language to Vision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.560323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.560323Z digest=sha256:d7e3962af686eb70871e379b2d1bf51c1003b0cde95a9da292127a420ffb0ae2

Observation c45523ae-bfe6-4356-b6b4-3f7ce6e92d7d · inbound

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM cites this paper.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Long Context Transfer from Language to Vision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.258703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.258703Z digest=sha256:0ebce9283074536df277777e16d90cad4e50a888757bee7676288dad165bd39c

Observation 094c0bcc-ac62-41ab-9ccc-b719cc5beb56 · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought Long Context Transfer from Language to Vision

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:15.746320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:15.746320Z digest=sha256:7309bd699e8ded63ca7843328c7e296b4ea80decdf19bcea290146a06130ce82

Observation 2108de4d-6b3d-48a7-8140-aff954899ef2 · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Long Context Transfer from Language to Vision

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:05.277632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:05.277632Z digest=sha256:20d369f1188c58703916d39b6fc29aba5c76b2c0d381ca17a64c3a00e632141b

Observation af4007e5-3f71-456d-80e7-b3c04993e6b8 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Long Context Transfer from Language to Vision

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:34:36.960131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:451e1d823ec1ccd84d298e05957a6b9db53e6eafe47ea5d7ea8669cfb699f546

Observation df459411-a9ef-4025-ba51-db9a32e0e78d · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Long Context Transfer from Language to Vision

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:00:51.367097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:d9976af00f6c1b8f31b96902e3d3bd1f6cea5d3c95e12834e14b701c7a94f51f

Observation d0f16f33-24b5-4af9-ad4e-4b8e07355c6d · inbound

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders cites this paper.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Long Context Transfer from Language to Vision

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:24.093163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.093163Z digest=sha256:ff89157454276407d60c772b6693961c101a22e512f4180a86d6b41154d75c79

Observation 708677d5-b6a1-4e7b-88c6-5af249e4da12 · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking Long Context Transfer from Language to Vision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:13.330350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:13.330350Z digest=sha256:85fcefb4ccd1681fea78518324b1c318315b5dc7e9d40e08cd2bffd76788f5f9

Observation 0efa8bca-4d4f-4abb-a51f-dcad74ffb272 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Long Context Transfer from Language to Vision

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:16.215765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:16.215765Z digest=sha256:4960fcd686ee433e8a3ddf661c5ebc34f3451650460cf395000ce5b0a8e06b0d

Observation 31892af6-6997-4b2c-9b72-0e9cc0709858 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Long Context Transfer from Language to Vision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.606703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.606703Z digest=sha256:fa97cbf1a3708f94ff20260ca82f2a4cd49235477d2d170bf72262c6d525368d

Observation 8ac67125-e712-4a7e-ac56-d9575d4c0163 · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding Long Context Transfer from Language to Vision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:29.901221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:29.901221Z digest=sha256:aeb05ad4df558c161cec9f7e9fc1d4549baad00655b7c486b12231a77381b3e9

Observation cc55fd6e-4219-447b-b325-f6fe6f62535c · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Long Context Transfer from Language to Vision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.916580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.916580Z digest=sha256:04a7e050c4db2db9133ed70b7f4bc281507203272564d344a28b4c265077d6a6

Observation b0a8dfb3-6cc0-40ee-a393-04b366ffd805 · inbound

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence cites this paper.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Long Context Transfer from Language to Vision

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:43.138721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:43.138721Z digest=sha256:36230982869e506fff8e94000cc0980c623f7e91b66255979ff8fabe430d56a9

Observation 7090f9e5-e826-48d2-a982-117d9ec7574c · inbound

Vid-SME: Membership Inference Attacks against Large Video Understanding Models cites this paper.

Vid-SME: Membership Inference Attacks against Large Video Understanding Models Long Context Transfer from Language to Vision

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:06.285354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:06.285354Z digest=sha256:c98e7579c8284efc6021c88b7fb780ff2742baa1f7f2e3f8a004702d78f6843d

Observation b10c4db0-dc68-4031-b0c2-9d98e407ecf4 · inbound

TextVidBench: A Benchmark for Long Video Scene Text Understanding cites this paper.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.836228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.836228Z digest=sha256:104cba96acd019456d7075dd6fc0e0892df5426403cc7f1b40ef2aa61ae29e93

Observation d56b795a-64cc-4291-86fd-07789b11d2f0 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Long Context Transfer from Language to Vision

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:51.563326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:51.563326Z digest=sha256:b9a367312018a4d22e85e8a7149914c593d24191ae1ee756067a12f913a87810

Observation f2a45ff1-3772-427d-a662-2d50614b02a3 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? Long Context Transfer from Language to Vision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:09.433526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:09.433526Z digest=sha256:ae9dfc8202d710872d914a8aea2567e56dd8bf0db5e798ad95c4c38f52a8494f

Observation d35a5bff-efa6-4b40-aa4c-bba752662cc4 · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos Long Context Transfer from Language to Vision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.343725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.343725Z digest=sha256:0d5c1a72064e2012c747a91f3b13e4398bcb3479ebf89bd8ba8445ffcd919aa3

Observation 372c0c6c-ddb6-46d8-8879-8a2903415c3f · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Long Context Transfer from Language to Vision

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.408552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.408552Z digest=sha256:b8935b1b76f047d03b655e2b73733ec2c7d1481ec95f494280bc8ce1cc055c0c

Observation cf0663c6-07a0-424d-a475-16f65e101b6b · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding Long Context Transfer from Language to Vision

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.299651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.299651Z digest=sha256:cc26634777f6d6ef7ea7c67cc2fe0a865b300d66f0ec58a1765ceec3badfc52d

Observation 48c77cfe-290f-4270-adb6-2b1adcf3f0d6 · inbound

Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? cites this paper.

Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? Long Context Transfer from Language to Vision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:36:57.778449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:36:57.778449Z digest=sha256:a9c3b9a2b86dc45972f8bcda651f6f519ad5f2ad1c02c3e7b79566facf410ade

Observation 780e01ce-6412-459e-9e7b-5403e03752ce · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Long Context Transfer from Language to Vision

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:56.105590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:56.105590Z digest=sha256:68f5309e8864424f398c1aa085f97f85ca6804786917a7ff6060fb6e107b6861

Observation 26b7f326-3691-4767-a429-eddfeaa91aef · inbound

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs cites this paper.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Long Context Transfer from Language to Vision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:33.079691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:33.079691Z digest=sha256:7a22bdc91ea8d973e50bfb52199c9bdea09e8215f8b516d6514ddc583b7c6e14

Observation d807b7f7-e6c9-475d-93ee-a9cd1cf256d3 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Long Context Transfer from Language to Vision

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:38.476955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:38.476955Z digest=sha256:646d70d9f4c5a634fe12c677789794f97594e1d235466df9833ce581ac2f84a4

Observation 81ae1300-a353-4646-9fc7-5186e7de2a7a · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Long Context Transfer from Language to Vision

Reference 143

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:51:15.954923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:6602292848dd3b12dedef53cdad3b3e40b645bf2f4db055d5b4fb24cde5d287d

Observation 0b8f8ce4-9400-4786-81bd-7c484317d5cc · inbound

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning cites this paper.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.814956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.814956Z digest=sha256:72eca2edb8b9c06f7bf3f7918090597bf66713bfa6a5b9ea7cea2d953937dce0

Observation 35bc73c8-054a-4132-919a-d3a9d7e7a3e6 · inbound

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation cites this paper.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Long Context Transfer from Language to Vision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.971332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.971332Z digest=sha256:697adcdeef6f0899e3b22196c0789717ce713f72170ec941dc0fa1146325c14c

Observation 38fdb9fe-d313-4570-8f2b-565d97b42853 · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Long Context Transfer from Language to Vision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.042420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.042420Z digest=sha256:3d848e974c78f9218197d8c7af124e14a8c9e02f947d9968abb227d30ef41e21

Observation 122a666a-3a86-4f3c-a925-119f5f0a1358 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.426259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.426259Z digest=sha256:b8a050cbe0470a07200d1703a1e5aec0324a90e913c93acc5a7721c90b5c3992

Observation ef8de3d9-0280-4974-bb2f-47853121def6 · inbound

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs cites this paper.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Long Context Transfer from Language to Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.749344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.749344Z digest=sha256:73eb103e44d8bef0ebbafaec34a30a839674d24bba771ecec655c0032c29130d

Observation ac77033f-768b-4197-b7c0-04d0923444c2 · inbound

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning cites this paper.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.903154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.903154Z digest=sha256:9e104273f1117fc43380ee668d14823b1bda2f6e882ffc3d0d2cfb27f7edb94b

Observation fafdc095-d006-47ba-8474-889537b0e4ff · inbound

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams cites this paper.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Long Context Transfer from Language to Vision

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.962828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.962828Z digest=sha256:a6e2bec3144f4a7761d6b44d80b43cf38d4bb1522e98da820835e5409050a7da

Observation 079be2da-dd1c-462c-b917-a51a2c231eee · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Long Context Transfer from Language to Vision

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:56.985190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:56.985190Z digest=sha256:0294cf48c2f3c0bd2262a636dda5d83f1d49f9b6d4b0efae7a5030f58d40c604

Observation 602199ed-14c7-4fbf-9815-939352dc8fa4 · inbound

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs cites this paper.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Long Context Transfer from Language to Vision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.185274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.185274Z digest=sha256:ddee5d22538b6de87a96065d9f9d0fedcc82e9e22612cf1a5c20bbef1031a55a

Observation 6118feca-54de-4a13-896f-92b60e129dd8 · inbound

RadEyeVideo: Enhancing general-domain Large Vision Language Model for chest X-ray analysis with video representations of eye gaze cites this paper.

RadEyeVideo: Enhancing general-domain Large Vision Language Model for chest X-ray analysis with video representations of eye gaze Long Context Transfer from Language to Vision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:25.859396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:25.859396Z digest=sha256:bef35882777e97cc8c4b32f67ad23d33affa1174cab7b73ed9be077b27675ae0

Observation 009ac0f9-9498-40c2-b9cd-b9a93ba48836 · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models Long Context Transfer from Language to Vision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.047880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.047880Z digest=sha256:d8aff36d4415d1bd11f620921e964bc5d3b219f279822f6cf0dba435ff2d6297

Observation fe22ba31-6ef0-422e-a89e-4b5188af112e · inbound

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding cites this paper.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Long Context Transfer from Language to Vision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:37.124321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:37.124321Z digest=sha256:f466e3ab15dd4d75cdcee76c8a7fe7fd167cabf05ad7a425191b2ecd0adf8e2c

Observation 3b5e3be6-f9fb-455b-8214-d751b065597f · inbound

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics cites this paper.

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics Long Context Transfer from Language to Vision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:32.507337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:05:32.507337Z digest=sha256:a1406732c2e3b295166915766aed1840ebb706bc4f475bed6ec93efd00a09632

Observation 3db7dd76-d71f-41af-b576-3b93ebe9c2ee · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Long Context Transfer from Language to Vision

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:35.312845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:35.312845Z digest=sha256:b6338d3371b41686116bec96014c043280a217ba1b1f780034d57e3825d7e782

Observation defef782-e812-4bf7-8dcb-9fda3c5ddc59 · inbound

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding cites this paper.

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding Long Context Transfer from Language to Vision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:16.823023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:10:16.823023Z digest=sha256:2aaac804310ada2bbc2f092d763c8fbe88fea21d69b152ac4945a0c9a028050c

Observation 217ed845-1e01-4a96-a6ec-079d028f42d4 · inbound

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment cites this paper.

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment Long Context Transfer from Language to Vision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T13:54:52.799037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:54:52.799037Z digest=sha256:5c8cb73352f034bf35627e30f784dd498707dc6784bd356fecf92ad72da9dd52

Observation a75667df-189b-40d5-a2e7-ab369304ba27 · inbound

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding cites this paper.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Long Context Transfer from Language to Vision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.150112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.150112Z digest=sha256:ccaf6e90a20fa98df13e6c87b4617edcd65cfe0d11b7bc7a0df902c0037bfcc7

Observation 7648395d-473f-4b5d-a0fc-713faf085d21 · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training Long Context Transfer from Language to Vision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:09.768750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:09.768750Z digest=sha256:82e6e285e98cef41fb11badb85d5e56a47187f8e27869e8a4889a670fa585d66

Observation 738943dd-7f6f-4f19-83b6-12b008fa9194 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Long Context Transfer from Language to Vision

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.764864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.764864Z digest=sha256:14df041237fc6fb2982df40a61dd25e3d088ef5243c87f56a22b64ff6304aac6

Observation d4504deb-bbe8-4d43-bf54-8eddaaf3a9c8 · inbound

Cambrian-S: Towards Spatial Supersensing in Video cites this paper.

Cambrian-S: Towards Spatial Supersensing in Video Long Context Transfer from Language to Vision

Reference 160

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:46:04.694927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T03:46:04.363500Z digest=sha256:31434ed2c672f491d47880c6f23d2ae99af412da6762a974535a7892a7ffa236

Observation b6d85b97-3627-4be6-b44e-c209e842c7f1 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning Long Context Transfer from Language to Vision

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:25:22.580751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:fa8375fd718856c981542fa971b0531f34a82af98b5e0d2faa579ff28d81b71b

Observation 5d0bb3f1-e85d-4ecf-bb20-9a521d582944 · inbound

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding cites this paper.

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding Long Context Transfer from Language to Vision

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:20:22.866300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T22:19:36.366837Z digest=sha256:dbd7237b6222b8d5bac646a76f9c6f42975a2a1832c4acb29cb8e8e8876ca303

Observation 7b972113-b370-45c8-9b32-5690607e6a70 · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning Long Context Transfer from Language to Vision

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:36.972775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:36.972775Z digest=sha256:c20d76ca81d2f695780d6c113e0e91710642e99d6cf81103448e4675c9c0b58d

Observation 82e609f1-d849-4451-89ca-dde64e8fbb9e · inbound

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding cites this paper.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Long Context Transfer from Language to Vision

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.646227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.646227Z digest=sha256:29b4c67593163c667325a2ab822d5921dea0274ca303ebe82c4790bfeebba388

Observation f964a623-93b3-404c-b769-131bde129cf3 · inbound

Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval cites this paper.

Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval Long Context Transfer from Language to Vision

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:11:23.285697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:10:01.596422Z digest=sha256:2d5bf63e73e03618487d0785cdbe869b86bf26a4002edfa18a325e2ada6332ce

Observation ed169029-a8de-464c-8137-bea9c7cca2a2 · inbound

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility cites this paper.

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility Long Context Transfer from Language to Vision

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:43:20.648500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T19:42:25.626448Z digest=sha256:7e8cc9b5b1e569d2daa0f1cf5d48365776ed1ba8ab96c1bcbad2ca2050c88bc7

Observation ff464fde-2c93-413f-85b0-02642e352f9f · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Long Context Transfer from Language to Vision

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:57:53.860496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:34b98e45ab23bd5c551140ca8a4425b9396d1291d61a81b8dc8cf1e5aaa92e78

Observation 334625a9-0400-42c3-8049-4db0e5b16ca5 · inbound

MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources cites this paper.

MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources Long Context Transfer from Language to Vision

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-03T06:50:29.849129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:50:29.849129Z digest=sha256:741424cd88181ec5c511ec482c99c63e35fc1172e465fd5728ddc52823e221f4

Observation 4780408a-6c60-4b1f-b636-5a1de8a8f82d · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets Long Context Transfer from Language to Vision

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:28.912117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:28.912117Z digest=sha256:d42fc9b2fc295dedecd142fc469f6b27c6603e773e67282d6b118d05f758618c

Observation 1862eb9f-6e6d-4eac-9e5b-43713013f906 · inbound

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents cites this paper.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Long Context Transfer from Language to Vision

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.040656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:5aef3746a6d17527c927b53212030337ba72a3335ae3be313609c8322eb7fca6

Observation 79630530-51c6-4ae5-b6c6-df80086870f8 · inbound

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation cites this paper.

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation Long Context Transfer from Language to Vision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T00:15:23.806078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T00:15:23.806078Z digest=sha256:a6ed766e268e1b53d88fa7d41565311a79c97676fec66fbe624e1716d2401aec

Observation e7d8a661-922b-44cd-81cc-fdea29c5d0af · inbound

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding cites this paper.

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding Long Context Transfer from Language to Vision

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T17:54:09.705716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:54:09.705716Z digest=sha256:1560d30deca5ed6a174a89cc6ffc4fe1a71cb92d1370e9e630d3f8fcfd843993

Observation 7b78802e-815f-412f-89dc-f2e33e3ba61a · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark Long Context Transfer from Language to Vision

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:08:04.465877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:a97a6fec723d581cecfa20891c5fab5f5b2f5dd4c24d5f176878d50359071a76

Observation bf500279-c1a0-4dad-b808-c81cae10a01d · inbound

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering cites this paper.

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering Long Context Transfer from Language to Vision

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T21:13:16.342523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T21:12:29.596207Z digest=sha256:9b06404e974502f0361b8e40c9bd87cb72929ac36fd3f21a7002bdc430d13836

Observation 4fb48766-c116-4ff5-ba90-e5f6da5eae34 · inbound

Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs cites this paper.

Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs Long Context Transfer from Language to Vision

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:32:59.381755Z digest=sha256:bf7cfe140311e87dbc7445e62402bc139db6684ea464409f82ffbb2bd72ec441

Observation fe8f2b3f-4c31-4421-8f8e-7b6b1f9de0ee · inbound

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models cites this paper.

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models Long Context Transfer from Language to Vision

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:20:19.653806Z digest=sha256:b55eced72b57586ba450d466dfe7218ccb8265fd92ae7fb5cda186271fc4ed0b

Observation c7f6e111-2efb-435c-b1fa-f3126e2d2c39 · inbound

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding cites this paper.

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding Long Context Transfer from Language to Vision

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:21:47.439019Z digest=sha256:37b22ef153bb36800f8731d6e625b5f3088887b4aa8f34d5960e82fa1f5a0e70

Observation 0d36e27f-b624-4232-8a0a-6456a6de7a6a · inbound

Small Vision-Language Models are Smart Compressors for Long Video Understanding cites this paper.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Long Context Transfer from Language to Vision

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:8b6f8a33d5bcb16ad88dc3ce2c99669e6073b73c279bacd63ced627add206ec4

Observation a2d7224b-1660-4b29-a8b2-0640f6d3c36e · inbound

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging cites this paper.

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging Long Context Transfer from Language to Vision

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:42:43.948462Z digest=sha256:6ca2e7908399080a5423ca4d3ac3880889ab2e39507c28e51ea10a23d72d1c0b

Observation dde20986-fddd-4353-96d0-49af6583fab0 · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Long Context Transfer from Language to Vision

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:0eef01ac82d87ff2462caca9e7b7179e1d0ccbc0f47534309be819612b20ab98

Observation edf47762-9789-499e-95b4-89092e306270 · inbound

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing cites this paper.

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing Long Context Transfer from Language to Vision

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:32:02.222528Z digest=sha256:aff73404c83dfc0881e5b9f386aae338b1786e6002cac3162f85d7384f5f8f1e

Observation 0d825d63-2330-46e5-beb2-60a149bad419 · inbound

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning cites this paper.

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Long Context Transfer from Language to Vision

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T06:51:52.861981Z digest=sha256:57eea6a8570a4700ed1bacc7fbc462b0959b843da536a9363dbf839482a054ca

Observation ed264919-1890-4210-a810-194c5ecf7f45 · inbound

EgoSelf: From Memory to Personalized Egocentric Assistant cites this paper.

EgoSelf: From Memory to Personalized Egocentric Assistant Long Context Transfer from Language to Vision

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:23:21.119521Z digest=sha256:a41e90bc9dcf2ea769ec008d60c63d2ce98289964163f23d0cb331d5503477b2

Observation 2ef0ece3-f894-4748-8c2a-b7124869c768 · inbound

Video-ToC: Video Tree-of-Cue Reasoning cites this paper.

Video-ToC: Video Tree-of-Cue Reasoning Long Context Transfer from Language to Vision

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T01:20:29.374012Z digest=sha256:c908a66d341441c692eac1f93f5f95c32b36107e2790a2a6387d0eeb07788a34

Observation fa674b76-523d-4ce2-96ac-602a24c065ff · inbound

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration cites this paper.

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration Long Context Transfer from Language to Vision

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T21:55:35.699057Z digest=sha256:87ad737378032d42f53b5826f7262af834cf322229c80487b69edf0e3db15dc4

Observation dfb2920d-67fc-491a-b61a-b28eb6e5d18b · inbound

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding cites this paper.

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding Long Context Transfer from Language to Vision

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T12:26:01.568507Z digest=sha256:7e1007f14d28a762ae78955074d0c6822807de39631e705166e06e60edd59ec3

Observation 5b998405-1ec4-4dd2-b88e-ad2fa3c56b75 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Long Context Transfer from Language to Vision

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T14:10:27.416341Z digest=sha256:69193dcfa1e69f21f4392de8b2c8b4fc99e7455d08a76d890472cf9cf8591515

Observation 441a83b3-bcfa-43fe-9b7c-3c71bebbe125 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Long Context Transfer from Language to Vision

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:45:34.772580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:ce928d91d533cd4f865823a84d8f3de6f69f7066a9763d31945a93cc56ac2cbe