Pith. sign in

Paper Citation Record · LEDGER

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

As of 4 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2605.07897.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.07897 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T02:12:20.651111Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T03:21:47.718375Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact25
  • verified fuzzy24
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ef130f4-3c4d-4368-a1d6-c2729db0baaf · outbound

This paper cites Qwen2.5-VL Technical Report.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.728343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:4786786db8b22a07fa8dba579b6374c8fb7f2ded89e2903ac70e6350106e6e2d

Observation aca26757-b3c8-4ee5-85ca-89377c2e702f · outbound

This paper cites Token merging: Your ViT but faster.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Token merging: Your ViT but faster

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.296353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:672e3fee39356f4d37894e982b9c0ac727f65833f2b1e9de0e743c31893d9708

Observation 11bb9116-a31f-4db5-9986-2dea5053d1e1 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Videollm-online: Online video large language model for streaming video

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.282447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:92cfd97619f6df8f8df92d7e74646262ebc266e4231a50dd7dba4769b12dc91a

Observation e39ad353-654d-49ad-8495-0f4fb42c4eea · outbound

This paper cites End-to-end autonomous driving: Challenges and frontiers.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding End-to-end autonomous driving: Challenges and frontiers.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.275673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:e9986b4ce29e369c53c1eaad7f8ce3162e7c85ec03cd09370700bae88860311d

Observation 67294575-ce6b-4214-b2d7-2e3d96f52576 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.335451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:55980ffb87f8d477e5dad2227ac66e6c33ede14ee0f6f5f824e26dfd8b5cb4ff

Observation 73cbad4a-5406-4008-a837-973b8350428e · outbound

This paper cites Streaming video question-answering with in-context video kv-cache retrieval.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Streaming video question-answering with in-context video kv-cache retrieval

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.279141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:1f7cd4d14fe479915d78ab81fc9d693c19fdbf8e3d6f43982f76d93e7627fd82

Observation 5a0e7e1d-e4e0-4693-bc82-fd3ad1ff0c3b · outbound

This paper cites Contextnav: Towards agentic multimodal in-context learning.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Contextnav: Towards agentic multimodal in-context learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.682389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:952a1364ddd0e07a421045685cf6e8d29fd564edf001b2c7309571916ca1fafe

Observation c3c5ac3f-d520-4fac-bc78-2f09cf970c3a · outbound

This paper cites Vispeak: Visual instruction feedback in streaming videos.ICCV.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Vispeak: Visual instruction feedback in streaming videos.ICCV

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.329770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:5f0808602ce3aadbc206df512f3b4b2e616dd120ebac0b63cd9ed3c29453f4ae

Observation 7c5c557c-67ee-4cfb-b736-c68cf4ec02c9 · outbound

This paper cites Framemind: Frame-interleaved chain-of-thought for video reasoning via reinforcement learning.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Framemind: Frame-interleaved chain-of-thought for video reasoning via reinforcement learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.358339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:65e4707482034254aa5e7217e4be6994ca7476d5a481ecaa051845eab74d49a4

Observation 08fafee2-43ab-4884-8e57-c6f90e5993d9 · outbound

This paper cites Online video understanding: Ovbench and videochat-online.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Online video understanding: Ovbench and videochat-online

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.321534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:9cff6b6a6abe0cef268fb723da020f8bce9e1e1d04c3623ca06d60c36f0f0607

Observation 164185d6-5052-4626-a537-f2f29a40f0bd · outbound

This paper cites GPT-4o System Card.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding GPT-4o System Card

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.600709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:6f40f337b94b47d92ad08421f3041386ce9067653a86b473ee123df51ed7367b

Observation 856863e4-259a-41bf-a02b-b435c1916619 · outbound

This paper cites Colbert: Efficient and effective passage search via contextualized late interaction over bert.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Colbert: Efficient and effective passage search via contextualized late interaction over bert

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.293819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:5b5b5ddfe94c0caee0502dab247df061ea9535abca665dde9b9691bb76fc151f

Observation f7674042-f15d-49a9-988d-8babf5d5f3af · outbound

This paper cites Interaction methods for smart glasses: A survey.IEEE access, 6:28712–28732.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Interaction methods for smart glasses: A survey.IEEE access, 6:28712–28732

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.310617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:e1019889af2b1a96cb902e7529e4e8ff951c274c63987499d72c4265afc4f4d1

Observation 12cfcf3d-8cf4-4513-bd55-77f4657c2a65 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.660336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:c56e5721e2221f3fbf380ec2c8aaa0687b676905b8cd15d509bc2638b75c4f41

Observation 66471926-6853-4272-a1f4-7383aab5a4e5 · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.707856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:8507b901945e70e4b0a574a3e696bd2ccc2f001a581b5d38d44b76d83ea020d5

Observation 930384d9-d9c0-4594-9ade-d05441dec74b · outbound

This paper cites Aligning cyber space with physical world: A comprehensive survey on embodied ai.IEEE/ASME Transactions on Mechatronics.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Aligning cyber space with physical world: A comprehensive survey on embodied ai.IEEE/ASME Transactions on Mechatronics

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.304421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:ab32685a870c7ebf4966d14cbd3c944ed450a43692365b7b7a48fd5a81885c96

Observation 2da1e5ae-2c09-4a3b-bac5-f7af9f1b4ce8 · outbound

This paper cites Thinking in streaming video.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Thinking in streaming video

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.723910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:3384ff4f32bc0a55347405e8a43afbc2559cdf256847d3c1df3dfbc1e1fd4a2b

Observation 33c38f6c-0c47-437a-a022-5ef812e39d04 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.306839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:060da94429e4911b0f32eabc694676fe8b69e3afb1bade06808a65775e9cef30

Observation 834c531f-c251-499a-86cd-301799397b71 · outbound

This paper cites A Survey of Context Engineering for Large Language Models.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding A Survey of Context Engineering for Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:45.795434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:367e5f2bd0a9593163ae3c21c9f677405dc0584c8227f5edf9bc2792a358fadd

Observation 4077ec18-0598-48a5-97b2-22f5156a1595 · outbound

This paper cites Gated differentiable working memory for long-context language modeling.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Gated differentiable working memory for long-context language modeling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.510940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:13172622afeba9451407e22a03fac7ba68f2af68869c582559a2ea7b4b725eb5

Observation 4721a065-4802-484a-8aa8-555228acd816 · outbound

This paper cites LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.573338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:a112e5ba93489ebc700824fb1764e274a10a856c098af67c36c106a771888d2a

Observation 8a3a5fa0-23f5-438e-8bd8-5f15150b23e4 · outbound

This paper cites Ovo-bench: How far is your video-llms from real-world online video understanding? InCVPR, pages 18902–18913.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Ovo-bench: How far is your video-llms from real-world online video understanding? InCVPR, pages 18902–18913

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.337915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:610e514974bbca1d55d8bbf0f1fb5ffb459966a3db61630587a1defca4d2b007

Observation 394ee2b2-0a37-46b4-9afe-d094c66f3237 · outbound

This paper cites Athresholdselectionmethodfromgray-levelhistograms.IEEETransactionsonSystems,Man,andCybernetics, 9(1):62–66.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Athresholdselectionmethodfromgray-levelhistograms.IEEETransactionsonSystems,Man,andCybernetics, 9(1):62–66

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.288427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:d98693ab01eb558e6a0db037a58d2fe254c8c15d8c5b7cfc18c094a0ea78a83e

Observation cf0aef41-8c40-469c-8371-14da1d47333b · outbound

This paper cites Streaming long video understanding with large language models.NeurIPS, 37:119336–119360.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Streaming long video understanding with large language models.NeurIPS, 37:119336–119360

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.285159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:36cfccdc735640d59690ef0efdc85cebeb72d27958432ab73565da240a9b1338

Observation 05bfd455-20aa-43be-ac4e-e18d0848ae8e · outbound

This paper cites Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.319311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:63ff603f981a589abc319670f724bb835aaf45585fd72ec2eaaa1ab14959cbd0

Observation a3b26720-147f-4ae9-a996-392e139c8168 · outbound

This paper cites Longvu: Spatiotemporal adaptive compression for long video-language understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Longvu: Spatiotemporal adaptive compression for long video-language understanding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.317309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:c45467057985238a724f202a1a6cce9a7a73c5b5fd532bda8ee7f5ef38edc5bb

Observation fdfbf315-9d46-48a6-b71f-c2678416f48d · outbound

This paper cites A Simple Baseline for Streaming Video Understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding A Simple Baseline for Streaming Video Understanding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.672127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:be3702ced1e71ebfec7e870962d4500a12e0f1f8a09cfbf1b1a674ebebeb7bf2

Observation e23b966d-1cb3-40e9-8228-87b3ede4758b · outbound

This paper cites Video understanding with large language models: A survey.IEEE Transactions on Circuits and Systems for Video Technology.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Video understanding with large language models: A survey.IEEE Transactions on Circuits and Systems for Video Technology

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.327400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:168905091db4c717dc7bb4346bcd90ebba67aaff47022c2468943b41fa36366e

Observation f2dbf2c8-d2ff-4443-b3ef-e4526cd558b5 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.640450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:7160ca35151174b5dcef9dbe636d7eab25fe2de1506b96e6c9f458887eaeff94

Observation 214a7323-1d43-4b65-a7c7-11d81a1ec7cc · outbound

This paper cites Streambridge: Turning your offline video large language model into a proactive streaming assistant.NeurIPS.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Streambridge: Turning your offline video large language model into a proactive streaming assistant.NeurIPS

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.301870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:9fc1ba45631759c5201a0b6806b66c73760efe07f0f0a87875868247845c4013

Observation 46f6a05a-b353-4bc2-8862-e30fd90d321a · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.473686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:715beac60a810c12cba518fb63c698fc332b28ab666bb4139589b9a65a287256

Observation 5ce4d7eb-816a-44b9-acb1-7698013bf443 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.484333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:b548eaab3df5f56deba7bb2a3adfad11fef09137bd44f6997610877261ebc7b5

Observation c2a8e64d-bb3c-445d-94b8-46bc820064af · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.415498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:3793b72c0374d41e2a7ed394a19b7def43b3467d905bbc9a05ccb8dd9563e4df

Observation f258556c-d103-4168-afd5-85ef3680fb2f · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:2b0fa5094160331d853a9fff19b0e094e815fde6feb666786416c46658a7e857

Observation 049a96ef-ffc7-4c3a-82ae-51cae100a626 · outbound

This paper cites CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.493339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:668cc17a8ddd0e08241f2872acb4f5fd2729191e1f90d9124a4f76853c8ccad2

Observation a1a9b116-c700-434b-bd40-d91c97c3ea38 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.NeurIPS.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding.NeurIPS

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.313099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:942b95ab1e7f542c8649361d00f3b7d66539c3ee9ec671b1cca68d231d774726

Observation ca0e2e08-160f-4b92-ae6b-17df6761d35d · outbound

This paper cites Fluxmem: Adaptive hierarchical memory for streaming video understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Fluxmem: Adaptive hierarchical memory for streaming video understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.434561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:dd5e265bdae322974d4653e8b210ba7c32456e1ccceb4f262d4682562dca254e

Observation ca912523-693d-4cc0-bbc8-1a4248e47317 · outbound

This paper cites StreamingVLM: Real-Time Understanding for Infinite Video Streams.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding StreamingVLM: Real-Time Understanding for Infinite Video Streams

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:51:33.451476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:91c06a6190f5454e7af009cbbac4ec8c3dd2b6f21d6225b9b12176c53d10f8a8

Observation 4ccd50be-7e92-49ec-9989-e43b071e86f1 · outbound

This paper cites StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.370849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:1ec3f4befed634d8cdcc8f85bbe58723b838be0c56df1f89f7c1a0cefe16bffd

Observation 71c3f810-d55d-475c-8b6d-829ad8fbdb2f · outbound

This paper cites Timechat-online: 80% visual tokens are naturally redundant in streaming videos.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Timechat-online: 80% visual tokens are naturally redundant in streaming videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.324135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:968453b0aa65e844822074b6260187c49725a59ad4bf61ade588a252695e36f3

Observation 7972e66d-952c-44ae-879d-643026da7ba4 · outbound

This paper cites Streamforest: Efficient online video understanding with persistent event memory.NeurIPS.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Streamforest: Efficient online video understanding with persistent event memory.NeurIPS

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.315244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:48920898be4a63511c28508c03647b26ff6a581db6858d03a8def8e25734d973

Observation 75411568-414b-4639-92f5-0aef2823a9bd · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:02:01.218694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:74745503ee0b7bb646761a4f18aaf2c2018d0af86c80586d8e8290c33a3720fe

Observation ed007fe1-729a-4b8d-83c9-ede42513274c · outbound

This paper cites Flash-vstream: Memory-based real-time understanding for long video streams.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Flash-vstream: Memory-based real-time understanding for long video streams

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.332328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:ebec053412a6f234bd8ba544d34a9aa682f407386a916cec261f3b3ac02865a4

Observation 9888cebf-3702-4571-b03c-8ec4b3f3ee2c · outbound

This paper cites HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.380773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:ea78e0ef7e2f87a48912aa5c78e1aa87c431ff0bb183bea1676b09150aee36bf

Observation b1dfc6ce-be28-46e3-aaa0-58a0929a12f8 · outbound

This paper cites Long Context Transfer from Language to Vision.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Long Context Transfer from Language to Vision

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:485588474c52100ae98e56f30753005b96773fad3d4356e7fe61a7f79be73573

Observation 02b79e42-e07e-4354-8171-c48e8b5cede1 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.454106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:8651fc5efd84ab655ad899f675de2b1ce876a28faeb3a97eceafdce56d1c42fb

Observation 8c780156-d464-416c-b990-d3138c9315e9 · outbound

This paper cites Weavetime: Stream from earlier frames into emergent memory in videollms.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Weavetime: Stream from earlier frames into emergent memory in videollms

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.530833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:e0bd8888c795cae4124580ef2ade1014de77b10b6d5613c93b34c2557174bf04

Observation 4d6215a8-9efa-4e93-b82e-05d7f8b4d6ec · outbound

This paper cites Whatobjectsarevisible?.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Whatobjectsarevisible?

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.298959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:6c122b4153331c206766ba4657a6ca9d2a7d5b34c581f011f44fc4f326dd7677

Observation 22a250f7-bf45-45ab-bbd4-efe61ed1d5f8 · outbound

This paper cites What objects are visible in the scene ?.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding What objects are visible in the scene ?

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.291504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:010fa65dfc23d47ae603d6e9407c78115e24a741c7729d3b4b1250719a53f78f

Pith citing papers

Observation 113d77bf-03e9-452a-8d1b-4a14e641c4ac · inbound

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding cites this paper.

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T03:21:47.718375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:21:47.718375Z digest=sha256:6c98dc618324604aab34598565718ddafbe1c1fa4829b7df7606c92422197ea0