Pith. sign in

Paper Citation Record · LEDGER

FOLIO: Focused Semantic Memory for Streaming Video Understanding

As of 21 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 0 inbound Pith citation observations for arXiv:2607.13298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13298 v1

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:40:51.856126Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 109 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92ac4837-9832-4e68-8262-dc1d69fa209a · outbound

This paper cites Goldfish: Vision- language understanding of arbitrarily long videos.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Goldfish: Vision- language understanding of arbitrarily long videos

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:40.247041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:40.247041Z digest=sha256:1efb96a4606e80480b35fe572e67326e41db57390d1255abff5ae181bcbb38d4

Observation 3a81e2e3-cfc2-4724-b6ac-39f05d2b2cdf · outbound

This paper cites Qwen3-VL Technical Report.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:40.306727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:40.306727Z digest=sha256:9bd70938a35b46561c3829521abca7ce363406c151b68389bd6921824fc388f0

Observation 7f6ed852-7b0d-4ada-b19c-0b13bb8d8368 · outbound

This paper cites Qwen2.5- vl technical report, 2025.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Qwen2.5- vl technical report, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:40.452140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:40.452140Z digest=sha256:f7f8a390a758997808bb37c1d2ee669485b9dbcaa6578ae6509920bfe76c6964

Observation 0479cafa-2f25-455b-a4d4-146c4e4fde04 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

FOLIO: Focused Semantic Memory for Streaming Video Understanding RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:40.605291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:40.605291Z digest=sha256:77b323e42b295d224577ed1288cf4e30d0ee3693bdebaa22e92c10ef7bdf29c7

Observation a688e2c5-3b35-408b-8bab-6b496859cd50 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Videollm-online: Online video large language model for streaming video

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:40.741228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:40.741228Z digest=sha256:80263913ce7298d262adde35de70a62f9dc20033e6ac543ca8111cce87bfcc3f

Observation 0ee16819-6ab5-44a0-8fa5-e1be06dc8d3c · outbound

This paper cites Livecc: Learning video llm with streaming speech transcription at scale.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Livecc: Learning video llm with streaming speech transcription at scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:40.914298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:40.914298Z digest=sha256:cb520e7519492699c359a21a9268a2b032c547785e8269ea27f0d28beabf5d7c

Observation 7a6aa7dc-0924-4bf6-b239-bfd95bb94cfe · outbound

This paper cites Stream- ingtom: Streaming token compression for efficient video understanding.arXiv preprint arXiv:2510.18269, 2025.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Stream- ingtom: Streaming token compression for efficient video understanding.arXiv preprint arXiv:2510.18269, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:41.029341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:41.029341Z digest=sha256:95f9bb79bcef87a9646adbfe2be3b514f7c4d57c59b13b424ef056027b7b6ae8

Observation 14909bf9-4e6a-49e3-9ca5-d4b22df09623 · outbound

This paper cites Streamkv: Streaming video question- answering with segment-based kv cache retrieval and com- pression.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Streamkv: Streaming video question- answering with segment-based kv cache retrieval and com- pression

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:41.155797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:41.155797Z digest=sha256:1cb31dddf0f3b2ccf3600a9ccc799e18fe0e4255dc5c398cdcb806b3060e74c3

Observation fb5ef3fe-aa98-483e-86ac-88d8538a82bf · outbound

This paper cites Streaming video question-answering with in-context video kv-cache retrieval.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Streaming video question-answering with in-context video kv-cache retrieval

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:41.318685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:41.318685Z digest=sha256:f8ad8d7ba17fdd1f5341f966a41458d150b8c70a88a6f920d606994954162caf

Observation 69f81289-f454-42ed-8285-5442a4917223 · outbound

This paper cites Streammind: Un- locking full frame rate streaming video dialogue through event-gated cognition.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Streammind: Un- locking full frame rate streaming video dialogue through event-gated cognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:41.512057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:41.512057Z digest=sha256:6d6e9d16e43d744cbf174b7fae1c2474025c13f02eb1e047020bc5a9a3a178f3

Observation aaa87bcc-a4da-4f84-97cf-5f91d7c1f9f1 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:41.698700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:41.698700Z digest=sha256:ad7398cacd1fc6455a1bc3ab33941072812e1d8ca68550b831b3bf57e9299a57

Observation ed48361d-44d9-4e21-8907-c207c71036bc · outbound

This paper cites Vispeak: Visual instruction feedback in streaming videos.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Vispeak: Visual instruction feedback in streaming videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:41.819932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:41.819932Z digest=sha256:67895a5dbb8ba8d560e4d8abbccb42a2be5fc12009e74109498b101237c28a12

Observation a501a3c6-1b7d-4d94-b769-85ad619547f7 · outbound

This paper cites an unresolved cited work.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:41.964153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:41.964153Z digest=sha256:f554985e706776081b84390d8211fb4f3a3354d08fff2f5588b624b25051c091

Observation 107746f8-e14a-43c1-b43d-b04548dc1879 · outbound

This paper cites Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:42.117347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:42.117347Z digest=sha256:bda73788450e360e7c1e1241ff65f0319e24c34da1327ba521efad6b9bbfef45

Observation 5bb57e00-b804-443d-8a37-9a2b7988ad93 · outbound

This paper cites Event-vstream: Event-driven real- time understanding for long video streams.arXiv preprint arXiv:2601.15655, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Event-vstream: Event-driven real- time understanding for long video streams.arXiv preprint arXiv:2601.15655, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:42.271382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:42.271382Z digest=sha256:eb4e6ce2a1a1dd042ebac15ecd6c10171e2bcf7bef9760683c2a6c51541434d5

Observation e490df33-f717-4c95-8f2b-3edb8a2da2b3 · outbound

This paper cites Wat: Online video understanding needs watching before thinking.arXiv preprint arXiv:2603.13412, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Wat: Online video understanding needs watching before thinking.arXiv preprint arXiv:2603.13412, 2026

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:42.449065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:42.449065Z digest=sha256:16e172f15b864c4a48d7a6a2db132523a77c7f6e16d47f25a1ded578c8eab74c

Observation 77f905a5-107d-437d-903d-fb63d30b704b · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:42.621001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:42.621001Z digest=sha256:2dfb04cbefc650cc0f14c054b6b09be323fd1a809bc6864e0c0bde868869b1c9

Observation 081bdcdd-a42e-4188-85be-6fdbbb6f31ee · outbound

This paper cites Online video understanding: Ovbench and videochat- online.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Online video understanding: Ovbench and videochat- online

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:42.775759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:42.775759Z digest=sha256:c7d6087a15c6b9af9ffdd0d49e1c91584a282a9642dddf2c59b60a8543b0edac

Observation 2df3c5ee-00cc-496d-94d1-c08c20afd595 · outbound

This paper cites Online video understanding: Ovbench and videochat- online.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Online video understanding: Ovbench and videochat- online

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:42.873549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:42.873549Z digest=sha256:864d18734e12a57045cd4723e64c25e4b9f856abd99ad3a24f34b60edcea0421

Observation 5807cf8c-7260-4925-a833-afc4b526ae5a · outbound

This paper cites Egospeak: learning when to speak for egocentric conver- sational agents in the wild.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Egospeak: learning when to speak for egocentric conver- sational agents in the wild

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:43.056189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:43.056189Z digest=sha256:2641d890eecbfec772e9ee779c771ada9c1aedc2e167a69c498a909041c63692

Observation b850061b-9fc1-4dc8-b962-d4d8ab59c382 · outbound

This paper cites Infinipot-v: Memory-constrained kv cache compres- sion for streaming video understanding.Advances in Neural Information Processing Systems, 38:138983–139013, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Infinipot-v: Memory-constrained kv cache compres- sion for streaming video understanding.Advances in Neural Information Processing Systems, 38:138983–139013, 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:43.199978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:43.199978Z digest=sha256:34bda96a29c8856f9514c03ee085734442d907cc535b938ead8a64c257d90d47

Observation 7d5c096b-6768-48af-949b-2462e88db09e · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:43.359412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:43.359412Z digest=sha256:9e260fcfdb980c7a9279d5d2360bcac7cf86f3449b73f8ef2fc728555803d0ee

Observation 1e6a4946-ae7c-4d80-afb7-c196369fcdbc · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:43.502119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:43.502119Z digest=sha256:8dcb37e2500ba7c7a4e0527add2ac03dba370deec191fedefaf9d15769a136f4

Observation bda3111c-9c49-4876-9324-1a8fb354f293 · outbound

This paper cites Videochat: Chat-centric video understanding.Science China Information Sciences, 68(10):200102, 2025.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Videochat: Chat-centric video understanding.Science China Information Sciences, 68(10):200102, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:43.627646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:43.627646Z digest=sha256:4e21c10c8136ccc2dec3493aafaf84e49109eb7c87c5ab8d803d364f10cee578

Observation 1ab66b5b-be95-41e6-903f-5da02a10e35a · outbound

This paper cites Freshmem: Brain-inspired frequency-space hybrid memory for streaming video understanding.arXiv preprint arXiv:2602.01683, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Freshmem: Brain-inspired frequency-space hybrid memory for streaming video understanding.arXiv preprint arXiv:2602.01683, 2026

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:43.826258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:43.826258Z digest=sha256:457c4bf0b932708cb534fe460207879c5ce4afce3282e77089f2900090bab18d

Observation 3d96b5dd-5104-4155-85cc-0630da8851e7 · outbound

This paper cites Lion-fs: Fast & slow video-language thinker as online video assistant.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Lion-fs: Fast & slow video-language thinker as online video assistant

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:44.115029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:44.115029Z digest=sha256:2c0a711e8d6f8eec60882766faf0b39ecf975487a680f7bd3decb102eca25787

Observation 8f77d249-e872-4ec2-bda3-c531a1e69f58 · outbound

This paper cites Llama-vid: An im- age is worth 2 tokens in large language models.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Llama-vid: An im- age is worth 2 tokens in large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:44.185626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:44.185626Z digest=sha256:2f28880c0cefbd9cde022a939f18ecdaa754af1b0fc140ac2e243c24c7fbe876

Observation 17f89242-abae-4e70-a51f-d88a8812ddd7 · outbound

This paper cites From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents.

FOLIO: Focused Semantic Memory for Streaming Video Understanding From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:44.305707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:44.305707Z digest=sha256:707997272795b87c55cb8c12644bcbdceb3071fa24804b4b3c529271e1347dde

Observation 1091435e-598c-427f-a6a2-4e5c27be402c · outbound

This paper cites OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning.

FOLIO: Focused Semantic Memory for Streaming Video Understanding OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:44.418468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:44.418468Z digest=sha256:298260fa4b1de6457ab300d157d9e6023f24e1902bc9286533a2baf01627a663

Observation 92f3fb23-0d68-4153-92d6-f3aa4d9fee15 · outbound

This paper cites Video-llava: Learning united visual represen- tation by alignment before projection.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Video-llava: Learning united visual represen- tation by alignment before projection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:44.553170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:44.553170Z digest=sha256:6f1debf542dca5847689629dfb2fa5dc91b2ccf32efc67f5958615428068645b

Observation 8587f1db-61ae-400d-bbb2-8a1b62ff0a7f · outbound

This paper cites Streamingbench: Assessing the gap for mllms to achieve streaming video understanding.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Streamingbench: Assessing the gap for mllms to achieve streaming video understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:44.705540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:44.705540Z digest=sha256:5471ab6f46bbbfd176c79ad315816c21c46ca80f71dd6f81356274553d2263a8

Observation f2642ebb-3cd7-411a-8d0b-35309ccac26f · outbound

This paper cites Speak while watch- ing: Unleashing true real-time video understanding capabil- ity of multimodal large language models.arXiv preprint arXiv:2601.06843, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Speak while watch- ing: Unleashing true real-time video understanding capabil- ity of multimodal large language models.arXiv preprint arXiv:2601.06843, 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:44.823539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:44.823539Z digest=sha256:ca18ab6cacccf919171d11b2f5d64e0c6b5d84f21143a0c6c4b675a2e8c66ce7

Observation ad9e9908-1673-475f-8184-7fe191510732 · outbound

This paper cites StreamChat: Chatting with Streaming Video.

FOLIO: Focused Semantic Memory for Streaming Video Understanding StreamChat: Chatting with Streaming Video

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:44.938616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:44.938616Z digest=sha256:000950b36c065e5ab4997dcbc5d9cd0d23586762567a2ae78270bc38a12a3511

Observation 2a3f4126-9c8d-4ef7-b8bf-5b3c671e7c91 · outbound

This paper cites Thinking in streaming video.arXiv preprint arXiv:2603.12938, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Thinking in streaming video.arXiv preprint arXiv:2603.12938, 2026

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:45.041545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:45.041545Z digest=sha256:fd30252771d89500dad55969e0bad48e8af65087e993cb2947f0831830d708ee

Observation 78f33efa-3b2b-40ba-8415-07c4ba057f3d · outbound

This paper cites Vista: Scene-aware optimization for streaming video question answering under post-hoc queries.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Vista: Scene-aware optimization for streaming video question answering under post-hoc queries

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:45.193669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:45.193669Z digest=sha256:1ed5a8e78824320cce17fde08cdd3d45026cc2092222446bc773e87af05eaf10

Observation 3a9ff75c-1b94-48bd-9934-4e3c38157c8c · outbound

This paper cites AURA: Always-On Understanding and Real-Time Assistance via Video Streams.

FOLIO: Focused Semantic Memory for Streaming Video Understanding AURA: Always-On Understanding and Real-Time Assistance via Video Streams

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:45.373894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:45.373894Z digest=sha256:04a83b1eba03d0e28d097290cf0bdaa9eea7c294c2b86eaf120b4648798d40e3

Observation 2ed3d7f9-7eb5-46c5-8867-e8c14836dea8 · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.Advances in Neural Information Processing Systems, 38:168008–168033, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Video-rag: Visually-aligned retrieval-augmented long video comprehension.Advances in Neural Information Processing Systems, 38:168008–168033, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:45.506793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:45.506793Z digest=sha256:8618606e3d60502cb5540aad049d6de0f3d2501f95f5dc3bc43bdec649e4fa87

Observation 05e9de73-c8f6-44e3-8062-9647d20c60fc · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:45.677511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:45.677511Z digest=sha256:21aca80e86214f812f8950bfff7a0e10e0d1f09451430fd0994524caa330fa2c

Observation f15bfcc4-c98f-42c7-9fd9-66f8cdccd5e3 · outbound

This paper cites LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval.

FOLIO: Focused Semantic Memory for Streaming Video Understanding LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:45.851326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:45.851326Z digest=sha256:58163d8efd5160328a2bfa9ca34445711d71741faad10fa401ed087952fb6267

Observation ded841e7-4165-450c-9c6e-0088580aadf0 · outbound

This paper cites Ovo-bench: How far is your video-llms from real-world online video understanding? InProceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Ovo-bench: How far is your video-llms from real-world online video understanding? InProceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:45.954821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:45.954821Z digest=sha256:b22c8a2613c57148bdeff6351788a42c715fd167bb775e3ea378d0615803c895

Observation 9a9088c7-d080-427d-a439-abbca51f0799 · outbound

This paper cites Streaming long video un- derstanding with large language models.Advances in Neural Information Processing Systems, 37:119336–119360, 2024.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Streaming long video un- derstanding with large language models.Advances in Neural Information Processing Systems, 37:119336–119360, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:46.189047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:46.189047Z digest=sha256:9ca33d94cadbbecc7e037616870928c2caea35ee62dc67272ed19fa8816c07d7

Observation 605428c9-ced5-4e26-b65c-746de9d130bc · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:46.344843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:46.344843Z digest=sha256:daf0ea26ade1881fa0b557d8456e52cd117e4ea0dc87c7fdd4bdd8cb118d3270

Observation 650776ae-51d1-4984-b7b1-88e9454365e2 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

FOLIO: Focused Semantic Memory for Streaming Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:46.409301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:46.409301Z digest=sha256:f53e3423d9d5f1600f1440016e0472bb99f5bf027c78e5afec6da853d1849848

Observation bc26dce3-1ee7-4bcf-98aa-63514a5f5157 · outbound

This paper cites A simple baseline for streaming video understanding.arXiv preprint arXiv:2604.02317, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding A simple baseline for streaming video understanding.arXiv preprint arXiv:2604.02317, 2026

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:46.495255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:46.495255Z digest=sha256:0970919e688fa55f9648f92cffd2ce2442d06c7660f4b6ae39cbfd2f0c6db090

Observation fc3ebf61-ac61-4d73-8c64-b465e1d70c48 · outbound

This paper cites Video- xl: Extra-long vision language model for hour-scale video understanding.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Video- xl: Extra-long vision language model for hour-scale video understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:46.570307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:46.570307Z digest=sha256:d520c246e301847cc01103ffed9af0d4abec9e52bebd3a47d6fcf5c7a8beec1e

Observation 6ad2b717-5d22-4ca6-bd10-fc8d39291973 · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

FOLIO: Focused Semantic Memory for Streaming Video Understanding DriveLM: Driving with Graph Visual Question Answering

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:46.637545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:46.637545Z digest=sha256:618171c62308ba9dde960a837b83539fe1c26f285705db3af6f9d2d8ba0c84e6

Observation 48beaaa8-0c4d-4a49-b98d-0f16b4059af5 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Moviechat: From dense token to sparse memory for long video understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:46.747360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:46.747360Z digest=sha256:45c9f9e40e3a65dc6243ae517e1c5999ccce79d5c1dbcf2c0c392e3afd6beaec

Observation 7b0d9cc7-cb53-473b-885c-d31b388862e0 · outbound

This paper cites Curvestream: Boosting streaming video understanding in mllms via curvature-aware hierarchical visual memory management.arXiv preprint arXiv:2603.19571, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Curvestream: Boosting streaming video understanding in mllms via curvature-aware hierarchical visual memory management.arXiv preprint arXiv:2603.19571, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:46.819805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:46.819805Z digest=sha256:42d2ab3f36177c4bb095cf9a8ca63849ccaf5d113bcec5bed2bcec3b74fe2028

Observation bd93c51b-bb06-4b14-b2db-d279c7a12004 · outbound

This paper cites Streambridge: Turning your offline video large language model into a proactive streaming assistant.Advances in Neural Information Processing Systems, 38:132332–132359,.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Streambridge: Turning your offline video large language model into a proactive streaming assistant.Advances in Neural Information Processing Systems, 38:132332–132359,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:46.933674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:46.933674Z digest=sha256:d79a0aa927cc31fc85dbeefd73ceffb2168cdb0b7579a2178ce0364572685e4d

Observation 4e230ec5-8c08-4fb5-84e0-4deef257fefa · outbound

This paper cites Think while watching: Online streaming segment-level memory for multi-turn video rea- soning in multimodal large language models.arXiv preprint arXiv:2603.11896, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Think while watching: Online streaming segment-level memory for multi-turn video rea- soning in multimodal large language models.arXiv preprint arXiv:2603.11896, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:47.033786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:47.033786Z digest=sha256:ec5e07c6832ba3018d4cba013ee514d113e9620e87ee97f20bf499cdffa23430

Observation e33e8f60-d39f-4d70-9a51-95a1f489eb80 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:47.124727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:47.124727Z digest=sha256:7affbd9e579c4e259aa69421b0350b2591b8c0d04c0564b0c423d8b93d26a5a1

Observation cb7b1314-5322-4540-b46f-d7430e60f5f1 · outbound

This paper cites Lvbench: An extreme long video understanding benchmark.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Lvbench: An extreme long video understanding benchmark

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:47.227901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:47.227901Z digest=sha256:927d8a6f81f3c14394b68faf3d52b19f9dbaea4319f1aa95b3faf89594767a44

Observation c97a63f8-117c-4ce7-bed3-8418963d3b18 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Videoagent: Long-form video understanding with large language model as agent

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:47.319098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:47.319098Z digest=sha256:e6d19d2bd6b0370253a310a197f646ad7847abf71d27e183e1eccb8482df735e

Observation d76a5321-1bb3-4b94-89cf-5c1bab3eba79 · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:47.388001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:47.388001Z digest=sha256:f32251e417a45aeed6e0be05e06722e19d16e8500ce42dd2ea96acd7805387b1

Observation ef2e1078-8004-42e4-be1d-af9ad1cf3ece · outbound

This paper cites Vide- ollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format.arXiv preprint arXiv:2411.17991, 1(3):5, 2024.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Vide- ollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format.arXiv preprint arXiv:2411.17991, 1(3):5, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:47.488282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:47.488282Z digest=sha256:2feded4a4f2aa52eb08b71713f7bb7329e4c2c251a673996d775be9e2c95ed22

Observation 67807932-6bcb-4079-bc9e-51b5674d8a18 · outbound

This paper cites Acceler- ating streaming video large language models via hierarchical token compression.arXiv preprint arXiv:2512.00891, 2025.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Acceler- ating streaming video large language models via hierarchical token compression.arXiv preprint arXiv:2512.00891, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:47.593361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:47.593361Z digest=sha256:1d6f626e0eef94d88502844ca0e0ebb5a8b32e35c536211cc0150292ae96b8a5

Observation 7e988e12-3a32-4484-bcf9-1a4d670c3047 · outbound

This paper cites Omnimmi: A comprehensive multi- modal interaction benchmark in streaming video contexts.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Omnimmi: A comprehensive multi- modal interaction benchmark in streaming video contexts

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:47.678186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:47.678186Z digest=sha256:1155a2f5de31837202b9459cf99d36bf6a57961a547db290ae56149e4d362928

Observation a27e5a17-8fd0-48d6-936b-6cb4d038a4e0 · outbound

This paper cites Episodic memory representation for long- form video understanding.arXiv preprint arXiv:2508.09486,.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Episodic memory representation for long- form video understanding.arXiv preprint arXiv:2508.09486,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:47.760523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:47.760523Z digest=sha256:ebf6792ce8bedb683da5e95b37e40a32ef1a9b66af5d09d49b54953d1e9fcb2f

Observation 19830a06-27c4-4b28-8a0d-bded57fd4607 · outbound

This paper cites 11 Videotree: Adaptive tree-based video representation for llm reasoning on long videos.

FOLIO: Focused Semantic Memory for Streaming Video Understanding 11 Videotree: Adaptive tree-based video representation for llm reasoning on long videos

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:47.792484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:47.792484Z digest=sha256:0cffe455c479913b5ab9fba38569a6a5106ad5b998c2a37e636f8ad85eb0c258

Observation 94b2a429-ce9e-4a0a-830f-b1825ec08421 · outbound

This paper cites Eventmemagent: Hierarchical event-centric memory for online video understanding with adaptive tool use.arXiv preprint arXiv:2602.15329, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Eventmemagent: Hierarchical event-centric memory for online video understanding with adaptive tool use.arXiv preprint arXiv:2602.15329, 2026

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:47.845688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:47.845688Z digest=sha256:3e770c649499c3f7bfde7502063921ef2fc744bda9bb82473507990fd1da5cea

Observation 191388df-15e1-42c8-8b09-b91be30caf2b · outbound

This paper cites Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation.Ad- vances in Neural Information Processing Systems, 37:109922– 109947, 2024.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation.Ad- vances in Neural Information Processing Systems, 37:109922– 109947, 2024

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:47.951763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:47.951763Z digest=sha256:afa8c9455a25b6dd343af8882cfb0d531a19ae9f648397ac521b807835c30859

Observation cfcffe0f-ea85-4ff8-8acc-0b9172dc8fde · outbound

This paper cites Next-qa: Next phase of question-answering to explaining tem- poral actions.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Next-qa: Next phase of question-answering to explaining tem- poral actions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:48.015629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:48.015629Z digest=sha256:d2ed98c41ec064d730176ecde663c9fa47cbe24e08169d37023551ae3fb16999

Observation 7102de25-5625-4e77-920e-5c55ba83be56 · outbound

This paper cites Fluxmem: Adaptive hierarchical memory for streaming video understanding.arXiv preprint arXiv:2603.02096, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Fluxmem: Adaptive hierarchical memory for streaming video understanding.arXiv preprint arXiv:2603.02096, 2026

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:48.051409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:48.051409Z digest=sha256:ee0cce5ceb309f356fea784890c23b3db339a14d6a1e1ac9e44462376732ef29

Observation a35e0a5b-0773-466f-8269-8709cf983134 · outbound

This paper cites Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:48.131445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:48.131445Z digest=sha256:9d1ec2c4af2fd63053e6a60b34bd6b18f91808e240b31b8b056d2bf0f5b37a3e

Observation 4aab3a6e-dce1-4db3-8ea6-e720f36e02c0 · outbound

This paper cites StreamingVLM: Real-Time Understanding for Infinite Video Streams.

FOLIO: Focused Semantic Memory for Streaming Video Understanding StreamingVLM: Real-Time Understanding for Infinite Video Streams

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:48.277130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:48.277130Z digest=sha256:3d670ca5883763c7c8ab8b691d4bcd13bebbf9c457e2a43777bd89e2b2b2ff3f

Observation 293e3802-b4f6-43b3-9d0b-d01a1a99fbb7 · outbound

This paper cites RTV-bench: Benchmarking MLLM continuous per- ception, understanding and reasoning through real-time video.

FOLIO: Focused Semantic Memory for Streaming Video Understanding RTV-bench: Benchmarking MLLM continuous per- ception, understanding and reasoning through real-time video

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:48.354846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:48.354846Z digest=sha256:061bf9ccafb4ec6fed28d15099057aa3d9047cd7bc3c5434575a13adbd6d3ad9

Observation 1ab9f8a5-2f51-45b5-8f2b-c11138fe01ea · outbound

This paper cites StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding.

FOLIO: Focused Semantic Memory for Streaming Video Understanding StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:48.484251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:48.484251Z digest=sha256:b07d4cb9da2b388353cf35b8442393da8244a9d2fa7adec86b48777b90049866

Observation ded4972f-d0b4-4d9b-8a7c-c283fb20302c · outbound

This paper cites Svbench: A benchmark with temporal multi-turn dialogues for streaming video understanding.arXiv preprint arXiv:2502.10810, 2025.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Svbench: A benchmark with temporal multi-turn dialogues for streaming video understanding.arXiv preprint arXiv:2502.10810, 2025

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:48.616626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:48.616626Z digest=sha256:d439c3c64187a021d6d6e6ac60874d50612ab41ec52d63905b26edd1edde3fd9

Observation 4b67abcb-9da9-4f3d-81ec-97d6128fec3a · outbound

This paper cites Livestar: Live streaming assis- tant for real-world online video understanding.Advances in Neural Information Processing Systems, 38:31266–31304,.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Livestar: Live streaming assis- tant for real-world online video understanding.Advances in Neural Information Processing Systems, 38:31266–31304,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:48.860850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:48.860850Z digest=sha256:2fcc390d96a890f8e550e2cac6183ec08ef80083c4f9b2dd31d6794c25c6389e

Observation f9e01330-bf06-451e-8948-e2acc84923db · outbound

This paper cites Timechat-online: 80% visual tokens are naturally redundant in streaming videos.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Timechat-online: 80% visual tokens are naturally redundant in streaming videos

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:48.999512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:48.999512Z digest=sha256:4f45e6eeef452286718fcd62b84529125244690ef397bc73a547ba6aeed68940

Observation fdd4536e-f4ed-4293-a0a7-b36f408ff6a1 · outbound

This paper cites Worldmm: Dynamic multimodal memory agent for long video reasoning.arXiv preprint arXiv:2512.02425, 2025.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Worldmm: Dynamic multimodal memory agent for long video reasoning.arXiv preprint arXiv:2512.02425, 2025

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:49.120572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:49.120572Z digest=sha256:4d0a3b1d4dfc73ea17d75312655301a030b26ed47b5f749ddb0ee2bafa9c7056

Observation 7256dec1-32ac-4ed5-8a1f-c63ce8f90d1d · outbound

This paper cites Eyes wide open: Ego proactive video-llm for streaming video.Ad- vances in Neural Information Processing Systems, 38:13420– 13463, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Eyes wide open: Ego proactive video-llm for streaming video.Ad- vances in Neural Information Processing Systems, 38:13420– 13463, 2026

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:49.196026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:49.196026Z digest=sha256:ccbf1c802cf947363715a1485374f2e0e538145e8a895c98718682a1f9b692e6

Observation b4f971bb-7ccf-40b8-bb42-e8117f951dc3 · outbound

This paper cites Streamforest: Efficient online video understand- ing with persistent event memory.Advances in Neural In- formation Processing Systems, 38:75804–75835, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Streamforest: Efficient online video understand- ing with persistent event memory.Advances in Neural In- formation Processing Systems, 38:75804–75835, 2026

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:49.276526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:49.276526Z digest=sha256:1d2baefb6f10cb8da71a98cadcb5e174bae954124337fd823ef667108d65354f

Observation 06b4bf85-672f-409a-b567-6da81f991a74 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:49.371480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:49.371480Z digest=sha256:ba364f19f4e6f2e6ede396aed06a057422e5cb01acba4dcf058601de07224dd8

Observation 1fab58bf-3224-4412-add9-4d877e13272f · outbound

This paper cites HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding.

FOLIO: Focused Semantic Memory for Streaming Video Understanding HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:49.468513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:49.468513Z digest=sha256:72ac54931ea8a1366fe697499cd4f956786d0af7a6427ed6c28a1d4aba7e4954

Observation 7114c5bc-4360-46a7-96e5-bb83427ff5ce · outbound

This paper cites Think-as-you-see: Stream- ing chain-of-thought reasoning for large vision-language mod- els.arXiv preprint arXiv:2603.02872, 2026.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Think-as-you-see: Stream- ing chain-of-thought reasoning for large vision-language mod- els.arXiv preprint arXiv:2603.02872, 2026

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:49.547863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:49.547863Z digest=sha256:98830c4dc47091e8f9fa7d0d9c51678059bde01c5b85624b4fce8728ffe14875

Observation c18d68e3-480d-42f8-9acd-72c48b903cfc · outbound

This paper cites Querystream: Advancing streaming video understanding with query-aware pruning and proactive response.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Querystream: Advancing streaming video understanding with query-aware pruning and proactive response

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:49.659308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:49.659308Z digest=sha256:ec6e3cfdd079ae92f8aa1db6a8acd4e7461921c12ea5be88115b55669d7e826b

Observation c7153a65-386a-4160-b939-63be0ae1703d · outbound

This paper cites InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions.

FOLIO: Focused Semantic Memory for Streaming Video Understanding InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:49.748388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:49.748388Z digest=sha256:c020f771c4d39d100c1f0b103812c572a4cdf3d908450735913e949972f444e3

Observation a27eb059-6f72-4afd-b2c5-a86da28f5db7 · outbound

This paper cites Long Context Transfer from Language to Vision.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Long Context Transfer from Language to Vision

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:49.833950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:49.833950Z digest=sha256:665dabbda802c33f8423d001fa2c6663d55a3306a45aaa178be1e8d2d5d192f7

Observation 809e9789-b714-4898-b197-2cda009b0a61 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

FOLIO: Focused Semantic Memory for Streaming Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:49.930725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:49.930725Z digest=sha256:6f462af0552b3dd2c28244f793a69deea2267b94763adbd23daa6bfd0c4b0ee9

Observation 95031e18-2fc1-4cc8-a19d-4d6ae93e4486 · outbound

This paper cites Proactive assistant dialogue generation from streaming egocentric videos.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Proactive assistant dialogue generation from streaming egocentric videos

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:50.012286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:50.012286Z digest=sha256:d3cd8dca0a692d1b3b5dd733f5a7e0b21e0293b470f798500e0cbc94284434b9

Observation 13383cae-d94a-463d-8591-6ad804ea5321 · outbound

This paper cites Proactive assistant dialogue generation from streaming egocentric videos.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Proactive assistant dialogue generation from streaming egocentric videos

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:50.124779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:50.124779Z digest=sha256:c1a87fc49757f683ff451d38627a266c680a836336a76023cfdedd1df58de637

Observation d8576c91-72a3-41fb-8150-9c972abf6679 · outbound

This paper cites Hierarchical event memory for accurate and low- latency online video temporal grounding.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Hierarchical event memory for accurate and low- latency online video temporal grounding

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:50.208860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:50.208860Z digest=sha256:709d08c56298f93fbcf59e7ddd6407a2a81978399c51dd8500d45265ca2dd3ce

Observation 48330fd9-008b-4fe8-ae01-535a18d86353 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

FOLIO: Focused Semantic Memory for Streaming Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 85

Resolution
malformed identifier
no resolver link, observed 2026-08-02T05:40:50.290445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:50.290445Z digest=sha256:9a017c63b0f01f754bc8cadeb4addcceef0eaef63e6311141519d31fd192dc00

Observation 82d93ecf-9e6c-4778-af21-84326ec4c09f · outbound

This paper cites an unresolved cited work.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:50.366425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:50.366425Z digest=sha256:c4139f0d504106eec1be9259fe54c183255fea2690e4166028ed7388d3ed2932

Observation 4069fdae-cf75-47e0-89c2-97ceb6a820c1 · outbound

This paper cites an unresolved cited work.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:50.513581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:50.513581Z digest=sha256:75c8634f9085a4cf658799a14715dcfc0fa632d0583c28f102b7bc4bd89203f8

Observation c98fb03d-546b-45e7-9516-0b445e7f34e9 · outbound

This paper cites Put them in compact_objects unless they are central to the action, then put them in detailed_objects.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Put them in compact_objects unless they are central to the action, then put them in detailed_objects

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:50.635364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:50.635364Z digest=sha256:23a2383adc2adb10377daca52f3c7143c2242e06effd99468a152370405ab81b

Observation f2dec515-99bd-4cc7-ad51-458e468fc612 · outbound

This paper cites an unresolved cited work.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:50.741826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:50.741826Z digest=sha256:2acb0217c5c5d954a97cb21c35e5c66d81b87f9601845d58ab85005ef7b3a387

Observation 83608dcd-6269-4091-b55d-78d8a32d3788 · outbound

This paper cites time": "{start_time:.1f}-{end_time:.1f}s.

FOLIO: Focused Semantic Memory for Streaming Video Understanding time": "{start_time:.1f}-{end_time:.1f}s

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:50.833842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:50.833842Z digest=sha256:b2772b38c828b80759fa61d96c3d79ebfba6f0fa21a24835ce261a6ad681e590

Observation ffc1c6af-df20-4005-8c35-0f72d37fcb68 · outbound

This paper cites an unresolved cited work.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:50.951018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:50.951018Z digest=sha256:2cf0ec9c52a91526764f803211e3ab20596d41e347caeb95787d6f1dab21867e

Observation bc10390b-e87e-4a0c-ad50-e20eb6c7c966 · outbound

This paper cites an unresolved cited work.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:51.064442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:51.064442Z digest=sha256:9885cea86edd2671fe418d2b6de5f78113b642536880739a03df1ebcca74b701

Observation 6464a1bf-3f1c-43af-aab1-153e931dcfdb · outbound

This paper cites an unresolved cited work.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:51.138752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:51.138752Z digest=sha256:47d9bcc159f366f154ea329fdb62fe99c24b48621999fa00d75e69d892c311f1

Observation 79f02603-b715-4002-a1b3-f34bbcd06826 · outbound

This paper cites an unresolved cited work.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Unresolved cited work

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:51.282189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:51.282189Z digest=sha256:19f3291276e72aa4cce82ce3da908ba2e63e069bad74fac66e830de72ef4cb6f

Observation 5435aa13-3c92-440d-a7ea-4a109c422fc1 · outbound

This paper cites avid reader.

FOLIO: Focused Semantic Memory for Streaming Video Understanding avid reader

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:51.395450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:51.395450Z digest=sha256:ab043a9a3ecc12f1af2a921a1b3e50e2efe1705e26b990c454359981ca221024

Observation 485f210d-02cb-4933-bf51-c5e92d47abe8 · outbound

This paper cites avid reader.

FOLIO: Focused Semantic Memory for Streaming Video Understanding avid reader

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:51.488681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:51.488681Z digest=sha256:4fa5417f688125c38824a0b246bcebbc03b100d2aa2a24d791d86df79efc8960

Observation f802922a-0e17-406d-8103-77b0793b7d95 · outbound

This paper cites an unresolved cited work.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:51.563637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:51.563637Z digest=sha256:81473aa43322347fb1d13b6056a6bf38c3e692a2f7e115c8d27211d16da9c022

Observation 8f667726-c5b7-4891-9258-58b38841f508 · outbound

This paper cites relevant_object_ids.

FOLIO: Focused Semantic Memory for Streaming Video Understanding relevant_object_ids

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:51.617908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:51.617908Z digest=sha256:e33c8de83d389f933fdc315555640b018096a31210646701e09a5186c439317e

Observation 4e65542d-9431-493c-9ca0-ed74aeebb15d · outbound

This paper cites an unresolved cited work.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:51.727831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:51.727831Z digest=sha256:d874f3ba10e12e7a5b5c45e0d38a04ca10616fa29b7fdb08b9f8afdca597632c

Observation b3009591-3b80-42e2-b8a2-0e75e1f18968 · outbound

This paper cites ## (!) CONCEPT QUESTION -- SPECIAL HANDLING REQUIRED.

FOLIO: Focused Semantic Memory for Streaming Video Understanding ## (!) CONCEPT QUESTION -- SPECIAL HANDLING REQUIRED

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:51.790099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:51.790099Z digest=sha256:41e87cc5889a712c0d42d46f133ee25e570c961383768d065f51847ccd916e13

Observation 2328a109-6093-4c42-81ac-12da28d09b51 · outbound

This paper cites Recent frames inform CURRENT state only.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Recent frames inform CURRENT state only

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:51.856126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:51.856126Z digest=sha256:4ff53ef87e4e7525b5c05baf2b47af41ec6927bcfe90280ba1f8dbb7902cbf22

Pith citing papers

No inbound Pith citation observations are available.