Pith. sign in

Paper Citation Record · LEDGER

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

As of 21 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 11 inbound Pith citation observations for arXiv:2506.23825.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23825 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:37:08.992465Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:14:01.995575Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy54
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6887147d-f18e-4ce0-be0b-649e9ee97769 · outbound

This paper cites GPT-4 Technical Report.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:01.869344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:01.869344Z digest=sha256:d48a5177596c0318c0170089100e32b334bbb012a39567685747dea47c951ee4

Observation f951860e-258b-40dd-9b3e-695ac42e79ff · outbound

This paper cites Self-calibrated clip for training-free open-vocabulary segmentation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Self-calibrated clip for training-free open-vocabulary segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:01.948138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:01.948138Z digest=sha256:0f2ffd9250b1be65a6ced8123c09e636ac8abf5090d5c317829aeb48b14339fc

Observation e8c54936-2933-4df9-a915-0306862a540a · outbound

This paper cites Memory consolidation enables long-context video understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Memory consolidation enables long-context video understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.070349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.070349Z digest=sha256:6e3627e86a0194a6a51e3da2a84433aebbbe0808b99ffd0dce6fd36e6002b7d0

Observation f1144657-4b06-4101-b70e-f10325e7f806 · outbound

This paper cites Language models are few-shot learners.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Language models are few-shot learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.199304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.199304Z digest=sha256:2d010ff4d4b63925f5deb72f431cac8e87990173da460aa9724460f63b691a6a

Observation ed047457-67d2-493f-b1bd-4a5a5b54c0da · outbound

This paper cites A Memory-Network Based Solution for Multivariate Time-Series Forecasting.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams A Memory-Network Based Solution for Multivariate Time-Series Forecasting

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:37:09.602810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:02.378276Z digest=sha256:8457211b1dcaa8d08ee475078fc83796aba22e0483ddf00f5b1c0a194e385b44

Observation caa6d757-081e-424c-9949-076661df0e62 · outbound

This paper cites Distributed deep learning model for intelligent video surveillance systems with edge computing.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Distributed deep learning model for intelligent video surveillance systems with edge computing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.499259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.499259Z digest=sha256:9ca356d8f13d4c458f5289176231105adc12b4169e4c8c1fff1c09e320901c0d

Observation aafc111a-b9fa-4326-966c-e495ab223121 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Videollm-online: Online video large language model for streaming video

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.595350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.595350Z digest=sha256:f8eb1359bd7655fd0003f0a099737409caf156175ce11ee2de6e80ccdd7a434c

Observation a3ad585c-95d6-4593-ae46-fc2760f3b7ac · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Sharegpt4video: Improving video understanding and generation with better captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:02.735952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:02.735952Z digest=sha256:604ac099ea5fe5fb603027e9a8b9cc315f1b686c4835847d783abd858bd5f737

Observation 7f9f6b6c-3478-4537-9fe5-05fe9a3a0293 · outbound

This paper cites Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.720770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:02.844901Z digest=sha256:2aa863da9a7df1cbc04500e213b26949177c562cec377e853b704a9e384412f0

Observation a8c5c172-dc6c-4e12-afaf-1b283327193c · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.023513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.023513Z digest=sha256:98b71e5b97ed4e9dee521d5b228ca21462122db87d22ecf6335190fe8b0ecea4

Observation 0fcbabcd-d5a3-40ab-b2f5-a01ba03c21a5 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.704419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.104975Z digest=sha256:d49247fd1931b2c2ac891099bc2f1d99c5255b33eb09460dd9be115c1d2834f0

Observation 522ef2db-b67c-4aa7-924f-78c744c54477 · outbound

This paper cites Flashattention-2: Faster attention with better paral- lelism and work partitioning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Flashattention-2: Faster attention with better paral- lelism and work partitioning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.689071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.209972Z digest=sha256:7f2561f278b89bfcd299efeafa3666e5999f6cdbb4a001e580ce9683937ac5af

Observation f8ab067d-b2d0-4b68-ace1-74339480aca4 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams An image is worth 16x16 words: Transformers for image recognition at scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.307705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.307705Z digest=sha256:de662db9c9db6400d5d883ac33d1d92f4753c1b4e7dca65590fb05a1dbe7edb5

Observation 2f202760-4744-4c6a-aa31-8f2171aa0c05 · outbound

This paper cites The Llama 3 Herd of Models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.374979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.374979Z digest=sha256:fd01615596ade957b6ab38842e0774e6ca185e9cba5f81f4e566d2d68becfd1e

Observation 53b89593-713c-4fb7-b545-9c19be088905 · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.662907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.468755Z digest=sha256:24dc915a5739eaf9b456d9224c6faabf42535aeb43344af4dcb06c872ff0c92c

Observation f0d1da68-38ef-429f-8509-457ba7bebf19 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.648958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.534615Z digest=sha256:5d921f7dba6de0439f96cd45023f28ccad7f7d680275a1dd8add966f34cc80fa

Observation 726d63d5-f4ff-4c22-bf44-8fd0f11dc009 · outbound

This paper cites Temporal sentence grounding in streaming videos.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Temporal sentence grounding in streaming videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.635328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.687221Z digest=sha256:c224770bafc7ea5161e78a10a1027f6bc14af607eba71232e2a8bbcf54e7d2b6

Observation 49749f40-590f-4bee-9800-da523fe23e36 · outbound

This paper cites Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answering.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.621253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.796093Z digest=sha256:56358420c2e5bc3dd164cac6b62aea982d49c38c858acbf8f15c699420b25bdb

Observation f38e2cb4-da5b-433f-92b1-b18c187df245 · outbound

This paper cites Clip- adapter: Better vision-language models with feature adapters.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Clip- adapter: Better vision-language models with feature adapters

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.607482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.874150Z digest=sha256:cac98a18d22aee9c2069642b35132ef49a0dd509911ada24aeb0e4f400b32912

Observation a7e01be7-e4e9-49f0-822a-ac95500b5f06 · outbound

This paper cites Frameexit: Conditional early exiting for efficient video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Frameexit: Conditional early exiting for efficient video recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.593252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:03.969979Z digest=sha256:fe3133153da5a6f3093b7ae481e80f9b38c55bf8cbcc63a6819e062d2b53a584

Observation 9b2576bf-59fc-4109-8f36-de56f08fe307 · outbound

This paper cites Dynamic neural networks: A survey.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Dynamic neural networks: A survey

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.580297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.043814Z digest=sha256:1adcf7760c47ef32a1b94c2f57acb3287df877b91dac8c355c46334e81e9278a

Observation 908cd346-dabb-4b80-93c1-c82eb7e70c04 · outbound

This paper cites A twofold siamese network for real-time object tracking.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams A twofold siamese network for real-time object tracking

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.566930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.091875Z digest=sha256:6198e3c4d9a9245fbd5f87543657590e90795ddc9d3d31bbac1418c35ff63ded

Observation e0f959ed-ecc4-46cc-b449-3d955ef8b133 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LoRA: Low-rank adaptation of large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.554209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.202477Z digest=sha256:c9316b545dd6270fcdf354b1add57aa8af60a2156c6ec840047a16278f5fad7a

Observation c2a9f090-f48f-4e98-b09c-0c5cd7f447a8 · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Movienet: A holistic dataset for movie understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.541398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.321669Z digest=sha256:7d5b3488dc2c33bebf290db04f1b31a3dae6bf7491f42a0a76dacc419bc9a544

Observation 461b260e-ed3d-44a6-9ecc-b9003b49afea · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.527716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.437980Z digest=sha256:b116ef4dff9e0c9b697a25bcba688864f786a3a8bda7d1325647f498b64dfc40

Observation 20d5ba49-83be-4a5b-9471-84ffaab9a37b · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Seed-bench: Benchmarking multimodal large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.512597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.531283Z digest=sha256:df147453254a6233614845e6659cbebcd66e27f1b38c341d0933e50629184a67

Observation 2bb99f58-390d-4cf3-a8b6-fd826d376847 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.625398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.625398Z digest=sha256:5b911164024337b79fb9bd1fd9f683c76f44ed364d4db895db7362697aed8e80

Observation c647090a-7381-4c06-bed0-c3a1c5d8f60c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.498813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.695915Z digest=sha256:15fcab425089f940f522c21db160cb8a8432782d7d56ccc13ea97b54290cad30

Observation 391d44d2-f142-4fd4-9d13-ea37546e6ca8 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.484102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.769774Z digest=sha256:671a7c6933efc44cbe1a913d51a6fb84b0f790f3b5f4d03d3b4f3353f1180715

Observation d3073307-a94a-44a6-b442-c23c61166907 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.470259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:04.874604Z digest=sha256:c4f93ecba790b20bb7879bf7694f66012e31b25d85e54fb6904df6ef813b13f7

Observation 82c186d8-45f2-428d-b58a-35d1f51e6d15 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Llama-vid: An image is worth 2 tokens in large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.456649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.041607Z digest=sha256:f432459441d368f36feb3891b47c010d5702d2f7832c8b7f571cdc8ea0acedc2

Observation 17555231-f16f-4331-b585-71124f0408f2 · outbound

This paper cites Visual instruction tuning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.444108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.162198Z digest=sha256:4f3a3eb2f02555c4f58fc11cb6372b964224054320b135f842fa69b4b130d1c7

Observation 1d17eb93-f65c-41fa-9216-1e97d9c997c0 · outbound

This paper cites Improved baselines with visual instruction tuning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Improved baselines with visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.430583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.248330Z digest=sha256:d7462e482f03bec540c16688cd411432aaea14a76239d13ef9e19bebb4aa3e15

Observation 0f4f6773-ee34-41bd-bd60-e24713ee2a90 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.338028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.338028Z digest=sha256:20593af0a2220211a167227b5522e4ff1e3ed8b8b1ce77fc347496422b5d84f6

Observation fe3eef6e-9ba6-4747-9420-16d2f1d9c9b1 · outbound

This paper cites Learning quality-aware dynamic mem- ory for video object segmentation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Learning quality-aware dynamic mem- ory for video object segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.415628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.434941Z digest=sha256:461f428df25439b4867ed125a20e09419a03bc3b5242ca4d53d6c088c3cd68bb

Observation 181b243a-8ba6-427b-a710-d3b95a117f09 · outbound

This paper cites Universal segmentation at arbitrary granularity with language instruction.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Universal segmentation at arbitrary granularity with language instruction

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.401854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.607752Z digest=sha256:876f857ad08de1af275a2d5d77ab064f203d8203773f7e5c278d48172a52e739

Observation 9a7864cc-bbaa-41c0-bd7b-dca40e139fa4 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.734732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.734732Z digest=sha256:85a629bd4811d8cefb6d76159645fbcf4637d251adcabaa40efcd1898ed656a2

Observation cc6f9ea1-d404-4abd-b1e3-0c16b116cd10 · outbound

This paper cites Soc: Semantic- assisted object cluster for referring video object segmentation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Soc: Semantic- assisted object cluster for referring video object segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.388612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.827409Z digest=sha256:314920f1d1a1a03ca336d6d90b905374cc8a78b707865b2853d03e976e031c51

Observation 8f68580d-5ba6-40b9-a3d8-7caff45da478 · outbound

This paper cites Multi-task deep learning for real-time 3d human pose estimation and action recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Multi-task deep learning for real-time 3d human pose estimation and action recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.375454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:05.934945Z digest=sha256:fc800b1943e6c49e4be1cc29f03cc466772d08f2ad9fe594f5e5f99b06a43c7d

Observation 22555c9a-deb7-4a70-8ce1-2f4a2a885b9d · outbound

This paper cites Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.362008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.023144Z digest=sha256:6ec24c060e65752f7dd610389b59a239781eff8d86c199314ac5179202a278fa

Observation be611fb7-0e21-471b-8211-369fb2e96bde · outbound

This paper cites Video-chatgpt: Towards detailed video under- standing via large vision and language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Video-chatgpt: Towards detailed video under- standing via large vision and language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.348389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.087437Z digest=sha256:7cd68f25a35f6f456dbebece67f5291e8e3266fa48f253c5555071e622d4a8d7

Observation 41fb913b-96b3-4fec-9d33-77a80d89d812 · outbound

This paper cites Some methods for classification and anal- ysis of multivariate observations.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Some methods for classification and anal- ysis of multivariate observations

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.334609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.189654Z digest=sha256:7506ac0abe0f1a9bdc080c8bc69886e20bd102a523b65157548d9103d39abc55

Observation 7f33954e-d803-4a93-a719-e53fc98c3706 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.320839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.280529Z digest=sha256:4492118708223d21976eeb90f4c763278b4bf56436e784f4ecaef901abc7223f

Observation ba0da123-c123-4aa5-b457-02ca53d8da66 · outbound

This paper cites Deepres: A deep learning-based video summarization strategy for resource- constrained industrial surveillance scenarios.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Deepres: A deep learning-based video summarization strategy for resource- constrained industrial surveillance scenarios

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.306067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.351008Z digest=sha256:fa9483dee4ed7411930b6d55c4a984683c1730bf0ca9f246dcae33331d5d99e4

Observation c70825d6-5f2e-4b0d-948d-316be74c920a · outbound

This paper cites Training language models to follow instructions with human feedback.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Training language models to follow instructions with human feedback

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.291436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.424661Z digest=sha256:8e0c24cc7002d6f4d84f87913f53c7ce8617ec8a0aed24f87549586786eb2181

Observation a26de3d5-160b-42bd-b8b9-9b830f97177c · outbound

This paper cites Streaming long video understanding with large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Streaming long video understanding with large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.276943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.479104Z digest=sha256:1fe0fe940152be6ed21bc6316be32bd52aefafb66b5fa65e0bc9d7e945f895ff

Observation 4b17bef0-607e-4494-8825-a72a1012c341 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.263120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.539428Z digest=sha256:67f65b54d2fa83506dd7861771816494d3039e8d2b3109e35ef60d80dc521151

Observation f2a3b187-e579-4d79-abc2-17de6c4e4e55 · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Robovqa: Multimodal long-horizon reasoning for robotics

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.248033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.601443Z digest=sha256:33ca6aa0ad6274200853e121d16463db21310fd1be21afcaaed58ee4e16becf0

Observation 964afe38-b9f3-4395-b522-4bdd7d6198ae · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Moviechat: From dense token to sparse memory for long video understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.232186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.651780Z digest=sha256:85a83d2466da061d87d9b2ba92b9d5890507f88fbe1c6e9b5ec3d9978194f915

Observation 53fc9ae4-2eb8-4020-acdd-603aa9cd8429 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Roformer: Enhanced transformer with rotary position embedding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.703267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.703267Z digest=sha256:288cff09d094e651ff1c9bf13c1a0c759556950af8b7e38729af7872161df732

Observation e2563ac0-4c65-4138-859c-d37055ffaebf · outbound

This paper cites Tracking as online decision-making: Learning a policy from streaming videos with reinforcement learning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Tracking as online decision-making: Learning a policy from streaming videos with reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.202961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.822196Z digest=sha256:a8df2a3a8e789f9de745556d026377802c30ed2ba6f0ee810095365d6c456488

Observation cfcecbf7-a784-46a9-b8e0-7978fc80d6fd · outbound

This paper cites Dynamic memory based attention network for sequential recommendation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Dynamic memory based attention network for sequential recommendation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.180383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:06.916665Z digest=sha256:cd968b377dbaea6cf184a6a372157d6c6e6483947306a4a408f981e2704b8896

Observation 6125cca0-ee24-438a-a6f3-d2b3b76defe7 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.008868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.008868Z digest=sha256:2fdad985c291d518f017b7a3f73f0a80cf730b90e21e50f1673a304ddce0d7ed

Observation b37c7254-c536-42ff-82bb-7afa328b9445 · outbound

This paper cites Kimi-VL Technical Report.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kimi-VL Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.074611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.074611Z digest=sha256:09c5731552182063b1c052fbbe246921e38f2260fc92edee157265855dc06b32

Observation cb49f620-fa9a-4472-9dfe-870c0cb1ac04 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaMA: Open and Efficient Foundation Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.127746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.127746Z digest=sha256:a79137316bc89039e286d9c8753b595b0aed111789687651ef844d025b2985b1

Observation 87ecba3e-14eb-4e4e-a8ff-8721d5970736 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.188899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.188899Z digest=sha256:95591efe86d75b6dc9c3691e57c1c5d45855a476787ed61a3da7821f61de1b46

Observation a350224f-da1f-4b50-90cf-84aaac305b3e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.239662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.239662Z digest=sha256:b008e92312d3d58584e07c67ef02c302cf84da50d0ca9a515defb9ccdfb5f58b

Observation 77574e63-3d4e-4b9b-adf3-d410f4e6881a · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LVBench: An Extreme Long Video Understanding Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.288161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.288161Z digest=sha256:92b2415c94e5a20a26fac06559b8aa28e512f73845dea56a9514dfadf7024442

Observation 242fa570-3d84-43c4-8f27-c01f37a67a4e · outbound

This paper cites ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.383398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.383398Z digest=sha256:968f7010ce9ae76a17fd358f7c9cdc431e4f789314b0304b25d6d84cae9a606a

Observation 5f9bbaf5-76ae-4e83-8855-cd7253172285 · outbound

This paper cites Adaptive focus for efficient video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Adaptive focus for efficient video recognition

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.162546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:07.498070Z digest=sha256:c04cfb4831db1dfc406391ffeec41deec6ef5f4f5c49b857a9549b62dccd570d

Observation f9a7a55a-0646-40da-a775-58ee96e5faa1 · outbound

This paper cites Adafocus v2: End-to-end training of spatial dynamic net- works for video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Adafocus v2: End-to-end training of spatial dynamic net- works for video recognition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.146070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:07.576870Z digest=sha256:6667a960e28c42dd98ce244d8efc2d758edddc5d924746b3a0cf5f978239a155

Observation 81c0ee5e-a38e-419d-a8b0-f15e86e45ecf · outbound

This paper cites Adafocusv3: On unified spatial-temporal dynamic video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Adafocusv3: On unified spatial-temporal dynamic video recognition

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.124808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:07.596032Z digest=sha256:fe7299b57b639d49327bdb8b6bb94dd42c0510e6cc8d6d9a61d9d447bfd10010

Observation bbdda88f-2070-44df-9029-f53700619c6f · outbound

This paper cites Hierarchical Memory for Long Video QA.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Hierarchical Memory for Long Video QA

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.699169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.699169Z digest=sha256:4f65343894428ef2745825e8d7f98023a9689d8cbdc5e3b33e376c1c0e441297

Observation 7c31022a-e811-44a1-adfe-b975d579c6d7 · outbound

This paper cites Ponder & Press: Advancing Visual GUI Agent towards General Computer Control.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Ponder & Press: Advancing Visual GUI Agent towards General Computer Control

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.848372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.848372Z digest=sha256:0087d42366c18ecfd579f1262a3cf176fec50b5dbc037b4e5b5bac666ec96b3b

Observation 983783b2-0c3d-4687-88f8-c2bab012ce25 · outbound

This paper cites Uni-adafocus: Spatial-temporal dynamic computation for video recognition.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Uni-adafocus: Spatial-temporal dynamic computation for video recognition

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.110255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.015864Z digest=sha256:435610b7ebd72e9927f2c90d062624c7a4150c5de0c7395530af814793402197

Observation 05e41abc-eb7f-4d65-90f7-facf52cbcefb · outbound

This paper cites Iterprime: Zero-shot referring image segmentation with iterative grad-cam refinement and primary word emphasis.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Iterprime: Zero-shot referring image segmentation with iterative grad-cam refinement and primary word emphasis

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.096179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.185558Z digest=sha256:9fc9050f798d47d3230f59974debe67f3c81dbb361ca0f5d76958c3606a48518

Observation 108b516f-0500-4305-a3ed-f0802abfa1fa · outbound

This paper cites Sam2-love: Segment anything model 2 in language- aided audio-visual scenes.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Sam2-love: Segment anything model 2 in language- aided audio-visual scenes

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.081701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.352293Z digest=sha256:733f3acb446dd432702f02ac8d87307a970c4f1b382f1244ca33d1ce7819ab85

Observation 9f00587a-b7bb-4c24-812d-f4a16f01545a · outbound

This paper cites Towards real-time multi-object tracking.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Towards real-time multi-object tracking

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.066009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.520067Z digest=sha256:65496786992d027df28dd7877b6c84c22efd9dac2a9eb316ead49d5db9ec0f4b

Observation 66990da4-9ee6-4cdd-894d-9844412dd7d1 · outbound

This paper cites Videollm-mod: Efficient video- language streaming with mixture-of-depths vision compu- tation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Videollm-mod: Efficient video- language streaming with mixture-of-depths vision compu- tation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.048899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.686159Z digest=sha256:dab34b2a8d6b887525d85aae1f60302f9cb71dc5db64489afeb24202eae3a02a

Observation f474a2e7-700c-40c7-beab-a09b604054f4 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Next-qa: Next phase of question-answering to explaining temporal actions

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.029636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.802065Z digest=sha256:c423c0907fe761e795ffa723258dc305303f6392a06c102e47443f351dfd4659

Observation 93c29c6b-51ce-4a07-9a67-99f57cf17501 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.873755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.873755Z digest=sha256:ef3c08a9e769b55d4df35e070a4d0541607ce7d0ad0eb7b61724db0e332bedb2

Observation 9eb08b61-673f-44b1-8016-d638a7fe5d3d · outbound

This paper cites Fine-grained video captioning via graph-based multi- granularity interaction learning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Fine-grained video captioning via graph-based multi- granularity interaction learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.012337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.879930Z digest=sha256:2703508e365d54efcc400cc6e3f66a41631d65dc4345ec646b173414704f587f

Observation b050451d-f53c-491f-950b-74967d66ffce · outbound

This paper cites Qwen2 Technical Report.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Qwen2 Technical Report

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.890274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.890274Z digest=sha256:57a208a2f1ab99126978c6175176158899821246fb25efa87e92dd8041d9578d

Observation fbe6b9ba-319a-4d92-9783-f9ee5860840b · outbound

This paper cites Language-aware vision transformer for referring segmentation.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Language-aware vision transformer for referring segmentation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.994052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.900750Z digest=sha256:3c577c40877ab519ebcd1f10434fc3d2983f40887a41f3eede6fe52f0fb3458b

Observation 1cf98fa4-906c-44af-af7c-e0594050c865 · outbound

This paper cites Atp-llava: Adaptive token pruning for large vision language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Atp-llava: Adaptive token pruning for large vision language models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.975487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.917306Z digest=sha256:763b892a09263c4b54d82f8610b48cd77b36d88e871b47b3003e5b4fc6f7d5d6

Observation d93939c1-ba4a-49c8-bdbf-69ed035ef2a7 · outbound

This paper cites V oco-llama: Towards vision compression with large language models.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams V oco-llama: Towards vision compression with large language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.958375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.923634Z digest=sha256:3841e2fd04e356ab7f33afaa60b761809781e3ac8b7875e89870d9eb83bef8dc

Observation f40af02a-f349-43ba-b593-e05875441c11 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Self-chained image-language model for video localization and question answering

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.933447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.934527Z digest=sha256:4122b66fb4db7030e134c43094df0d07345562a4b0198a2d538b55ef48be03c0

Observation 5708aa3f-d747-4b72-9ae2-314e41a9b0a1 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.913146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.942219Z digest=sha256:f626b16e2d1aa0809acf1ad4646c65b5e79e09bde7d4524ffdb129eab52fedee

Observation 3b095665-9420-4c13-8203-70c5b0f8f2b4 · outbound

This paper cites Real-time action recognition with enhanced motion vector cnns.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Real-time action recognition with enhanced motion vector cnns

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.892300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.956859Z digest=sha256:2d597434659f8a139dc9ce9030538e2cb08ae39b2fe7f9e25ebcac0ccec5bca7

Observation fafdc095-d006-47ba-8474-889537b0e4ff · outbound

This paper cites Long Context Transfer from Language to Vision.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Long Context Transfer from Language to Vision

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.962828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.962828Z digest=sha256:7bcccfb5dd2a31a7ba01abce1bd2576387657b07df49b5a6bde6ba5559fab682

Observation 113072f7-b149-48b4-98a8-8e9f4e1254e5 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.967515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.967515Z digest=sha256:31bc672253daa2549878801602877bb0bf9410c3abc73433e4159c8c09eee183

Observation 52e61739-f8b5-43b7-a14c-62e6067bddfc · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams MLVU: Benchmarking Multi-task Long Video Understanding

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.972286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.972286Z digest=sha256:5f85de197d8d704eece507c3a22cba7c69cebfea1e98c1da1d2b3996d45eb628

Observation 2c8153d6-b46f-4322-b8f9-89626a8f23b0 · outbound

This paper cites Streaming dense video captioning.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Streaming dense video captioning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.875831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.977508Z digest=sha256:9a4c05079a912f15a4fb4b88c3433fdcf1cfb9dcb708484ae1a344e833e66815

Observation def39505-06cc-4ccf-baed-4dff0a300a59 · outbound

This paper cites InstaRevive: One-Step Image Enhancement via Dynamic Score Matching.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams InstaRevive: One-Step Image Enhancement via Dynamic Score Matching

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:37:09.053145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.984190Z digest=sha256:eb1ec28598d3d9d49de39c6a42b00fd6ab989c63650c89e9a5550ce72ecaad9a

Observation e10025c3-8e4e-4ce7-8f23-a47b69e63cce · outbound

This paper cites an unresolved cited work.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Unresolved cited work

Reference 1000

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:37:09.856397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:37:08.992465Z digest=sha256:ba7430edcee95c8f53ab6a3c3a127b8cda971c225a04cbd95dbb2a49369195bb

Pith citing papers

Observation 5d958aba-4e18-443b-bce7-01c2b8959940 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:35.208400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:35.208400Z digest=sha256:008353b57d835877f411e846389027d28e2068e317f3ae86f3984e69a6573d7c

Observation c89077fc-d844-4efe-80f0-208470118693 · inbound

ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion cites this paper.

ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T16:14:01.995575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:14:01.995575Z digest=sha256:3fa8f2bcc825296fb5f264da8d88a4a6707eed4b5ea4ea87107b7fb9ee5db9ca

Observation 011ffc0d-af10-4df5-b428-909d0d7ea02e · inbound

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation cites this paper.

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T14:32:00.456065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:32:00.456065Z digest=sha256:f1f7b8da876dcda302c46e498683da8e173a3b5f686b5cbd7f846aa4b880e41e

Observation 6f3dff6f-4043-41f6-b592-79e133eaebbc · inbound

Mosaic: Cross-Modal Clustering for Efficient Video Understanding cites this paper.

Mosaic: Cross-Modal Clustering for Efficient Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:10:34.388472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:07:36.133404Z digest=sha256:f9e13170f303f9099eec82eb9a4195d157c276e2a6c1e7b4bd7aad84f8ae3521

Observation ce12786c-d594-42aa-8810-0d56971e0a6c · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:03.709919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:1d778c7813291ddaf2781ada45eee478e123e15d85b9fe6c3e3e0548687609a8

Observation 78041cb7-9628-4162-a2e3-c47795275266 · inbound

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning cites this paper.

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.780231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T06:51:52.861981Z digest=sha256:65e0254604af671398abf4341c9e43053336cbec4682b96bdb052c628b2ecee8

Observation 44ceb307-4b36-4750-bc72-23d3b53ade81 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:20:56.393925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T02:30:55.939351Z digest=sha256:238ad1e167b0f8e856c82af1930e03ca924719ccdf4e8e3c7eb95e15274a3520

Observation cf4b5ff5-efc3-4be8-997b-d36a18201cae · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:17.733473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:b20e457d1189e8f424c7dab7a6f2f3d0686bb5854f8996ea82c8c219b94bc163

Observation 54e3e141-b2f2-4945-be55-c0e5ffa8023f · inbound

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering cites this paper.

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-06-28T02:01:29.198189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T01:52:33.768494Z digest=sha256:3f7f98800725618ab0d24504a81f009dfc841a4db5e1a08c19912081b0919705

Observation 061c197b-c067-4d9a-bdf5-0e5971b2fd3a · inbound

Harnessing Streaming Video in the Wild cites this paper.

Harnessing Streaming Video in the Wild Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:37:25.733342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T18:47:55.910417Z digest=sha256:5638c2fa87a2b1facd0c4ecb3b4f2f1e6ac3aa499a9ff6ae2ed1d099b08ce490

Observation e4615993-e9ff-4785-81b1-0e0ca671466b · inbound

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation cites this paper.

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-14T04:39:27.358904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:39:27.358904Z digest=sha256:7abc919a8f3b84b3698e774fdd2ccb48d5775512576bed2aae1c0f1b582b0d14