Pith. sign in

Paper Citation Record · LEDGER

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding

As of 6 August 2026, this Paper Citation Record lists 100 of 120 outbound references and 1 inbound Pith citation observation for arXiv:2606.06991.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.06991 v1

Coverage vector

measured 100 of 120 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T22:11:01.690237Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:12:46.295455Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T20:38:56.159073Z

Reference resolution

100 of 120 outbound references displayed

  • verified exact43
  • verified fuzzy0
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e44ff72d-fe3d-44eb-90f9-764ea821fa25 · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.876383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:48c778a48131b704836b3126a3a259c1703d0e27810b088764a911944f8b7e97

Observation 1a891852-1a2d-428e-816b-60dd59de663f · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.864217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:0c24eeab55f2ca1efd4dc319a947d31cb8ce35d12df5095bb63bf98f619e06d2

Observation 51ee5fd0-7ae1-41e2-a6f0-8d1600b3246a · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.795455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:7d51dad474799466e5f29b0e55940248ef11a2704cd684885e26dc152af3174a

Observation f217b3a6-fd3e-4fae-8576-c32e802f9c1f · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:0f2614b02cfec51db0eca9dccbf4058bfaa78d02b09c16f90729b253f55b8ac9

Observation 783c631f-84aa-4d63-91a5-94c250fdfd21 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.866731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:cc257ed0a6e89e9bc51c41b1b6e6dff22e2ba968de898ed4ac6d687374d50dd5

Observation f041428e-de8c-40d3-bb42-9bdb76454b86 · outbound

This paper cites Cat+: Investigating and enhancing audio-visual understanding in large language models,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Cat+: Investigating and enhancing audio-visual understanding in large language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:b986776beb17709be132734d1c1e2b4c4cabd0a91cb62ce46076cdb89894556a

Observation 347ae490-2311-4983-b3b0-10409177c9df · outbound

This paper cites Motionllm: Understanding human behaviors from human motions and videos,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Motionllm: Understanding human behaviors from human motions and videos,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:f2a4ab7b2b38b801dec85ab333e3881520384262b2be5c359f638ed325ca6f3c

Observation bca48a04-5b63-4239-8ebb-2a1e2178ad89 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.809524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:1e35fc30b37a0825f9e8d4022b2a7cde6846ece631ebb59bf8037a0ae034c4f1

Observation 687ae146-f056-42c1-bd57-d3b9df6a016b · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.842034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:abc8e8a5b51632c527dd98c99e525fc57f1ae6319ae124f430643c41d5cc42d1

Observation 56177e63-364b-4896-b1c1-32a93ade46e9 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:9a7313b45318ac8175a5c9eaf61cbdcf10cc59d08294ce4ceeb6c632a8da800e

Observation 9c83236c-1277-4bf2-91ae-6d15731c4a4e · outbound

This paper cites Hier-egopack: Hierarchical egocentric video understanding with diverse task perspectives,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Hier-egopack: Hierarchical egocentric video understanding with diverse task perspectives,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:6a3c0d8c77e64d181d1b417b102608f355164b10f75fb62f312596985d74fed2

Observation f17c9624-dcc6-4b32-9028-2268977d50fc · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.826114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:a0343ffa53418db863fac2fbaf9c122deb8c4b0145355e6f43e85766687546ba

Observation f5be6b5e-1180-4d3a-8fdf-bff484473869 · outbound

This paper cites MovieChat+: Question-aware Sparse Memory for Long Video Question Answering.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding MovieChat+: Question-aware Sparse Memory for Long Video Question Answering

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.889758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:0fa26e272efb5feb19a1bfa2dfa45796cad543dc347829708a095696035d31ce

Observation d363a032-1ae2-48ab-bfaf-c9fb289d84dc · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:666c59d580b96aa6bad8d8ec4744cf5c451700dd7ccc0dfba12de789c13f7c9d

Observation 8173de8a-3b4c-4445-84ba-8ab655a6aa64 · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.855632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:33ea0a49b43a382c6f715bb6f62c37303017b0a4ce588c77f66610079e4304ed

Observation 81fc4cfe-e144-4d78-b1f4-5742af59720c · outbound

This paper cites Longvila: Scaling long-context visual language models for long videos,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Longvila: Scaling long-context visual language models for long videos,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:434bc144ec6b1d1578c646f45e8438157e134cee8da182994a577abc8c5555bf

Observation e27fd12c-64af-4ba7-88b3-6b57edf869bb · outbound

This paper cites Long Context Transfer from Language to Vision.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Long Context Transfer from Language to Vision

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.876722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:0e771e6e0e277ef503edaf7c97e624e66b84abfe7508c0c18703f075ce53ecd9

Observation a8bb0fa2-d1eb-412d-8eff-688c1ea46bc7 · outbound

This paper cites Momentor++: Advancing video large language models with fine-grained long video reasoning,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Momentor++: Advancing video large language models with fine-grained long video reasoning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:1e7f4708ff24ad117fd264a888979f850202005db094e8a3b5f41183a3dfb6a5

Observation be19c702-97c4-4b6c-8744-afab3935855f · outbound

This paper cites Selongvlm: Empowering long video language models with self-corrective clip selection,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Selongvlm: Empowering long video language models with self-corrective clip selection,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:3d6919c9deefb0a325571d65013ba86723b728128e99f3a3981df32a45da4460

Observation 08d4febe-2674-4aed-86e3-12008e935641 · outbound

This paper cites Ego-r1: Agentic chain-of-tool-thought for ultra- long egocentric video reasoning,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Ego-r1: Agentic chain-of-tool-thought for ultra- long egocentric video reasoning,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:dfe118c02bc61e87f60fe5dd2e02ec59d793f3e2ca6496429a12dfd8be506ac9

Observation 22a8f6d5-0ef1-4327-95fe-cdab7aaffdd9 · outbound

This paper cites Interaction methods for smart glasses: A survey,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Interaction methods for smart glasses: A survey,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:fee796b9e45ad9ad8f7e78bbf787ea6ed1b30dcc49ba585e1c3ddf9171fea5f7

Observation 3502ee20-345a-4410-84ac-c55a95aaf25f · outbound

This paper cites A head-mounted three dimensional display,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding A head-mounted three dimensional display,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:96b7eee36175ea0933700797bd53ed3ddf1c298ae422f117752793958d06bbde

Observation 1e52514e-baf1-4fd6-a96a-91555d5b7deb · outbound

This paper cites Horn,Robot vision.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Horn,Robot vision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:0a0f3389daef9d4171e0d886c5220ec8d079b5e5b7e8a80bdb6a1158c55d56dc

Observation c56ad348-8324-4a9c-b77a-9d7e04a6650c · outbound

This paper cites Videollm-online: Online video large language model for streaming video,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Videollm-online: Online video large language model for streaming video,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:78304c885849f4c3538d02cc581ba39cd88a571303826414f12b36812fc4cd27

Observation f4bf46b1-3444-464b-9d89-526ff4284aea · outbound

This paper cites Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:7b75969e1eded840114887e546bf6d98d3b25179ea45dedcd6aa35141b50144c

Observation d7174b84-30a2-46e4-9cb7-35dc0e6ed94e · outbound

This paper cites LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.842378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:11061b71df643aef24650705fcf6da0b7b2d84c629a6205ec996d21ca4b49061

Observation 1230205f-9be4-4540-aa55-7e786b5f8d94 · outbound

This paper cites StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.869265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:91b034d6e794ff9d43134a9a00ae87cdcdcabcc50274b95049ea875320aba997

Observation d66e9752-c647-4405-a5d7-dac30b4e6cbe · outbound

This paper cites Videollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Videollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.858844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:019dc8e5c51d2edbaa84ff32468967e9214c5c7e8386e66abf148d237cbbcc21

Observation 46820132-3a17-4a48-adab-3482f65d276d · outbound

This paper cites Streambridge: Turning your offline video large language model into a proactive streaming assistant.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Streambridge: Turning your offline video large language model into a proactive streaming assistant

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.894856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:97534ea9c1a615ae0cfc585bc93b6e165fbc620ff9e5cd1a6912c423f6387f3f

Observation c8ac6f70-b6fa-405c-b10c-9647e0ccc7a4 · outbound

This paper cites Streaming long video understanding with large language models,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Streaming long video understanding with large language models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:202d1e6137091f60e96a6b77d8880571794f6a68da5fcd532c0c0779a1cd84bb

Observation f55d66f6-6600-4c4f-bba3-e2673b0c26db · outbound

This paper cites Livestar: Live streaming assistant for real-world online video understanding,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Livestar: Live streaming assistant for real-world online video understanding,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:9e4d3328db602b355bef99ef707be5263c0d1613708701faecdd6ce861d98225

Observation 6227ce77-3ec2-40de-9470-32c1250b84ee · outbound

This paper cites SpeakStream: Streaming Text-to-Speech with Interleaved Data.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding SpeakStream: Streaming Text-to-Speech with Interleaved Data

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.868675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:a422423841cc86a3034a9abe7786b062186c0f8dddcff72aa9f0d10e4ff2e645

Observation 958d7d40-c1d3-4ae5-bc13-ddf1ee987a11 · outbound

This paper cites Soccernet: A scalable dataset for action spotting in soccer videos,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Soccernet: A scalable dataset for action spotting in soccer videos,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:ed8745e9183f3d6c92b771aa10c2d4cfba3c4cb70f096fd255f7d7673404180e

Observation 7c3403eb-813a-4ae9-88b6-6fb758406a17 · outbound

This paper cites Generating live soccer-match commentary from play data,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Generating live soccer-match commentary from play data,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:8249a59534caca0549bdf834781d11d1b6ee7f4c214567490918b99c93d19188

Observation 16658cd1-52b2-4661-9f34-52011b235bcd · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Sharegpt4video: Improving video understanding and generation with better captions,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:3b00c656ee812a4435df7f4fac93308ed27233d37b81838721def5042d4990d9

Observation ddc228c1-894e-424f-87e6-7733a20abc00 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.811669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:1c00866f9cba11a09f2f22a543ee18e3f00698d80e2b877a940290c82caf9e9d

Observation 8c1932a5-5ea0-49b5-ac6b-e7123ff82066 · outbound

This paper cites Video recap: Recursive captioning of hour-long videos,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Video recap: Recursive captioning of hour-long videos,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:65bf605a3e89dd70c61abcf3ae0626629d57166ac1aafd684798fad98b39d46d

Observation 9fa00143-7797-4935-bf59-024f7f501136 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.874439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:43914568bfdfcca3fcfc17d9b31cb4d6a282ccd68aaadb4012a92cd0c81c4dc3

Observation 5cc267bd-176e-4e9d-b154-c291d58b9451 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.850489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:54efe7635299adc54438695ad8ea4d16844e80bdaec02ddb550f2966cb808de9

Observation 5f1f1c72-766d-48f2-a376-cab0b09d49f0 · outbound

This paper cites GPT-4 Technical Report.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding GPT-4 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.861724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:0569d569f9730c6f09ce07477f1dd7d780bd170048a860f5e4307f21856f8888

Observation e07efb0e-149e-450e-8edd-a60cfaac64fc · outbound

This paper cites Training language models to follow instructions with human feedback,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Training language models to follow instructions with human feedback,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:3063b3bd81b6b46dc6c830428c7213adfc99ae7f0b6636800f1e53e902c3e779

Observation cfcbfc03-9e1e-4175-a08e-de1d81a5fcac · outbound

This paper cites Improving language understanding by generative pre-training,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Improving language understanding by generative pre-training,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:6d9fe6849f662095e816d315c86fd8affc894195b4dbca8a8b368ced306c97d5

Observation 41c8b56e-334b-40ee-9663-dfe43f28164a · outbound

This paper cites Ldre: Llm-based divergent reasoning and ensemble for zero-shot composed image retrieval,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Ldre: Llm-based divergent reasoning and ensemble for zero-shot composed image retrieval,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:99588f8969b7c3842922d754048b1d17a7f30d42eb3fee38bb0a28b388491860

Observation 4e915c48-e109-4e87-b4c7-b4f0bf5bfbac · outbound

This paper cites Vila: On pre-training for visual language models,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Vila: On pre-training for visual language models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:d7dd08718c72dbcf1f97f319df7274dc6d61074adcc39305ed47f5f33dac792c

Observation e14d055e-86f4-4011-8ddf-4f8ab84db104 · outbound

This paper cites Seman- tic editing increment benefits zero-shot composed image retrieval,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Seman- tic editing increment benefits zero-shot composed image retrieval,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:50f3bd1723fe8f725c9268f2267b88bddd556bf3651001207d41b29523889a1e

Observation 1c113b12-9f8a-4a48-957e-1f4b72ce12e9 · outbound

This paper cites Large Language Models are Temporal and Causal Reasoners for Video Question Answering.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Large Language Models are Temporal and Causal Reasoners for Video Question Answering

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.804217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:1cf16654be832cd1ab388e8452b50976bc000f0347817d7f8b1738f2081dd78b

Observation 0fe125f9-d2ae-4f88-95bf-91beea473d4b · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Mvbench: A comprehensive multi-modal video understanding benchmark,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:c3db211cda2781639f2bd4883c14792c36887a5a1c15650bc25c978ba97bc98a

Observation 23c6ebff-b944-4eb8-94a0-615f5a6f9d25 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.831217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:2ee7694e61ed476f1f443066dbbf9ac4d247683b7ac6128e0568dee4cab6eb61

Observation 1f48bc89-8f8f-4cb3-8dd8-cab0f38f4898 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Llava-next: A strong zero-shot video understanding model,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:614ff65ee22ae0f2a5a8428de3bceeb74bc6d148a7f1492f8eb7a700c301ef8e

Observation 46d4f6c4-d652-4d9f-84d6-d751bdb8cdd6 · outbound

This paper cites Learning to answer visual questions from web videos,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Learning to answer visual questions from web videos,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:8f8909ad795f2b008479147aa82e7812e3ae242bef52fed880d11432e5a7a778

Observation 98afd22a-9b1f-4963-9c92-fa20c5be6321 · outbound

This paper cites Transformer-empowered invariant grounding for video question answering,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Transformer-empowered invariant grounding for video question answering,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:19662a3cd012f67e28ec1872e49710471a569deec060ac25caf81e22778225da

Observation 49492a6e-29e4-4435-aab6-47ff51297220 · outbound

This paper cites Intentqa: Intent question answering in videos by cognitive context reasoning,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Intentqa: Intent question answering in videos by cognitive context reasoning,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:408a33643fce07bf0a04cd5f5d1ef862f281d2d9ade72c613e32907b095560e5

Observation 5698a124-ebf2-4180-87f2-898289dfe6ad · outbound

This paper cites Parse, align and aggregate: Graph-driven compositional reasoning for video question answering,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Parse, align and aggregate: Graph-driven compositional reasoning for video question answering,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:b0cba11e2bd4d2a7fd13d342cdeedc328c40550339d0437100c9984d5b5e8d42

Observation 6de7a588-f3bb-4e14-b3b3-5bfb8a2dfec8 · outbound

This paper cites Mecd+: Unlocking event-level causal graph discovery for video reasoning,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Mecd+: Unlocking event-level causal graph discovery for video reasoning,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:cc0980c856963965f77296fe9e6fae81013a6f05c024b19921666b569149b52e

Observation 40fc762a-d2ef-4453-ac9b-3b88c89b558f · outbound

This paper cites Adversa: Abductive driving accident video understanding,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Adversa: Abductive driving accident video understanding,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:efb3e5eef88a71abd06bc378f033e82fc1b3a5d234f455941417702425847c1a

Observation 35083534-a016-4ef5-b87c-50f74c1333e8 · outbound

This paper cites Moviechat+: Question-aware sparse memory for long video question answering,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Moviechat+: Question-aware sparse memory for long video question answering,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:fd3cf6141ca86e4eb5f467522515b50c36b42fa8c78cee0f9d51bc5ac85985cf

Observation 85b7a303-a2cd-4634-8cf3-311c71fb059b · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:c61aae6fccfad23b82083db48daf5f6a8248dd49a52d3b4fcb993b38c9d31f93

Observation a747fe2a-7028-4349-a0c4-b1a75ea58ca7 · outbound

This paper cites Vtg-gpt: Tuning-free zero- shot video temporal grounding with gpt,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Vtg-gpt: Tuning-free zero- shot video temporal grounding with gpt,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:cf1d964d40cb50df87502ec4496d496f58395012fd1cf52c1c2811c6fbb3b101

Observation 00a19d4c-3a9f-4e3b-a7f2-bfc08c45edc1 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.834984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:e390d7db0c80530bb1a9655e3b231dd6552aa67aa1a50ae5e0b5aff65e9c731d

Observation cec50cc2-e48c-4eac-bb6c-b309438f64c1 · outbound

This paper cites A survey on video temporal grounding with multimodal large language model,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding A survey on video temporal grounding with multimodal large language model,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:a470d9fd0db3a507832741c1e4aaf212d7b905387fb6354064ff9f11d82dc943

Observation 7cd82c12-2165-4bd5-9b8d-f56eb15859e9 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.873755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:86614ef168fa2f463127ae153138493a6f95cfa6882fb8daba5f23b3339dc81e

Observation 6fedf0cf-600e-492c-97bd-cef25aae9926 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.884170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:422eaf2c892893e3899237223f93a2d89ba08dac82c0f6e70ade9f5bf1dbb337

Observation 20bb6c12-8282-4aeb-a318-a41cbfcf2743 · outbound

This paper cites Valor: Vision-audio-language omni-perception pretraining model and dataset,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Valor: Vision-audio-language omni-perception pretraining model and dataset,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:146e34701435c23f68a0b3b890598632f136cd62cd36388643dc04214a8125ec

Observation 04f1bc10-b5a7-4c01-b5c6-f2a25f2bcadb · outbound

This paper cites Cap4video++: Enhancing video understanding with auxiliary captions,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Cap4video++: Enhancing video understanding with auxiliary captions,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:06f0676d4ae6265d024589848ec354da73d1e047cefce1d7234a0612b4136644

Observation f0ff9915-3ee8-4d2e-9c08-b01d8b655049 · outbound

This paper cites Hierarchical banzhaf interaction for general video-language representation learning,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Hierarchical banzhaf interaction for general video-language representation learning,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:d0d70ba9ff8d0dbcaefbdaeae7c20d71af63ed624fd1b7d5e119db00d84d2dc5

Observation fe3559d6-e628-47ea-880f-a3e500390de5 · outbound

This paper cites Video dataflywheel: Resolving the impossible data trinity in video-language understanding,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Video dataflywheel: Resolving the impossible data trinity in video-language understanding,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:eb854d194c25ba831c03e3d72e625242dd8e7191540176308bb407da3540f2fd

Observation 307bbdf0-3cc4-4161-b222-78059325826a · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.839847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:eb7a03ec1a8cc9aba4bee19384180704cb46bed6d22f1e72debca0840e8ad5b7

Observation 68141cf3-11bd-49eb-98f9-eb56b4648e18 · outbound

This paper cites Streaming dense video captioning,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Streaming dense video captioning,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:1955eb8804235bf4bb2d495ec43f0aa1b6535e94d8b8533e12b56cc598da8ac5

Observation cc8386c5-477b-468e-8966-e97c33317c4c · outbound

This paper cites Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.826651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:eaa153132e0e3129ce37e7dd3918be66d469be57026faf8ff7208e30c1e106f7

Observation da088f27-ff86-4e43-bf20-b4a5e92fd45e · outbound

This paper cites Merlot reserve: Neural script knowledge through vision and language and sound,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Merlot reserve: Neural script knowledge through vision and language and sound,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:2c11d5d03ce9c511ea7640163856461e40638eae15c3071b1b1a63c48733928a

Observation beaa738a-70ab-47c2-a3de-2f75721b34f1 · outbound

This paper cites LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.833884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:6fe0adac01e3d9aaf6f5f96c61aec12ad8c9a73d003814a76877d573f363aff1

Observation c49cf114-8c63-410b-a161-ca270edd949c · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:c629aef66c63a76c4b1561dbd906518e2ebf7fc1e70c24620a2731dc5558209f

Observation 52daadf1-6945-4010-b379-dc782656f36c · outbound

This paper cites ViSpeak: Visual Instruction Feedback in Streaming Videos.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding ViSpeak: Visual Instruction Feedback in Streaming Videos

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.853502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:a2e37755f845f29daec22e690d4c889a4a953eb349eeb106a0d67e2be9bf6920

Observation 3481cff8-ebc7-4199-8d03-3b6f1077b17d · outbound

This paper cites Open-ended hierarchical streaming video understanding with vision language models,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Open-ended hierarchical streaming video understanding with vision language models,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:9fd45365f18bbf4289386218b93fee5e0cc0ede72badb69788221735d36da1c4

Observation 72d567aa-8864-4a4b-b84a-49523d869cec · outbound

This paper cites StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.886774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:a22886cf7598901d17f307b895044c3f590b41fee98ec8bf461acaed854d19dd

Observation bf29b0ec-3341-4261-93fd-305ba69ab72a · outbound

This paper cites Learning to respond: A large-scale benchmark and progressive learning framework for trigger-centric online video understanding.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Learning to respond: A large-scale benchmark and progressive learning framework for trigger-centric online video understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:1c7ca43d53094cf8625211466414a94176a90b82c5249da9321d9eaf1366b5dc

Observation fb1e9b31-2784-4889-9a37-db3130d6c89c · outbound

This paper cites Egospeak: learning when to speak for egocentric conversational agents in the wild,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Egospeak: learning when to speak for egocentric conversational agents in the wild,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:b398fc9fa09227c3aa5a26f09cdc156447d7bed045efa888102dd3493749d0a7

Observation 201df5d6-37ad-478c-aa4a-a827e73da964 · outbound

This paper cites Streamlined dense video captioning,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Streamlined dense video captioning,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:aa3a721fae817de38ae1326545545c9dc95b8d0d4be18ae94943d57025be6a34

Observation 591348b0-58bf-44a0-bab9-92564619ecc5 · outbound

This paper cites Proact-VL: A Proactive VideoLLM for Real-Time AI Companions.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Proact-VL: A Proactive VideoLLM for Real-Time AI Companions

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.845142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:ccb0392c4446a0c419b49d697020efc4a6a88e065fcab0801c81b32a50e99cfe

Observation 08702598-86a7-4a93-8be2-466b567e768f · outbound

This paper cites Streamready: Learning what to answer and when in long streaming videos,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Streamready: Learning what to answer and when in long streaming videos,

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.768417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:7995e914d92df472711eabaa62a48231ff857e008d7c567b9c2c754ae32f2568

Observation 8aba4993-2c88-4504-998f-06d72778a934 · outbound

This paper cites Roma: Real-time omni-multimodal assistant with interactive streaming under- standing,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Roma: Real-time omni-multimodal assistant with interactive streaming under- standing,

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.881699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:846824a2e1c1f7e11e972cbcfb97aeb3d13935b2a82f908be6854f4f803ddb53

Observation 6c7b5c68-8c65-44f2-9aaf-f5cba3b59cf2 · outbound

This paper cites Rehg, Minsu Kim, and Yong Man Ro.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Rehg, Minsu Kim, and Yong Man Ro

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.781759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:2e3a60529cac1caf117e66d138f28fb933c94f0d9c17e11eaa686f70854c6b99

Observation 3820b706-82f9-4baf-9d83-a37a3bc5a889 · outbound

This paper cites Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.865947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:9b4a68c93c153f98220b0410ced938d195bbe2cfdf6f194b5954784b0c110ec1

Observation fbfffffb-f11f-406d-a276-5a27f8c652c1 · outbound

This paper cites Timechat-online: 80% visual tokens are naturally redundant in streaming videos,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Timechat-online: 80% visual tokens are naturally redundant in streaming videos,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:44a4a02a6fc419684c2ff165975dee9584b3b8d448f34998dd241b6d764a42bf

Observation 19e98f7b-c8bd-4cf1-a610-ef2657e38882 · outbound

This paper cites Livecc: Learning video llm with streaming speech transcription at scale,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Livecc: Learning video llm with streaming speech transcription at scale,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:cf0ea00c26610add1345798c870415553252859a52a8fb51a18592f325ed3a29

Observation fde229d1-7ce0-44d1-8369-a33998627afd · outbound

This paper cites Eyes wide open: Ego proactive video-llm for streaming video.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Eyes wide open: Ego proactive video-llm for streaming video

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.818679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:f6b40bf364512a5b82e64fe1982247d5a9ffd233368087c12ddc6a2326a3073f

Observation 2f57766a-d5a3-48f5-a834-5e95d888ffed · outbound

This paper cites Proactive Assistant Dialogue Generation from Streaming Egocentric Videos.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Proactive Assistant Dialogue Generation from Streaming Egocentric Videos

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.871955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:cad3664f4c7d92e92538ec04319c22b4ca4216ad5f531ededa4002086b45df7e

Observation 3389fc50-b9c4-4ea3-97f1-b50c8be4f684 · outbound

This paper cites AssistPDA: An Online Video Surveillance Assistant for Video Anomaly Prediction, Detection, and Analysis.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding AssistPDA: An Online Video Surveillance Assistant for Video Anomaly Prediction, Detection, and Analysis

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.881882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:c0ccc01ae5c01440103aaafb02ff27574f63d1508362a856b7a44d70c5b93705

Observation 4c7ddc3f-bec3-44a8-a7ce-90f8d129c2a5 · outbound

This paper cites Streaming Video Instruction Tuning.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Streaming Video Instruction Tuning

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.823683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:b304ae26df0d2149fe5fb3a98b4631c6f78674ff8e02ca51b61c20f87c276205

Observation 8fa9d949-2f76-4ff4-83ea-401707a18e0e · outbound

This paper cites Mmduet2: Enhancing proactive interaction of video mllms with multi-turn reinforcement learning.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Mmduet2: Enhancing proactive interaction of video mllms with multi-turn reinforcement learning

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.839489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:0d32e71319b50a1c435d85353e07837ec1c1f51dba394a34579a22a570807e3d

Observation 913de0f0-f8c2-4060-b632-aba4ed5a6a36 · outbound

This paper cites Thinking in streaming video.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Thinking in streaming video

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.892281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:64248b27e31203fc8528371f9b9db7eaa0ffa9ee9d40d4f55464d89247b9af23

Observation 6c894e2e-acde-421f-ab67-e992241ade2e · outbound

This paper cites Streamingclaw technical report.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Streamingclaw technical report

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.884040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:203c98b54e5221dbd75443f19ec2306ea8b90a87ea33aec533b686c8a63c682f

Observation b10940fc-465e-44b2-8354-eee9e74ef737 · outbound

This paper cites Querystream: Advancing streaming video understanding with query-aware pruning and proactive response,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Querystream: Advancing streaming video understanding with query-aware pruning and proactive response,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:308ac6cca1f15ef3df928727970b0e184a6957c00a6a3de5b1968a498c543d16

Observation 3f0e521a-1a94-439d-9a8b-4638881663ad · outbound

This paper cites Attention is all you need,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Attention is all you need,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:1c332e42c51b40168231c3c295ff8e6b29f96777f1665ce8982e3576d0222a5e

Observation 947f3dfa-9509-4b99-ad6f-88e084244afa · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.879057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:d1def0e0cd27573c5376cbdf39c950acf25751cd5da012a1c75e2cc58cc8e433

Observation b62462ba-6f47-43e3-948c-df820a424819 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.785043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:abddd97c04c432950a03566a691c1aceceffc2c459001a155e4476ca9f6d66e7

Observation 1b47b3d4-dc9d-473a-a501-7d4530b185f8 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.812326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:8e65095b3f0c0e44d30bb7deb5cfb4d5064948fffc34cce5fba646d25a015b97

Observation 4fdfc87b-a66e-46e7-8164-46bd3ab84990 · outbound

This paper cites Qwen2.5-VL Technical Report.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Qwen2.5-VL Technical Report

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.798286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:8314502fe5d96cdc54578697db3e03d58e35fa8b248f84279aa241a81019f615

Observation 6942a4ba-250e-4a9a-a0f8-dd6cd29fc844 · outbound

This paper cites Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction,.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction,

Reference 100

Resolution
unresolved
no resolver link, observed 2026-06-27T22:11:01.690237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:17ca16dd3ed6e9d28490dbb5f65e7a3e8128aa419d45a74901c5da30b5442501

Observation b275d5f1-12b8-4bb1-a872-6cacc65ccf9d · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.871284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:d4769c57763408d1c788d5f7724210341ff9ffae42d4413e49072e1c7d010302

Pith citing papers

Observation dd0c6591-e0ee-474c-b682-973138cd64dd · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:56.160538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:d281b451ce4468bcc5517d0d61bd8b0c9b63ee5afbf4da5f83ebed7e16bc2937