Pith. sign in

Paper Citation Record · LEDGER

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

As of 13 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.05703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05703 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:50:24.333801Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3903d1a2-3c6e-4527-846d-a9016c6c01a5 · outbound

This paper cites A simple baseline for streaming video under- standing.arXiv preprint arXiv:2604.02317, 2026.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding A simple baseline for streaming video under- standing.arXiv preprint arXiv:2604.02317, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.227021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.227021Z digest=sha256:03c0f8de739465a957866f0fdc77091a89f7474968df240d3caad786fed3f312

Observation 4a577aba-f4ab-4436-a7bd-e48ce53a87ba · outbound

This paper cites Improving patient safety through video monitoring.Rehabilitation Nursing, 2016.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Improving patient safety through video monitoring.Rehabilitation Nursing, 2016

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.898118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.231022Z digest=sha256:e853de233a862d3c0fb44839eec16872e86ccecec884abe212225197651c198f

Observation 3d2d6482-5087-4e89-a260-2cc3170b3b15 · outbound

This paper cites Safefac: Video- based smart safety monitoring for preventing industrial work accidents.Expert Systems with Applications, 215:119397, 2023.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Safefac: Video- based smart safety monitoring for preventing industrial work accidents.Expert Systems with Applications, 215:119397, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.888592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.234541Z digest=sha256:1deb0cac6f3fd463adde7d39dc4e704f452f594ebee2d43575f46192c5a2b00e

Observation e43d4235-558d-4ad6-a9e2-38a61c0926f0 · outbound

This paper cites When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis.Accident Analysis & Prevention, 219:108077, 2025.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis.Accident Analysis & Prevention, 219:108077, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.878978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.238171Z digest=sha256:491f8300045c04356dbcbcd515c5e97efe96092132931f7c36262d22d30d9f1a

Observation 59bbdf9c-5194-48aa-83eb-8cf564613e3a · outbound

This paper cites Interactive language: Talking to robots in real time.IEEE Robotics and Automation Letters, 2023.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Interactive language: Talking to robots in real time.IEEE Robotics and Automation Letters, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.241392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.241392Z digest=sha256:317f111f0fa217c39fd4e44fbb3909ed6e8269f90599b17a8ef575686fd7b842

Observation d3f2bba2-f884-4ca0-b96e-b336c4464c22 · outbound

This paper cites Robix: A Unified Model for Robot Interaction, Reasoning and Planning.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Robix: A Unified Model for Robot Interaction, Reasoning and Planning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.245044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.245044Z digest=sha256:0848f2139316456d3d09c87fee24732ae63ca5b535fcc1b2060e0798d374f9cb

Observation cf9a048d-fca0-417e-9c42-503fd413c479 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.248995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.248995Z digest=sha256:51e8ca92f6e023c715c962bbd07f67d7cdb93be2e5d29c5d091b6bc53c941ee6

Observation 16ca9a20-18a1-4408-97c1-00812ad7e5aa · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Video-llava: Learning united visual representation by alignment before projection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.252135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.252135Z digest=sha256:317d96e46d69211cc9e5faace5dc6dce6602ee8c61cbc0a9ef101fc7016fc99a

Observation 4a18752b-f8f3-476c-8fc3-44a0b062ecc1 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.255218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.255218Z digest=sha256:66ddf52a509d01d45b1f0902b4459d2f053471ee2001996ad958da127cbf6b9c

Observation 82799247-ef52-46e2-b753-2b9232383c0f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.258340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.258340Z digest=sha256:f8bddda8520972a6ec41139ec62632c2a76986f311a199464138ba67f1f77ea2

Observation 5664c91d-ae8f-429d-80cc-c56d7150b537 · outbound

This paper cites Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.261792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.261792Z digest=sha256:54498cad01dd0828ebb7f0d154056267fbe5d9507588e382cdd4fb89c678171f

Observation 2cf61a10-1d53-4a5e-8ab6-a20354a97fd1 · outbound

This paper cites GPT-Realtime-2: A multimodal speech-to-speech large language model.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding GPT-Realtime-2: A multimodal speech-to-speech large language model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.845727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.265553Z digest=sha256:f9b34a00a160767d5d1f8ed4b4f6a13e97100269bf3bd104a0efd7bcb2f9910c

Observation d37bf162-a569-4e01-a7bb-dff7f6b76b50 · outbound

This paper cites AURA: Always-On Understanding and Real-Time Assistance via Video Streams.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding AURA: Always-On Understanding and Real-Time Assistance via Video Streams

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.268633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.268633Z digest=sha256:1115d576c89e346f37f3532e57838416c1308051da32bb4d4154103f9e86d0f5

Observation c7eba5f1-b69a-48e8-b106-1c41ab5f84d9 · outbound

This paper cites Interaction models: A scalable approach to human-ai collaboration.Thinking Machines Lab: Connectionism, May 2026.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Interaction models: A scalable approach to human-ai collaboration.Thinking Machines Lab: Connectionism, May 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.272031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.272031Z digest=sha256:dae5947a3fd7c37d9b14ad621aecac08cf3f0e6c3a78cd86982d27bcdc2a4093

Observation 25ff5857-ae51-497e-8706-aeb77515ae2b · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Videollm-online: Online video large language model for streaming video

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.830601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.275109Z digest=sha256:660b2cbdd10b983388b91986056b5d70afbdd0984a090245e03991b8c4932b64

Observation 58b7d634-87c6-41e9-8e3e-957ea1589e48 · outbound

This paper cites Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.278098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.278098Z digest=sha256:04ea6da53a240ed101067a70b5c4c157c70cf7a28c2ce64a0a3da05d24b0c58e

Observation 9c899d0e-caee-47dd-b09e-668c6d2f4c6b · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.281054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.281054Z digest=sha256:b004fa17075aaa77f926086a67cd7199c828ebf71e9d6177907b9fad3af4d139

Observation 31843355-7f06-403b-bf30-c4f36e4b8725 · outbound

This paper cites Streamingbench: Assessing the gap for mllms to achieve streaming video understanding.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Streamingbench: Assessing the gap for mllms to achieve streaming video understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.284148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.284148Z digest=sha256:9eba1f89a46ade12443ba80608cfbc7e1b1ceaf12872c83f1365b4363b5d5fec

Observation 95b6604d-2dc4-4b75-9603-591c921618e6 · outbound

This paper cites Omnimmi: A comprehensive multi-modal interaction benchmark in streaming video contexts.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Omnimmi: A comprehensive multi-modal interaction benchmark in streaming video contexts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.802600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.286701Z digest=sha256:4254f6fe64f7301c3cb1e924fa547f7c61f9ec3c1adaac4771f31accd6f91af0

Observation e09a70cb-c01a-4f3b-9c2e-4c94e180ab8d · outbound

This paper cites MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.289509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.289509Z digest=sha256:34b15f52014c1543415ac3b3dcf796414d27502dc6ec950efa1ffa71d4c4355e

Observation ce204b9c-a39a-47c6-85d5-c11a4fc3e4c7 · outbound

This paper cites Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.292417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.292417Z digest=sha256:d155d590d4f6cbe198988a97f3ddf04031d4fcd0f17893aa8a02b2d7360b7806

Observation a5a6a0bf-3362-4624-b25f-4bafb160593d · outbound

This paper cites Streamforest: Efficient online video understanding with persistent event memory.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Streamforest: Efficient online video understanding with persistent event memory

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.793276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.295249Z digest=sha256:7ba23594fd4e82249016420e612c4f39fd8066a4178a2d561ea73d0794b47c4d

Observation 38847f93-255c-4b9f-bca7-8479eb977004 · outbound

This paper cites Thinking in streaming video.arXiv preprint arXiv:2603.12938, 2026.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Thinking in streaming video.arXiv preprint arXiv:2603.12938, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.297756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.297756Z digest=sha256:6d72b55e99ad95d9605e2aadbb1be881309d84d74bd08c176f6dfb7a545628c0

Observation 0ac28c44-b74b-405f-8d10-12b3a7cb4cf1 · outbound

This paper cites MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-08T00:50:24.531503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.300279Z digest=sha256:67dbe6d3caaf5b3dd447ec433fdd881d6064b32a00c5f5261900d7f5e653634f

Observation 5ea62cd8-0878-4201-814c-69b47c0be8cf · outbound

This paper cites Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory.arXiv preprint arXiv:2508.09736, 2025.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory.arXiv preprint arXiv:2508.09736, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.303426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.303426Z digest=sha256:c233da30f991a2f3fa828e68e2a76e68e7281e821aebf348aea906d1b17cd238

Observation b3ee4ad1-a74e-4d7e-8b8f-4b866dcd6611 · outbound

This paper cites Egolife: Towards egocentric life assistant.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Egolife: Towards egocentric life assistant

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.783779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.306535Z digest=sha256:b14fbe0ee70d9d1d01f7a6bfbd9a9b1f4317a42c763ecc1ae19f0412ac14b5bc

Observation 875285ee-8527-4e50-93f8-63a78aee8821 · outbound

This paper cites Agentic Very Long Video Understanding.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Agentic Very Long Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.309568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.309568Z digest=sha256:9d23fefac8f8a13ff1ae522222c7df93fec98e592a25dff8eb8ac7577cbb148c

Observation 8c9e68c5-d467-4f3f-ab8a-4febd009af69 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828– 28857, 2024.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828– 28857, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.774905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.312926Z digest=sha256:b2670da60a51bdd49f43ba0217c678d84d3c1f970b9f6a7fb6e4a611aa6cb39c

Observation ed3b5740-1506-43fb-91e7-877a7c3bfa18 · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Mlvu: Benchmarking multi-task long video understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.765627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.315905Z digest=sha256:d1b52359edb5c29a8d3acaa3b01325d9ae405a987a24a83b3e86d4cfad93eeef

Observation 6fb8e054-2ff6-4dee-b9bd-0d52cc08a61f · outbound

This paper cites Ovo-bench: How far is your video-llms from real-world online video understanding? InProceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Ovo-bench: How far is your video-llms from real-world online video understanding? InProceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.318968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.318968Z digest=sha256:be24ae6a03f8275170dc4d865346d099f9cc1f2df5538beb25826a691816f1aa

Observation a4ed0abe-f650-4c8f-b88d-7ba8f15f0366 · outbound

This paper cites Online video understanding: Ovbench and videochat-online.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Online video understanding: Ovbench and videochat-online

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.751138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.322017Z digest=sha256:92c2b00d0e7ae9218e7f86a7692a36d2ca5dc5ba93e32ff6fffeb3a2d06215d2

Observation 8f433c35-7286-413f-b921-3c0e9a8e6cd5 · outbound

This paper cites Rtv-bench: Benchmarking mllm continuous perception, understanding and reasoning through real-time video.Advances in Neural Information Processing Systems, 38, 2026.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Rtv-bench: Benchmarking mllm continuous perception, understanding and reasoning through real-time video.Advances in Neural Information Processing Systems, 38, 2026

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.741064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.324946Z digest=sha256:1f91c54a86c0fdfbbfc6c30cb0de2272e94b6f7a71508a0ba940b6fefa51163b

Observation 5560658a-e2d9-403b-87c0-a50731452018 · outbound

This paper cites Ost- bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding.Advances in Neural Information Processing Systems, 38, 2026.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Ost- bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding.Advances in Neural Information Processing Systems, 38, 2026

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.729246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T00:50:24.327981Z digest=sha256:09397ece51982f6ab6e38fd4d63091a42aa79a325d516425590153c29b7905e5

Observation 9b00560a-cfbb-425a-bc4f-714bc25329e6 · outbound

This paper cites Can vision-language models answer face to face questions in the real-world?arXiv preprint arXiv:2503.19356, 2025.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Can vision-language models answer face to face questions in the real-world?arXiv preprint arXiv:2503.19356, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.330915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.330915Z digest=sha256:51c7035aad8abf76e62545cfacace870cddd05015e4e71302c84b55b755b7f90

Observation e53c0b9f-8478-4534-8b78-99aab0de20cf · outbound

This paper cites HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents

Reference 35

Resolution
malformed identifier
no resolver link, observed 2026-08-08T00:50:24.333801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.333801Z digest=sha256:976dd3b2c6516c744cf06f5bb41c93bfdcb5b3179c93d85d85d47d0dc3da2544

Pith citing papers

No inbound Pith citation observations are available.