Pith. sign in

Paper Citation Record · LEDGER

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

As of 14 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.05703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05703 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:50:24.333801Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3903d1a2-3c6e-4527-846d-a9016c6c01a5 · outbound

This paper cites A simple baseline for streaming video under- standing.arXiv preprint arXiv:2604.02317, 2026.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding A simple baseline for streaming video under- standing.arXiv preprint arXiv:2604.02317, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.227021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.227021Z digest=sha256:c386b8237227d7fe0cc0a0c836ca1d8aa4713e8f733f3ebda96a08f6a55f6169

Observation 4a577aba-f4ab-4436-a7bd-e48ce53a87ba · outbound

This paper cites Improving patient safety through video monitoring.Rehabilitation Nursing, 2016.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Improving patient safety through video monitoring.Rehabilitation Nursing, 2016

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.898118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.231022Z digest=sha256:ead7698aa84568d44f60a71ed8fe5202d80d7417c38cf0bd948c7859c2571be3

Observation 3d2d6482-5087-4e89-a260-2cc3170b3b15 · outbound

This paper cites Safefac: Video- based smart safety monitoring for preventing industrial work accidents.Expert Systems with Applications, 215:119397, 2023.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Safefac: Video- based smart safety monitoring for preventing industrial work accidents.Expert Systems with Applications, 215:119397, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.888592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.234541Z digest=sha256:36a850e604eaa64a5065c3d568ff1ee408385288c4fb2398dbc52f078fe50283

Observation e43d4235-558d-4ad6-a9e2-38a61c0926f0 · outbound

This paper cites When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis.Accident Analysis & Prevention, 219:108077, 2025.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis.Accident Analysis & Prevention, 219:108077, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.878978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.238171Z digest=sha256:04a4496ef71dbcc5c6c723aaf0bc4a93d415de3049b45973c64a12286d36d9cc

Observation 59bbdf9c-5194-48aa-83eb-8cf564613e3a · outbound

This paper cites Interactive language: Talking to robots in real time.IEEE Robotics and Automation Letters, 2023.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Interactive language: Talking to robots in real time.IEEE Robotics and Automation Letters, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.241392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.241392Z digest=sha256:822bea57897764ac6422912bb967233aea690aa498e2ccc35c6f2d11f47196a9

Observation d3f2bba2-f884-4ca0-b96e-b336c4464c22 · outbound

This paper cites Robix: A Unified Model for Robot Interaction, Reasoning and Planning.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Robix: A Unified Model for Robot Interaction, Reasoning and Planning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.245044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.245044Z digest=sha256:a3bd75c808be88786cad160ea2c61fe3ce9c6fc3e56f49ae51c8705495fba5da

Observation cf9a048d-fca0-417e-9c42-503fd413c479 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.248995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.248995Z digest=sha256:d290680fefa7fd8e75b90cf092bb119b510f35e94b504c7fbbf22cc1bf7e111d

Observation 16ca9a20-18a1-4408-97c1-00812ad7e5aa · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Video-llava: Learning united visual representation by alignment before projection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.252135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.252135Z digest=sha256:d03ccc20109749cedb3686896fe72f0a9e838aa27b9ee151294e32178a88d4f3

Observation 4a18752b-f8f3-476c-8fc3-44a0b062ecc1 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.255218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.255218Z digest=sha256:28ee59a9bd50cb1dcfd8aa23faea4534d1ba395a4dda38128974f3af0173b3c8

Observation 82799247-ef52-46e2-b753-2b9232383c0f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.258340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.258340Z digest=sha256:d8f15fc36098c1ce6de41c324a473f4a4e36a3e39cb5341858c9527e9c14754d

Observation 5664c91d-ae8f-429d-80cc-c56d7150b537 · outbound

This paper cites Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.261792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.261792Z digest=sha256:92e272103cb61e108f8e64010d6df744ead852310b401b86c1e8e7ff0ee5769f

Observation 2cf61a10-1d53-4a5e-8ab6-a20354a97fd1 · outbound

This paper cites GPT-Realtime-2: A multimodal speech-to-speech large language model.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding GPT-Realtime-2: A multimodal speech-to-speech large language model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.845727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.265553Z digest=sha256:92c3a4b6b04d551379b552c826856aacc21dbcdd95a6bb2aeeb42819cd399507

Observation d37bf162-a569-4e01-a7bb-dff7f6b76b50 · outbound

This paper cites AURA: Always-On Understanding and Real-Time Assistance via Video Streams.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding AURA: Always-On Understanding and Real-Time Assistance via Video Streams

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.268633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.268633Z digest=sha256:b7ef875cf36d05add7ebcdd014e2af976767d2c4f273d302b5f59ad5949fa51f

Observation c7eba5f1-b69a-48e8-b106-1c41ab5f84d9 · outbound

This paper cites Interaction models: A scalable approach to human-ai collaboration.Thinking Machines Lab: Connectionism, May 2026.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Interaction models: A scalable approach to human-ai collaboration.Thinking Machines Lab: Connectionism, May 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.272031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.272031Z digest=sha256:b1028980bddcb1d2df02b246a586ac49ce429526fdeadd59bf8d9e98fb9474c5

Observation 25ff5857-ae51-497e-8706-aeb77515ae2b · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Videollm-online: Online video large language model for streaming video

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.830601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.275109Z digest=sha256:926818aa695beb0f34ecf1166d9777596f9226c38d9b3afde5921959aa417945

Observation 58b7d634-87c6-41e9-8e3e-957ea1589e48 · outbound

This paper cites Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.278098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.278098Z digest=sha256:ea0d34d4bddf00a0189d3db53d4e9e2c5cafa56f9c38f199271e4b78e9e44c8b

Observation 9c899d0e-caee-47dd-b09e-668c6d2f4c6b · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.281054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.281054Z digest=sha256:4ee261d32a64dbc51bcb6b8fe60da22d94d9ed5849c12202767001e3b24bbc00

Observation 31843355-7f06-403b-bf30-c4f36e4b8725 · outbound

This paper cites Streamingbench: Assessing the gap for mllms to achieve streaming video understanding.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Streamingbench: Assessing the gap for mllms to achieve streaming video understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.284148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.284148Z digest=sha256:0a25be510799081d1e4921fb8ea788c4c75813dceb857aa134cf09de8d96d296

Observation 95b6604d-2dc4-4b75-9603-591c921618e6 · outbound

This paper cites Omnimmi: A comprehensive multi-modal interaction benchmark in streaming video contexts.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Omnimmi: A comprehensive multi-modal interaction benchmark in streaming video contexts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.802600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.286701Z digest=sha256:90c91dadc4475927314c404a895394f6523f61b24b7d6c55e9b0e35a22fda2e9

Observation e09a70cb-c01a-4f3b-9c2e-4c94e180ab8d · outbound

This paper cites MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.289509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.289509Z digest=sha256:472477269c992cbda51d2a722ac5a62e2f32b83011247e44dad91146d685f70e

Observation ce204b9c-a39a-47c6-85d5-c11a4fc3e4c7 · outbound

This paper cites Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.292417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.292417Z digest=sha256:d3e1040abdb91ed62ad74c2d289a6f58ef2144ea3445618732ab249c60bfbd40

Observation a5a6a0bf-3362-4624-b25f-4bafb160593d · outbound

This paper cites Streamforest: Efficient online video understanding with persistent event memory.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Streamforest: Efficient online video understanding with persistent event memory

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.793276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.295249Z digest=sha256:1793f9343015c45e92534b74d7c9234aaad0e49c15ee92562d82c0fb5a3bd76e

Observation 38847f93-255c-4b9f-bca7-8479eb977004 · outbound

This paper cites Thinking in streaming video.arXiv preprint arXiv:2603.12938, 2026.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Thinking in streaming video.arXiv preprint arXiv:2603.12938, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.297756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.297756Z digest=sha256:889c930370dd277834309195b5cfd61ad724a191720c12072b885d01d0a94f64

Observation 0ac28c44-b74b-405f-8d10-12b3a7cb4cf1 · outbound

This paper cites MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-08T00:50:24.531503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.300279Z digest=sha256:ff7cc53e662054c594c4a954ee0be008dca58edd25c1e3cf9344259491862cb7

Observation 5ea62cd8-0878-4201-814c-69b47c0be8cf · outbound

This paper cites Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory.arXiv preprint arXiv:2508.09736, 2025.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory.arXiv preprint arXiv:2508.09736, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.303426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.303426Z digest=sha256:546e6fc6c84bb6c0e96320e113c260089d85076e7d8dd73ac3144d8990cd8d73

Observation b3ee4ad1-a74e-4d7e-8b8f-4b866dcd6611 · outbound

This paper cites Egolife: Towards egocentric life assistant.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Egolife: Towards egocentric life assistant

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.783779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.306535Z digest=sha256:c4d77ac8a4bc82b4aeb3cebf0eecdeca3e705f764ce06a8fb04827bd2147ecaa

Observation 875285ee-8527-4e50-93f8-63a78aee8821 · outbound

This paper cites Agentic Very Long Video Understanding.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Agentic Very Long Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.309568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.309568Z digest=sha256:206dcb8645bd4192a6f1a2f0304d27c2063a6527966a05efbd34b558389aa614

Observation 8c9e68c5-d467-4f3f-ab8a-4febd009af69 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828– 28857, 2024.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828– 28857, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.774905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.312926Z digest=sha256:221963334cd0fd403f83e052679f51fbfae0cd69fb68b80c14996601e4968836

Observation ed3b5740-1506-43fb-91e7-877a7c3bfa18 · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Mlvu: Benchmarking multi-task long video understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.765627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.315905Z digest=sha256:c72929ce27799d99faf1853027386478d797e3166dc86d9b9a79287216d5188e

Observation 6fb8e054-2ff6-4dee-b9bd-0d52cc08a61f · outbound

This paper cites Ovo-bench: How far is your video-llms from real-world online video understanding? InProceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Ovo-bench: How far is your video-llms from real-world online video understanding? InProceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.318968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.318968Z digest=sha256:ccca8ae7d4ae1e33286688f7cd33b4a1168fc3aec67e569391d49fa6c22d18d2

Observation a4ed0abe-f650-4c8f-b88d-7ba8f15f0366 · outbound

This paper cites Online video understanding: Ovbench and videochat-online.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Online video understanding: Ovbench and videochat-online

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.751138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.322017Z digest=sha256:3f33c86b644b39dd79afe8d5fb387cb73457e8217694218fcb3e3cf3b667ad0f

Observation 8f433c35-7286-413f-b921-3c0e9a8e6cd5 · outbound

This paper cites Rtv-bench: Benchmarking mllm continuous perception, understanding and reasoning through real-time video.Advances in Neural Information Processing Systems, 38, 2026.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Rtv-bench: Benchmarking mllm continuous perception, understanding and reasoning through real-time video.Advances in Neural Information Processing Systems, 38, 2026

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.741064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.324946Z digest=sha256:481ef0d8ed6a5297d8af6f6c84bd47a31aa6d0620e3c205a4902b160d5509127

Observation 5560658a-e2d9-403b-87c0-a50731452018 · outbound

This paper cites Ost- bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding.Advances in Neural Information Processing Systems, 38, 2026.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Ost- bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding.Advances in Neural Information Processing Systems, 38, 2026

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:50:24.729246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T00:50:24.327981Z digest=sha256:c5f3125e14397d1863edea575012ceabf68a9061d9248d5f1ef48a0d3659da8c

Observation 9b00560a-cfbb-425a-bc4f-714bc25329e6 · outbound

This paper cites Can vision-language models answer face to face questions in the real-world?arXiv preprint arXiv:2503.19356, 2025.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Can vision-language models answer face to face questions in the real-world?arXiv preprint arXiv:2503.19356, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:24.330915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.330915Z digest=sha256:fac6945a0477bb91bac0960d8574070bcbf24aa1f9898ad8b0f717c63410a603

Observation e53c0b9f-8478-4534-8b78-99aab0de20cf · outbound

This paper cites HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents.

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents

Reference 35

Resolution
malformed identifier
no resolver link, observed 2026-08-08T00:50:24.333801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:24.333801Z digest=sha256:843e42b3532344059060028525298ea3247e65057316ff69b41730ef8bef2fa3

Pith citing papers

No inbound Pith citation observations are available.