Pith. sign in

Paper Citation Record · LEDGER

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding

As of 11 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2605.07575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.07575 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:00:34.728880Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact3
  • verified fuzzy37
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch23

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 34d708c9-0e1e-4c52-b0b9-c1330abfb6dc · outbound

This paper cites Aho and Jeffrey D.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Aho and Jeffrey D

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.313377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:0fd67d3dd5bd7e7d00d2fdc68b68bc9f88bf70a5ee5c2729c0ad1199a132b04d

Observation 3f5c29a1-97dd-494e-b10e-ae560996ffab · outbound

This paper cites an unresolved cited work.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-12T21:01:53.309075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:714ec11e55571ff90b791f933602f9fa4d66497478e01b02f404c09657da1073

Observation fd9e72fd-7fcf-45b8-9041-faeaf5420c4a · outbound

This paper cites Chandra and Dexter C.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Chandra and Dexter C

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:17.209168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:3efe7f9937660f809a612f8a83b65edd52458d5835e8a84a6765dbe9308cffed

Observation 173dfd9f-9e8d-4850-94e6-3c9e2e54a2ea · outbound

This paper cites Scalable training of.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Scalable training of

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.405813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:15f7a77662a00898b8762066a0dfde58d63f16fb903908c8210100bf837428fc

Observation 5c05f806-b81b-47d2-942d-8fcd1b4c8f2a · outbound

This paper cites an unresolved cited work.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-12T21:01:53.409696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:f05d7d374e37a4f6282654af24ca1663c772903d4a62f965c8a9f4fc6c308cde

Observation e35ea1a3-5e98-4b56-8352-9f5bf5b96f75 · outbound

This paper cites Tetreault , title =.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Tetreault , title =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.401198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:524bbda61585d9e4bdefaa6a865b4af4d304ddbfc7d2abd656a32d350c7df285

Observation cdb73abf-2946-4c04-acb6-a4c357871cd9 · outbound

This paper cites A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.385857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:cacbe9afc49f203cfeaf51a1f85479070c2152082369867ca0b6341589003bf3

Observation 8a7c5190-de14-4713-b03c-229393445c14 · outbound

This paper cites Proceedings of the 33rd ACM International Conference on Multimedia , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the 33rd ACM International Conference on Multimedia , pages=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.389484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:a42f7df3580ff8cbb8435c8cf1b20c3b9887e68aa1a8e8070c6934dea9930d19

Observation 04595ad3-7d6f-4aa1-9719-3f25bbe35201 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.382305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:acd37c58778a2434d4b6965840a5c8cb3740296d67140dafb3ece1265964705f

Observation c740b639-209b-411f-a730-fdfc9fe76937 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.366527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:e5096135efef4784475a5782db74970b3480c19e939dafa1636082f027d4d05b

Observation 1c40b3ce-d0b5-46d9-bfbe-1cdbfd4f38a2 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.370854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:0858409061b149ef3c75c610f3d37be2edb78d7a82ae8130f97876497bfcfae0

Observation dfd62a81-e7fc-4a76-bbdb-28126ee4f5d2 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.397076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:da5b9bcb71bfa7d486204f32d183a105ad816a47df1d4ea532ed524462c741b2

Observation 5a2141aa-f525-4652-ae75-c3bb7be1ed26 · outbound

This paper cites Streambridge: Turning your offline video large language model into a proactive streaming assistant.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Streambridge: Turning your offline video large language model into a proactive streaming assistant

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.722174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:d969afdfdb4976cb9c2c20ebf138f31d4477debf737b490c384f06c7c1dea5e2

Observation f1f166c1-fde2-46e7-a129-033fb3bb5cea · outbound

This paper cites StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:01:17.737478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:b432e8bbb26ad2deb611e75bcb4631202140809fa2a9ac2d02d9475778aa9f66

Observation 701eb792-3946-4f86-92e4-90277abdf5c2 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.354816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:63c0e475b260c8dd3baf2bf1d4b6abd4e925f1bfebaeda3fde682875df6936c1

Observation 821f8b9c-55dd-4e55-85c9-2580248dc621 · outbound

This paper cites Neurocomputing , volume=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Neurocomputing , volume=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.357972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:a018dd50871d102aa61dd8504b734a6001c3723cc36fbe5e31889e24dd93670f

Observation c8a55867-b880-45ff-b791-b6868aa3e4fd · outbound

This paper cites Ai for service: Proactive assistance with ai glasses.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Ai for service: Proactive assistance with ai glasses

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.696044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:eceaf157c6eebaee1e3497d38be7195a6baa1571f1adf3a6c8c9b85e2407a048

Observation 5d3545d8-7eee-4058-8dc2-95c3d607eae1 · outbound

This paper cites Eyes wide open: Ego proactive video-llm for streaming video.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Eyes wide open: Ego proactive video-llm for streaming video

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.755819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:1c41a7df6a00dc72acb5037fe9ad8aa3bdd711f27ec947442f300b41aa1ce6e1

Observation f25e8a62-10a7-4610-990b-271d0fb836fb · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:01:17.685074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:2afe381fc86f3ee0c6d9fce2d27937d2d6bc8fed52850ee480eb759111c27b51

Observation e68a8438-83f1-4109-b5cd-36c9b9750cb4 · outbound

This paper cites Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.745300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:0274a751ee2215f55dfc69803cd6d79707bfc30ec4a1d7390ae147f5d93daf6e

Observation 9cd86e45-bb99-4530-99c2-be1b0a5d9849 · outbound

This paper cites Streamforest: Efficient online video understanding with persistent event memory.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Streamforest: Efficient online video understanding with persistent event memory

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.691519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:b8cb0c99d14c6c4a08b0f0ee979d9e7caed3af9aca8fe7971d044da7824fae4b

Observation 9fa7ea87-cb35-4f6d-a548-3bf6e5e0d9f6 · outbound

This paper cites Streaming Video Question-Answering with In-context Video KV-Cache Retrieval.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Streaming Video Question-Answering with In-context Video KV-Cache Retrieval

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.752967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:fbed15cbd32df8307bd08488db67f8fb9c3d22a5d017f7373ddb268f1b0721ac

Observation cf4b5ff5-efc3-4be8-997b-d36a18201cae · outbound

This paper cites Flash-VStream: Efficient Real-Time Understanding for Long Video Streams.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:17.733473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:e303003f0122862291d31ff2a5addfe67b8cab6f7bff81d7c797c1f0da486960

Observation 94461aac-4ebe-4e10-8421-5fcab3d2fd1a · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Advances in Neural Information Processing Systems , volume=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.340864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:c75366c751fffe40281e9b649f82b87299906992b79f49f67baea67baa0f32de

Observation b1e47bed-b978-42ab-bbe1-1b1066e8a397 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:01:17.758939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:8e12a1737b06a882c0aef5022c67e36e07170206a98f2ef6649c68aba3694258

Observation 23e19ffc-1302-4fbd-a97f-eb0905c49462 · outbound

This paper cites Qwen3 Technical Report.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Qwen3 Technical Report

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:01:17.714612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:92524e23e1a0f51ba1c8ac2a1e0034387a1904ea93c134952f9da57696533112

Observation 8500b5a6-a38e-4324-a067-4887a00535ad · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.718872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:dd15136076d0c6dad558a2cd1bd4693f7118ec1a823b0b9ddb4dedcf1b29365f

Observation 98f0a2b6-0692-43e7-883f-8d6077adb6f5 · outbound

This paper cites European Conference on Computer Vision , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding European Conference on Computer Vision , pages=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.374910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:5c73c84a2afc3cf7d6ef9fd48bd72affa5fdce8060aef09211295777a38a47c2

Observation 0553acf7-302b-4332-8a5c-a91a7a6f117a · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.412869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:e233e2734982af26191d6c9dbb3caaca92e00d8edaa2319afe29442357ce2049

Observation 8bf157e1-b807-47d1-873e-c440a8114efc · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.436236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:7e9d63891daccfb3490d945d11bbee08662b128bdd64cecb994ffa2587a5e28e

Observation 6d7a20fe-55de-48f3-8a35-d06f5a73bb6a · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.442194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:75368173b70f7b639ffc9186acfb3224efee249d9d3e71642c407d744461d77f

Observation 9658a757-a29c-4972-98cc-75ae8dc93a6b · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence , volume=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding IEEE Transactions on Pattern Analysis and Machine Intelligence , volume=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.416579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:a63fce282a1ec51d46f5e60029bae1f95d9247331e3dd3592d16ac4a94218433

Observation 384c3800-8835-480b-90de-e9c540ad8ef2 · outbound

This paper cites Forty-second International Conference on Machine Learning , year=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Forty-second International Conference on Machine Learning , year=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.419939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:721eb6f494ff07972b8d68140b511d94718d7acf2825b5b64541cf10c4515df3

Observation 6cba1d7b-8403-43ce-adad-3d3814085a2a · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.423378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:bc26b97983027448ddafd845267f744e224ca56d8cd9ca0c2a1945a8dafcdbd8

Observation ae9e39d8-68ce-4f67-bb5b-3e9a14907c1f · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops , pages=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.445528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:f322c14a83df3cbd5c386beae08751e8ff5bbe9a466a249c521484694561dd47

Observation a78fe43b-f489-4526-bdd0-ff9e02016602 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.393413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:22717a7cdb2480ae5310ab70e2181b44f75c41cff61b4c17d013fdc19213c952

Observation 789f9eb3-e2de-47b0-b23e-e0b12cba276f · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.336986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:802565d746d11dfc799a62c9b677e31aec64763e7e1298faf07f41a22d07f5a7

Observation fa83b644-f921-4527-a781-cc8bfa22215c · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.345022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:06ba47c52dbc682f675a1cb53d6d1f5d8dc741f796b45bcd79329f5a562e82ef

Observation 0349df26-09c1-4f6b-9e59-e02046a11bf3 · outbound

This paper cites Advances in neural information processing systems , volume=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Advances in neural information processing systems , volume=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.348509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:af994a8ea1cd99a2203fbffb03fdd4bc468a883b970ea826fcec2a3d4a3cff32

Observation c0ba247f-b4cb-4aeb-b3cf-d2c66be3398d · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.328610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:38d31ede742f600a86cb4b8cc7c0e23d7735c22c8851b57dc317c64868a13f05

Observation 3690bb62-e57f-4330-9c40-c7c3a95a9db2 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.320190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:c120673054b7f7cd5da9d257cfa9ce9b7788a8affcfc9d53ef369b3c64e473c9

Observation 0411c91d-92dc-429d-9bdf-7fab140237e9 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.316563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:b1ac443089b7ab888c037831b5ec4f78182eaffa7c4cf81748ec389b30a6d8bb

Observation 8fdec2a3-5971-405e-a208-11964a11bf1b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Advances in Neural Information Processing Systems , volume=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.323771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:7a562674bd0902c512ec5ca1feb0f1e4796f3a720992f0c2284dcdbec26a9d47

Observation edd0ba34-8583-4c54-935b-8c0ace8db1aa · outbound

This paper cites Qwen3-VL Technical Report.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Qwen3-VL Technical Report

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:01:17.765558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:b02b4d432f776cd47bda23d7c06f6c3d8aac6c91bee765b841e67a8cafea8e3d

Observation 30c162f4-8dd4-4eec-9413-79d13fe04088 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:01:17.772036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:a7c30ecd22ecd66154bbc0ae03710a7e4f256d94d117a2ce5877a55da3b88b2e

Observation 23984aed-351d-4f4d-a7ed-1975f0fe16a5 · outbound

This paper cites GPT-4o System Card.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding GPT-4o System Card

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:01:17.762419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:a95f19d6aaff722787b7de201bf5bba102e6a21d7b6498917200c0d3e03c9190

Observation 6e727fc9-1374-4aaa-be82-ef3aa364db11 · outbound

This paper cites 2024 , howpublished =.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding 2024 , howpublished =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.332774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:03726651fec3fdb5dc618f5a6f5626449f72c1898bba28da090f704777d1d9fe

Observation f3f3a713-1c5e-49d0-a521-f79bd625483c · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:01:17.687828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:12a32c49057767c976a5dcfd69911cac1db39b0e3254525d8ba461b22e7c2964

Observation 9d67ab4f-234e-47dc-97bf-d9c28edd2117 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.351724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:48b39c313300fc41956c72212715d55e49860df8609c56f7ba5c16d686a8de64

Observation 0a27ff8e-79be-4bdd-9476-ea62c810ea37 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.707528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:a915d9d4ec94370ddd885b9afd91b24917a105f94972cdc4e3f952fd3e362d92

Observation 12772f70-4170-48a9-9bcd-d3b325ebd2f0 · outbound

This paper cites Long Context Transfer from Language to Vision.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Long Context Transfer from Language to Vision

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:214f56b627a01207a88f2e0e7c2990c91b7bf6a469f9174bb1ff5dd5c58d0d15

Observation 38c21101-2298-42c5-ac41-ae76d63e5fee · outbound

This paper cites Science China Information Sciences , volume=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Science China Information Sciences , volume=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.363034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:3d6b2bc4b2c333bf04edea5e91eb1f04ed3626e733144957ed97eef5dc99a2ce

Observation c1005e51-57e9-4c04-9b1d-862c7d6bfbde · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.700487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:72396d93e480bd8a495ca6e1df74384f30a6284e1540d7cf3e9eb71ab8d7cc82

Observation 93afac63-3fe7-4a56-b370-51a5aebe35c0 · outbound

This paper cites an unresolved cited work.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-12T21:01:53.379122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:7e611b245610a7ce871aa417cb2efb8cf308cd5bae05aed6ed2d915ce475de3f

Observation e67d6f04-b80b-4621-aa18-43ec66c5fc5f · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:00:53.610624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:c947f1169bae11ddfa54f6386d994c9442e8605111658c748636c0cc5d71dc9f

Observation b51ca90e-0617-48a3-98a6-82827f44cf4a · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:01:17.748999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:d4329533956152b719e150bac24957fb7dc407581bf7329241769a6fe6e20ea0

Observation 2818515e-def5-403b-9de9-aa01c9d1ba7c · outbound

This paper cites Qwen2.5-VL Technical Report.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Qwen2.5-VL Technical Report

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:01:17.703748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:4ab11f3b498ddffc0e7fc7aa6e98398e322f5774008c49982fbce2efcc0b81e6

Observation 60ed806d-def4-459c-ace4-3b940f257edb · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:53:33.725971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:bd3f4211d0c9d484466a27d35a0726e234ef326c64460d0f37716fe9f05e5851

Observation 615bdaeb-9701-41ef-a4fb-d8bd47d7de9f · outbound

This paper cites TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.726003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:442bff9116e4a7f46be0b16f50300908ea7f64693b6ccb46470268906e830a16

Observation cb524be8-5383-4d7e-91e6-fba881acc382 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.453106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:8b6781897b7c3d68dceecdf66bcbe7d94a5925df209e941259a652e832c7c58b

Observation 60b42ab6-7b73-4b89-8079-c915c0e712ff · outbound

This paper cites Frontiers of Computer Science , volume=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Frontiers of Computer Science , volume=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.439125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:ebc65293b2b25e2fd869039051808d4a0c7d841d31f18a025ca12fae084ead7f

Observation 7450a099-166b-44d4-a798-85954d1e0bec · outbound

This paper cites ACM Transactions on Information Systems , volume=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding ACM Transactions on Information Systems , volume=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.449003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:459a4133ea44b27e30a66b1fc89da86dc5423cd91f03a4b7c33f8c5dc8b38dd4

Observation 4db74dba-565f-4273-86e7-9f36c0502e76 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Advances in Neural Information Processing Systems , volume=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.433414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:f9bf34a22b51426ddb1d2fb5dceaba4854189786f92e33c92ec6a7d3dae007a2

Observation d434120a-7374-4c9d-a5e6-a6f338312af5 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.427194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:fcb0c3f0aab66b31f8303fb90a44e516593b349b744326a8ade09fb958b564fa

Observation db8864db-16a7-4ae4-bf04-c5a62c109e3c · outbound

This paper cites IEEE Transactions on Mobile Computing , year=.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding IEEE Transactions on Mobile Computing , year=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:01:53.430488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:57b1ac7b4167b884eca225df8b9381f984532294e052622f30c479a60dbcc0e4

Observation 54b888d4-6da8-46bb-85a6-2b2ad46d75a3 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:59ffb2ec11189e2972f36be22806a8679c677b10b82c7f43589cd980283749a0

Pith citing papers

No inbound Pith citation observations are available.