Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:50:44.081050Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2505.03581.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:50:44.081050Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2f0771a3-f1ca-4c86-8055-a9b79bb5f917 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4c7416d5-173b-4e37-be71-32794146e0e3 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Beyond bare queries: Open-vocabulary object retrieval with 3d scene graph,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 84232814-3510-4da1-a4f1-dd96325f2251 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Search3d: Hierarchical open-vocabulary 3d segmenta- tion,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d7074e35-bd7b-4e54-99cb-ece5e1f239d7 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7d94e18-e615-42c0-b8b0-51c029add4f6 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Clio: Real-time task-driven open-set 3d scene graphs,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9494be8-c440-42a5-b3bf-6204d5c977cd · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes 4d panoptic scene graph generation,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 46c15d20-7df7-4f50-a137-3da68dc4bdce · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes G-retriever: Retrieval-augmented generation for textual graph understanding and question answering,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 645bd11b-dc08-4e04-8761-17006d2928db · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Let Your Graph Do the Talking: Encoding Structured Data for LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8cc9849-0fb3-46d9-bcf3-73b8e8b9bccb · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Can llms enhance performance prediction for deep learning models?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8128cca7-1eee-44b0-a647-76d99cacdfdf · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes STAR: A Benchmark for Situated Reasoning in Real-World Videos
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c57e5aa-525e-46ac-9490-95e8720d9733 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Agqa 2.0: An updated benchmark for compositional spatio-temporal reasoning,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1885515b-95f2-4100-8240-6e09a118c67a · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Visual genome: Connecting language and vision using crowdsourced dense image annotations,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f758b0d9-18c6-47ca-b6e9-be1bb9780696 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Gqa: A new dataset for real- world visual reasoning and compositional question answering,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a767cad-a860-4189-86b5-bef289d6efa8 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Panoptic scene graph generation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 65521ffd-71c5-42e0-b747-d37d6c798f85 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Action genome: Actions as compositions of spatio-temporal scene graphs,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0fed5c8e-ac80-492e-a37b-dd2684cfbfa4 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Panoptic video scene graph generation,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a090021c-ddbf-4edd-b9a9-393dcf6de988 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Egtr: Extracting graph from transformer for scene graph generation,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30a67c2d-1411-459d-a581-f32df200fe56 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Reltr: Relation transformer for scene graph generation,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6806e38-1f43-436e-a43a-bf1604f5a723 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Oed: towards one-stage end-to- end dynamic scene graph generation,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1605903d-9619-4b17-b3b2-2cfbf2bf3ee1 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes SAM 2: Segment Anything in Images and Videos
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e316375-8fa3-404a-83ff-1832345ffe1f · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Llavanext: Improved reasoning, ocr, and world knowledge,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22188316-e050-47d6-9f6f-be776b314683 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Yolo-world: Real-time open-vocabulary object detection,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3418715-8ace-498d-bcf7-ef9b3d605286 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes GPT-4 Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c725269-06a3-4ead-ae5e-0a3f44765277 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Yi: Open Foundation Models by 01.AI
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f7d309-02d3-4c7c-b892-64d46f8a505c · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes NVILA: Efficient Frontier Visual Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740d89b9-7117-4b46-a85b-ff8648391f89 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Qwen2.5-VL Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a8e3370-8c8f-4ecf-ad26-162d9a30876b · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04a36963-6c86-4e53-87ab-7c9073ebd39a · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6eeaa08-33b7-48ca-b31f-e80b1c3944c1 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Llm4sgg: large language models for weakly supervised scene graph generation,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 93d64783-f5cc-43c8-b905-59f1660eccb1 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db1551b9-ffdd-478d-b7fb-7678a1a290c6 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9818144f-18fa-476e-a85b-1e19b57e3fa5 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Visual Large Language Models for Generalized and Specialized Applications
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d0cf3e8-42d0-4d39-a3f8-79333589329e · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes (2.5+ 1) d spatio- temporal scene graphs for video question answering,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4cc7d048-e096-4191-b2f3-254a8c9b05f7 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Action scene graphs for long-form understanding of egocentric videos,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1b858b1f-bfeb-47f1-b0bd-5e31f8284cb9 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d0f2412a-688e-4361-b79c-5c67e17a3d8c · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1abcacc-c6b9-435a-a869-a612ea80fcd2 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4250c15-2b0a-4a16-ae60-b28dee74ec5e · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Benchmarking graph neural networks,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d938cc35-c39c-4172-ac51-f9f679921791 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Masked Label Prediction: Unified Message Passing Model for Semi-Supervised Classification
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd3c880a-c483-4c30-9183-6ac3d7a94fd0 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Roformer: En- hanced transformer with rotary position embedding,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a12445ac-7fcb-4ea8-88e4-d207e0a54489 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd2c11b-fd18-4e51-806c-f2e36ee7cb5b · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Agqa: A benchmark for compositional spatio-temporal reasoning,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c168e277-342d-4817-9a93-37fe007677cd · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Lora: Low-rank adaptation of large language models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f94f6c-05ae-411a-8abb-bfb10691e8ad · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes The Llama 3 Herd of Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85b94c4e-2a7b-44a0-9418-7c43271d08d7 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Do We Really Need Complicated Model Architectures For Temporal Networks?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc2ae432-b86d-49e9-82a7-11240a495aa5 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Attention is all you need,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea23b79a-d636-4a5f-9585-088fc0a12db6 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3d640f6-5953-462d-85d4-55aaab99866f · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 41a135df-9a6f-469b-b794-3ed9cdc277a5 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Self-chained image-language model for video localization and question answering,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 219cc9f7-2e91-40e7-a4f4-2e2d071e4456 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Vila: Efficient video-language alignment for video question answering,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 17ca962d-91c5-44f7-ae3e-ab8089b6aac4 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c57fd8c-d894-4d96-baa6-3b39df3c6587 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Look, Remember and Reason: Grounded reasoning in videos with language models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9fe86359-9402-4674-81ae-90c404ae6f0a · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Glance and focus: Memory prompt- ing for multi-event video question answering,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 204bab2a-2e42-4642-97bb-685fc96fff48 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Learning to reason iteratively and parallelly for complex visual reasoning scenarios,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7e94c167-1768-4650-bcc6-10a6cde4d138 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec0ce2d-c820-44ef-b6b5-a7500e15d60f · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Learning Fine-Grained Visual Understanding for Video Question Answering via Decoupling Spatial-Temporal Modeling
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e527ce8f-8f68-4236-8d92-491033c3c6e2 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49832ac9-9274-4e89-9d63-778c489cd185 · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes A note on the prize collecting traveling salesman problem,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 825f9578-8f73-459c-9c8b-16f2f5cc1acb · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c137e4b7-3635-4d63-91ba-aceac1a5e0ed · outbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes AGQA 2.0: An Updated Benchmark for Compositional Spatio-Temporal Reasoning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.