Pith. sign in

Paper Citation Record · LEDGER

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes

As of 16 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2505.03581.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03581 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:50:44.081050Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy21
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f0771a3-f1ca-4c86-8055-a9b79bb5f917 · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.709445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.879396Z digest=sha256:be99940bbf898613368ad27ecd46e5de20e72a525d2d6e10833634fdab40624e

Observation 4c7416d5-173b-4e37-be71-32794146e0e3 · outbound

This paper cites Beyond bare queries: Open-vocabulary object retrieval with 3d scene graph,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Beyond bare queries: Open-vocabulary object retrieval with 3d scene graph,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.700268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.883967Z digest=sha256:409a5c33360ab0b6864c6780466425dae302c487a12b164e9e8dff860e9ebaeb

Observation 84232814-3510-4da1-a4f1-dd96325f2251 · outbound

This paper cites Search3d: Hierarchical open-vocabulary 3d segmenta- tion,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Search3d: Hierarchical open-vocabulary 3d segmenta- tion,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.689789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.888014Z digest=sha256:9708a3bdd1eac9aab229443880e02fac275fb4bd9e33981615a88332a3da6d5e

Observation d7074e35-bd7b-4e54-99cb-ece5e1f239d7 · outbound

This paper cites Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.891225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.891225Z digest=sha256:966f92cd604dd07814e1e492b0878946f63706832cf6f7e0713f46af0e5bf10d

Observation b7d94e18-e615-42c0-b8b0-51c029add4f6 · outbound

This paper cites Clio: Real-time task-driven open-set 3d scene graphs,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Clio: Real-time task-driven open-set 3d scene graphs,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.895182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.895182Z digest=sha256:8a962f6434aad18d75424f294a76b70e9b91c6fc7702f4476605608400d93b7e

Observation b9494be8-c440-42a5-b3bf-6204d5c977cd · outbound

This paper cites 4d panoptic scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes 4d panoptic scene graph generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.668082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.898419Z digest=sha256:4cf72d61a0d93bd033c1ba5ec5f403cb90744366592960a171662acea3dd278b

Observation 46c15d20-7df7-4f50-a137-3da68dc4bdce · outbound

This paper cites G-retriever: Retrieval-augmented generation for textual graph understanding and question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes G-retriever: Retrieval-augmented generation for textual graph understanding and question answering,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.658325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.902087Z digest=sha256:d8736aa42817d635c30a501332c3539be0794238e487738d5d4966a89b7d112e

Observation 645bd11b-dc08-4e04-8761-17006d2928db · outbound

This paper cites Let Your Graph Do the Talking: Encoding Structured Data for LLMs.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Let Your Graph Do the Talking: Encoding Structured Data for LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.905097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.905097Z digest=sha256:0e7ed0af2e8c32e0c7232dde1251c684bb840a5cf6688e62fae0904ada8ad140

Observation e8cc9849-0fb3-46d9-bcf3-73b8e8b9bccb · outbound

This paper cites Can llms enhance performance prediction for deep learning models?.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Can llms enhance performance prediction for deep learning models?

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.647941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.908611Z digest=sha256:527ae614ec18e7f4727b6d1582bf1adf46e2575ed7a8234d604c396b2eefa731

Observation 8128cca7-1eee-44b0-a647-76d99cacdfdf · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.912140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.912140Z digest=sha256:b0de76a4713b4677b59ab9ee70574f79132ca6faa88ea2a74c1c4e08e3a3e901

Observation 8c57e5aa-525e-46ac-9490-95e8720d9733 · outbound

This paper cites Agqa 2.0: An updated benchmark for compositional spatio-temporal reasoning,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Agqa 2.0: An updated benchmark for compositional spatio-temporal reasoning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.637619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.915839Z digest=sha256:eecfaf9eb4f3475a260dcce4c4667797465a3d131502ebc8e1bc9e66517a7454

Observation 1885515b-95f2-4100-8240-6e09a118c67a · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Visual genome: Connecting language and vision using crowdsourced dense image annotations,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.922505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.922505Z digest=sha256:6ee9b4885410bb98a511ac6907a56b7ab01665e34156a4b7954526fd3bfaefdd

Observation f758b0d9-18c6-47ca-b6e9-be1bb9780696 · outbound

This paper cites Gqa: A new dataset for real- world visual reasoning and compositional question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Gqa: A new dataset for real- world visual reasoning and compositional question answering,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.925544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.925544Z digest=sha256:c1e6db1a869d84bf6d545c75b1affdb1e1827f697214bba21a96b240c8ad29e6

Observation 9a767cad-a860-4189-86b5-bef289d6efa8 · outbound

This paper cites Panoptic scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Panoptic scene graph generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.618448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.928530Z digest=sha256:5116c6d72065d84365cadf9f8cb4ba2ac68314d84a22c4e8494a967335dd628d

Observation 65521ffd-71c5-42e0-b747-d37d6c798f85 · outbound

This paper cites Action genome: Actions as compositions of spatio-temporal scene graphs,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Action genome: Actions as compositions of spatio-temporal scene graphs,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.609510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.931740Z digest=sha256:cd83fb8e304a4f9f5529922f73c25acb154aa58d5fc176e05e1a4c9f14458dea

Observation 0fed5c8e-ac80-492e-a37b-dd2684cfbfa4 · outbound

This paper cites Panoptic video scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Panoptic video scene graph generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.600160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.934810Z digest=sha256:3b06795dc2f7696f6083745889f243f8e281f74ebbf35c826b75b8dbe9559773

Observation a090021c-ddbf-4edd-b9a9-393dcf6de988 · outbound

This paper cites Egtr: Extracting graph from transformer for scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Egtr: Extracting graph from transformer for scene graph generation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.938041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.938041Z digest=sha256:7c6c9523ceff0b889194c8c520634b760546a5fbf2fab01ba1732ef31c2fcc12

Observation 30a67c2d-1411-459d-a581-f32df200fe56 · outbound

This paper cites Reltr: Relation transformer for scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Reltr: Relation transformer for scene graph generation,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.941400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.941400Z digest=sha256:6e6345ec6732b1092185bc42f827ab148ee27019aca7e66802d95eae5e327b31

Observation a6806e38-1f43-436e-a43a-bf1604f5a723 · outbound

This paper cites Oed: towards one-stage end-to- end dynamic scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Oed: towards one-stage end-to- end dynamic scene graph generation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.580376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.945053Z digest=sha256:df4a6d00290bda2ffe6587ec4e5650f3aaf8254657ffc1a5bd74be85a0522654

Observation 1605903d-9619-4b17-b3b2-2cfbf2bf3ee1 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes SAM 2: Segment Anything in Images and Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.948634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.948634Z digest=sha256:27db8eff5cc557bef0e2f0f0732fa06d243da110480b9057320520f51ab1230c

Observation 7e316375-8fa3-404a-83ff-1832345ffe1f · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Llavanext: Improved reasoning, ocr, and world knowledge,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.952811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.952811Z digest=sha256:64d55d48637f671248e4d92305007ea68cf7efe397088ea914a3453630f7c4c7

Observation 22188316-e050-47d6-9f6f-be776b314683 · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Yolo-world: Real-time open-vocabulary object detection,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.956865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.956865Z digest=sha256:76e36d84cc3c06ac034a936f3d1ae126ea0539f867fd2153b679df7ba349eed4

Observation e3418715-8ace-498d-bcf7-ef9b3d605286 · outbound

This paper cites GPT-4 Technical Report.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.960371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.960371Z digest=sha256:5df0114e52001a4a6ad440dc388389c6c28a575a2380d6ee6a637b6d508c6459

Observation 5c725269-06a3-4ead-ae5e-0a3f44765277 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Yi: Open Foundation Models by 01.AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.963321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.963321Z digest=sha256:5bf4a34cc0f82995f45a6d7349635985a7f5f601513b92d8c6617e78ea07d9ac

Observation 77f7d309-02d3-4c7c-b892-64d46f8a505c · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes NVILA: Efficient Frontier Visual Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.967414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.967414Z digest=sha256:69c580a6f73de2299e5c7db9986ad2396f331fda181cc1a56cce961a099196dd

Observation 740d89b9-7117-4b46-a85b-ff8648391f89 · outbound

This paper cites Qwen2.5-VL Technical Report.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Qwen2.5-VL Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.970524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.970524Z digest=sha256:18f7347ce0a047c40e10bd6d6702302cae4122ff1bda16d739c262acd379cfa7

Observation 8a8e3370-8c8f-4ecf-ad26-162d9a30876b · outbound

This paper cites Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.973960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.973960Z digest=sha256:243f8d55feb6bce166c318d0463cd67abae61c6a5dbc680713a50a62370d7f67

Observation 04a36963-6c86-4e53-87ab-7c9073ebd39a · outbound

This paper cites ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.977587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.977587Z digest=sha256:cb0ce11a8af8b686e5a2efbd1692fd4416a50e67fad84018f91179ba9e7e7c04

Observation f6eeaa08-33b7-48ca-b31f-e80b1c3944c1 · outbound

This paper cites Llm4sgg: large language models for weakly supervised scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Llm4sgg: large language models for weakly supervised scene graph generation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.560483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.980913Z digest=sha256:bfb35e31bba2143e86becc4f12764206c313147c3603cbf95360874b0f5108fe

Observation 93d64783-f5cc-43c8-b905-59f1660eccb1 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.984048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.984048Z digest=sha256:367a70e9dcde3e63e92010361a050907283959424d6d96a8819868ac79ca92b5

Observation db1551b9-ffdd-478d-b7fb-7678a1a290c6 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.987304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.987304Z digest=sha256:b9a22317ba46f23c112cf826bad96047ef119535e952f12da6da1c4d47e34b01

Observation 9818144f-18fa-476e-a85b-1e19b57e3fa5 · outbound

This paper cites Visual Large Language Models for Generalized and Specialized Applications.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Visual Large Language Models for Generalized and Specialized Applications

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.990891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.990891Z digest=sha256:27c373f935b9c1c2b2aeee578a8acb78dc354cc6846506dab7466eb5b8e0eb9d

Observation 3d0cf3e8-42d0-4d39-a3f8-79333589329e · outbound

This paper cites (2.5+ 1) d spatio- temporal scene graphs for video question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes (2.5+ 1) d spatio- temporal scene graphs for video question answering,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.551272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.993936Z digest=sha256:aa0e3f6fe323ede995e3eed0eda80c9b2605db40b5850802cdb3b99a38fdb5e1

Observation 4cc7d048-e096-4191-b2f3-254a8c9b05f7 · outbound

This paper cites Action scene graphs for long-form understanding of egocentric videos,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Action scene graphs for long-form understanding of egocentric videos,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.541815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:43.997532Z digest=sha256:95e1312824da5a210a329d674be98f2536b386d20bc9989ae1b84d35ef0813a3

Observation 1b858b1f-bfeb-47f1-b0bd-5e31f8284cb9 · outbound

This paper cites HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:50:44.318336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:44.000451Z digest=sha256:cbe8ea31d319ad0512458541cd37386df96984dc9c360bc539f23e53ee06b152

Observation d0f2412a-688e-4361-b79c-5c67e17a3d8c · outbound

This paper cites STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.003979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.003979Z digest=sha256:895261cd74b5c0ea03b3ebace49ed68e1fafcb56f42c28de3655915323319298

Observation a1abcacc-c6b9-435a-a869-a612ea80fcd2 · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.007182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.007182Z digest=sha256:92d9ef9ef842b4214521601d5d2f1dd5ace659d98ff276bf1ff3715d33d6f001

Observation d4250c15-2b0a-4a16-ae60-b28dee74ec5e · outbound

This paper cites Benchmarking graph neural networks,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Benchmarking graph neural networks,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.010411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.010411Z digest=sha256:81e897ed93ddd3a6165d67804c46597ac5e1736bab06639d05a625558630e607

Observation d938cc35-c39c-4172-ac51-f9f679921791 · outbound

This paper cites Masked Label Prediction: Unified Message Passing Model for Semi-Supervised Classification.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Masked Label Prediction: Unified Message Passing Model for Semi-Supervised Classification

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.013299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.013299Z digest=sha256:284ca352650a821e9b5ee409a088b3de94e94161c90e0a95894fbd4148d6e86c

Observation bd3c880a-c483-4c30-9183-6ac3d7a94fd0 · outbound

This paper cites Roformer: En- hanced transformer with rotary position embedding,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Roformer: En- hanced transformer with rotary position embedding,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.016593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.016593Z digest=sha256:1faea35f0eecc40919d00b562d377857c1e289d806830d6a23482e580f385b82

Observation a12445ac-7fcb-4ea8-88e4-d207e0a54489 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.019556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.019556Z digest=sha256:5f40dc0ff67beb2a7528b45e86ca404bd1064932c8fddfa63e18670f23eac638

Observation 8cd2c11b-fd18-4e51-806c-f2e36ee7cb5b · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Agqa: A benchmark for compositional spatio-temporal reasoning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.516269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:44.022635Z digest=sha256:f39d9a2051fe59635479b85d11f442c774a0a1c2fc35e632412182650794f424

Observation c168e277-342d-4817-9a93-37fe007677cd · outbound

This paper cites Lora: Low-rank adaptation of large language models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Lora: Low-rank adaptation of large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.025383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.025383Z digest=sha256:6951fbfe4bd149b9a506d39c1f3a70f16e30c63d869655d0897b8558225bf9b3

Observation c8f94f6c-05ae-411a-8abb-bfb10691e8ad · outbound

This paper cites The Llama 3 Herd of Models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes The Llama 3 Herd of Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.029059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.029059Z digest=sha256:465025f6ef080cdbaf3da2fe7e038f8e10a4948442f7828b0121d70755649fe7

Observation 85b94c4e-2a7b-44a0-9418-7c43271d08d7 · outbound

This paper cites Do We Really Need Complicated Model Architectures For Temporal Networks?.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Do We Really Need Complicated Model Architectures For Temporal Networks?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.032856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.032856Z digest=sha256:55d5463fd787e474afa806402643a3c56b7dcfd3dd4041015edf360acb819ec2

Observation fc2ae432-b86d-49e9-82a7-11240a495aa5 · outbound

This paper cites Attention is all you need,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Attention is all you need,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.036476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.036476Z digest=sha256:7b17fd606fdeed22bf4d7eb061fdfdd1e26e6a741fc67b19177b232c875549c4

Observation ea23b79a-d636-4a5f-9585-088fc0a12db6 · outbound

This paper cites Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.040102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.040102Z digest=sha256:3d6f351906abc7fd5ceac16973e4ba107b669cdf885f6ffc38e4954c16fc5db4

Observation f3d640f6-5953-462d-85d4-55aaab99866f · outbound

This paper cites Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.495603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:44.044120Z digest=sha256:44e5cd91dbb108d2e58911a3b5c4204a35e29a9b3c5fe38f97c9d833d5ce47c1

Observation 41a135df-9a6f-469b-b794-3ed9cdc277a5 · outbound

This paper cites Self-chained image-language model for video localization and question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Self-chained image-language model for video localization and question answering,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.485545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:44.047370Z digest=sha256:10b0690c3acdd87ba4ca561e6aad213602858db79b93f9eeb941480b59511f25

Observation 219cc9f7-2e91-40e7-a4f4-2e2d071e4456 · outbound

This paper cites Vila: Efficient video-language alignment for video question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Vila: Efficient video-language alignment for video question answering,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.476300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:44.051426Z digest=sha256:49c990473197d7d49bad4cd9e415fc2be15d42fd7c0a6cb3467ebd048a73ae5b

Observation 17ca962d-91c5-44f7-ae3e-ab8089b6aac4 · outbound

This paper cites End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.054394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.054394Z digest=sha256:afdaa4474ad72e10e97a7dacaffa822cd396dc0e2e64d55b76fe65c59ebb16cb

Observation 9c57fd8c-d894-4d96-baa6-3b39df3c6587 · outbound

This paper cites Look, Remember and Reason: Grounded reasoning in videos with language models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Look, Remember and Reason: Grounded reasoning in videos with language models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:50:44.244898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:44.057798Z digest=sha256:f8faf8ac29cb13354965392e7b25ce56b34e181d45764a64789eba5e907a5860

Observation 9fe86359-9402-4674-81ae-90c404ae6f0a · outbound

This paper cites Glance and focus: Memory prompt- ing for multi-event video question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Glance and focus: Memory prompt- ing for multi-event video question answering,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.465775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:44.060839Z digest=sha256:7cec473404a48b3f81891b0c4fe36fbbb634aa9334fe06c15b1756fe40467eab

Observation 204bab2a-2e42-4642-97bb-685fc96fff48 · outbound

This paper cites Learning to reason iteratively and parallelly for complex visual reasoning scenarios,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Learning to reason iteratively and parallelly for complex visual reasoning scenarios,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.456171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:44.064372Z digest=sha256:d5e1f2878e8bd8fe6c625d3d73d0fdae1d0f3e9e262f591e8538b51ebbb4bf04

Observation 7e94c167-1768-4650-bcc6-10a6cde4d138 · outbound

This paper cites Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.067588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.067588Z digest=sha256:0d9b63ccf515051e731739ffa87fbbd51386bdb0c3dabe15d20b828ab0902c44

Observation fec0ce2d-c820-44ef-b6b5-a7500e15d60f · outbound

This paper cites Learning Fine-Grained Visual Understanding for Video Question Answering via Decoupling Spatial-Temporal Modeling.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Learning Fine-Grained Visual Understanding for Video Question Answering via Decoupling Spatial-Temporal Modeling

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:50:44.134813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:44.071264Z digest=sha256:57d135a3543c30ff7fae08653a877d507e4cb9e21a92c2eeae2f1064292839d8

Observation e527ce8f-8f68-4236-8d92-491033c3c6e2 · outbound

This paper cites G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.074558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.074558Z digest=sha256:2bb1ef66b32ae6e336524a549e85f802da92289636fb17e3d66748a03c74193f

Observation 49832ac9-9274-4e89-9d63-778c489cd185 · outbound

This paper cites A note on the prize collecting traveling salesman problem,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes A note on the prize collecting traveling salesman problem,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.445522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:50:44.078054Z digest=sha256:5d6f9d564bf77b6d287079299cecf196602b0128ed000ac2f4b3c90e5019a71b

Observation 825f9578-8f73-459c-9c8b-16f2f5cc1acb · outbound

This paper cites FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.081050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.081050Z digest=sha256:8cf8c814ee0f514afc5aa75cd658828356f550a1bd5ac74531c1f58653b161a3

Observation c137e4b7-3635-4d63-91ba-aceac1a5e0ed · outbound

This paper cites AGQA 2.0: An Updated Benchmark for Compositional Spatio-Temporal Reasoning.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes AGQA 2.0: An Updated Benchmark for Compositional Spatio-Temporal Reasoning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.919104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.919104Z digest=sha256:5afe357b33b6ef0a5875d1cc5123e886ee77af8d5fe28ecc98eb30d05ff89c55

Pith citing papers

No inbound Pith citation observations are available.