Pith. sign in

Paper Citation Record · LEDGER

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

As of 16 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 3 inbound Pith citation observations for arXiv:2412.00161.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00161 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T06:02:34.065885Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:50:44.003979Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:46:56.770108Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved39
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3119cec1-1208-49af-b60a-e42b1c201082 · outbound

This paper cites GPT-4 Technical Report.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.833859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.833859Z digest=sha256:26a7ebb8c47be8994cedbd78ca9a0a871dd73557bffa36d29d965b265b52329d

Observation b35a6944-8fb1-4233-a2b9-731f362e2e59 · outbound

This paper cites Self-Training: A Survey.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Self-Training: A Survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.838855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.838855Z digest=sha256:1dd8806e9c440d1b91d7d90ef3202b34a28a5219654f24658d59b69cda672d7a

Observation 60a94add-ab46-4473-9ecc-8462cfd3663c · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.843350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.843350Z digest=sha256:31e20db8fe072bcab2d055567fcb1ae66e2fc0eeb5967ab1216fc264f9fcdb29

Observation a11ef4d5-c5a8-4a32-a889-57bf3e8126b4 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Activitynet: A large-scale video benchmark for human activity understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.847470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.847470Z digest=sha256:bf23f5ee5f2835f3b106d8ce5d8305dedffb51b292fa78ba4dab9f1754d60515

Observation 016fa148-3736-4f07-b712-9bbc0b167869 · outbound

This paper cites Castellano.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Castellano

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.792722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.851365Z digest=sha256:43b7741c274f22616d1133f7cf731879c171e795f718dda1d4c361a26afd79da

Observation 0606bf1f-d35a-4860-bf87-df77798b889e · outbound

This paper cites Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.855535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.855535Z digest=sha256:daa2398b5d04ea6a50252c57b4b99977b57b329bbc14a695a1cc3bd8c2c41d05

Observation af7384e6-ea8d-4b67-ba30-5e87a8850081 · outbound

This paper cites Palm: Scaling language modeling with pathways.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Palm: Scaling language modeling with pathways

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.859624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.859624Z digest=sha256:02b0831fdaea46a217400866c1c6a244611e4eba0832ffb6964708d7da5dc4ce

Observation 1903718e-b72f-4fb8-8a4e-b9814314d933 · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training VILA$^2$: VILA Augmented VILA

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.863324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.863324Z digest=sha256:f1e1f952f7feecfa06683a4c867a2a8c850cb48e0a6d49d440a5d2d277dbba03

Observation 5fc4f296-b068-4ef5-95b3-70180deba854 · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.774606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.867033Z digest=sha256:cddcb938764473bf38eec1a822d49adad4c92937091465269198a14f4f498385

Observation 9999ec75-bd08-4bf5-8d35-e1901626d8a8 · outbound

This paper cites Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.870633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.870633Z digest=sha256:871bf5f47a932ed4cfd5d7363475e1cb5ad16840fbf058c8236953de12f82d5d

Observation c780f082-c36e-4600-8633-41bed9d3a0a5 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.762297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.874334Z digest=sha256:b7f23a67473d8330f80f6c7d01350f10d999da9c41d140e1657201f8bec9de50

Observation f837daa1-14ec-4911-bde4-68be39403107 · outbound

This paper cites Spatio-temporal action graph networks.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Spatio-temporal action graph networks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.751134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.877977Z digest=sha256:da30b7691d2417006d38faadb7466cdc77a85b41872d2e666553c8a5543cce0a

Observation b3ac5e22-a313-469a-8b02-f3a1678b48b4 · outbound

This paper cites V2xum-llm: Cross-modal video summarization with temporal prompt instruction tuning.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training V2xum-llm: Cross-modal video summarization with temporal prompt instruction tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.881327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.881327Z digest=sha256:accae029fd67b39d6e3d94a3b7f8883768ffd3a9b484ed7c384d8e8b995b0632

Observation e4f429cc-745e-49b7-aa45-5cd8a78a5a7d · outbound

This paper cites Large Language Models Can Self-Improve.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Large Language Models Can Self-Improve

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.884736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.884736Z digest=sha256:10cf62779a525f46bb40456d754ceac481ff441adbc00f7c07e0204e55000ce3

Observation 052da5bc-60a5-49f4-bbe6-c6d4eb75ff15 · outbound

This paper cites Action genome: Actions as compositions of spatio- temporal scene graphs.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Action genome: Actions as compositions of spatio- temporal scene graphs

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.739334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.888550Z digest=sha256:80683b6b66677746c9133dda43f3eda0fb063d1f933c4ded78f6fbedd555f22f

Observation 8b78d2f8-f0d3-4d92-9903-42b43f168d9b · outbound

This paper cites Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.891781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.891781Z digest=sha256:802701f70e127a705a325758bc667ba47c654dd1316c4f27d59bbfaa64ab80c9

Observation 87138f8c-a0eb-4861-9a4a-40ed3b4d5510 · outbound

This paper cites Segmentation in the perception and memory of events.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Segmentation in the perception and memory of events

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.727275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.896306Z digest=sha256:7a2fe22522b08dc3ad011ac021f33a5ca4cdc10ce06828e53c97ebc55228a799

Observation e4e52cce-912a-4608-bbed-a54658110083 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training TVQA: Localized, Compositional Video Question Answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.899867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.899867Z digest=sha256:0ed6c26d1bb3229958b3dc783a6e3424c617f9a60fe5dec798ae8f224f72e6ce

Observation f663fcd5-988a-4bc7-89f6-ba7b9bae87d7 · outbound

This paper cites Adap- tive hierarchical graph reasoning with semantic coherence for video-and-language inference.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Adap- tive hierarchical graph reasoning with semantic coherence for video-and-language inference

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.715685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.903559Z digest=sha256:14f5793496b8f5d100024f428376915b64166244388475370dccf473cf0a0521

Observation cad64fe0-6a44-43ce-816e-af1fe177f464 · outbound

This paper cites Fine-grained semantically aligned vision-language pre-training.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Fine-grained semantically aligned vision-language pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.907066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.907066Z digest=sha256:71d4afda1a4a594b8e3050b3a0ebd1609cfd0ffcef7f1d6778678205bad87e1f

Observation be4addca-cfb1-4d05-aac2-420fbea3fc52 · outbound

This paper cites Fine-tuning multimodal llms to follow zero-shot demonstrative instructions.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Fine-tuning multimodal llms to follow zero-shot demonstrative instructions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.696809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.910618Z digest=sha256:e8b4cd28a3673fe33c0d9e243fd97336a522ed5f25b5c4cd977db1de5b1a062b

Observation c7ca8f38-bb6a-47e3-a050-77ed83fc3a52 · outbound

This paper cites Variational cross- graph reasoning and adaptive structured semantics learning for compositional temporal grounding.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Variational cross- graph reasoning and adaptive structured semantics learning for compositional temporal grounding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.685586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.914166Z digest=sha256:227b209793bd8ee74169b91c14fbb37d0178852e3f95f9625999902f24aac201

Observation f014b3fb-87cb-430b-82e2-4bcdf615f6da · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training VideoChat: Chat-Centric Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.918689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.918689Z digest=sha256:847f30c615e8b0dabe8860d61f4b2f8ff7b76bf6cbb4b28a3d548bda0b13c1fd

Observation f48a6de6-93a5-48ca-8065-f91952b6a129 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.674337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.922417Z digest=sha256:d18a517dfd04fbb7259bb7938d315e42520c6eefd3ea97426484e3598ebdb13d

Observation ade81513-35b2-418e-b934-2e0ea1566107 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.926111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.926111Z digest=sha256:c09e60ea77807b6ec4a47e0d072449a9548ff8c46177705ce7b80e14d8f7d534

Observation 1d93cbb1-038d-41ae-8cce-bdb35297f1ae · outbound

This paper cites VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.930189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.930189Z digest=sha256:f513d431d7dc9c912decac3011832f3b60174f03f8e9a416dff3a8f8ddf727a6

Observation c235d881-39db-4be3-8812-6133045c31d5 · outbound

This paper cites Discrim- inative hierarchical modeling of spatio-temporally compos- able human activities.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Discrim- inative hierarchical modeling of spatio-temporally compos- able human activities

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.663897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.933748Z digest=sha256:5058d97c0b959ad8f9a3d778c7901d67df4229e1b298cc55c3ab7b5ffa77f296

Observation a0fbc49d-51dc-4167-bb93-7b2c46708b5c · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.937143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.937143Z digest=sha256:45a82506d754a2338c3c68cb03549f4c82c88d29c36c97a364720a1010f73587

Observation 7d9229ea-2905-4b51-abfc-e99b5a57c469 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Vila: On pre-training for vi- sual language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.652790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.940729Z digest=sha256:758bfe0e306f3cb97b44523d4f26cadadf7003b7560859607ca49cae123add41

Observation a07a0184-8a19-4924-a76d-810fa316febf · outbound

This paper cites Improved baselines with visual instruction tuning.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Improved baselines with visual instruction tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.943897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.943897Z digest=sha256:fd7e6af771a45490d339b1aed65d6d494ef824fadec0528ff233872017418228

Observation 5d5a31d2-fb4f-4a83-810a-cedf0f645a44 · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.947128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.947128Z digest=sha256:5fd32641d3220045bb3b6ce8bf0635efae0006a913b3434a76d6f8748f338510

Observation 1423de4b-e459-4740-8288-e018e3f737e7 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training TempCompass: Do Video LLMs Really Understand Videos?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.951029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.951029Z digest=sha256:c673a6aadd5a020821e81e960d29b31624faa79a0bbbe4fa429fed3ebaa14e27

Observation 2e460265-76b2-48ae-bda7-a498738e491a · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.954963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.954963Z digest=sha256:9c6a328f86d39951692ec5385cc4537dff82b5b11d4ff2981d103c12518b8c3c

Observation 8673a5a6-6b84-40ce-b136-0e5358d4aa28 · outbound

This paper cites Orca 2: Teaching Small Language Models How to Reason.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Orca 2: Teaching Small Language Models How to Reason

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.959256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.959256Z digest=sha256:17ea05fe9a480d6d356cc5da275ba746119f610de6a69c54d4a6b7964038f8e5

Observation cc060ef6-6f9e-4f50-990b-90eb605fe3f0 · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Compositional chain-of-thought prompting for large multimodal models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.634922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.962968Z digest=sha256:89ae9cefcf106ad21d966ea211645f423f7a11b30557f3ebaa709609525fa125

Observation 90db1c2d-792c-4e59-85c6-1bbac8a11746 · outbound

This paper cites Hig: Hierarchical interlacement graph approach to scene graph generation in video understanding.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Hig: Hierarchical interlacement graph approach to scene graph generation in video understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.623933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.966164Z digest=sha256:6e3229e873bd8b79d54e625c750f0fa127691fe3444f1b6ef4ccdc526ec61d14

Observation dfd69555-a6ef-471d-9590-a883f7a767da · outbound

This paper cites Chatgpt, 2023.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Chatgpt, 2023

Reference 37

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T06:02:34.612958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.969545Z digest=sha256:428f621613f09c35e5e9aa320e8a89f3abbd911f7431b4833e5781f93315a2ad

Observation 98d830c6-afdf-487d-8f3e-e4679b3cb855 · outbound

This paper cites Gpt-4v(ision) system card, 2023.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Gpt-4v(ision) system card, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.972715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.972715Z digest=sha256:f756dad212e9ae4298e33e49c937905d67aa84094871e805aad5b0fe56635410

Observation 01b87ade-8e47-4abb-a7de-c04542646600 · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.976196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.976196Z digest=sha256:4e943440829530c24f9da3456ae57f8460c46da498cf44dca22f74071b1f6aa6

Observation 017f2444-abdc-4b9d-aa8f-1e80c492723b · outbound

This paper cites A computational model of event segmentation from perceptual prediction.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training A computational model of event segmentation from perceptual prediction

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.595225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.979838Z digest=sha256:79314393534ba47a514373611f4fce3e524a4eedadeca3bade2f80aee1437e2f

Observation 49821944-a8a2-4995-a7ad-99ff2517c84e · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Moviechat: From dense token to sparse memory for long video understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.982993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.982993Z digest=sha256:4f8acfc024ec3853e14370a111aaae952775d962ebf3fd5fc104c966caff703f

Observation 4a6fa4bc-09ce-4b9a-8349-4247ba8c861d · outbound

This paper cites To click or not to click: Automatic selection of beautiful thumbnails from videos.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training To click or not to click: Automatic selection of beautiful thumbnails from videos

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.577309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.986253Z digest=sha256:1ef92f2c090931214826fb721eeb77f0abd5d4a03a1014cf2c1b315299878372

Observation 2e89331b-fa67-4d18-bca5-5b502a901f8f · outbound

This paper cites Human brain activity time-locked to narrative event bound- aries.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Human brain activity time-locked to narrative event bound- aries

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.566379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:33.989612Z digest=sha256:b7c9853d247a467134e7f691de1088270006494fe6bc2fdc83f75a46974e6a42

Observation 0837432b-bbe0-416f-95bb-3540e3e07d63 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Finetuned Language Models Are Zero-Shot Learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.992918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.992918Z digest=sha256:cf4b46bd49fc5255dec037b5f9e38710c6d25f459256fae254e8ce18e0bb5ef1

Observation 7447d830-5eb2-4bbe-bd35-503c5c759b4d · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:33.996637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:33.996637Z digest=sha256:54f1d113f91ac69451e82368029b820cdf1c4ab3276a82b080e8835b561d7820

Observation 606647ec-df26-4bb0-945a-b4abb75cc09c · outbound

This paper cites Graph information bottleneck.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Graph information bottleneck

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.555150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:34.000404Z digest=sha256:3bcb21789ba2d4702f44f23154fedf09918d3d12471316d0e49f654aa6232129

Observation 993c8a96-1962-4fc3-b3af-18704ee7b246 · outbound

This paper cites FreeVA: Offline MLLM as Training-Free Video Assistant.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training FreeVA: Offline MLLM as Training-Free Video Assistant

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.003740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.003740Z digest=sha256:f1587c9b9893ba9999596ed786f6b39394464a48be90d97ba270e4914b6c7031

Observation be0f4d1d-94ae-4742-8e75-eb1edeaeea3f · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Next-qa: Next phase of question-answering to explaining temporal actions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.543673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:34.007393Z digest=sha256:7db5f9647d46e614f85408b35db264612ecc9db976468d4e7d1a8f3433dc9443

Observation 8a5256a6-c33a-41e4-88fb-b43819056c86 · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.010959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.010959Z digest=sha256:8ea4d05aa08bcfe949edc80c54ad8b78818bc5c1e2532f75158f8a12c59e1b3b

Observation 8f07c334-1dad-4d0a-8602-dd6e1a41508c · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Msr-vtt: A large video description dataset for bridging video and language

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.014379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.014379Z digest=sha256:304b123ecbf5572422d2f8fffcb340d58fd3ebd0d0fd0205f6dabe15359b3da4

Observation ee8ced4e-69f7-4030-b6d7-fbe96905bc17 · outbound

This paper cites Stronger Models are NOT Stronger Teachers for Instruction Tuning.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Stronger Models are NOT Stronger Teachers for Instruction Tuning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.017797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.017797Z digest=sha256:f82b09096b10b1a43506779caea45948347c472b79794e71a8d15005c4e61093

Observation 51825da9-c282-49f4-8125-7012d9621d0c · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.021601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.021601Z digest=sha256:32ea95165a9ccaa31181c7c0ccb7f1b64c2dcb1234687aca779f84978b06bf02

Observation 11c50a23-74f1-4a80-abb5-f93e9835e63c · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.025631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.025631Z digest=sha256:b5b48fbb6b6a1ee6e2f902bfac21e6d97262176300335a5567d79ddb642b2381

Observation 4e92cc0a-c592-4f55-937f-8fe7167c3e7f · outbound

This paper cites Visually-prompted language model for fine-grained scene graph generation in an open world.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Visually-prompted language model for fine-grained scene graph generation in an open world

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.518871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:34.029996Z digest=sha256:984b4fe30db2f07b882ed16a986e86af6d6859b19155e91c62e61d4999d63df3

Observation 5efd4230-307a-4c2e-aad5-8fb09f17c7ba · outbound

This paper cites Anetqa: A large-scale benchmark for fine-grained compositional reasoning over untrimmed videos.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Anetqa: A large-scale benchmark for fine-grained compositional reasoning over untrimmed videos

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.507129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:34.033277Z digest=sha256:29523d9c3fe2f99400e0d95ae6aefb865415fbd089dc659f63afe494cccec348

Observation bc2ee29b-2ad4-424c-a429-96c49583a8cd · outbound

This paper cites Compositional video understanding with spatiotem- poral structure-based transformers.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Compositional video understanding with spatiotem- poral structure-based transformers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.495582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:34.036556Z digest=sha256:4fee2b4ebc15bb82ff285d147d3e0fdb32be57c1299ede3ce56c6cbff44c18bf

Observation 154fb1f8-d3ee-486a-a701-94359e54f8da · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Star: Bootstrapping reasoning with reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.040000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.040000Z digest=sha256:5a1f72df52d243b22abd9b8343f195cfbae1937bba08238cf54d3101bc0e2549

Observation d0c859ae-c2f1-4f22-826c-669789dc5852 · outbound

This paper cites Star: Self-taught reasoner bootstrapping reasoning with reasoning.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Star: Self-taught reasoner bootstrapping reasoning with reasoning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T06:02:34.478180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:34.044237Z digest=sha256:a939cf7dfa08d0975fcb90beadd035c8206f410c499de8d1232d784403a6335a

Observation ae857cb0-6a0e-4769-bf8e-5326fb206eb6 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.047689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.047689Z digest=sha256:2485908c2c6a33515e3984537a913ecb0eca0bff2b7b0636e7d6830c4e754728

Observation 5d1e833d-3cbe-4e0e-b38f-906af9c8b3e6 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video instruction tuning with synthetic data, 2024

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.051366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.051366Z digest=sha256:c94c6a9e9150d41be11f763e701f55c61f14f74bd49aeb12cefbdeeada53ca68

Observation 7dfedadf-564d-40ba-a5f1-a2195ac7ed9e · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Instruction-Following Evaluation for Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.054933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.054933Z digest=sha256:72b15fd6f6571405ffeb502d1f2037077162fd83cb185d500f4a1d5dd182125e

Observation c06b9e89-7c5a-42ce-b0e4-e0e881a5be83 · outbound

This paper cites Enhancing Logical Reasoning in Large Language Models through Graph-based Synthetic Data.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Enhancing Logical Reasoning in Large Language Models through Graph-based Synthetic Data

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-12T06:02:34.122126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T06:02:34.058568Z digest=sha256:409a6f2f1e0a8131652c3369c95b262144f1eb1571189de0e04576356f317fc4

Observation 579ee0a6-7a6d-4d79-90ef-a653b2bc9e7b · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training MLVU: Benchmarking Multi-task Long Video Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.062155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.062155Z digest=sha256:dd86f21796c0361e2c08a2d7a7e7567befc578e818c31953a5751a58d0317b85

Observation 950c5de4-7ac2-446d-abc0-7c45252c75f1 · outbound

This paper cites Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.065885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.065885Z digest=sha256:35dd33be3fa44173ade8f4565086067f06d7b7bcbb8eb9997f6cf35cb93c7242

Pith citing papers

Observation d0f2412a-688e-4361-b79c-5c67e17a3d8c · inbound

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes cites this paper.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.003979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.003979Z digest=sha256:5fbf55db864ec35e0ff9b53cae580900e98bec276d97ff5d78efd6f8bfbe5a58

Observation c139af4a-b154-427e-9360-3ae217bb71dd · inbound

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL cites this paper.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.555613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.555613Z digest=sha256:ddef95f8effa1e952a75b3af559cf679d5923299995feef8ffbb7c09d44d7967

Observation ad8db802-3408-4c6c-b710-e779351b9078 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.771532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:41dc016c96df8750c4e959f3db5dc8b0d35e53290310c3419ed5b70dc498c0e1