Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T06:02:34.065885Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 3 inbound Pith citation observations for arXiv:2412.00161.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T06:02:34.065885Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:50:44.003979Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T12:46:56.770108Z
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3119cec1-1208-49af-b60a-e42b1c201082 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b35a6944-8fb1-4233-a2b9-731f362e2e59 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Self-Training: A Survey
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60a94add-ab46-4473-9ecc-8462cfd3663c · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a11ef4d5-c5a8-4a32-a889-57bf3e8126b4 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Activitynet: A large-scale video benchmark for human activity understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 016fa148-3736-4f07-b712-9bbc0b167869 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Castellano
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0606bf1f-d35a-4860-bf87-df77798b889e · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af7384e6-ea8d-4b67-ba30-5e87a8850081 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Palm: Scaling language modeling with pathways
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1903718e-b72f-4fb8-8a4e-b9814314d933 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training VILA$^2$: VILA Augmented VILA
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc4f296-b068-4ef5-95b3-70180deba854 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video-of-thought: Step-by-step video reasoning from perception to cognition
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9999ec75-bd08-4bf5-8d35-e1901626d8a8 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c780f082-c36e-4600-8633-41bed9d3a0a5 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Agqa: A benchmark for compositional spatio-temporal reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f837daa1-14ec-4911-bde4-68be39403107 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Spatio-temporal action graph networks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b3ac5e22-a313-469a-8b02-f3a1678b48b4 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training V2xum-llm: Cross-modal video summarization with temporal prompt instruction tuning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f429cc-745e-49b7-aa45-5cd8a78a5a7d · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Large Language Models Can Self-Improve
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 052da5bc-60a5-49f4-bbe6-c6d4eb75ff15 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Action genome: Actions as compositions of spatio- temporal scene graphs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8b78d2f8-f0d3-4d92-9903-42b43f168d9b · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87138f8c-a0eb-4861-9a4a-40ed3b4d5510 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Segmentation in the perception and memory of events
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e4e52cce-912a-4608-bbed-a54658110083 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training TVQA: Localized, Compositional Video Question Answering
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f663fcd5-988a-4bc7-89f6-ba7b9bae87d7 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Adap- tive hierarchical graph reasoning with semantic coherence for video-and-language inference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cad64fe0-6a44-43ce-816e-af1fe177f464 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Fine-grained semantically aligned vision-language pre-training
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be4addca-cfb1-4d05-aac2-420fbea3fc52 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Fine-tuning multimodal llms to follow zero-shot demonstrative instructions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c7ca8f38-bb6a-47e3-a050-77ed83fc3a52 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Variational cross- graph reasoning and adaptive structured semantics learning for compositional temporal grounding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f014b3fb-87cb-430b-82e2-4bcdf615f6da · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training VideoChat: Chat-Centric Video Understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f48a6de6-93a5-48ca-8065-f91952b6a129 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ade81513-35b2-418e-b934-2e0ea1566107 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d93cbb1-038d-41ae-8cce-bdb35297f1ae · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c235d881-39db-4be3-8812-6133045c31d5 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Discrim- inative hierarchical modeling of spatio-temporally compos- able human activities
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a0fbc49d-51dc-4167-bb93-7b2c46708b5c · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d9229ea-2905-4b51-abfc-e99b5a57c469 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Vila: On pre-training for vi- sual language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a07a0184-8a19-4924-a76d-810fa316febf · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Improved baselines with visual instruction tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d5a31d2-fb4f-4a83-810a-cedf0f645a44 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training ST-LLM: Large Language Models Are Effective Temporal Learners
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1423de4b-e459-4740-8288-e018e3f737e7 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training TempCompass: Do Video LLMs Really Understand Videos?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e460265-76b2-48ae-bda7-a498738e491a · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8673a5a6-6b84-40ce-b136-0e5358d4aa28 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Orca 2: Teaching Small Language Models How to Reason
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc060ef6-6f9e-4f50-990b-90eb605fe3f0 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Compositional chain-of-thought prompting for large multimodal models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 90db1c2d-792c-4e59-85c6-1bbac8a11746 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Hig: Hierarchical interlacement graph approach to scene graph generation in video understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dfd69555-a6ef-471d-9590-a883f7a767da · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Chatgpt, 2023
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 98d830c6-afdf-487d-8f3e-e4679b3cb855 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Gpt-4v(ision) system card, 2023
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01b87ade-8e47-4abb-a7de-c04542646600 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 017f2444-abdc-4b9d-aa8f-1e80c492723b · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training A computational model of event segmentation from perceptual prediction
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 49821944-a8a2-4995-a7ad-99ff2517c84e · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Moviechat: From dense token to sparse memory for long video understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a6fa4bc-09ce-4b9a-8349-4247ba8c861d · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training To click or not to click: Automatic selection of beautiful thumbnails from videos
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2e89331b-fa67-4d18-bca5-5b502a901f8f · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Human brain activity time-locked to narrative event bound- aries
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0837432b-bbe0-416f-95bb-3540e3e07d63 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Finetuned Language Models Are Zero-Shot Learners
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7447d830-5eb2-4bbe-bd35-503c5c759b4d · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training STAR: A Benchmark for Situated Reasoning in Real-World Videos
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 606647ec-df26-4bb0-945a-b4abb75cc09c · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Graph information bottleneck
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 993c8a96-1962-4fc3-b3af-18704ee7b246 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training FreeVA: Offline MLLM as Training-Free Video Assistant
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be0f4d1d-94ae-4742-8e75-eb1edeaeea3f · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Next-qa: Next phase of question-answering to explaining temporal actions
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8a5256a6-c33a-41e4-88fb-b43819056c86 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video question answer- ing via gradually refined attention over appearance and mo- tion
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f07c334-1dad-4d0a-8602-dd6e1a41508c · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Msr-vtt: A large video description dataset for bridging video and language
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee8ced4e-69f7-4030-b6d7-fbe96905bc17 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Stronger Models are NOT Stronger Teachers for Instruction Tuning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51825da9-c282-49f4-8125-7012d9621d0c · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c50a23-74f1-4a80-abb5-f93e9835e63c · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training CLEVRER: CoLlision Events for Video REpresentation and Reasoning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e92cc0a-c592-4f55-937f-8fe7167c3e7f · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Visually-prompted language model for fine-grained scene graph generation in an open world
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5efd4230-307a-4c2e-aad5-8fb09f17c7ba · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Anetqa: A large-scale benchmark for fine-grained compositional reasoning over untrimmed videos
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bc2ee29b-2ad4-424c-a429-96c49583a8cd · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Compositional video understanding with spatiotem- poral structure-based transformers
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 154fb1f8-d3ee-486a-a701-94359e54f8da · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Star: Bootstrapping reasoning with reasoning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0c859ae-c2f1-4f22-826c-669789dc5852 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Star: Self-taught reasoner bootstrapping reasoning with reasoning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ae857cb0-6a0e-4769-bf8e-5326fb206eb6 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d1e833d-3cbe-4e0e-b38f-906af9c8b3e6 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video instruction tuning with synthetic data, 2024
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dfedadf-564d-40ba-a5f1-a2195ac7ed9e · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Instruction-Following Evaluation for Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c06b9e89-7c5a-42ce-b0e4-e0e881a5be83 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Enhancing Logical Reasoning in Large Language Models through Graph-based Synthetic Data
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 579ee0a6-7a6d-4d79-90ef-a653b2bc9e7b · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training MLVU: Benchmarking Multi-task Long Video Understanding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950c5de4-7ac2-446d-abc0-7c45252c75f1 · outbound
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0f2412a-688e-4361-b79c-5c67e17a3d8c · inbound
DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c139af4a-b154-427e-9360-3ae217bb71dd · inbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad8db802-3408-4c6c-b710-e779351b9078 · inbound
VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.