Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:36:29.862040Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2507.18531.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:36:29.862040Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-21T04:35:39.356919Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T04:39:35.268570Z
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dc6a7e31-8201-4180-ba38-9dafbbc109b2 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42204628-102c-42c1-b5f6-45ebfafbd812 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f414133-f829-442a-b0ed-51acb9ff197f · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bbb85bc4-a452-4b8e-8689-a4428efd383a · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLM: Modeling Video Sequence with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13215b2a-2602-4f18-bf25-eeb3a07d5c9b · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 731f6ea5-715e-42e1-b475-8aae446a1e72 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a16528c4-38d0-4fba-ac9f-3c2da59e467c · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0c156231-190f-4f73-82ca-c18507dc9569 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ad8fe91-e6ad-465f-808c-5951059f9066 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76d80ee7-7cae-45e7-a7ff-4b17429fda7c · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0523fc10-059e-42a5-8757-946baf492dc1 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5abe577e-90ef-4cc7-bd88-a015cf6a7550 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 308f1e94-137a-42e2-afa2-859c0e6185df · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30c5393b-afda-49a0-8f06-22ab38060af6 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Kastner, Yasutomo Kawanishi, Trung Thanh Nguyen, and Junan Chen
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b4dad6c9-5a3b-4527-bb40-5c391699ec5b · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe291c8-0c5b-44b1-a2eb-1298ca80235e · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 93078549-2e74-4600-8872-1a6b7649fb2a · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 288323a4-0bae-499f-a36d-8327cb51eeba · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning The First Place Solution of WSDM Cup 2024: Leveraging Large Language Models for Conversational Multi-Doc QA
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6149dfaf-44e7-4ebe-91ac-a649344baf6b · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning GroundingGPT:Language Enhanced Multi-modal Grounding Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e131865-637e-4ff8-b0d5-85303bf7a887 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8577305e-dd05-4d2f-afa0-f5367d9c21b6 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0bd1384c-21e0-418e-a515-4dca5363be3d · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a5d55e2-8499-4997-8b8f-7e832017e2e5 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfb91417-6894-4509-bd7d-a56b5435c770 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae3ba893-d498-469e-b526-bc5454769ffc · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e52972a-9ce2-4cb0-94fb-a9ebb190d69b · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Reinforced Video Captioning with Entailment Rewards
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bac88a13-b8bd-4d48-a736-82ca1ff6b729 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0481b339-1842-4dcd-a94c-bd18a6cd7b7a · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc1ed0eb-b562-42ae-a323-31a5a96752f1 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Audio-Visual LLM for Video Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0a0e0de-24b5-4f19-a70d-d297d9ff9c47 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4374199a-da92-4eb0-95ac-ec8cef97716d · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dc758d5-3d19-4cd6-847a-be6cc3d62e1e · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07d0839b-d37b-44fd-9e28-5f60c3c7f7dd · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e59a2a-f5cb-47b1-8749-73b6fb7f2721 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In Proceedings of the 31st ACM International Conference on Multimedia
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad3ad4e7-b1ae-4c7d-b6bc-9f6e5a73b46a · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 095751ea-1c5c-419e-8643-040b6bf1c3d9 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 319a4095-870c-416d-a64c-ca4636aa7a69 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31188037-4dab-4cc2-ad2d-635ee4f64788 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f38be541-5772-4fe8-9f6b-4b15eb7bf32b · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 19b5d2ad-b7c1-4368-8a6c-07d279cc7f99 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c973e061-80aa-4b03-9b71-d641c42539c6 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e461aa4-bd38-4e8b-a4bb-e5313aa4cb95 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9b73b3b-3978-44c9-8fe7-ea8da1860696 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c5cf68d7-08d0-45c3-9875-4330fd4e649a · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 53ec807f-a90b-4c6d-a069-f22338da95a8 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68df01c5-db36-4869-a1ba-cc133ae62b8f · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5835d953-a9c2-44f1-84de-92bb33d1d9e7 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d7067fad-f68f-460a-a4fe-d845ebd89253 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31dcf7d2-1e4f-48d5-88ad-189e61538283 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29c554df-1207-4a9c-80f4-2f8247fbcc36 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 03d931d2-76c4-40a9-b6b1-26f88676df9e · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a70c7932-59e7-4ac5-a0af-b4d76ac58e99 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Qwen3 Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcf9863a-fb78-41c6-92da-21c52c478f7c · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be0ef333-e251-4713-8cf3-b698da8e4b71 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 918efbcc-c026-423b-8c60-ec56077fd46b · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e30fd84-e35e-4b66-8987-d6987b31df39 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 87f6d969-dcfa-47c6-a344-a8830c003976 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 535d5528-c211-40a6-89a1-41072bdde2d6 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97411bd4-bb74-4ae5-b95c-028126faebd9 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c16232-4046-467a-8847-08fb25ec7307 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3d196621-58ea-4326-ad9c-5b540ec82e8b · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning OVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1c897235-9b2f-48d4-afe8-7c2a0937e9fc · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In Proceedings of the IEEE conference on computer vision and pattern recognition
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e3a30e2-4a94-4ba2-abf5-c0628384966f · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In Proceedings of the IEEE/CVF international conference on computer vision
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4d657601-617a-4e66-80b6-9ff221e4d053 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In 2020 IEEE International Conference on Multimedia and Expo (ICME)
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a3898842-383a-46dc-bcdc-848142eec563 · outbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc631b3f-aa4e-4c98-9e70-26fcb13e0df4 · inbound
RoadTones: Tone Controllable Text Generation from Road Event Videos IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.