Pith. sign in

Paper Citation Record · LEDGER

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning

As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2507.18531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18531 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:36:29.862040Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T04:35:39.356919Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T04:39:35.268570Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc6a7e31-8201-4180-ba38-9dafbbc109b2 · outbound

This paper cites Qwen2.5-VL Technical Report.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.756643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.756643Z digest=sha256:28927fbd612f49388a4c41ffa1167491f7553f6d77ea3ce66d0fe91055cbb594

Observation 42204628-102c-42c1-b5f6-45ebfafbd812 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.774002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.774002Z digest=sha256:0715767c99428f41f4c2144afc7107a858f15024b1aeb4340eabba1a4868b53c

Observation 2f414133-f829-442a-b0ed-51acb9ff197f · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.630954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:28.804732Z digest=sha256:9370616ecc91eed09f81c24ec7c782439537cd811245f5fa514e66d701cb181e

Observation bbb85bc4-a452-4b8e-8689-a4428efd383a · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLM: Modeling Video Sequence with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.824818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.824818Z digest=sha256:cd32df0c36e11df163fbf68f49605c5e86c3231ae89058505c169402b6b438ad

Observation 13215b2a-2602-4f18-bf25-eeb3a07d5c9b · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.830571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.830571Z digest=sha256:ab25de27d6e186865cf58c8f60a72ef600a83bffac3f0d8ad633310d0193693f

Observation 731f6ea5-715e-42e1-b475-8aae446a1e72 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.598539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:28.862647Z digest=sha256:faecccdcfd3af36d18bf80bedd789f415e4479db906fe8bd5d9d0d0903da1c6b

Observation a16528c4-38d0-4fba-ac9f-3c2da59e467c · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.567933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:28.870648Z digest=sha256:605c5b28e6e0a71167575929a151f36c40113b0c7a0a70fdc0693067532c084d

Observation 0c156231-190f-4f73-82ca-c18507dc9569 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.875907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.875907Z digest=sha256:e3d58df972329c66236c4c4f0a58cbf70f3622c762b15acf086353c00db0b049

Observation 4ad8fe91-e6ad-465f-808c-5951059f9066 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.519400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:28.884001Z digest=sha256:33591fd090d9e81a1b15e94764c1b03b81c8a7d9bd7c4d9b1476c39479503e70

Observation 76d80ee7-7cae-45e7-a7ff-4b17429fda7c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.890640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.890640Z digest=sha256:91ad5c86d132e2693b0105fc7a0347bbed3a273df8c85611f22981004121f3d5

Observation 0523fc10-059e-42a5-8757-946baf492dc1 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.484662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:28.896730Z digest=sha256:315ebf539297a1bdba05d653aa1219b4ecec78184ebaaff99233d5193e058cb1

Observation 5abe577e-90ef-4cc7-bd88-a015cf6a7550 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.902907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.902907Z digest=sha256:25b63b7ab7a606c6e370b378efea5cb0f42dea3b5e7e75682b958c10e73030ba

Observation 308f1e94-137a-42e2-afa2-859c0e6185df · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.911252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.911252Z digest=sha256:d2861f58aad0a027e7def37f3abfb64780f34c3a5d082a07c96024b85c9b3c46

Observation 30c5393b-afda-49a0-8f06-22ab38060af6 · outbound

This paper cites Kastner, Yasutomo Kawanishi, Trung Thanh Nguyen, and Junan Chen.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Kastner, Yasutomo Kawanishi, Trung Thanh Nguyen, and Junan Chen

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:33.409054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:28.918071Z digest=sha256:3ee24f84c4ba1c41c04891c4f6da27687f31cd02e86ab2a7adffa57c48d785d8

Observation b4dad6c9-5a3b-4527-bb40-5c391699ec5b · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.935806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.935806Z digest=sha256:e1a78be53833aaed625b447e038d93106676cab31239e9473c6f0d8f6b5a6d64

Observation dfe291c8-0c5b-44b1-a2eb-1298ca80235e · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.330172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:28.950400Z digest=sha256:24e1af124e1c09a4fd583e691d6dbd63e526cbc5cad80fe82f3e94df84c3181c

Observation 93078549-2e74-4600-8872-1a6b7649fb2a · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.136958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:28.957212Z digest=sha256:8cb31787119ed1682289e035a0e174c3a639edc799f002175c9fb620ac7e4498

Observation 288323a4-0bae-499f-a36d-8327cb51eeba · outbound

This paper cites The First Place Solution of WSDM Cup 2024: Leveraging Large Language Models for Conversational Multi-Doc QA.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning The First Place Solution of WSDM Cup 2024: Leveraging Large Language Models for Conversational Multi-Doc QA

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:36:30.683604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:28.963157Z digest=sha256:2fb22757362940be5c489194d4b3eb32508cced67c4601f5a6150a35c57d5172

Observation 6149dfaf-44e7-4ebe-91ac-a649344baf6b · outbound

This paper cites GroundingGPT:Language Enhanced Multi-modal Grounding Model.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning GroundingGPT:Language Enhanced Multi-modal Grounding Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.980544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.980544Z digest=sha256:01bf113d2b244cd4e62735931346f6b73ab4550ae7649fe194b9a35c485f2573

Observation 6e131865-637e-4ff8-b0d5-85303bf7a887 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.987835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.987835Z digest=sha256:43e8f6f84705edb36db72520c4523dea70387c682f28a2a6acd2cf0085e89e1e

Observation 8577305e-dd05-4d2f-afa0-f5367d9c21b6 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:32.907929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.038967Z digest=sha256:cfac0f9389882da645cb3cf5aadf7a86ed9b12e1d7ddc1d92a9e4b106018fb53

Observation 0bd1384c-21e0-418e-a515-4dca5363be3d · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.074740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.074740Z digest=sha256:59ffcfab2ac495247f137d450d537627fa995909f58894e217644dbd01803afb

Observation 9a5d55e2-8499-4997-8b8f-7e832017e2e5 · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.101479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.101479Z digest=sha256:bd56db2a8cfb5059fcf1edf81c1c5d0a46ff8e084dc372989940e4765390e2b6

Observation cfb91417-6894-4509-bd7d-a56b5435c770 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.124744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.124744Z digest=sha256:02d2149bd6e19b646e9dbbdc2745a36cde317365c321491ba2ca5d78cd5f0ebb

Observation ae3ba893-d498-469e-b526-bc5454769ffc · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.149139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.149139Z digest=sha256:8b2ae8d3f425de50b598ef4a8cbdfef677d749687a620ae004e607ac8ddcd124

Observation 5e52972a-9ce2-4cb0-94fb-a9ebb190d69b · outbound

This paper cites Reinforced Video Captioning with Entailment Rewards.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Reinforced Video Captioning with Entailment Rewards

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:36:30.540454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.160509Z digest=sha256:3b1ae50dd183bb1c2f0308a4e8370485a8b0054982bca9a172246d49cd8327ba

Observation bac88a13-b8bd-4d48-a736-82ca1ff6b729 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.179246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.179246Z digest=sha256:66de7139467240a4dcc0d7e72fa4ab449616bf96d93d1c7d6971f8c5b1b45b93

Observation 0481b339-1842-4dcd-a94c-bd18a6cd7b7a · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:32.645965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.204885Z digest=sha256:e7ddb8b721cf261318364dbe545919724d26340a75c5f405426b8dd3ea26d57f

Observation bc1ed0eb-b562-42ae-a323-31a5a96752f1 · outbound

This paper cites Audio-Visual LLM for Video Understanding.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Audio-Visual LLM for Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.216841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.216841Z digest=sha256:49996e42f3aa51a9fe9a13efb298ae3576984b87eff283aba41feac3fa09ff70

Observation f0a0e0de-24b5-4f19-a70d-d297d9ff9c47 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.236144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.236144Z digest=sha256:1fd911cb006872e18e344c0f6ec54e1634900913b451bc75ccd555951d89d1b1

Observation 4374199a-da92-4eb0-95ac-ec8cef97716d · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.279165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.279165Z digest=sha256:197adc3d14f4f342309424ca34b12ff78e35a6f93b43be1789e7719704d07327

Observation 3dc758d5-3d19-4cd6-847a-be6cc3d62e1e · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.296001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.296001Z digest=sha256:849699fbea75304f546bc1e1d63f32450a9bc441c89b5fc64a817d409ece8dd5

Observation 07d0839b-d37b-44fd-9e28-5f60c3c7f7dd · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.332967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.332967Z digest=sha256:3c87fb869eada0dff762823a8f24b397ca1281512adbefe2b64561ec9a164fe5

Observation 06e59a2a-f5cb-47b1-8749-73b6fb7f2721 · outbound

This paper cites In Proceedings of the 31st ACM International Conference on Multimedia.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In Proceedings of the 31st ACM International Conference on Multimedia

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.326969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.326969Z digest=sha256:842eefd32da564f65b3b047349796c0cf2e289195c2e137475c2c9e573fe0beb

Observation ad3ad4e7-b1ae-4c7d-b6bc-9f6e5a73b46a · outbound

This paper cites Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.394779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.394779Z digest=sha256:0a3e21d5ca8f79fc47c16eba407fe3659d18811c0fcafe4d075b81e71f4c93b4

Observation 095751ea-1c5c-419e-8643-040b6bf1c3d9 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.447467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.447467Z digest=sha256:82e608d66bebf271551fead34e44d6cac7e78473310fdaa2c4459254131ec5d2

Observation 319a4095-870c-416d-a64c-ca4636aa7a69 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:32.040258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.365989Z digest=sha256:fa3cc0ce5acaf4754031129f66d939ecce2de8883afd106793e3d2157ab69d72

Observation 31188037-4dab-4cc2-ad2d-635ee4f64788 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.694759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.491095Z digest=sha256:ab12eb6e8d9b4c81e18317f7c40e2ccf0ff7b4315ef01128b03c711e19738dec

Observation f38be541-5772-4fe8-9f6b-4b15eb7bf32b · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.549968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.499869Z digest=sha256:c8da294786a9fbca847fb950735299906c5302fd0dfb5c20807086eda75707cf

Observation 19b5d2ad-b7c1-4368-8a6c-07d279cc7f99 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.474796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.474796Z digest=sha256:1bda50e61839d3d7e8cdb725fbc7f1d6029d015f126b926f12c1f4b99a356453

Observation c973e061-80aa-4b03-9b71-d641c42539c6 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.138936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.532008Z digest=sha256:30b1849f5c5e18ce378d4f8e08a9e57a5e23af045d00d0f756866d4dffab75a7

Observation 9e461aa4-bd38-4e8b-a4bb-e5313aa4cb95 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.583224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.583224Z digest=sha256:1ed6d815f747aa8f073c459b752653d56d3bc108d0c5e42d295c8079c7212efc

Observation c9b73b3b-3978-44c9-8fe7-ea8da1860696 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.450436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.509421Z digest=sha256:2800d9f15ef440cb85fb08221b8e53a8ac1ce0e2664493bc5765005962794919

Observation c5cf68d7-08d0-45c3-9875-4330fd4e649a · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.038003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.654739Z digest=sha256:ca13f328c269a20e094166d39d7c796b37a4384e0dcba18f870edd50ae97d045

Observation 53ec807f-a90b-4c6d-a069-f22338da95a8 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.673490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.673490Z digest=sha256:1a33dcccee04879fabed9a129c33cb2f0f17999a6b34ed8f792861d438d651b6

Observation 68df01c5-db36-4869-a1ba-cc133ae62b8f · outbound

This paper cites CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.684895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.684895Z digest=sha256:d4adbe1201966051872a4df9699dc7aa9ec38243ba632843bf03da2c1a4d9426

Observation 5835d953-a9c2-44f1-84de-92bb33d1d9e7 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.110929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.607606Z digest=sha256:7ebcb0fb3aee4e066b48ff20dcd946f876558f8814f6c1ab4ca36eaa769d0f3c

Observation d7067fad-f68f-460a-a4fe-d845ebd89253 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.709137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.709137Z digest=sha256:0b4cc90deba6d4c9ebdc2035ece74a74af1f626d58552b2f790b1aa3d5213290

Observation 31dcf7d2-1e4f-48d5-88ad-189e61538283 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.976741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.726957Z digest=sha256:8dec8307b317402ac7436a87d89e476591db973c5a4527cb2c835d3f11b0b3c4

Observation 29c554df-1207-4a9c-80f4-2f8247fbcc36 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.946168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.741989Z digest=sha256:bb2a21b02cc21779a661541d1a853c00f9e6c72e9d28f01c94c6529015bb873a

Observation 03d931d2-76c4-40a9-b6b1-26f88676df9e · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.906135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.748938Z digest=sha256:b46004e61e436fd07e4b1ca6c4d14e86665c484169ff011772d4863a1c560a3c

Observation a70c7932-59e7-4ac5-a0af-b4d76ac58e99 · outbound

This paper cites Qwen3 Technical Report.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Qwen3 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.695562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.695562Z digest=sha256:79c052178a850711c415a2fb6e2c50bba54f114bb686b449bd9d0815d4af3555

Observation fcf9863a-fb78-41c6-92da-21c52c478f7c · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.767987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.767987Z digest=sha256:82033b7afe6882fb59a74dd1ec74ad90bb2607a0a94e9e3335a70e5e2ca415e0

Observation be0ef333-e251-4713-8cf3-b698da8e4b71 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.714217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.714217Z digest=sha256:19120d20e63aac55c2f069d67b1d12769b1568f20868c760a7cfccdd353cb9dc

Observation 918efbcc-c026-423b-8c60-ec56077fd46b · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.787485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.787485Z digest=sha256:e8db5ad6c7c42e04bb286f5aac4a0ce7b71f306b91f2fea504781744a83b2c9c

Observation 5e30fd84-e35e-4b66-8987-d6987b31df39 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.825970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.797201Z digest=sha256:172ecc57cb2fd203401c47e96914458aab0955d1b80ecb0eed2a646c9e37519f

Observation 87f6d969-dcfa-47c6-a344-a8830c003976 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.809595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.809595Z digest=sha256:809eaf10822c39d50b2bc404782cf19c4e0b1907a2374bd83a3a5ffbc80808fc

Observation 535d5528-c211-40a6-89a1-41072bdde2d6 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.762895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.762895Z digest=sha256:a14c3d1a9cbc210b16602918bb72d5f25d26ca76004cc6e6c9fe3cd9337ec651

Observation 97411bd4-bb74-4ae5-b95c-028126faebd9 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.862040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.862040Z digest=sha256:3cd461df8123616dd1e206b16cc136c5eccd31d0f8491a6c2a38f01c79f75c1c

Observation 03c16232-4046-467a-8847-08fb25ec7307 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.864740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.780390Z digest=sha256:239f0666b2ac6c88b10d12e12c57029c5990ce2d9fe9fe68c2516983a98e7d47

Observation 3d196621-58ea-4326-ad9c-5b540ec82e8b · outbound

This paper cites OVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning OVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:36:30.025967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.835315Z digest=sha256:fd2465d6790a178411aea9f4adb9ad0cc1779019ad1e7f961aa651604667b5f7

Observation 1c897235-9b2f-48d4-afe8-7c2a0937e9fc · outbound

This paper cites In Proceedings of the IEEE conference on computer vision and pattern recognition.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:31.241487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.515564Z digest=sha256:97a310cb4a4f1fe184406b4ecc395fe2e02d820508b5836bf9b33c795d8419eb

Observation 9e3a30e2-4a94-4ba2-abf5-c0628384966f · outbound

This paper cites In Proceedings of the IEEE/CVF international conference on computer vision.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In Proceedings of the IEEE/CVF international conference on computer vision

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:32.304827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.356331Z digest=sha256:5e09fe10d23192223c3ba553b961ee1f26a32a30557e99531d51bd4d99f54d35

Observation 4d657601-617a-4e66-80b6-9ff221e4d053 · outbound

This paper cites In 2020 IEEE International Conference on Multimedia and Expo (ICME).

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In 2020 IEEE International Conference on Multimedia and Expo (ICME)

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:31.068892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:36:29.625877Z digest=sha256:14a66cbd222f5ede86ea633b5faa987c2808a1a69b170b6e63d68dc4afde1a78

Observation a3898842-383a-46dc-bcdc-848142eec563 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.852339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.852339Z digest=sha256:c6dbc91e6f15b30bb58360eea7759354489ae7284273ec4f88c0f3dd14e7542e

Pith citing papers

Observation dc631b3f-aa4e-4c98-9e70-26fcb13e0df4 · inbound

RoadTones: Tone Controllable Text Generation from Road Event Videos cites this paper.

RoadTones: Tone Controllable Text Generation from Road Event Videos IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:39:35.270214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:35:39.356919Z digest=sha256:d1488f6a159478dd4af538729cf6fc9190006c88263fb319c348c1bf8bff8db1