Pith. sign in

Paper Citation Record · LEDGER

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding

As of 8 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2506.19288.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19288 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:59.323308Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:19:24.701282Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy30
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3528a83-a7ee-40c9-8223-310854cfe123 · outbound

This paper cites A data set for airborne maritime surveillance environments,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding A data set for airborne maritime surveillance environments,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:06.929626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:54.380544Z digest=sha256:1cd7f37ea2ed18d601e55e3f9cb0bac7dbbdd3aabb123099db09ed54bf3b9338

Observation 27844cef-f1f5-49bf-89ae-bd96dc7abad6 · outbound

This paper cites Asy-vrnet: Waterway panoptic driving perception model based on asymmetric fair fusion of vision and 4d mmwave radar,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Asy-vrnet: Waterway panoptic driving perception model based on asymmetric fair fusion of vision and 4d mmwave radar,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:06.623178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:54.500267Z digest=sha256:f49aa33d6f91e05ace1fed5194b9f37a3a7d921df98711a41f015303431e6c35

Observation d187f314-9867-4cd8-92e4-20129d9884a0 · outbound

This paper cites Usvtrack: Usv-based 4d radar-camera tracking dataset for autonomous driving in inland waterways,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Usvtrack: Usv-based 4d radar-camera tracking dataset for autonomous driving in inland waterways,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:06.412323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:54.565772Z digest=sha256:7922f1639748710bf13ddfc43791fad6883158632f1090ff266546e4d8e9c74b

Observation 58f25f48-1449-4749-baeb-add1f0d6227c · outbound

This paper cites Are we ready for unmanned surface vehicles in inland waterways? the usvinland multisensor dataset and benchmark,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Are we ready for unmanned surface vehicles in inland waterways? the usvinland multisensor dataset and benchmark,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:06.214649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:54.693238Z digest=sha256:b3923047de412b9596e0ff4693669f04b1c26a5b5eb483d98fb73d8af79869d2

Observation 2c3b71b9-e01b-4816-bcc3-40a397e49e88 · outbound

This paper cites Achelous: A fast unified water-surface panoptic perception framework based on fusion of monocular camera and 4d mmwave radar,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Achelous: A fast unified water-surface panoptic perception framework based on fusion of monocular camera and 4d mmwave radar,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:06.026304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:54.824633Z digest=sha256:36b7cd8c8f0068b8c780c95acac682e24c2d468bd93ca6fc4b25d79d3ae6ae0f

Observation 79cf704c-20eb-4558-a91b-4e448572d1cc · outbound

This paper cites Achelous++: Power-Oriented Water-Surface Panoptic Perception Framework on Edge Devices based on Vision-Radar Fusion and Pruning of Heterogeneous Modalities.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Achelous++: Power-Oriented Water-Surface Panoptic Perception Framework on Edge Devices based on Vision-Radar Fusion and Pruning of Heterogeneous Modalities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:54.910034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:54.910034Z digest=sha256:51050bd2ea6dd79f9c6e5b8de2043d341bc83612ffb80011c19085bc7620b7a2

Observation 0576f8b1-8e07-4bea-8e14-96544a8f7a1f · outbound

This paper cites Watervg: Waterway visual grounding based on text-guided vision and mmwave radar,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Watervg: Waterway visual grounding based on text-guided vision and mmwave radar,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:05.860607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:55.019750Z digest=sha256:9831009918e9fff6a5c62cad9bf409b595a10232e00718e0a1b193131a01c45c

Observation 7f149857-30d3-4942-a92c-43a369d6c990 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:55.097495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:55.097495Z digest=sha256:0a3519c993acaf47561f4c0958db6159d5aa6aec8ae38a03244026a49cddead8

Observation 431507a3-2ede-4617-b2b5-ea98e04d71f4 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:05.709641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:55.164644Z digest=sha256:f3c477e061acf94c086ed4d62abcf476ee9533db71945fe616b502039122b198

Observation 7334db13-52ec-418f-87ab-77411f5f0147 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:05.546474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:55.220892Z digest=sha256:ca2edf7f1f3337761c40792269cfbcd6cdd2dee5de43f585dab31f7d5e1d915e

Observation 601eddcf-5424-4c4d-a648-0289034f962a · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Im2text: Describing images using 1 million captioned photographs,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:05.365422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:55.270503Z digest=sha256:029b53801e685f303ad34d2dbe97cff8003da1ffa8d22f0e66c7431c52699c5d

Observation da03011a-f677-40e5-b2df-e04dad77a848 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Vizwiz grand challenge: Answering visual questions from blind people,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:05.181603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:55.337293Z digest=sha256:b38514f445aeeaee6772303c73b893b9e8ba6de57e53dfef93aaf899765de840

Observation e1d16594-950c-45b2-9659-bccd1f4ee483 · outbound

This paper cites Learning deep represen- tations of fine-grained visual descriptions,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Learning deep represen- tations of fine-grained visual descriptions,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:04.957616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:55.390881Z digest=sha256:48adc279720262ebd4649bde93a35238330de68a98e4f832972c7c94cd12b7d7

Observation 2cb29d9f-7359-49d0-8fde-146b32fce0fe · outbound

This paper cites Fashion captioning: Towards generating accurate descrip- tions with semantic rewards,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Fashion captioning: Towards generating accurate descrip- tions with semantic rewards,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:04.757576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:55.446462Z digest=sha256:faf8d7859586968b17d82c955c126533fab64accf1f24dfa568a5e31a7c5796a

Observation 4022af13-e19d-43dc-918c-2f1726fbd04d · outbound

This paper cites Break- ingnews: Article annotation by image and text processing,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Break- ingnews: Article annotation by image and text processing,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:04.514738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:55.500107Z digest=sha256:ac4f9601de5919923782370b632be1ad706603dfcacb9b25f25754a14d720ce2

Observation 75037dd3-a52a-4951-b344-c6a3f57d8f03 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Textcaps: a dataset for image captioning with reading comprehension,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:04.264772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:55.578720Z digest=sha256:ae08b9a032f342da3dfb8fec2a125893aeafda6392b98631b527f2ac0c2d9d36

Observation 329a93c2-2979-4e66-b914-a3310f0f000e · outbound

This paper cites Deep learning approaches on image captioning: A review,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Deep learning approaches on image captioning: A review,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:04.048418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:55.637589Z digest=sha256:362982668781150497dc6ce0755de945f752408f5c941ae68911a55d4e18816a

Observation 5ebd6dec-ff85-48d5-ba1d-f98c01e19e20 · outbound

This paper cites Waterscenes: A multi-task 4d radar-camera fusion dataset and benchmarks for autonomous driving on water surfaces,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Waterscenes: A multi-task 4d radar-camera fusion dataset and benchmarks for autonomous driving on water surfaces,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:03.821704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:55.692726Z digest=sha256:e52a88d700c29335c8696e757b4aea9550ac021a8b9fb32db0d1776a80ad8a54

Observation 8a1acd54-b7bc-4ea2-8881-6048d843805c · outbound

This paper cites an unresolved cited work.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:13:03.672962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:55.735704Z digest=sha256:e727727938be27cf45b533dd067b37e4a0c73151f745688ec4e737bcb7791793

Observation c198008a-e337-447b-8511-0a07ceeeca1c · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:55.793428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:55.793428Z digest=sha256:2d971454ea947d19c58ee43817acbc555b58e97baf0087e8e2590f4ddbc61a8d

Observation dfaff453-8f72-497a-862d-28849e747998 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Learning transferable visual models from natural language supervision,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:55.845170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:55.845170Z digest=sha256:31c7adfd928f6b3affde6fd6b2c7d6bfc938a2e78323d063a71187d30f524675

Observation a0c4bf0c-1a5d-4af1-aaf4-e86de4796e50 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Align before fuse: Vision and language representation learning with momentum distillation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:55.902203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:55.902203Z digest=sha256:5a9f24764a15755e07e7fd6f453584ce8f74576a51199ec1558feb8203646835

Observation 615558af-e3db-48dd-94ec-a9b5986425d3 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:55.972188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:55.972188Z digest=sha256:dd2a3b055783f7c2ad60a10325f06aade5b4ad9574d0f57e1f07b5adc016ce83

Observation 8d91ef8b-87e7-4f03-b628-b2e580385b65 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:56.022808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:56.022808Z digest=sha256:880b2259bc934ee9b6ef6efcb48f2d7d6d8754717e868180854d96fbe3ce34da

Observation 811bea4f-c2ec-423b-9955-40aacfb6ac4f · outbound

This paper cites Visual instruction tuning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Visual instruction tuning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:56.086648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:56.086648Z digest=sha256:84c9baa56cd43886d55074eb20bb907099f0cfbae5e8dbb07d734e5d49b9fcae

Observation 7e759b75-38a1-4828-a366-2cae7131c08d · outbound

This paper cites Qwen2. 5 technical report,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Qwen2. 5 technical report,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:03.389808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:56.155414Z digest=sha256:89c4d0e7615c0835bf3b78a8fc2c61f22d2c3beed00e0627b6871a44bd613155

Observation 3174bdf5-d3e6-42cf-a494-b862fee32ba3 · outbound

This paper cites Task-adaptive attention for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Task-adaptive attention for image captioning,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:56.258705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:56.258705Z digest=sha256:70bf9aa75f946e6542240158e192bf194dc29c554bd24b9cda496dedb141a13c

Observation f5901940-4ad3-43da-97f9-3d76d4d4e981 · outbound

This paper cites Vision-enhanced and consensus- aware transformer for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Vision-enhanced and consensus- aware transformer for image captioning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:03.193848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:56.355939Z digest=sha256:708e26e0156dd172dc376f850a10b4ef118c8aee6881f7f6f5d107894f19d110

Observation b1cb9332-8d86-436e-8410-4221571f69a7 · outbound

This paper cites Adaptive path selection for dynamic image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Adaptive path selection for dynamic image captioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:02.958702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:56.434598Z digest=sha256:dbd3cab3395cdb71a01344957317c0681f23f1e7e5883e4c244890cf9ec67222

Observation c81c2a21-60cd-4941-b0fa-f8d52be3c83e · outbound

This paper cites Double-stream position learning trans- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 14 former network for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Double-stream position learning trans- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 14 former network for image captioning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:02.744958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:56.534168Z digest=sha256:5b67e6babd2ebd69fa064941fa0f2109fd7b05e78293202b65eb329915c622f2

Observation 4ebe8d1e-bf2b-449c-8e2e-150ce11c4437 · outbound

This paper cites A comprehen- sive survey of 3d dense captioning: Localizing and describing objects in 3d scenes,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding A comprehen- sive survey of 3d dense captioning: Localizing and describing objects in 3d scenes,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:02.556920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:56.603663Z digest=sha256:8e20e37b77537d7db85bb414999647a12567daa56b6d704313b826ef7c9d0fdf

Observation 31978f6c-b1d4-4a7c-a8c2-79ae2febe81b · outbound

This paper cites Spt: Spatial pyramid transformer for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Spt: Spatial pyramid transformer for image captioning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:02.375032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:56.743713Z digest=sha256:7e6670aa2571f9c82b5bac4bb3f49d8e5020bf6f9be532a8df5829f99dff9735

Observation c05fa789-e8b7-46ba-b3e6-d75d53bcf8d7 · outbound

This paper cites Multimodal transformer with multi- view visual representation for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Multimodal transformer with multi- view visual representation for image captioning,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:56.928074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:56.928074Z digest=sha256:81e9a6fa40d8d1ac051004187a1ae0a945ffecf03695bff0deb8a8e2afac928c

Observation ccabe1f6-81d6-4c7e-9296-bd3d434b7c5b · outbound

This paper cites MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:56.998779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:56.998779Z digest=sha256:77df7c9870245317208b614a554dc5e414081848eb643eb3c29863a5f8db7b35

Observation c5e489b2-72b8-42ec-b315-0eda9e0dcf3a · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:02.224750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:57.061781Z digest=sha256:8e8c822ca665054a37244d30e0bd06d95083d72be8ce3c0a5a87f604b8d3ad33

Observation 500fdbab-b469-46d2-af42-cf3f814ed487 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:57.142952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:57.142952Z digest=sha256:647dd153bab8a8065006ff47650c01ab900ad836a35340729f64c6c781ed1c7a

Observation edbf4cf7-badf-40a5-b9ba-4172a7c575d7 · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:57.273153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:57.273153Z digest=sha256:d941b4ca67b8a1f4d7c6f5f8afccb44265488645f5edbff7eb9006da80cd364c

Observation aa0ae569-7531-44af-8e29-9d3b379ef496 · outbound

This paper cites Mask-vrdet: A robust riverway panoptic perception model based on dual graph fusion of vision and 4d mmwave radar,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Mask-vrdet: A robust riverway panoptic perception model based on dual graph fusion of vision and 4d mmwave radar,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:01.997200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:57.364159Z digest=sha256:683d0565d76723f3d245ccfc389bbbe0ca21f856277d7e4bb8e1c8e1d422098a

Observation af2b816b-3b25-4b3c-b476-0dcaae6dfda0 · outbound

This paper cites NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave Radar.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave Radar

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:59.949657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:57.454821Z digest=sha256:07c2a729ac90eda72bb15adaedc7d00a16156cfdb2d11c2119fb84f4f2209c2a

Observation 6cab96c3-08cf-4e27-9026-62cc3e78ca61 · outbound

This paper cites GPT-4 Technical Report.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding GPT-4 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:57.527835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:57.527835Z digest=sha256:5732923ff96ed6efb36c842eed6b04e1654d46cc700cd825f87bbaa31080d6c4

Observation 081bb90d-1ec1-4d9d-bdec-b68e315506ea · outbound

This paper cites DeepSeek-V3 Technical Report.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding DeepSeek-V3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:57.627616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:57.627616Z digest=sha256:258b03ed13d3b467c93faf38d87b21c6ebc494a39a65186906f4b50ae52a8083

Observation 2111c1d1-a8db-441d-b219-6cc38c9c150c · outbound

This paper cites Mobileclip: Fast image-text models through multi-modal reinforced training,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Mobileclip: Fast image-text models through multi-modal reinforced training,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:01.851137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:57.757659Z digest=sha256:169f02cc07308e51171ea9340eded4c658b7e76d45a5e299231a3e826660b83c

Observation 0c5fce76-134f-4e97-a4e0-9e70b75f8a98 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Coca: Contrastive captioners are image-text foundation models,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:57.830163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:57.830163Z digest=sha256:5bf2900dd0cbecb8a4a0bec4846374f37c117f299d82171de8969a3755a34be0

Observation d61b9442-439c-4671-96b3-98a58e1265df · outbound

This paper cites Sigmoid loss for language image pre-training,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Sigmoid loss for language image pre-training,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:57.922031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:57.922031Z digest=sha256:5afa47a0fa24986cc218eee2850a3594a9b2a0308f10212b98fc6c9fcc368de8

Observation 7bd0ff1f-2783-425c-8f03-d20c7b35d236 · outbound

This paper cites Agent attention: On the integration of softmax and linear attention,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Agent attention: On the integration of softmax and linear attention,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:01.649115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:58.027415Z digest=sha256:0239b46b73c71c902c230960035c5a714ef7fdd6ca872377f2c102e530719a47

Observation abb2ff5e-e04c-496d-9c7a-99f5e45d3fe4 · outbound

This paper cites Dual-level collaborative transformer for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Dual-level collaborative transformer for image captioning,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:01.360157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:58.117598Z digest=sha256:49a59b3e3e18a03a13f095ad413cc8b11175af9d7033c8e60a5c5a85fbf17691

Observation 810a7100-652f-4612-bdfb-ef681593103d · outbound

This paper cites Comprehending and ordering seman- tics for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Comprehending and ordering seman- tics for image captioning,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:01.148678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:58.239788Z digest=sha256:b6c23be52ebad7a0bcd295d33530043e2b7926c48b688bd97fa0d62dc4b58230

Observation c1a5c21e-a9df-45cd-b725-08b75466a7fa · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Oscar: Object-semantics aligned pre-training for vision-language tasks,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:58.335734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:58.335734Z digest=sha256:c4e666c7c59e01d5c2fa022e516c96bd0aa71d4ad16398f584d2a5850d3fad4b

Observation 145b0466-f1ea-4ed3-9f35-537488884268 · outbound

This paper cites an unresolved cited work.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:58.477687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:58.477687Z digest=sha256:5dc1128b637ca2b8350ead670ec7975839bc3cd11acae7ed9f58f0c52dcfd73a

Observation 666e3eec-e2e9-46bd-8dd8-258b15ae58fe · outbound

This paper cites Qwen2.5-VL Technical Report.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Qwen2.5-VL Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:58.578001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:58.578001Z digest=sha256:09870c182c6835a5e8ee0b2ac097673961f73fad3316f04b6ff065d59298f599

Observation 427794b7-6134-4495-99f0-437837990883 · outbound

This paper cites MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:59.596017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:58.675967Z digest=sha256:33259cb587f9658cdffc5903a39fdd4b39d74e820feb13f1b4152b1270415886

Observation 44a722cf-0c9f-4a35-94af-64a5b07bfb53 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Rouge: A package for automatic evaluation of summaries,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:58.785881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:58.785881Z digest=sha256:df3adb9b5f639fb0639f944c114f0180c3536269c2ca145890b104dfd5c1ba5c

Observation 976e26ed-e61f-4926-a8a6-318f66543f8d · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Bleu: a method for automatic evaluation of machine translation,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:58.879043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:58.879043Z digest=sha256:e8a6baf72acf0e634e17130208e27bb50673fd4c4498cbb774584f99f4d17fa3

Observation 410c64ab-1944-44a5-91c3-6d9a08ea2364 · outbound

This paper cites Lavie and A.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Lavie and A

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:00.851698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:59.001834Z digest=sha256:eb4f5dc78424206d2520e4a5b1639f32c85ec9735e151ef9da439407eb24ef2d

Observation 7157aa19-ca8c-463a-8a93-0f58f8e2020a · outbound

This paper cites Cider: Consensus- based image description evaluation,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Cider: Consensus- based image description evaluation,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:59.106904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:59.106904Z digest=sha256:c8551f31c43e53d6fd40e4171e43c5c99097297b31ce7e1eed2c8afad4b82347

Observation 1f34d844-ea05-4991-95c6-f0c3e2dfafa7 · outbound

This paper cites Drivelm: Driving with graph visual question answering,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Drivelm: Driving with graph visual question answering,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:00.559148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:59.233791Z digest=sha256:1c8ed250bcaaa31ffd7ec42679e4a3102901db6000bb20334e639535ac2a034d

Observation 53883036-0972-4f8f-85de-8dcd25d81229 · outbound

This paper cites an unresolved cited work.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:13:00.327800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:12:59.323308Z digest=sha256:93cb26890b90037484d63550ffcddd29853db982575059593190999e6306793a

Pith citing papers

Observation 9fdddfd6-4ace-4f4d-9d43-1eb7af353c64 · inbound

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey cites this paper.

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T19:19:24.701282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:19:24.701282Z digest=sha256:315bdcf2152a048ecad68c9cf57c927f096d93a47fdd764ffa9f0229c9b42ca1