Pith. sign in

Paper Citation Record · LEDGER

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism

As of 22 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2504.16761.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16761 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:59:50.316255Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e1cb311-37ca-4359-b927-c7ab66e21405 · outbound

This paper cites Deep learning approaches on image captioning: A review,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Deep learning approaches on image captioning: A review,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.120841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.120841Z digest=sha256:ffa6098269b091364e1080ce3fa96651be001bb8e66b4bffcc6fc703756dcd27

Observation b7b0b336-bd7a-41e4-a778-0eb9cd830dfb · outbound

This paper cites From methods to datasets: A survey on image-caption generators,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism From methods to datasets: A survey on image-caption generators,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.893753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.127279Z digest=sha256:887f78ced7af4b08cf19d86bd58fc9eecfbf1ed68aa53025daace933e14b98c3

Observation 87fdbd8f-dce6-4cf8-a912-db90c75a14e4 · outbound

This paper cites A survey on vision transformer,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism A survey on vision transformer,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.132579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.132579Z digest=sha256:a6606f37320f0959f5402b7164e50f379750564237c60b45e191a0f3fb0d55a8

Observation ffe5c0a4-8251-4e05-87b5-52ec63548431 · outbound

This paper cites Text augmentation using bert for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Text augmentation using bert for image captioning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.870814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.137872Z digest=sha256:f29a71c9c246a3200d1852e4f587224e8c94c744cdd22d57bfc6a7bfe4accc58

Observation c73c811d-153d-4a05-9a24-71dcd9445b7f · outbound

This paper cites Contrastive language- image pre-training with knowledge graphs,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Contrastive language- image pre-training with knowledge graphs,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.856515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.143478Z digest=sha256:fa8e63c4de1e31fc8cc85f9372653cece66da64097650bac0367c48b04c254cc

Observation 796b4b33-25ed-4c19-875f-c9efd9dcf15b · outbound

This paper cites Entangled transformer for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Entangled transformer for image captioning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.841837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.148733Z digest=sha256:9ec8a6c3b352be26fa20fc9409ddc028fa54b77257c242ccac125b72d3b6f581

Observation 50f91054-d0dc-4760-8b62-d836717ef205 · outbound

This paper cites Meshed-memory transformer for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Meshed-memory transformer for image captioning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.154037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.154037Z digest=sha256:5561d3b37f82d7a5cbd88de72cd3750e78df66dd0d7370e2906c44a9d7d30cb0

Observation 2ae568fc-702d-49ea-a29f-d76e962bfea7 · outbound

This paper cites Multimodal transformer with multi- view visual representation for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Multimodal transformer with multi- view visual representation for image captioning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.159043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.159043Z digest=sha256:580e55f3c5608d0c9cb17f5f4f5f7b4ce0bd23f2c921428c2fc6a0984a894a96

Observation 01d9fa9e-2d74-4b48-9bff-46ff8475adf6 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Learning transferable visual models from natural language supervision,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.163495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.163495Z digest=sha256:e8e130142c49c247cc1cbf1edede777d81763d55fc7dfd0c6e7b4609d7b5e01f

Observation 6f91ecc6-80b0-4a38-9fe3-dbf7a1a2c04a · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.168472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.168472Z digest=sha256:970104d56e51d77e79e4f502f0a16bd69eaba43d081ef68439affa14324363a6

Observation 9df9c8cd-5337-4a79-860e-83625bfa41cb · outbound

This paper cites Samt- generator: A second-attention for image captioning based on multi-stage transformer network,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Samt- generator: A second-attention for image captioning based on multi-stage transformer network,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.799177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.174099Z digest=sha256:c370900d327ee20a83c85e15a52b416f102ee419cca2d1b331f9048f08a750e6

Observation c7c0c7ab-4bfc-4e79-b76c-8d6c95c89611 · outbound

This paper cites S2 transformer for image captioning.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism S2 transformer for image captioning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.785470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.178706Z digest=sha256:6d915a605bbcbb0e7cea40957c691e90454365334295ba690e27292e1269cccc

Observation 835c1840-d3b7-49a5-a270-64fb83d1497e · outbound

This paper cites Ca-captioner: A novel concentrated attention for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Ca-captioner: A novel concentrated attention for image captioning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.770415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.182811Z digest=sha256:b49818d14f506aad5a119d4613fdc314af7757858583c3d5fb14dc53ddb7b57c

Observation 53e8609a-4a5e-44a0-a4a2-b47043953969 · outbound

This paper cites Image cap- tioning using transformer-based double attention network,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Image cap- tioning using transformer-based double attention network,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.755184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.187255Z digest=sha256:fd6b80c12da705eb40e4b3302f6303f8ca36a781531a041cf578187d5e414257

Observation f6d96a6a-8c9c-40b5-89c7-3410ac9c3490 · outbound

This paper cites With a little help from your own past: Prototypical memory networks for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism With a little help from your own past: Prototypical memory networks for image captioning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.739343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.192488Z digest=sha256:013adb00d8041f8d769e3f6ac6ad043369b220e26d3eac5e0c753b6c4d934e76

Observation a9809b41-cca4-45cf-9dd5-5876bed97e91 · outbound

This paper cites Haav: Hierarchical aggregation of augmented views for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Haav: Hierarchical aggregation of augmented views for image captioning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.725022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.198070Z digest=sha256:a7e03fba89f0454800ce2d61f48e89effda736481bb9277affb83a200a5c404f

Observation 9a3bed3e-b3c6-46b0-bdc9-31a73a0dd0b1 · outbound

This paper cites Dual vision transformer,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Dual vision transformer,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.709713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.202445Z digest=sha256:1ea2306dea7adf2d60c31798f72002b48cf18a47784abec4b41406e7c645d6c5

Observation fe0f43ed-15bd-42ac-b9c9-6adbc4859024 · outbound

This paper cites Spt: Spatial pyramid transformer for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Spt: Spatial pyramid transformer for image captioning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.693549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.206883Z digest=sha256:4717c4853218339202e1290c815a1b93d81e1a39d9c66f9a1b2d639af0c73ccd

Observation e15e0415-71c3-4b20-b798-97944ec6a639 · outbound

This paper cites Show and tell: A neural image caption generator,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Show and tell: A neural image caption generator,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.211563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.211563Z digest=sha256:56c5e9e472a702674afa11a11cb96571da174b4443d43f660e296095e22f29ad

Observation 430aab14-aa6e-4b49-b049-e57d8578222d · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.216074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.216074Z digest=sha256:96bfcec14973b1eee3967dfc5213e1d829f7048a41bb35465bcaacbe0cc32c85

Observation a037e398-a076-4bfd-8ecf-67c5544adaec · outbound

This paper cites A topic-based multi-channel attention model under hybrid mode for image caption,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism A topic-based multi-channel attention model under hybrid mode for image caption,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.659361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.220542Z digest=sha256:aaf3b5fd45d598d923282c9e070a2109a9671938e726d878b7a0ea30cffe8928

Observation 9bfd8312-5085-4675-b998-0214f19a2c6e · outbound

This paper cites Transformer model incorporating local graph semantic attention for image caption,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Transformer model incorporating local graph semantic attention for image caption,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.644315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.224647Z digest=sha256:00ed4c9867e2c73ee72f8cfad09e2f12500e847733b770bdd3aeccc7b6ee04fc

Observation 343b0f64-5407-40fa-b452-aae5d9a44804 · outbound

This paper cites Dynamic-balanced double-attention fusion for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Dynamic-balanced double-attention fusion for image captioning,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.629265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.228795Z digest=sha256:5d9bc1e7f4b275f6b2b01867f5822af1fa1bf241c41611a8bbaea1b107cb19df

Observation 74635107-53e3-4e22-8f17-3d180e6f0165 · outbound

This paper cites Improving image captioning by leveraging intra-and inter-layer global representation in transformer network,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Improving image captioning by leveraging intra-and inter-layer global representation in transformer network,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.614455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.233363Z digest=sha256:35d3d71fe8f51be20c7440799be2f48369fe3dd564560d07e27797f08f803635

Observation 0a3f1b45-9da4-40b4-a68a-9f61f8e7dd69 · outbound

This paper cites Geometry attention transformer with position-aware lstms for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Geometry attention transformer with position-aware lstms for image captioning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.599268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.238055Z digest=sha256:368645d48bfcfd69149282ae202d32cffddda314d7fa65130176538e19a4bb6c

Observation 540ef87b-320b-4f75-85fd-47163a910185 · outbound

This paper cites Vision-enhanced and consensus- aware transformer for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Vision-enhanced and consensus- aware transformer for image captioning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.243481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.243481Z digest=sha256:fcdc3bdd9ccc8009ab4b0255439d087be43d1854c60066f96e94cdbefca93437

Observation f2e2bed9-5fae-4ef8-a674-cdf1488301c1 · outbound

This paper cites A novel cross-fusion method of different types of features for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism A novel cross-fusion method of different types of features for image captioning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.576910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.248490Z digest=sha256:0fc6ee4f2a9aec5ab5b313cc622ff094be07c24227bc7204165043c039e384ab

Observation 19673f21-c8db-439d-bf9b-9379c1a07100 · outbound

This paper cites Transformer- based local-global guidance for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Transformer- based local-global guidance for image captioning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.562334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.253464Z digest=sha256:e9ba50d7d2cb3aa9dcff44f8c1ef4ff541d76bb7950a950b59f45e35553d8c5a

Observation 1f05575c-ef93-461c-92bb-6fa746ef6e96 · outbound

This paper cites Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN).

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.258035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.258035Z digest=sha256:34b2efc329a19c7d115f916b5b157f35181530cf0248888304fc29f7e60efbc4

Observation 047cc0eb-30a7-45a8-8844-a72b6c9746ff · outbound

This paper cites Re- view networks for caption generation,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Re- view networks for caption generation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.547899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.262490Z digest=sha256:d01eac1af5de7910824c756362f3e28f2738a48ca27cedaa45398f6e910bf06e

Observation 2539ab21-cf4c-4897-af52-1c6f36889711 · outbound

This paper cites Semantic compositional networks for visual captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Semantic compositional networks for visual captioning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.532672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.266726Z digest=sha256:ccccd68ca88f4dea36748ba88e4b0ed03237bdfec9eaca826a32f92f020a6e49

Observation d3c9ac04-1a96-4481-91db-d5ecbfc99201 · outbound

This paper cites Knowing when to look: Adaptive attention via a visual sentinel for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Knowing when to look: Adaptive attention via a visual sentinel for image captioning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.271127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.271127Z digest=sha256:a223c6afbeaae78644eb6622a026a4d7ff727e89d6846158eb082c3e7a49a578

Observation 595b87ea-c349-4863-a98d-25d380a7d535 · outbound

This paper cites Self- critical sequence training for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Self- critical sequence training for image captioning,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.275779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.275779Z digest=sha256:6cfe7d2b48d3253e2b0edba2331a6a81db62b16530fe76fed3b1ffb57c2331a0

Observation ad6fd4ab-682f-4b6a-9040-8e04d72780f4 · outbound

This paper cites Gatecap: Gated spatial and semantic attention model for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Gatecap: Gated spatial and semantic attention model for image captioning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.499479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.281535Z digest=sha256:0b3c7db04feb890333f2cfe83722d802dae115dd0cecc01f8fa5be453991f5c1

Observation 8df0a377-a171-43a8-bd92-ff17314d8f73 · outbound

This paper cites Boosting image captioning with attributes,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Boosting image captioning with attributes,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.484054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.285739Z digest=sha256:035be2276cf60a9d3e52908610e417a2a28e1f0457ffbabc9fcc7828a100fc96

Observation 60dbf035-3441-4005-a563-33243def2958 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Bottom-up and top-down attention for image captioning and visual question answering,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.290298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:50.290298Z digest=sha256:c3c0744461ee35ebc9405effe41f9fa124cf709ddf780fb9397d5a2b83c8dbe0

Observation 212a3d8d-5a4a-4043-8042-a318f2f89b10 · outbound

This paper cites Recurrent fusion network for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Recurrent fusion network for image captioning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.459200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.294473Z digest=sha256:cec6d5049c0aa677937f838991d79cc30e73bd2a562e8e5a94b01c2d11a46ebc

Observation 42127f5c-5451-4bfd-b3d8-0973b8a6c1b3 · outbound

This paper cites Exploring visual relationship for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Exploring visual relationship for image captioning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.443247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.298918Z digest=sha256:5b288c9fef9b0ee52cc10cc403a65ebaec1311466a007381a1c2ce8582c7f04e

Observation f280bec8-dabd-4509-b03a-5d18894cfa55 · outbound

This paper cites Auto-encoding scene graphs for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Auto-encoding scene graphs for image captioning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.428184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.303322Z digest=sha256:40c14054f82aa6f202f2e9c707945d442b640caab92bcc80951e62f90442259d

Observation 1c866516-02e6-4314-b6e8-65afb39f6aad · outbound

This paper cites Attention on attention for image captioning,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Attention on attention for image captioning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.413370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.307714Z digest=sha256:19520dc380177bb8e9bb993efaac51ca97733d83341be192639fa71e713e018c

Observation d2ab1097-61cb-4866-aae6-f7430c5de033 · outbound

This paper cites Image captioning using vision encoder decoder model,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Image captioning using vision encoder decoder model,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.398123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.312099Z digest=sha256:5d4219661fc01634f714a0d9aabedb9a78b42a2f0f6e37a8163b61aa944c30bb

Observation 4a49fa3d-0689-4799-897d-a967c9f2188b · outbound

This paper cites Optimal trans- formers based image captioning using beam search,.

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism Optimal trans- formers based image captioning using beam search,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:59:50.383025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:59:50.316255Z digest=sha256:4360f51315284eb293a322c30a0b922fcd60a33bc45ec360e1c205219df8e5a4

Pith citing papers

No inbound Pith citation observations are available.