REVIEW 4 major objections 4 minor 130 references
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A survey of attention-based image captioning across languages finds English systems still lead, while multilingual captioning is held back by scarce data and imperfect metrics.
desk verdict A genuinely useful multilingual survey with a load-bearing reliability problem in its central table. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery that carries the survey is its classification scheme rather than a new algorithm. The central object is the attention mechanism: scaled dot-product attention and its multi-head extension from the transformer [46], soft versus hard attention, and bottom-up versus top-down attention. The survey uses this scheme to organize dozens of models, then adds two comparison tools—Table 2, which records state-of-the-art BLEU, METEOR, ROUGE, CIDEr, and SPICE scores per dataset and language, and Table 6, which summarizes multilingual captioning models and their reported improvements. These tables are what allow the authors to argue that English and non-English captioning are at different stages of maturity.
What would settle it
Re-reading each row of Table 2 against the cited papers would settle the data side: if even a handful of BLEU or CIDEr values differ from the originals, the comparative analysis in Section 4.6 loses its factual foundation. For the coverage side, re-running the stated keyword search on additional literature databases and listing major attention-based captioning papers it misses would test whether the survey's sample is representative.
Extended reading notes
Core claim
The paper's central claim is that the field of image captioning has consolidated around attention-based transformer models, and that this consolidation has not yet benefited all languages equally. On its own terms, the survey establishes a taxonomy: handcrafted approaches, deep learning encoder-decoder models (CNN-RNN, CNN-LSTM, CNN-GRU), transformer-based models, attention-based mechanisms (soft, hard, bottom-up, top-down, multi-head), and graph-based relational models. It then reads the comparative results in Table 2 as showing that English systems evaluated on MS COCO reach the highest BLEU and CIDEr scores, while Arabic, Indonesian, Bengali, Vietnamese, and Myanmar systems post lower scores on smaller or translated datasets. The authors conclude that semantic inconsistency, limited reasoning ability, and data scarcity in non-English languages are the main bottlenecks, and they point to multimodal learning, multilingual transfer, and real-time applications as the next research targets.
Load-bearing premise
The survey's conclusions rest on the assumption that the performance numbers in its main comparison table are faithful transcriptions of the cited papers and that its keyword search of the scholarly literature from 2018 to 2024 returned a representative sample of the field.
Editorial extensions
If this is right
- New captioning systems should build on transformer encoders with multi-head or bottom-up attention rather than plain CNN-LSTM decoders, because the surveyed state-of-the-art methods all rely on attention-based alignment.
- Machine translation of English captions is not a reliable way to build non-English captioning corpora; the survey reports that translated captions produce poor sentence structure, so native-speaker annotation is the safer path.
- Evaluation of captioning models should report several metrics together, since BLEU, METEOR, ROUGE, CIDEr, and SPICE each capture different aspects and no single score tracks human judgment well.
- Non-English image captioning will need larger, culturally grounded datasets before transformer models can close the gap with English results.
- Graph-based representation of object relationships is a promising route to more detailed captions, because it encodes interactions between objects rather than only their presence.
Reading between the lines
- Editorial inference: the cross-language score gaps in Table 2 likely mix model quality with dataset difficulty, since caption counts, vocabulary size, and annotation protocols differ; a fair comparison would need matched multilingual benchmarks.
- Editorial inference: a testable next step is to run one fixed transformer backbone on parallel captions in several languages to isolate how much of the performance gap comes from language-specific data rather than architecture.
- Editorial inference: the survey's emphasis on reasoning limits suggests that object-relationship graphs or memory-augmented transformers, which it reviews only briefly, may matter more for caption quality than larger attention heads.
- Editorial inference: multilingual captioning could borrow cross-lingual pretraining from machine translation to reduce the need for large native caption sets.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of attention-based and transformer-based image captioning, with a particular focus on multilingual captioning. It proposes a taxonomy (hand-crafted, deep-learning, transformer-based, attention-based, graph-based), reviews benchmark datasets and evaluation metrics (BLEU, METEOR, ROUGE, CIDEr, SPICE), and discusses limitations and future directions. The paper also provides a comparative table of state-of-the-art methods (Table 2) and a separate summary of multilingual models (Table 6). The central claim is that the survey serves as a comprehensive, reliable reference for attention-based image captioning across languages.
Significance. If the compilation were reliable, the survey would fill a useful niche: most prior surveys focus on English-centric image captioning, whereas this one explicitly targets Arabic, Bengali, Indonesian, Myanmar, and Vietnamese, among others. The taxonomy, dataset descriptions, and discussion of evaluation metrics are broadly useful for researchers entering the area. The paper does not present new experiments or machine-checked derivations, so its value depends entirely on the accuracy and completeness of its literature compilation. The multilingual scope is a genuine differentiator, but the current internal inconsistencies in Table 2 and the limited search protocol substantially reduce the confidence a reader can place in the survey as a reference.
major comments (4)
- [§1.3, §4.2.3, Table 2, Table 6] The paper claims in §1.3 that the survey covers 'including, but not limited to, English, Arabic, Vietnamese, Myanmar, and Indonesian,' and §4.2.3 describes the Vietnamese model of [3], which also appears in Table 6. However, Table 2 contains no Vietnamese row at all, so a reader relying on the main comparative table would incorrectly conclude that no Vietnamese image-captioning work met the inclusion criteria. This directly undermines the comprehensiveness claim that is central to the survey. Either add the Vietnamese entry to Table 2 or revise the scope statement to match what is actually tabulated.
- [Table 2, §4.6, §5.2] Several dataset labels in Table 2 contradict the text. Row [56] is labeled 'Arabic Flickr616,' but §5.2 explains that the model of [56] was trained and tested on images from MS COCO and Flickr8k, with a test split of 616 images; there is no dataset named 'Flickr616.' Row [88] is labeled 'Bengali Flickr4k-Bn,' while §4.6 states that the method of [88] was applied to the Bengali Flickr8k dataset, and §4.2.3 also describes Flickr8k translated with Google Translator. In addition, §4.6 says the table includes 'Arabic datasets such as Flickr8k and Flickr30k,' but no Arabic Flickr30k row exists in Table 2. Since Table 2 is the central quantitative synthesis of the survey, these inconsistencies are load-bearing and must be corrected or explained.
- [§4.4.1, References [118] and [51]] The text states that Show, Attend, and Tell was the first attentive deep paradigm for image captioning and cites [118]. However, reference [118] in the bibliography is 'Show, attend and read: A simple and strong baseline for irregular text recognition' by Li et al., not Xu et al.'s 'Show, attend and tell' paper, which is reference [51]. This is not a mere formatting slip: a survey that aims to be a reliable reference must correctly attribute foundational work. The citation should be fixed, and the reference list should be checked for similar mismatches across the manuscript.
- [§1.3] The search protocol is too thin to support the claim of a comprehensive multilingual survey. It reports only five English keywords searched in Google Scholar for 2018–2024, with no language-specific search terms, no inclusion/exclusion criteria, no screening counts, and no validation against known multilingual benchmarks or seed papers. Given that the paper's distinctive contribution is multilingual coverage, the protocol should be substantially expanded (e.g., searches in the target languages or using known multilingual datasets as seeds) and reported in enough detail for the search to be reproducible.
minor comments (4)
- [§1, first paragraph] The sentence 'the coverage, inventiveness, and complexity of the generated sentences coverage, inventiveness, and complexity are limited' contains a duplicated phrase and should be rewritten.
- [Table 5 and §6.3] The METEOR expansion in Table 5 is 'Metric for Evaluation of Translation with Explicit ORdering,' while §6.3 calls it 'The Metric for Explicit Ordering Translation Evaluation.' These should be reconciled, and the odd capitalization in 'ORdering' should be fixed.
- [References [75] and [101]] References [75] and [101] are the same paper ('Geometry attention transformer with position-aware lstms for image captioning') and should be consolidated, especially because both appear in Table 2 and Table 6.
- [§2.2] The sentence 'The study categorized supervised learning-based techniques into encoder-decoder architecture-based, compositional architecture-based, attention-based, semantic concept-based, stylized captions, dense image captioning, and novel object-based image captioning, as outlined by [2]' is confusing: the list mixes categories from the cited study with those from [2], and the attribution is unclear.
Circularity Check
Survey is a literature review; no derivation chain reduces to its own inputs or to load-bearing self-citations.
full rationale
The paper is a survey of attention-based and transformer-based image captioning methods. Its central claims are summaries of external papers, benchmark datasets, and evaluation metrics; there is no derivation chain in which a quantity is defined in terms of the thing it is supposed to predict. The authors cite several of their own prior works ([8], [9] for texture features in Section 1; [92] for deep belief networks in Section 4.1; [129], [130] for GAN loss and brain MR-CT synthesis in Section 7), but these are background or motivational citations and are not load-bearing for the survey's conclusions about the state of the field. Table 2 transcribes reported performance numbers from the cited literature, and such transcription is a reporting act, not a fitted-parameter-then-predicted-on-the-same-data loop. Internal inconsistencies between Table 2 and Sections 4.6/5.2/5.3 (e.g., no Vietnamese row despite the claimed Vietnamese coverage, and dataset-label mismatches for rows [56] and [88]) are accuracy and reporting concerns, not circularity: they do not show that any result equals its input by construction. Per the scoring rules, the absence of self-consistent derivation or statistically forced prediction yields a non-finding with score 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The Google Scholar keyword search from 2018 to 2024 captures the relevant literature for image captioning, particularly multilingual work.
- domain assumption The numbers in Table 2 are accurately reproduced from the cited papers.
- domain assumption Standard evaluation metrics (BLEU, METEOR, CIDEr, ROUGE, SPICE) are valid indicators of caption quality as used in the field.
Cite this review
Pith. "Pith review of Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation." pith.science (2026). https://pith.science/paper/74XWTFH7
@misc{pith2026250605399,
author = {Pith},
title = {Pith review of: Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/74XWTFH7}},
note = {Machine review of arXiv:2506.05399}
}
read the original abstract
Image captioning involves generating textual descriptions from input images, bridging the gap between computer vision and natural language processing. Recent advancements in transformer-based models have significantly improved caption generation by leveraging attention mechanisms for better scene understanding. While various surveys have explored deep learning-based approaches for image captioning, few have comprehensively analyzed attention-based transformer models across multiple languages. This survey reviews attention-based image captioning models, categorizing them into transformer-based, deep learning-based, and hybrid approaches. It explores benchmark datasets, discusses evaluation metrics such as BLEU, METEOR, CIDEr, and ROUGE, and highlights challenges in multilingual captioning. Additionally, this paper identifies key limitations in current models, including semantic inconsistencies, data scarcity in non-English languages, and limitations in reasoning ability. Finally, we outline future research directions, such as multimodal learning, real-time applications in AI-powered assistants, healthcare, and forensic analysis. This survey serves as a comprehensive reference for researchers aiming to advance the field of attention-based image captioning.
Figures
Figures from the paper (13 more)
Reference graph
Works this paper leans on
-
[88]
M. Humaira, P. Shimul, M. A. R. K. Jim, A. S. Ami, F. M. Shah, A hybridized deep learning method for bengali image captioning, Interna- tional Journal of Advanced Computer Science and Applications 12 (2) (2021)
work page 2021
-
[56]
H. A. Al-Muzaini, T. N. Al-Yahya, H. Benhidour, Automatic arabic im- age captioning using rnn-lst m-based language model and cnn, Interna- tional Journal of Advanced Computer Science and Applications 9 (6) (2018)
work page 2018
-
[118]
H. Li, P. Wang, C. Shen, G. Zhang, Show, attend and read: A simple and strong baseline for irregular text recognition, in: Proceedings of the AAAI conference on artificial intelligence, V ol. 33, 2019, pp. 8610– 8617
work page 2019
-
[3]
H. N. Tien, T.-H. Do, V .-A. Nguyen, Image captioning in vietnamese language based on deep learning network, in: International Conference on Computational Collective Intelligence, Springer, 2020, pp. 789–800
2020
-
[51]
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, Y . Bengio, Show, attend and tell: Neural image caption generation with visual attention, in: International conference on machine learning, PMLR, 2015, pp. 2048–2057
2015
-
[1]
S. Bai, S. An, A survey on automatic image caption generation, Neuro- computing 311 (2018) 291–304
2018
-
[2]
M. Z. Hossain, F. Sohel, M. F. Shiratuddin, H. Laga, A comprehensive survey of deep learning for image captioning, ACM Computing Surveys (CsUR) 51 (6) (2019) 1–36
2019
-
[4]
Cheikh, M
M. Cheikh, M. Zrigui, Active learning based framework for image cap- tioning corpus creation, in: International Conference on Learning and Intelligent Optimization, Springer, 2020, pp. 128–142
2020
Show all 130 references
-
[5]
J. Yu, J. Li, Z. Yu, Q. Huang, Multimodal transformer with multi-view visual representation for image captioning, IEEE transactions on circuits and systems for video technology 30 (12) (2019) 4467–4480
2019
-
[6]
Biswas, M
R. Biswas, M. Barz, D. Sonntag, Towards explanatory interactive image captioning using top-down and bottom-up features, beam search and re- ranking, KI-K¨unstliche Intelligenz 34 (4) (2020) 571–584
2020
-
[7]
Ghandi, H
T. Ghandi, H. Pourreza, H. Mahyar, Deep learning approaches on image captioning: A review, ACM Comput. Surv. 56 (3) (oct 2023)
2023
-
[8]
O. S. Al-Kadi, Combined statistical and model based texture features for improved image classification, in: 4th IET International Conference on Advances in Medical, Signal and Information Processing-MEDSIP 2008, IET, 2008, pp. 1–4
2008
-
[9]
O. S. Al-Kadi, Supervised texture segmentation: a comparative study, in: 2011 IEEE Jordan Conference on Applied Electrical Engineering and Computing Technologies (AEECT), IEEE, 2011, pp. 1–5
2011
-
[10]
Ayesha, S
H. Ayesha, S. Iqbal, M. Tariq, M. Abrar, M. Sanaullah, I. Abbas, A. Rehman, M. F. K. Niazi, S. Hussain, Automatic medical image in- terpretation: State of the art and future directions, Pattern Recognition (2021) 107856
2021
-
[11]
Chendake, P
P. Chendake, P. Korpal, S. Bhor, R. Bansal, S. Patil, D. Deshpande, Learning system for kids, International Journal of Recent Advances in Multidisciplinary Topics 2 (6) (2021) 71–75
2021
-
[12]
Ogura, N
A. Ogura, N. Hayashi, T. Negishi, H. Watanabe, E ffectiveness of an e- learning platform for image interpretation education of medical staff and students, Journal of digital imaging 31 (5) (2018) 622–627
2018
-
[13]
A. K. Muhammed Kunju, S. Baskar, S. Zafar, B. AR, A transformer based real-time photo captioning framework for visually impaired peo- ple with visual attention, Multimedia Tools and Applications (2024) 1– 20
2024
-
[14]
D. H. Fudholi, Y . Windiatmoko, N. Afrianto, P. E. Susanto, M. Suyuti, A. F. Hidayatullah, R. Rahmadi, Image captioning with attention for smart local tourism using e fficientnet, in: IOP Conference Series: Ma- terials Science and Engineering, V ol. 1077, IOP Publishing, 2021,...
2021
-
[15]
Nivedita, P
M. Nivedita, P. Chandrashekar, S. Mahapatra, Y . A. V . Phamila, S. K. Selvaperumal, Image captioning for video surveillance system using neural networks, International Journal of Image and Graphics (2021) 2150044
2021
-
[16]
Hoxha, F
G. Hoxha, F. Melgani, B. Demir, Toward remote sensing image retrieval under a deep image captioning perspective, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 13 (2020) 4462–4475
2020
-
[17]
Z. Wang, Z. Huang, Y . Luo, Paic: Parallelised attentive image caption- ing, in: Australasian Database Conference, Springer, 2020, pp. 16–28
2020
-
[18]
Shuster, S
K. Shuster, S. Humeau, H. Hu, A. Bordes, J. Weston, Engaging image captioning via personality, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 12516–12526
2019
-
[19]
Fujiyoshi, T
H. Fujiyoshi, T. Hirakawa, T. Yamashita, Deep learning-based image recognition for autonomous driving, IATSS research 43 (4) (2019) 244– 252
2019
-
[20]
A. P. Shah, J.-B. Lamare, T. Nguyen-Anh, A. Hauptmann, Cadp: A novel dataset for cctv tra ffic camera based accident analysis, in: 2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (A VSS), IEEE, 2018, pp. 1–9
2018
-
[21]
Guinness, E
D. Guinness, E. Cutrell, M. R. Morris, Caption crawler: Enabling reusable alternative text descriptions using reverse image search, in: Pro- ceedings of the 2018 CHI Conference on Human Factors in Computing Systems, 2018, pp. 1–11
2018
-
[22]
Huang, L
Q. Huang, L. Yang, H. Huang, T. Wu, D. Lin, Caption-supervised face recognition: Training a state-of-the-art face model without manual an- notation, in: European Conference on Computer Vision, Springer, 2020, pp. 139–155
2020
-
[23]
P. P. Khaing, et al., Attention-based deep learning model for image cap- tioning: a comparative study, International Journal of Image, Graphics and Signal Processing 11 (6) (2019) 1
2019
-
[24]
F. Chen, X. Li, J. Tang, S. Li, T. Wang, A survey on recent advances in image captioning, in: Journal of Physics: Conference Series, V ol. 1914, IOP Publishing, 2021, p. 012053
1914
-
[25]
Stefanini, M
M. Stefanini, M. Cornia, L. Baraldi, S. Cascianelli, G. Fiameni, R. Cuc- chiara, From show to tell: a survey on deep learning-based image cap- tioning, IEEE transactions on pattern analysis and machine intelligence 45 (1) (2022) 539–559
2022
-
[26]
G. Luo, L. Cheng, C. Jing, C. Zhao, G. Song, A thorough review of models, evaluation metrics, and datasets on image captioning, IET Im- age Processing 16 (2) (2022) 311–332
2022
-
[27]
Zohourianshahzadi, J
Z. Zohourianshahzadi, J. K. Kalita, Neural attention for image cap- tioning: review of outstanding methods, Artificial Intelligence Review (2021) 1–30
2021
-
[28]
Senior, G
H. Senior, G. Slabaugh, S. Yuan, L. Rossi, Graph neural networks in vision-language image understanding: a survey, The Visual Computer (2024) 1–26
2024
-
[29]
T. Pang, P. Li, L. Zhao, A survey on automatic generation of medical imaging reports based on deep learning, BioMedical Engineering On- Line 22 (1) (2023) 48
2023
-
[30]
C. Wohlin, Guidelines for snowballing in systematic literature studies and a replication in software engineering, in: Proceedings of the 18th International Conference on Evaluation and Assessment in Software En- gineering, EASE ’14, Association for Computing Machinery, New Yor...
2014
-
[31]
Zohourianshahzadi, J
Z. Zohourianshahzadi, J. K. Kalita, Neural attention for image caption- ing: review of outstanding methods, Artificial Intelligence Review 55 (5) (2022) 3833–3862
2022
-
[32]
Sharma, D
H. Sharma, D. Padha, A comprehensive survey on image caption- ing: from handcrafted to deep learning-based techniques, a taxonomy and open research issues, Artificial Intelligence Review 56 (11) (2023) 13619–13661
2023
-
[33]
Sharma, D
H. Sharma, D. Padha, Domain-specific image captioning: a comprehen- sive review, International Journal of Multimedia Information Retrieval 13 (2) (2024) 20
2024
-
[34]
Sharma, D
H. Sharma, D. Padha, A. Selwal, A survey on attention-based image captioning: Taxonomy, challenges, and future perspectives, in: Inter- national Conference on Machine Intelligence and Signal Processing, Springer, 2022, pp. 681–694
2022
-
[35]
Oluwasammi, M
A. Oluwasammi, M. U. Aftab, Z. Qin, S. T. Ngo, T. V . Doan, S. B. Nguyen, S. H. Nguyen, G. H. Nguyen, Features to text: a comprehensive survey of deep learning on semantic segmentation and image captioning, Complexity 2021 (2021)
2021
-
[36]
Stani ¯ut˙e, D
R. Stani ¯ut˙e, D. ˇSeˇsok, A systematic literature review on image caption- ing, Applied Sciences 9 (10) (2019) 2024
2019
-
[37]
Sharma, D
H. Sharma, D. Padha, From templates to transformers: a survey of mul- timodal image captioning decoders, in: 2023 International Conference on Computer, Electronics & Electrical Engineering & their Applications (IC2E3), IEEE, 2023, pp. 1–6
2023
-
[38]
X. Liu, Q. Xu, N. Wang, A survey on deep neural network-based image captioning, The Visual Computer 35 (3) (2019) 445–470
2019
-
[39]
S.-H. Choi, S. Y . Jo, S. H. Jung, Component based comparative analysis of each module in image captioning, ICT Express 7 (1) (2021) 121–125
2021
-
[40]
Thirunavukarasu, E
R. Thirunavukarasu, E. Kotei, A comprehensive review on transformer network for natural and medical image analysis, Computer Science Review 53 (2024) 100648. doi:https://doi.org/10.1016/j.cosrev.2024.100648. URL https://www.sciencedirect.com/science/article/pii/S1574013724000327
2024
-
[41]
Kotei, R
E. Kotei, R. Thirunavukarasu, Medical image analysis with vision trans- formers for downstream tasks and clinical report generation, in: Intelli- gent Systems and Sustainable Computational Models, Auerbach Publi- cations, pp. 288–307
-
[42]
Zhao, A systematic survey of remote sensing image captioning, IEEE Access 9 (2021) 154086–154111
B. Zhao, A systematic survey of remote sensing image captioning, IEEE Access 9 (2021) 154086–154111
2021
-
[43]
X. Lu, B. Wang, X. Zheng, X. Li, Exploring models and data for remote 28 sensing image caption generation, IEEE Transactions on Geoscience and Remote Sensing 56 (4) (2017) 2183–2195
2017
-
[44]
M. Z. Hossain, F. Sohel, M. F. Shiratuddin, H. Laga, M. Bennamoun, Text to image synthesis for improved image captioning, IEEE Access 9 (2021) 64918–64928
2021
-
[45]
Zhang, H
W. Zhang, H. Shi, S. Tang, J. Xiao, Q. Yu, Y . Zhuang, Consensus graph representation learning for better grounded image captioning, in: Proc 35 AAAI Conf on Artificial Intelligence, 2021, pp. 3394–3402
2021
-
[46]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, in: Advances in neu- ral information processing systems, 2017, pp. 5998–6008
2017
-
[47]
S. He, W. Liao, H. R. Tavakoli, M. Yang, B. Rosenhahn, N. Pugeault, Image captioning through image transformer, in: Proceedings of the Asian Conference on Computer Vision, 2020, pp. 153–169
2020
-
[48]
T. Wolf, J. Chaumond, L. Debut, V . Sanh, C. Delangue, A. Moi, P. Cis- tac, M. Funtowicz, J. Davison, S. Shleifer, et al., Transformers: State- of-the-art natural language processing, in: Proceedings of the 2020 Con- ference on Empirical Methods in Natural Language Processing:...
2020
-
[49]
T. Lu, J. Wang, F. Min, Full-memory transformer for image captioning, Symmetry 15 (1) (2023) 190
2023
-
[50]
D. Wang, H. Hu, D. Chen, Transformer with sparse self-attention mech- anism for image captioning, Electronics Letters 56 (15) (2020) 764–766
2020
-
[52]
H. Wang, Y . Zhang, X. Yu, An overview of image caption generation methods, Computational intelligence and neuroscience 2020 (2020)
2020
-
[53]
Pedersoli, T
M. Pedersoli, T. Lucas, C. Schmid, J. Verbeek, Areas of attention for image captioning, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 1242–1250
2017
-
[54]
Jindal, A deep learning approach for arabic caption generation us- ing roots-words, in: Thirty-First AAAI Conference on Artificial Intelli- gence, 2017, pp
V . Jindal, A deep learning approach for arabic caption generation us- ing roots-words, in: Thirty-First AAAI Conference on Artificial Intelli- gence, 2017, pp. 4941–4942
2017
-
[55]
V . Jindal, Generating image captions in arabic using root-word based recurrent neural networks and deep neural networks, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 32, 2018, p. 144–151
2018
-
[57]
Mualla, J
R. Mualla, J. Alkheir, Development of an arabic image description sys- tem, International Journal of Computer Science Trends and Technology (IJCST)–6 (3) (2018) 205–213
2018
-
[58]
ElJundi, M
O. ElJundi, M. Dhaybi, K. Mokadam, H. M. Hajj, D. C. Asmar, Re- sources and end-to-end neural network models for arabic image cap- tioning., in: VISIGRAPP (5: VISAPP), 2020, pp. 233–241
2020
-
[59]
Hejazi, K
H. Hejazi, K. Shaalan, Deep learning for arabic image captioning: A comparative study of main factors and preprocessing recommendations, International Journal of Advanced Computer Science and Applications 12 (11) (2021)
2021
-
[60]
S. M. Sabri, Arabic image captioning using deep learning with attention, Ph.D. thesis, University of Georgia (2021)
2021
-
[61]
Emami, P
J. Emami, P. Nugues, A. Elnagar, I. Afyouni, Arabic image captioning using pre-training of deep bidirectional transformers, in: Proceedings of the 15th International Conference on Natural Language Generation, 2022, pp. 40–51
2022
-
[62]
M. T. Lasheen, N. H. Barakat, Arabic image captioning: the e ffect of text pre-processing on the attention weights and the bleu-n scores, Int J Adv Comput Sci Appl 13 (7) (2022) 11
2022
-
[63]
Alsayed, T
A. Alsayed, T. M. Qadah, M. Arif, A performance analysis of transformer-based deep learning models for arabic image captioning, Journal of King Saud University-Computer and Information Sciences 35 (9) (2023) 101750
2023
-
[64]
Elbedwehy, T
S. Elbedwehy, T. Medhat, Improved arabic image captioning model using feature concatenation with pre-trained word embedding, Neural Computing and Applications 35 (2023) 1–17. doi:10.1007/s00521-023- 08744-1
2023 doi
-
[65]
Karpathy, L
A. Karpathy, L. Fei-Fei, Deep visual-semantic alignments for generating image descriptions, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3128–3137
2015
-
[66]
J. Bineeshia, Image caption generation using cnn-lstm based approach, in: Proceedings of the First International Conference on Combinatorial and Optimization, ICCAP 2021, December 7-8 2021, Chennai, India, 2021, pp. 1–9
2021
-
[67]
Y . Ma, J. Ji, X. Sun, Y . Zhou, R. Ji, Towards local visual modeling for image captioning, Pattern Recognition 138 (2023) 109420
2023
-
[68]
Jiang, Z
T. Jiang, Z. Zhang, Y . Yang, Modeling coverage with semantic embed- ding for image caption generation, The Visual Computer 35 (11) (2019) 1655–1665
2019
-
[69]
do Carmo Nogueira, C
T. do Carmo Nogueira, C. D. N. Vinhal, G. da Cruz J´unior, M. R. D. Ull- mann, Reference-based model using multimodal gated recurrent units for image captioning, Multimedia Tools and Applications 79 (41) (2020) 30615–30635
2020
-
[70]
J. Lu, C. Xiong, D. Parikh, R. Socher, Knowing when to look: Adaptive attention via a visual sentinel for image captioning, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 375–383
2017
-
[71]
Kalimuthu, A
M. Kalimuthu, A. Mogadala, M. Mosbach, D. Klakow, Fusion models for improved image captioning, in: Pattern Recognition. ICPR Interna- tional Workshops and Challenges: Virtual Event, January 10–15, 2021, Proceedings, Part VI, Springer, 2021, pp. 381–395
2021
-
[72]
Abdussalam, Z
A. Abdussalam, Z. Ye, A. Hawbani, M. Al-Qatf, R. Khan, Numcap: a number-controlled multi-caption image captioning network, ACM Transactions on Multimedia Computing, Communications and Appli- cations 19 (4) (2023) 1–24
2023
-
[73]
Shrimal, T
A. Shrimal, T. Chakraborty, Attention beam: An image captioning ap- proach (student abstract), in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 35, 2021, pp. 15887–15888
2021
-
[74]
W. Zhao, X. Wu, J. Luo, Cross-domain image captioning via cross- modal retrieval and model adaptation, IEEE Transactions on Image Pro- cessing 30 (2020) 1180–1192
2020
-
[76]
H. Zhu, R. Wang, X. Zhang, Image captioning with dense fusion con- nection and improved stacked attention module, Neural Process. Lett. 53 (2) (2021) 1101–1118
2021
-
[77]
Fei, Attention-aligned transformer for image captioning, in: Proceed- ings of the AAAI Conference on Artificial Intelligence, V ol
Z. Fei, Attention-aligned transformer for image captioning, in: Proceed- ings of the AAAI Conference on Artificial Intelligence, V ol. 36, 2022, pp. 607–615
2022
-
[78]
Y . Wang, J. Xu, Y . Sun, End-to-end transformer based model for im- age captioning, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 36, 2022, pp. 2585–2594
2022
-
[79]
X. Yang, Y . Liu, X. Wang, Reformer: The relational transformer for image captioning, in: Proceedings of the 30th ACM International Con- ference on Multimedia, 2022, pp. 5398–5406
2022
-
[80]
Mulyawan, A
R. Mulyawan, A. Sunyoto, A. H. Muhammad, Automatic indonesian image captioning using cnn and transformer-based model approach, in: 2022 5th International Conference on Information and Communications Technology (ICOIACT), 2022, pp. 355–360
2022
-
[81]
Mulyanto, E
E. Mulyanto, E. I. Setiawan, E. M. Yuniarno, M. H. Purnomo, Auto- matic indonesian image caption generation using cnn-lstm model and feeh-id dataset, in: 2019 IEEE International Conference on Computa- tional Intelligence and Virtual Environments for Measurement Systems and App...
2019
-
[82]
J. J. Wijadi, B. Ghi ffar, N. N. Qomariyah, Indonesian language im- age captioning using encoder-decoder with attention approach, in: 2024 2nd International Conference on Software Engineering and Information Technology (ICoSEIT), 2024, pp. 342–347
2024
-
[83]
A. A. Nugraha, A. Arifianto, Suyanto, Generating image description on indonesian language using convolutional neural network and gated recurrent unit, 2019 7th International Conference on Information and Communication Technology (ICoICT) (2019) 1–6
2019
-
[84]
W. P. Pa, T. L. Nwe, et al., Automatic myanmar image captioning us- ing cnn and lstm-based language model, in: Proceedings of the 1st Joint Workshop on Spoken Language Technologies for Under-resourced lan- guages (SLTU) and Collaboration and Computing for Under-Resourced Langu...
2020
-
[85]
S. P. P. Aung, W. P. Pa, T. L. Nwe, Improving myanmar image cap- 29 tion generation using nasnetlarge and bi-directional lstm, in: 2023 IEEE Conference on Computer Applications (ICCA), 2023, pp. 1–6
2023
-
[86]
Muhammad Shah, M
F. Muhammad Shah, M. Humaira, M. A. R. K. Jim, A. Saha Ami, S. Paul, Bornon: Bengali image captioning with transformer-based deep learning approach, SN Computer Science 3 (2022) 1–16
2022
-
[87]
M. F. Khan, S. Sadiq-Ur-Rahman, M. S. Islam, Improved bengali im- age captioning via deep convolutional neural network based encoder- decoder model, in: Proceedings of International Joint Conference on Advances in Computational Intelligence, Springer, 2021, pp. 217–229
2021
-
[89]
Kulkarni, V
G. Kulkarni, V . Premraj, V . Ordonez, S. Dhar, S. Li, Y . Choi, A. C. Berg, T. L. Berg, Babytalk: Understanding and generating simple image descriptions, IEEE transactions on pattern analysis and machine intelli- gence 35 (12) (2013) 2891–2903
2013
-
[90]
Kpalma, J
K. Kpalma, J. Ronsin, An overview of advances of pattern recognition systems in computer vision, Vision Systems (2007) 26
2007
-
[91]
Sezgin, B
M. Sezgin, B. l. Sankur, Survey over image thresholding techniques and quantitative performance evaluation, Journal of Electronic imaging 13 (1) (2004) 146–168
2004
-
[92]
Almanaseer, M
W. Almanaseer, M. Alshraideh, O. Al-Kadi, A deep belief network clas- sification approach for automatic diacritization of arabic text, Applied Sciences 11 (11) (2021) 5228
2021
-
[93]
J. Chen, H. Zhuge, A news image captioning approach based on multi- modal pointer-generator network, Concurrency and Computation: Prac- tice and Experience (2019) e5721
2019
-
[94]
Zakraoui, S
J. Zakraoui, S. Elloumi, J. M. Alja’am, S. B. Yahia, Improving arabic text to image mapping using a robust machine learning technique, IEEE Access 7 (2019) 18772–18782
2019
-
[95]
Saleh, J
M. Saleh, J. M. Alja’am, Towards adaptive multimedia system for as- sisting children with arabic learning di fficulties, in: 2019 IEEE Jordan International Joint Conference on Electrical Engineering and Informa- tion Technology (JEEIT), IEEE, 2019, pp. 794–799
2019
-
[96]
Alzubaidi, J
L. Alzubaidi, J. Zhang, A. J. Humaidi, A. Al-Dujaili, Y . Duan, O. Al- Shamma, J. Santamar ´ıa, M. A. Fadhel, M. Al-Amidie, L. Farhan, Re- view of deep learning: Concepts, cnn architectures, challenges, applica- tions, future directions, Journal of big Data 8 (1) (2021) 1–74
2021
-
[97]
Zhang, W
W. Zhang, W. Nie, X. Li, Y . Yu, Image caption generation with adap- tive transformer, in: 2019 34rd Youth Academic Annual Conference of Chinese Association of Automation (Y AC), IEEE, 2019, pp. 521–526
2019
-
[98]
L. Guo, J. Liu, X. Zhu, P. Yao, S. Lu, H. Lu, Normalized and geometry- aware self-attention network for image captioning, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 10327–10336
2020
-
[99]
Kumar, V
D. Kumar, V . Srivastava, D. E. Popescu, J. D. Hemanth, Dual-modal transformer with enhanced inter-and intra-modality interactions for im- age captioning, Applied Sciences 12 (13) (2022) 6733
2022
-
[100]
Nguyen, M
V .-Q. Nguyen, M. Suganuma, T. Okatani, Grit: Faster and better image captioning transformer using dual visual features, in: Computer Vision– ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23– 27, 2022, Proceedings, Part XXXVI, Springer, 2022, pp. 167–184
2022
-
[101]
C. Wang, Y . Shen, L. Ji, Geometry attention transformer with position- aware lstms for image captioning, Expert Systems with Applications 201 (2022) 117174
2022
-
[102]
C. Liu, R. Zhao, Z. Shi, Remote-sensing image captioning based on multilayer aggregated transformer, IEEE Geoscience and Remote Sens- ing Letters 19 (2022) 1–5
2022
-
[103]
G. Li, L. Zhu, P. Liu, Y . Yang, Entangled transformer for image cap- tioning, in: Proceedings of the IEEE /CVF International Conference on Computer Vision, 2019, pp. 8928–8937
2019
-
[104]
Y . Wei, C. Wu, G. Li, H. Shi, Sequential transformer via an outside-in attention for image captioning, Engineering Applications of Artificial Intelligence 108 (2022) 104574
2022
-
[105]
Dubey, F
S. Dubey, F. Olimov, M. A. Rafique, J. Kim, M. Jeon, Label-attention transformer with geometrically coherent objects for image captioning, Information Sciences 623 (2023) 812–831
2023
-
[106]
Jiang, X
W. Jiang, X. Li, H. Hu, Q. Lu, B. Liu, Multi-gate attention network for image captioning, IEEE Access 9 (2021) 69700–69709
2021
-
[107]
Kandala, S
H. Kandala, S. Saha, B. Banerjee, X. X. Zhu, Exploring transformer and multilabel classification for remote sensing image captioning, IEEE Geoscience and Remote Sensing Letters 19 (2022) 1–5
2022
-
[108]
J. H. Tan, Y . H. Tan, C. S. Chan, J. H. Chuah, Acort: A compact ob- ject relation transformer for parameter e fficient image captioning, Neu- rocomputing 482 (2022) 60–72
2022
-
[109]
X. Wang, X. Fang, Y . Yang, Dm-catn: Deep modular co-attention trans- former networks for image captioning, in: International Conference on Artificial Intelligence and Intelligent Information Processing (AIIIP 2022), V ol. 12456, SPIE, 2022, pp. 600–606
2022
-
[110]
Cornia, M
M. Cornia, M. Stefanini, L. Baraldi, R. Cucchiara, Meshed-memory transformer for image captioning, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2020, pp. 10578–10587
2020
-
[111]
Cornia, L
M. Cornia, L. Baraldi, R. Cucchiara, Explaining transformer-based im- age captioning models: An empirical analysis, AI Communications 35 (2) (2022) 111–129
2022
-
[112]
Y . Pan, T. Yao, Y . Li, T. Mei, X-linear attention networks for image captioning, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2020, pp. 10971–10980
2020
-
[113]
Huang, W
L. Huang, W. Wang, J. Chen, X.-Y . Wei, Attention on attention for image captioning, in: Proceedings of the IEEE /CVF international conference on computer vision, 2019, pp. 4634–4643
2019
-
[114]
L. Ke, W. Pei, R. Li, X. Shen, Y .-W. Tai, Reflective decoding network for image captioning, in: Proceedings of the IEEE /CVF international conference on computer vision, 2019, pp. 8888–8897
2019
-
[115]
C. Yan, Y . Hao, L. Li, J. Yin, A. Liu, Z. Mao, Z. Chen, X. Gao, Task- adaptive attention for image captioning, IEEE Transactions on Circuits and Systems for Video Technology 32 (1) (2022) 43–51
2022
-
[116]
Y . Wang, X. Sun, X. Li, W. Zhang, X. Gao, Reasoning like humans: on dynamic attention prior in image captioning, Knowledge-Based Systems 228 (2021) 107313
2021
-
[117]
J. Liu, K. Cheng, H. Jin, Z. Wu, An image captioning algorithm based on combination attention mechanism, Electronics 11 (9) (2022) 1397
2022
-
[119]
Anderson, X
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, L. Zhang, Bottom-up and top-down attention for image captioning and visual question answering, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6077–6086
2018
-
[120]
T. Yao, Y . Pan, Y . Li, T. Mei, Exploring visual relationship for image captioning, in: Proceedings of the European conference on computer vision (ECCV), 2018, pp. 684–699
2018
-
[121]
Hodosh, P
M. Hodosh, P. Young, J. Hockenmaier, Framing image description as a ranking task: Data, models and evaluation metrics, Journal of Artificial Intelligence Research 47 (2013) 853–899
2013
-
[122]
Young, A
P. Young, A. Lai, M. Hodosh, J. Hockenmaier, From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions, Transactions of the Association for Computational Linguistics 2 (2014) 67–78
2014
-
[123]
T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll´ar, C. L. Zitnick, Microsoft coco: Common objects in context, in: European conference on computer vision, Springer, 2014, pp. 740– 755
2014
-
[124]
Papineni, S
K. Papineni, S. Roukos, T. Ward, W.-J. Zhu, Bleu: a method for au- tomatic evaluation of machine translation, in: Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 2002, pp. 311–318
2002
-
[125]
Lin, Rouge: A package for automatic evaluation of summaries, in: Text summarization branches out, 2004, pp
C.-Y . Lin, Rouge: A package for automatic evaluation of summaries, in: Text summarization branches out, 2004, pp. 74–81
2004
-
[126]
Banerjee, A
S. Banerjee, A. Lavie, Meteor: An automatic metric for mt evaluation with improved correlation with human judgments, in: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, 2005, pp. 65–72
2005
-
[127]
Vedantam, C
R. Vedantam, C. Lawrence Zitnick, D. Parikh, Cider: Consensus-based image description evaluation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 4566–4575
2015
-
[128]
Anderson, B
P. Anderson, B. Fernando, M. Johnson, S. Gould, Spice: Semantic propositional image caption evaluation, in: European conference on computer vision, Springer, 2016, pp. 382–398. 30
2016
-
[129]
Abu-Srhan, M
A. Abu-Srhan, M. A. Abushariah, O. S. Al-Kadi, The e ffect of loss function on conditional generative adversarial networks, Journal of King Saud University-Computer and Information Sciences 34 (9) (2022) 6977–6988
2022
-
[130]
Abu-Srhan, I
A. Abu-Srhan, I. Almallahi, M. A. Abushariah, W. Mahafza, O. S. Al- Kadi, Paired-unpaired unsupervised attention guided gan with transfer learning for bidirectional brain mr-ct synthesis, Computers in Biology and Medicine 136 (2021) 104763
2021
-
[131]
L. Gong, J. M. Crego, J. Senellart, Enhanced transformer model for data- to-text generation, in: Proceedings of the 3rd Workshop on Neural Gen- eration and Translation, 2019, pp. 148–156. 31
2019
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.