REVIEW 4 major objections 4 minor 93 references
An Ensemble Model with Attention Based Mechanism for Image Captioning
T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper claims that an ensemble of eight pre-trained CNN feature extractors feeding an attention-based transformer decoder, with voting that selects the highest-BLEU caption, outperforms published image-captioning results on Flickr8k…
desk verdict The headline numbers are an artifact of oracle selection: the ensemble picks the caption with the highest BLEU against test references and then reports that BLEU. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the ensemble-voting pipeline: eight pre-trained CNN feature extractors each produce a visual feature map, a transformer encoder-decoder with multi-head scaled dot-product attention turns each feature map into a caption using beam search of width 10, and a voting module compares the candidate captions against reference captions and keeps the one with the highest BLEU-1 score. The attention mechanism is what aligns image regions with each generated word, and the voting step is what the paper credits for the final gain over individual models.
What would settle it
Run the same eight CNN/transformer configurations on Flickr30k with the voting step replaced by a fixed aggregator or by a scorer that does not see the reference captions. If BLEU-1 and BLEU-4 drop to the level of the best single model, the reported gain depends on knowing the ground-truth caption during selection.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that combining many CNN visual encoders before a transformer decoder, then selecting the candidate caption with the highest BLEU-1 score, makes generated captions richer and more accurate than any single model in the ensemble. On Flickr8K the proposed model reports BLEU-1/2/3 of 0.728, 0.495, and 0.323; on Flickr30K it reports BLEU-1 through BLEU-4 of 0.798, 0.561, 0.387, and 0.269, plus SPICE of 0.387, and states these outperform the latest methods. The ablation study shows the full ensemble beating every reduced ensemble on all reported metrics, with paired t-tests at p < 0.05 separating the configurations.
Load-bearing premise
The voting step assumes that at selection time the system can compare each candidate caption with human-written reference captions and pick the one with the highest BLEU-1 score; without such references, the selection procedure cannot run as described.
Editorial extensions
If this is right
- A captioning system can be built without recurrent language models: pretrained CNNs plus an attention-based transformer produce the reported gains.
- Adding more diverse CNN encoders monotonically raises BLEU and SPICE scores on Flickr8k, according to the ablation study.
- Voting by BLEU-1 outperforms every single base model in the ensemble, so ensemble aggregation is the reported source of the improvement.
- The reported Flickr30k BLEU-1 through BLEU-4 scores (0.798, 0.561, 0.387, 0.269) exceed all cited comparison methods on those four n-gram levels.
Reading between the lines
- If the voting step indeed compares candidates with ground-truth captions, the reported numbers are an upper bound; a deployed system would need a reference-free selector, leaving the practical gain over a single model unmeasured.
- The ablation's monotone improvements suggest diversity among encoders drives the result; a fair comparison would hold computation constant and test a single transformer trained longer.
- A testable extension is to replace BLEU-1 selection with a semantic scorer such as SPICE, which may reward a different winner and change the ensemble's reported trade-offs.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an image captioning system that combines eight CNN feature extractors with a transformer encoder-decoder and an ensemble voting step. The final caption is selected as the one with the highest BLEU-1 score among the eight models' outputs, and the model is evaluated on Flickr8k and Flickr30k, where it reports state-of-the-art BLEU, METEOR, CIDEr, and SPICE scores. The central technical content is a survey-style exposition of transformers, attention mechanisms, and ensemble learning, followed by experiments with several CNN backbones.
Significance. If valid, the proposed ensemble and the reported gains over prior work would be a useful contribution to image captioning. The paper does provide a reasonably broad review of related work, an ablation study, and qualitative examples, and it reports a range of metrics. However, the evaluation protocol is fundamentally circular: the final caption is chosen by comparing candidate captions against the same reference captions that are later used to compute the reported BLEU scores. This makes the headline numbers oracle-selected maxima rather than the outputs of a deployable system, and the comparison with prior methods is therefore not meaningful. The absence of a reference-free selection rule means the central claim of state-of-the-art performance is not established.
major comments (4)
- [§3.3.5, Algorithm 1 Steps 23–26] The ensemble voting step selects the caption with the highest BLEU-1 score computed against the reference captions, and the same references are then used to compute the BLEU scores reported in Tables 2 and 3. At inference on a new image, reference captions are not available, so the method as specified cannot select a caption. The reported numbers are therefore a maximum over the candidate set with respect to the evaluation references, not the output of a deployable ensemble, and the comparison with prior methods is invalid because no competing method is allowed access to test references during generation. The paper offers no reference-free proxy for the voting step, so this is a load-bearing flaw in the central claim.
- [§3.3.5 vs. §4.6] The ensemble architecture is specified inconsistently: Algorithm 1 (line 12) lists eight CNN feature extractors (ResNet50, ResNet101, EfficientNetV2, VGG16, VGG19, EfficientNetB4, ResNet152, RegNetX120), while §3.3.5 states that the voting model combines the results of 'each of the eight transformer models,' and §4.6 says the model 'combines eight CNN models via a voting process.' Without a unique statement of what is being ensembled, the method is not reproducible and the attribution of the reported gains is unclear.
- [Tables 2 and 3, §4.3] The quantitative presentation contains a direct contradiction: the text in §4.3 says 'Our model obtained the highest result of the METEOR score, 0.604,' but Table 2 reports METEOR=0.235 and CIDEr=0.604, and the same paragraph later lists 'ROUGE L (0.432) and CIDEr (0.604)' as competitive results. This misreporting of the metric values undermines confidence in the numerical claims, although it is secondary to the circularity issue above.
- [§4.5, Table 4] The ablation study does not cleanly isolate the contribution of the proposed ensemble. The ablation baselines include MobileNetV2, which is absent from the full ensemble model in Algorithm 1 (line 12), and no explanation is given for how the full model relates to the baselines. Consequently, the apparent monotonic improvement from Baseline 1 to the full ensemble cannot be attributed to the specific CNN set or to the voting mechanism.
minor comments (4)
- [Throughout] The dataset name is spelled inconsistently as both 'Flickr8K' and 'Flicker8k'; the latter spelling appears in several places, including Tables 2 and 4.
- [Table 3] The row for the proposed model lists '0.2690.443' without a separator, making it difficult to distinguish the BLEU-4 value from the ROUGE-L value.
- [§3.3.5] The phrase 'voting-on' in Algorithm 1 (line 25) is never defined; the surrounding text describes BLEU-based selection, but the algorithm and the prose should use a single, precise term.
- [§4.6] The claim that paired t-tests between the full ensemble and each ablation variant were statistically significant (p < 0.05) is not accompanied by the test statistic, the number of samples, or the exact p-values, so the claim cannot be verified.
Circularity Check
The ensemble's voting step selects the caption with the highest BLEU-1 score against test reference captions, then the paper reports BLEU-[1-4] on those same references; the headline SOTA scores are oracle-selected maxima, not deployable predictions.
-
self definitional
[Section 3.3.5 'Ensemble learning'; Algorithm 1 Steps 23–26; results in Tables 2–3]
"The BLEU score-1 was considered for this purpose, and the prediction result will be accepted from the model that gains the highest BLEU score."
The final output caption is defined as the candidate that maximizes BLEU-1, a metric that by definition (Section 4.2.1) compares generated captions to human reference captions. Those same reference captions are then used in Section 4.3 to compute the reported BLEU-[1-4] values in Tables 2 and 3. Selecting the maximum of a set and then reporting that maximum is not an empirical evaluation of an ensemble; it is an oracle-selection artifact. At inference on a new image no reference captions exist, so the method as written cannot produce the reported outputs. The paper never states that references are withheld during voting or replaced by a reference-free proxy, so the claimed state-of-the-art scores reduce by construction to the best available candidate rather than to a predictive ensemble.
full rationale
The central benchmark claim is self-definitional rather than independently derived. The paper's own voting rule in Section 3.3.5 accepts the caption 'from the model that gains the highest BLEU score,' and BLEU is computed against ground-truth reference captions. The same references are then used to report the headline BLEU-[1-4] results in Tables 2 and 3, so the reported scores are maxima over the candidate set with respect to the evaluation references. This invalidates the comparison with prior methods, none of which receive reference captions at generation time. No reference-free voting rule, held-out model selection, or deployment-time alternative is described. Apart from this central issue, the transformer/attention architecture is standard and the author self-citations are not load-bearing; the circularity is concentrated in the evaluation loop, where the output is selected by the metric used to score it.
Assumptions & free parameters
free parameters (4)
- learning_rate =
0.00001
- beam_width_k =
10
- ensemble_size =
8
- batch_size =
64
assumptions (5)
- domain assumption Pre-trained CNN models (ResNet, VGG, EfficientNet, RegNetX) provide useful image feature representations for captioning.
- domain assumption A transformer encoder-decoder can generate a caption from CNN feature maps.
- domain assumption BLEU, ROUGE, METEOR, CIDEr, and SPICE are valid proxies for caption quality.
- ad hoc to paper Reference captions are available to select the best candidate caption during voting.
- domain assumption Paired t-test assumptions hold for the ablation comparisons.
Cite this review
Pith. "Pith review of An Ensemble Model with Attention Based Mechanism for Image Captioning." pith.science (2026). https://pith.science/paper/3PXAOLCF
@misc{pith2026250114828,
author = {Pith},
title = {Pith review of: An Ensemble Model with Attention Based Mechanism for Image Captioning},
year = {2026},
howpublished = {\url{https://pith.science/paper/3PXAOLCF}},
note = {Machine review of arXiv:2501.14828}
}
read the original abstract
Image captioning creates informative text from an input image by creating a relationship between the words and the actual content of an image. Recently, deep learning models that utilize transformers have been the most successful in automatically generating image captions. The capabilities of transformer networks have led to notable progress in several activities related to vision. In this paper, we thoroughly examine transformer models, emphasizing the critical role that attention mechanisms play. The proposed model uses a transformer encoder-decoder architecture to create textual captions and a deep learning convolutional neural network to extract features from the images. To create the captions, we present a novel ensemble learning framework that improves the richness of the generated captions by utilizing several deep neural network architectures based on a voting mechanism that chooses the caption with the highest bilingual evaluation understudy (BLEU) score. The proposed model was evaluated using publicly available datasets. Using the Flickr8K dataset, the proposed model achieved the highest BLEU-[1-3] scores with rates of 0.728, 0.495, and 0.323, respectively. The suggested model outperformed the latest methods in Flickr30k datasets, determined by BLEU-[1-4] scores with rates of 0.798, 0.561, 0.387, and 0.269, respectively. The model efficacy was also obtained by the Semantic propositional image caption evaluation (SPICE) metric with a scoring rate of 0.164 for the Flicker8k dataset and 0.387 for the Flicker30k. Finally, ensemble learning significantly advances the process of image captioning and, hence, can be leveraged in various applications across different domains.
Reference graph
Works this paper leans on
-
[1]
ACM Computing Surveys (CsUR)51(6), 1–36 (2019)
Hossain, M.Z., Sohel, F., Shiratuddin, M.F., Laga, H.: A comprehensive survey of deep learning for image captioning. ACM Computing Surveys (CsUR)51(6), 1–36 (2019)
2019
-
[2]
In: International Conference on Learning and Intelligent Optimization, pp
Cheikh, M., Zrigui, M.: Active learning based framework for image caption- ing corpus creation. In: International Conference on Learning and Intelligent Optimization, pp. 128–142 (2020). Springer
2020
-
[3]
IEEE transactions on circuits and systems for video technology30(12), 4467–4480 (2019)
Yu, J., Li, J., Yu, Z., Huang, Q.: Multimodal transformer with multi-view visual representation for image captioning. IEEE transactions on circuits and systems for video technology30(12), 4467–4480 (2019)
2019
-
[4]
ACM Comput
Ghandi, T., Pourreza, H., Mahyar, H.: Deep learning approaches on image cap- tioning: A review. ACM Comput. Surv.56(3) (2023) https://doi.org/10.1145/ 3617592
2023
-
[5]
Pattern Recognition, 107856 (2021)
Ayesha, H., Iqbal, S., Tariq, M., Abrar, M., Sanaullah, M., Abbas, I., Rehman, A., Niazi, M.F.K., Hussain, S.: Automatic medical image interpretation: State of the art and future directions. Pattern Recognition, 107856 (2021)
2021
-
[6]
Journal of digital imaging31(5), 622–627 (2018)
Ogura, A., Hayashi, N., Negishi, T., Watanabe, H.: Effectiveness of an e-learning platform for image interpretation education of medical staff and students. Journal of digital imaging31(5), 622–627 (2018)
2018
-
[7]
Academic Press, ??? (2017)
Depeursinge, A., Al-Kadi, O.S., Mitchell, J.R.: Biomedical Texture Analysis: Fundamentals, Tools and Challenges. Academic Press, ??? (2017)
2017
-
[8]
PhD thesis, University of Sussex (2010)
Al-Kadi, O.S.: Tumour grading and discrimination based on class assignment and quantitative texture analysis techniques. PhD thesis, University of Sussex (2010)
2010
Show all 93 references
-
[9]
International Journal of Recent Advances in Multidisciplinary Topics 2(6), 71–75 (2021) 25
Chendake, P., Korpal, P., Bhor, S., Bansal, R., Patil, S., Deshpande, D.: Learning system for kids. International Journal of Recent Advances in Multidisciplinary Topics 2(6), 71–75 (2021) 25
2021
-
[10]
IEEE Transactions on Learning Technologies (2024)
Ayyoub, H.Y., Al-Kadi, O.S.: Learning style identification using semi-supervised self-taught labeling. IEEE Transactions on Learning Technologies (2024)
2024
-
[11]
In: Durmus, E., Gupta, V., Liu, N., Peng, N., Su, Y
Ahsan, H., Bhatt, D., Shah, K., Bhalla, N.: Multi-modal image captioning for the visually impaired. In: Durmus, E., Gupta, V., Liu, N., Peng, N., Su, Y. (eds.) Pro- ceedingsofthe2021ConferenceoftheNorthAmericanChapteroftheAssociation for Computational Linguistics: Student Rese...
2021 doi
-
[12]
International Journal of Image and Graphics, 2150044 (2021)
Nivedita, M., Chandrashekar, P., Mahapatra, S., Phamila, Y.A.V., Selvaperu- mal, S.K.: Image captioning for video surveillance system using neural networks. International Journal of Image and Graphics, 2150044 (2021)
2021
-
[13]
In: Australasian Database Conference, pp
Wang, Z., Huang, Z., Luo, Y.: Paic: Parallelised attentive image captioning. In: Australasian Database Conference, pp. 16–28 (2020). Springer
2020
-
[14]
Comput- ers & Electrical Engineering 78, 108–119 (2019) https://doi.org/10.1016/j
Saleem, S., Dilawari, A., Khan, U.G., Iqbal, R., Wan, S., Umer, T.: Stateful human-centered visual captioning system to aid video surveillance. Comput- ers & Electrical Engineering 78, 108–119 (2019) https://doi.org/10.1016/j. compeleceng.2019.07.009
2019 doi
-
[15]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp
Shuster, K., Humeau, S., Hu, H., Bordes, A., Weston, J.: Engaging image caption- ing via personality. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12516–12526 (2019)
2019
-
[16]
IATSS research43(4), 244–252 (2019)
Fujiyoshi, H., Hirakawa, T., Yamashita, T.: Deep learning-based image recogni- tion for autonomous driving. IATSS research43(4), 244–252 (2019)
2019
-
[17]
In: Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, pp
Guinness, D., Cutrell, E., Morris, M.R.: Caption crawler: Enabling reusable alter- native text descriptions using reverse image search. In: Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, pp. 1–11 (2018)
2018
-
[18]
Computers & Electrical Engineering119, 109626 (2024) https://doi.org/10.1016/j.compeleceng.2024.109626
Xiao, F., Zhang, N., Xue, W., Gao, X.: Sentinel mechanism for visual semantic graph-based image captioning. Computers & Electrical Engineering119, 109626 (2024) https://doi.org/10.1016/j.compeleceng.2024.109626
2024
-
[19]
Computers and Electrical Engineering 104, 108429 (2022) https://doi.org/10
Zhang, Z., Zhang, H., Wang, J., Sun, Z., Yang, Z.: Generating news image cap- tions with semantic discourse extraction and contrastive style-coherent learning. Computers and Electrical Engineering 104, 108429 (2022) https://doi.org/10. 1016/j.compeleceng.2022.108429
2022
-
[20]
Neurocomputing 311, 291–304 (2018)
Bai, S., An, S.: A survey on automatic image caption generation. Neurocomputing 311, 291–304 (2018)
2018
-
[21]
In: 4th IET International Conference on Advances in 26 Medical, Signal and Information Processing-MEDSIP 2008, pp
Al-Kadi, O.S.: Combined statistical and model based texture features for im- proved image classification. In: 4th IET International Conference on Advances in 26 Medical, Signal and Information Processing-MEDSIP 2008, pp. 1–4 (2008). IET
2008
-
[22]
In: 2011 IEEE Jordan Conference on Applied Electrical Engineering and Computing Technologies (AEECT), pp
Al-Kadi, O.S.: Supervised texture segmentation: a comparative study. In: 2011 IEEE Jordan Conference on Applied Electrical Engineering and Computing Technologies (AEECT), pp. 1–5 (2011). IEEE
2011
-
[23]
Frontiers of Computer Science14, 241–258 (2020)
Dong, X., Yu, Z., Cao, W., Shi, Y., Ma, Q.: A survey on ensemble learning. Frontiers of Computer Science14, 241–258 (2020)
2020
-
[24]
Multimedia Tools and Applications83(2), 5309–5325 (2024)
Verma, A., Yadav, A.K., Kumar, M., Yadav, D.: Automatic image caption gener- ation using deep learning. Multimedia Tools and Applications83(2), 5309–5325 (2024)
2024
-
[25]
Journal of Xi’an Shiyou University, Natural Science Edition9, 1088–1095 (2023)
Dahri, F.H., Chandio, A.A., Dahri, N.A., Soomro, M.A.: Image caption genera- tor using convolutional recurrent neural network feature fusion. Journal of Xi’an Shiyou University, Natural Science Edition9, 1088–1095 (2023)
2023
-
[26]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp
Karpathy, A., Fei-Fei, L.: Deep visual-semantic alignments for generating image descriptions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3128–3137 (2015)
2015
-
[27]
Wireless Communications and Mobile Computing 2020, 1–7 (2020)
Chu, Y., Yue, X., Yu, L., Sergei, M., Wang, Z.: Automatic image captioning based on resnet50 and lstm with soft attention. Wireless Communications and Mobile Computing 2020, 1–7 (2020)
2020
-
[28]
In: Proceedings of the AAAI Conference on Artificial Intelligence, vol
Fei, Z.: Attention-aligned transformer for image captioning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, pp. 607–615 (2022)
2022
-
[29]
Information Sciences 623, 812–831 (2023)
Dubey, S., Olimov, F., Rafique, M.A., Kim, J., Jeon, M.: Label-attention trans- former with geometrically coherent objects for image captioning. Information Sciences 623, 812–831 (2023)
2023
-
[30]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp
Guo, L., Liu, J., Zhu, X., Yao, P., Lu, S., Lu, H.: Normalized and geometry-aware self-attention network for image captioning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10327–10336 (2020)
2020
-
[31]
In: Proceedings of the Asian Conference on Computer Vision (2020)
He, S., Liao, W., Tavakoli, H.R., Yang, M., Rosenhahn, B., Pugeault, N.: Image captioning through image transformer. In: Proceedings of the Asian Conference on Computer Vision (2020)
2020
-
[32]
arXiv preprint arXiv:2012.12975 (2020)
Velioglu, R., Rose, J.: Detecting hate speech in memes using multimodal deep learning approaches: Prize-winning solution to hateful memes challenge. arXiv preprint arXiv:2012.12975 (2020)
2020 arXiv
-
[33]
Information Sciences567, 23–41 (2021) 27
Meel, P., Vishwakarma, D.K.: Han, image captioning, and forensics ensemble multimodal fake news detection. Information Sciences567, 23–41 (2021) 27
2021
-
[34]
The Visual Computer, 1–18 (2022)
Zhong, J., Cao, Y., Zhu, Y., Gong, J., Chen, Q.: Multi-channel weighted fusion for image captioning. The Visual Computer, 1–18 (2022)
2022
-
[35]
Dalla Serra, F., Deligianni, F., Dalton, J., O’Neil, A.Q.: Cmre-uog team at imageclefmedical caption 2022: Concept detection and image captioning (2022)
2022
-
[36]
Neural Computing and Applications 34(21), 18391–18406 (2022)
Salur, M.U., Aydın, İ.: A soft voting ensemble learning-based approach for multimodal sentiment analysis. Neural Computing and Applications 34(21), 18391–18406 (2022)
2022
-
[37]
IEEE Journal of Biomedical and Health Informatics (2022)
Singh, D., Kaur, M., Alanazi, J.M., AlZubi, A.A., Lee, H.-N.: Efficient evolving deep ensemble medical image captioning network. IEEE Journal of Biomedical and Health Informatics (2022)
2022
-
[38]
Sensors 22(12), 4392 (2022)
Kim, B.C., Kim, H.C., Han, S., Park, D.K.: Inspection of underwater hull surface condition using the soft voting ensemble of the transfer-learned models. Sensors 22(12), 4392 (2022)
2022
-
[39]
Journal of King Saud University- Computer and Information Sciences34(9), 6977–6988 (2022)
Abu-Srhan, A., Abushariah, M.A., Al-Kadi, O.S.: The effect of loss function on conditional generative adversarial networks. Journal of King Saud University- Computer and Information Sciences34(9), 6977–6988 (2022)
2022
-
[40]
Computers in Biology and Medicine136, 104763 (2021)
Abu-Srhan, A., Almallahi, I., Abushariah, M.A., Mahafza, W., Al-Kadi, O.S.: Paired-unpairedunsupervisedattentionguidedganwithtransferlearningforbidi- rectional brain mr-ct synthesis. Computers in Biology and Medicine136, 104763 (2021)
2021
-
[41]
IEEE Internet of Things Journal8(12), 9463–9472 (2020)
Alkadi, O., Moustafa, N., Turnbull, B., Choo, K.-K.R.: A deep blockchain framework-enabled collaborative intrusion detection for protecting iot and cloud networks. IEEE Internet of Things Journal8(12), 9463–9472 (2020)
2020
-
[42]
Journal of Artificial Intelligence Research 47, 853–899 (2013)
Hodosh, M., Young, P., Hockenmaier, J.: Framing image description as a rank- ing task: Data, models and evaluation metrics. Journal of Artificial Intelligence Research 47, 853–899 (2013)
2013
-
[43]
Transactions of the Association for Computational Linguistics2, 67–78 (2014)
Young, P., Lai, A., Hodosh, M., Hockenmaier, J.: From image descriptions to visual denotations: New similarity metrics for semantic inference over event de- scriptions. Transactions of the Association for Computational Linguistics2, 67–78 (2014)
2014
-
[44]
PhD thesis, Vrije Universiteit Amsterdam (October 2019)
van Miltenburg, E.: Pragmatic factors in (automatic) image description. PhD thesis, Vrije Universiteit Amsterdam (October 2019)
2019
-
[45]
Complexity 2021 (2021) 28
Oluwasammi, A., Aftab, M.U., Qin, Z., Ngo, S.T., Doan, T.V., Nguyen, S.B., Nguyen, S.H., Nguyen, G.H.: Features to text: a comprehensive survey of deep learning on semantic segmentation and image captioning. Complexity 2021 (2021) 28
2021
-
[46]
Applied Sciences 11(11), 5228 (2021)
Almanaseer, W., Alshraideh, M., Alkadi, O.: A deep belief network classification approach for automatic diacritization of arabic text. Applied Sciences 11(11), 5228 (2021)
2021
-
[47]
KI-Künstliche Intelligenz 34(4), 571–584 (2020)
Biswas, R., Barz, M., Sonntag, D.: Towards explanatory interactive image cap- tioning using top-down and bottom-up features, beam search and re-ranking. KI-Künstliche Intelligenz 34(4), 571–584 (2020)
2020
-
[48]
Concurrency and Computation: Practice and Experience 34(7), 5721 (2022)
Chen, J., Zhuge, H.: A news image captioning approach based on multi- modal pointer-generator network. Concurrency and Computation: Practice and Experience 34(7), 5721 (2022)
2022
-
[49]
Journal of big Data8(1), 1–74 (2021)
Alzubaidi, L., Zhang, J., Humaidi, A.J., Al-Dujaili, A., Duan, Y., Al-Shamma, O., Santamaría, J., Fadhel, M.A., Al-Amidie, M., Farhan, L.: Review of deep learning: Concepts, cnn architectures, challenges, applications, future directions. Journal of big Data8(1), 1–74 (2021)
2021
-
[50]
In: Proceedings of International Joint Conference on Advances in Computational Intelligence, pp
Faiyaz Khan, M., Sadiq-Ur-Rahman, S., Islam, S.,et al.: Improved bengali image captioning via deep convolutional neural network based encoder-decoder model. In: Proceedings of International Joint Conference on Advances in Computational Intelligence, pp. 217–229 (2021). Springer
2021
-
[51]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L.: Bottom-up and top-down attention for image captioning and visual question answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6077–6086 (2018)
2018
-
[52]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp
Huang, G., Liu, Z., Van Der Maaten, L., Weinberger, K.Q.: Densely connected convolutional networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4700–4708 (2017)
2017
-
[53]
Applied Sciences 9(10), 2024 (2019)
Stani¯ ut˙ e, R., Šešok, D.: A systematic literature review on image captioning. Applied Sciences 9(10), 2024 (2019)
2019
-
[54]
In: International Conference on Machine Learning, pp
Tan, M., Le, Q.: Efficientnet: Rethinking model scaling for convolutional neu- ral networks. In: International Conference on Machine Learning, pp. 6105–6114 (2019). PMLR
2019
-
[55]
Applied soft computing96, 106691 (2020)
Marques, G., Agarwal, D., Torre Díez, I.: Automated medical diagnosis of covid- 19 through efficientnet convolutional neural network. Applied soft computing96, 106691 (2020)
2020
-
[56]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., Chen, L.-C.: Mobilenetv2: Inverted residuals and linear bottlenecks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4510–4520 (2018) 29
2018
-
[57]
Procedia Computer Science197, 198–207 (2022) https://doi.org/10.1016/j.procs.2021.12.132
Indraswari, R., Rokhana, R., Herulambang, W.: Melanoma image classifica- tion based on mobilenetv2 network. Procedia Computer Science197, 198–207 (2022) https://doi.org/10.1016/j.procs.2021.12.132 . Sixth Information Systems International Conference (ISICO 2021)
2022 doi
-
[58]
Multimedia Tools and Applications79(41), 30615–30635 (2020)
Carmo Nogueira, T., Vinhal, C.D.N., Cruz Júnior, G., Ullmann, M.R.D.: Reference-based model using multimodal gated recurrent units for image caption- ing. Multimedia Tools and Applications79(41), 30615–30635 (2020)
2020
-
[59]
In: Advances in Neural Information Processing Systems, pp
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, ., Polosukhin, I.: Attention is all you need. In: Advances in Neural Information Processing Systems, pp. 5998–6008 (2017)
2017
-
[60]
In: Proceedings of the 3rd Workshop on Neural Generation and Translation, pp
Gong, L., Crego, J.M., Senellart, J.: Enhanced transformer model for data-to- text generation. In: Proceedings of the 3rd Workshop on Neural Generation and Translation, pp. 148–156 (2019)
2019
-
[61]
In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp
Wolf, T., Chaumond, J., Debut, L., Sanh, V., Delangue, C., Moi, A., Cistac, P., Funtowicz, M., Davison, J., Shleifer, S.,et al.: Transformers: State-of-the-art natural language processing. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processi...
2020
-
[62]
Neurocomputing 311, 291–304 (2018) https://doi.org/10.1016/j.neucom.2018.05.080
Bai, S., An, S.: A survey on automatic image caption generation. Neurocomputing 311, 291–304 (2018) https://doi.org/10.1016/j.neucom.2018.05.080
2018 doi
-
[63]
Concurrency and Computation: Practice and Experience, 5721 (2019)
Chen, J., Zhuge, H.: A news image captioning approach based on multi- modal pointer-generator network. Concurrency and Computation: Practice and Experience, 5721 (2019)
2019
-
[64]
In: Proceedings of the IEEE International Conference on Computer Vision, pp
Pedersoli, M., Lucas, T., Schmid, C., Verbeek, J.: Areas of attention for image captioning. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 1242–1250 (2017)
2017
-
[65]
Computers & Electrical Engineer- ing 92, 107114 (2021) https://doi.org/10.1016/j.compeleceng.2021.107114
Mishra,S.K.,Dhir,R.,Saha,S.,Bhattacharyya,P.,Singh,A.K.:Imagecaptioning in hindi language using transformer networks. Computers & Electrical Engineer- ing 92, 107114 (2021) https://doi.org/10.1016/j.compeleceng.2021.107114
2021
-
[66]
Computational intelligence and neuroscience2020 (2020)
Wang, H., Zhang, Y., Yu, X.: An overview of image caption generation methods. Computational intelligence and neuroscience2020 (2020)
2020
-
[67]
IEEE transactions on pattern analysis and machine intelligence45(1), 539–559 (2022)
Stefanini, M., Cornia, M., Baraldi, L., Cascianelli, S., Fiameni, G., Cucchiara, R.: From show to tell: A survey on deep learning-based image captioning. IEEE transactions on pattern analysis and machine intelligence45(1), 539–559 (2022)
2022
-
[68]
ArXivabs/1611.08562 (2016) 30
Li, J., Monroe, W., Jurafsky, D.: A simple, fast diverse decoding algorithm for neural generation. ArXivabs/1611.08562 (2016) 30
2016 arXiv
-
[69]
IEEE Access10, 99129–99149 (2022)
Mienye, I.D., Sun, Y.: A survey of ensemble learning: Concepts, algorithms, applications, and prospects. IEEE Access10, 99129–99149 (2022)
2022
-
[70]
Machine learning24, 123–140 (1996)
Breiman, L.: Bagging predictors. Machine learning24, 123–140 (1996)
1996
-
[71]
Machine learning5, 197–227 (1990)
Schapire, R.E.: The strength of weak learnability. Machine learning5, 197–227 (1990)
1990
-
[72]
Neural Networks 5(2), 241–259 (1992) https://doi.org/10.1016/S0893-6080(05)80023-1
Wolpert, D.H.: Stacked generalization. Neural Networks 5(2), 241–259 (1992) https://doi.org/10.1016/S0893-6080(05)80023-1
1992 doi
-
[73]
In: Cocchi, M
Ballabio, D., Todeschini, R., Consonni, V.: Chapter 5 - recent advances in high- level fusion methods to classify multiple analytical chemical data. In: Cocchi, M. (ed.) Data Fusion Methodology and Applications. Data Handling in Sci- ence and Technology, vol. 31, pp. 129–155. ...
2019 doi
-
[74]
Sakarya University Journal of Science 25 (2021) https://doi.org/10.16984/saufenbilder.901960
Ahad, A.: Vote-based: Ensemble approach. Sakarya University Journal of Science 25 (2021) https://doi.org/10.16984/saufenbilder.901960
2021 doi
-
[75]
International Conference on Learning Representations (2014)
Kingma, D., Ba, J.: Adam: A method for stochastic optimization. International Conference on Learning Representations (2014)
2014
-
[76]
In: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pp
Papineni, K., Roukos, S., Ward, T., Zhu, W.-J.: Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pp. 311–318. Association for Computational Linguistics, Philadelphia, Pe...
2002
-
[77]
In: Text Summarization Branches Out, pp
Lin, C.-Y.: Rouge: A package for automatic evaluation of summaries. In: Text Summarization Branches Out, pp. 74–81 (2004)
2004
-
[78]
65–72 (2005)
Banerjee, S., Lavie, A.: Meteor: An automatic metric for mt evaluation with improvedcorrelationwithhumanjudgments.In:ProceedingsoftheAclWorkshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation And/or Summarization, pp. 65–72 (2005)
2005
-
[79]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp
Vedantam, R., Lawrence Zitnick, C., Parikh, D.: Cider: Consensus-based image description evaluation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4566–4575 (2015)
2015
-
[80]
In: European Conference on Computer Vision, pp
Anderson,P.,Fernando,B.,Johnson,M.,Gould,S.:Spice:Semanticpropositional image caption evaluation. In: European Conference on Computer Vision, pp. 382– 398 (2016). Springer
2016
-
[81]
10 (2004) 31
Lin, C.-Y.: Rouge: A package for automatic evaluation of summaries, p. 10 (2004) 31
2004
-
[82]
The Visual Computer35(11), 1655–1665 (2019)
Jiang, T., Zhang, Z., Yang, Y.: Modeling coverage with semantic embedding for image caption generation. The Visual Computer35(11), 1655–1665 (2019)
2019
-
[83]
arXiv preprint arXiv:2006.10923 (2020)
Patel, A., Varier, A.: Hyperparameter analysis for image captioning. arXiv preprint arXiv:2006.10923 (2020)
2020 arXiv
-
[84]
In: 2020 IEEE 14th International Conference on Semantic Computing (ICSC), pp
Katpally, H., Bansal, A.: Ensemble learning on deep neural networks for image caption generation. In: 2020 IEEE 14th International Conference on Semantic Computing (ICSC), pp. 61–68 (2020). IEEE
2020
-
[85]
In: Proceedings of the First International Conference on Combinatorial and Opti- mization, ICCAP 2021, December 7-8 2021, Chennai, India (2021)
Bineeshia, J.: Image caption generation using cnn-lstm based approach. In: Proceedings of the First International Conference on Combinatorial and Opti- mization, ICCAP 2021, December 7-8 2021, Chennai, India (2021)
2021
-
[86]
Pattern Recognition138, 109420 (2023)
Ma, Y., Ji, J., Sun, X., Zhou, Y., Ji, R.: Towards local visual modeling for image captioning. Pattern Recognition138, 109420 (2023)
2023
-
[87]
In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp
You, Q., Jin, H., Wang, Z., Fang, C., Luo, J.: Image captioning with semantic at- tention. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4651–4659 (2016). https://doi.org/10.1109/CVPR.2016.503
2016 doi
-
[88]
IEEE transactions on pattern analysis and machine intelligence39(12), 2321– 2334 (2016)
Fu, K., Jin, J., Cui, R., Sha, F., Zhang, C.: Aligning where to see and what to tell: Image captioning with region-based attention and scene-specific contexts. IEEE transactions on pattern analysis and machine intelligence39(12), 2321– 2334 (2016)
2016
-
[89]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp
Lu, J., Xiong, C., Parikh, D., Socher, R.: Knowing when to look: Adaptive at- tention via a visual sentinel for image captioning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 375–383 (2017)
2017
-
[90]
Neural Processing Letters 49, 177–185 (2019)
He, C., Hu, H.: Image captioning with text-based visual attention. Neural Processing Letters 49, 177–185 (2019)
2019
-
[91]
In: Pattern Recognition
Kalimuthu, M., Mogadala, A., Mosbach, M., Klakow, D.: Fusion models for im- proved image captioning. In: Pattern Recognition. ICPR International Workshops and Challenges: Virtual Event, January 10–15, 2021, Proceedings, Part VI, pp. 381–395 (2021). Springer
2021
-
[92]
ACM Transactions onMultimediaComputing,CommunicationsandApplications 19(4),1–24(2023)
Abdussalam, A., Ye, Z., Hawbani, A., Al-Qatf, M., Khan, R.: Numcap: a number-controlled multi-caption image captioning network. ACM Transactions onMultimediaComputing,CommunicationsandApplications 19(4),1–24(2023)
2023
-
[93]
Yang, M., Liu, J., Shen, Y., Zhao, Z., Chen, X., Wu, Q., Li, C.: An ensemble of generation-and retrieval-based image captioning with dual generator genera- tive adversarial network. IEEE Transactions on Image Processing29, 9627–9640 (2020) 32 Figure 7: Samples of correct capti...
2020
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.