REVIEW 4 major objections 5 minor 92 references
Hierarchical Deep Multi-modal Network for Medical Visual Question Answering
T0 review · 4 major / 5 minor · reviewed 2026-08-27 · deepseek-v4-flash
Pith's one-line read Medical VQA improves when questions are routed by answer type.
desk verdict Useful routing idea and a clean ablation, but the headline baseline comparison falls apart against the paper's own Table 6. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the two-level question-routing hierarchy. At the root, a linear SVM classifies each question using a binary presence vector for ten question-identifier words ('is', 'was', 'are', 'how', 'can', 'does', 'which', 'what', 'type', 'there') concatenated with a tf-idf vector over the top 500 vocabulary terms. The output class chooses the learning path: Yes/No questions go to a two-class softmax, while Others questions go to a multi-label decoder that emits answer words from a vocabulary. Images are encoded with Inception-Resnet-v2 into 1000 features, questions with a Bi-LSTM over concatenated word and subword embeddings, and the modalities are fused by concatenation followed by batch normalization. The routing works by reducing the output space for each question to the only plausible answer type.
What would settle it
Run the original released checkpoints or official challenge submissions of the four prior systems on the RAD and CLEF18 test sets with the paper's evaluation script. If they reproduce their originally reported scores, for example CLEF18 BLEU 0.161 and 0.134 for the two best-scoring prior systems, the paper's significant margins over those baselines would not hold.
Extended reading notes
Core claim
The central claim is that question segregation is itself a performance lever, not just a convenience. A linear SVM that marks a small set of question-identifier words and a tf-idf vector separates Yes/No from Others questions with F1 near 0.99 on RAD and 0.91 on CLEF18; routing on that decision into two specialist answer models — a two-class softmax over {Yes, No} and a word-by-word sequence predictor — improves BLEU and WBBS scores on RAD, CLEF18, and their combination, compared with the same multimodal network left unsegmented. The paper reports that the router raises Yes/No precision, recall, and F1 by as much as 0.5 points and keeps descriptive questions from being collapsed into Yes/No answers.
Load-bearing premise
The claimed margins over prior systems rest on the assumption that the authors' re-implementations of those systems are as strong as the original published systems; if the originals score higher on the same test sets, the reported margins shrink or vanish.
Editorial extensions
If this is right
- Yes/No questions can be answered with a two-word output vocabulary, removing the dominant error mode in which a monolithic model answers them with descriptive phrases.
- A question-segregation layer can be placed on top of an existing multimodal encoder and answer generator, and the paper shows that adding it raises scores on both RAD and CLEF18.
- The segregation benefit survives training on RAD and CLEF18 combined, so it is not tied to one dataset's particular mix of question types.
- On small medical datasets, simpler fusion without elaborate attention can beat more complex fusion mechanisms, because dedicated branches are easier to train and less prone to overfitting.
- The error categories identified in the paper — semantic mismatch, modality or plane confusion, specification gaps, boundary loss, and miscellaneous reasoning errors — give concrete targets for the next generation of medical VQA systems.
Reading between the lines
- If the routing gain comes from shrinking the output space, then a finer taxonomy than the binary Yes/No vs Others split, such as RAD's eleven native question categories, should yield further gains; the paper does not test that extension.
- The router is cheap and model-agnostic, so a clean test of its value would freeze the answer generators and toggle only the routing module across several general-domain VQA datasets.
- The 'Others' branch still mixes one-word answers, short phrases, and long descriptions; separating those by expected answer length may improve sequence metrics further.
- Evaluation with BLEU under-rewards synonymous medical terms, so a medical semantic-similarity metric could change the measured size of the routing benefit.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes HQS-VQA, a hierarchical model for medical visual question answering that first uses an SVM with hand-crafted features to segregate questions into Yes/No and Others, then routes each question to a dedicated answer-prediction subnetwork. The Yes/No branch is a two-class softmax classifier; the Others branch is a sequence-generation model; image features come from Inception-ResNet-v2, question features from Bi-LSTM, and modalities are fused by concatenation. The model is evaluated on RAD and CLEF18, with and without the question segregation module. The authors report that adding QS improves performance and claim that the full model outperforms existing baselines by significant margins.
Significance. The question-routing idea is simple and potentially transferable, and the internal with/without-QS comparison is a clean ablation that shows consistent gains across both datasets. The paper also provides reproducible code links and a detailed error analysis, which are strengths. However, the headline claim of outperforming baselines is not established: the baseline comparison in Table 6 is uncalibrated, and the paper's own reported numbers contradict the abstract. If the authors reframe the contribution as a QS ablation study, the core result would be of interest to the medical VQA community.
major comments (4)
- [Abstract; Section 5.2, Table 6] The abstract's claim that the proposed HQS-VQA technique 'outperforms the baseline models with significant margins' is contradicted by the paper's own Table 6. On CLEF18, the proposed model obtains BLEU 0.132, which is below the officially reported Peng et al. (2018) score of 0.161 and essentially tied with the official Zhou et al. (2018) score of 0.134. The text after Table 6 explicitly states that the authors cannot directly compare their re-implementations with the proposed approach because the participants did not use the same evaluation setup. Therefore the headline result is not supported by the evidence presented.
- [Section 5.2] The baseline re-implementations are not validated as faithful reproductions. For Zhou et al. (2018), even the authors' run of the official code yields BLEU 0.072, far below the reported 0.134. For Peng et al. (2018) and Abacha et al. (2018), the re-implementations drop to 0.023 and 0.051, respectively, from reported scores of 0.161 and 0.121. These gaps suggest substantial differences in preprocessing, postprocessing, or evaluation protocol. The paper's conclusion that the proposed model outperforms these baselines is therefore based on an uncalibrated comparison, and the 'significant margins' claim is not justified.
- [Section 5.2.1, Tables 5 and 10] The internal with/without-QS comparison is the most credible result in the paper, but its interpretation is limited by the QS classifier's low recall for Yes/No questions on CLEF18 (0.28, F1-score 0.44 in Table 5). To separate routing accuracy from answer-generation quality, the authors should report an oracle-routing condition in which questions are assigned to the correct branch, and should quantify how much of the gain in Table 8 is lost to routing errors. This would strengthen the otherwise plausible claim that QS itself is responsible for the observed improvements.
- [Conclusion] The conclusion repeats the unsupported claim that the proposed model 'outperformed all the stated baseline models.' This sentence should either be removed or replaced with a statement that clearly distinguishes the QS ablation result from the uncalibrated baseline comparison, consistent with the caveat already acknowledged in Section 5.2.
minor comments (5)
- [Section 3.2.4] The metric name is misspelled as 'BiLingual' instead of 'Bilingual', and the paragraph following Eq. (12) refers to 'BLUE score' instead of 'BLEU score'.
- [Abstract] The sentence 'The existing techniques in VQA-Med fail to distinguish between the different question types sometimes complicates the simpler problems, or over-simplifies the complicated ones' is grammatically awkward and should be rewritten for clarity.
- [Section 3.2.3] The description of selecting the 10 question identifier words should clarify whether the selection was made using the training set, validation set, or external knowledge, to avoid the appearance of training-set peeking.
- [Table 11] The caption states 'question-type Yes/No' but the table shows Others-type questions; the caption should be corrected to avoid confusion.
- [Section 4.1] The list of 11 RAD question categories enumerates only 10 items; the category list should be completed or corrected.
Circularity Check
No circularity: the QS benefit and baseline comparisons are empirical measurements, not reductions to the paper's own inputs.
full rationale
The paper's central claims are empirical: the QS module is an SVM trained on question features, and its effect is measured by comparing the same answer-prediction architecture with and without the segregation module (Table 8). No equation defines the improvement into existence; the with-QS and without-QS models are genuinely different training setups and are evaluated on held-out test data. The abstract's 'significant margins' claim rests on comparisons against re-implemented baselines, and the paper explicitly concedes in Section 5.2 that direct comparison with the official ImageCLEF2018 scores is not possible because 'the participants did not use their own evaluation setup/script.' That is an evaluation-validity problem, not a circularity problem. The hand-selected question identifier words and the tf-idf top-500 feature choice are tuned on training data, but they do not predetermine the downstream answer-generation results, and the QS accuracy itself is reported on test splits. The self-citations to Gupta et al. appear only in the introduction and related work as background on question answering and are not load-bearing for the proposed architecture or its evaluation. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no known result is repackaged as new under a new coordinate system. The derivation chain is therefore self-contained with respect to circularity, and the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (8)
- Question identifier word set =
{is, was, are, how, can, does, which, what, type, there}
- tf-idf top-k feature count =
500 (from vocabulary of 2000)
- Question dictionary size =
1050
- Max question length =
21
- Max answer length (Others) =
11
- Bi-LSTM hidden units per direction =
128
- Training hyperparameters =
dropout 0.5, batch size 256, epochs 251, Adam
- Answer dictionary size =
number of unique answer words in training data
assumptions (5)
- domain assumption Medical VQA questions can be partitioned into two useful types: Yes/No and Others.
- domain assumption A dedicated per-type model outperforms a single joint model for VQA-Med.
- domain assumption ImageNet-pretrained Inception-ResNet-v2 features transfer usefully to radiology images.
- domain assumption SVM with linear kernel and the chosen hand-crafted features is an adequate question segregator.
- ad hoc to paper The re-implementations of baseline systems in Table 6 faithfully reproduce the original systems.
Cite this review
Pith. "Pith review of Hierarchical Deep Multi-modal Network for Medical Visual Question Answering." pith.science (2026). https://pith.science/paper/3MTXNWO6
@misc{pith2026200912770,
author = {Pith},
title = {Pith review of: Hierarchical Deep Multi-modal Network for Medical Visual Question Answering},
year = {2026},
howpublished = {\url{https://pith.science/paper/3MTXNWO6}},
note = {Machine review of arXiv:2009.12770}
}
read the original abstract
Visual Question Answering in Medical domain (VQA-Med) plays an important role in providing medical assistance to the end-users. These users are expected to raise either a straightforward question with a Yes/No answer or a challenging question that requires a detailed and descriptive answer. The existing techniques in VQA-Med fail to distinguish between the different question types sometimes complicates the simpler problems, or over-simplifies the complicated ones. It is certainly true that for different question types, several distinct systems can lead to confusion and discomfort for the end-users. To address this issue, we propose a hierarchical deep multi-modal network that analyzes and classifies end-user questions/queries and then incorporates a query-specific approach for answer prediction. We refer our proposed approach as Hierarchical Question Segregation based Visual Question Answering, in short HQS-VQA. Our contributions are three-fold, viz. firstly, we propose a question segregation (QS) technique for VQAMed; secondly, we integrate the QS model to the hierarchical deep multi-modal neural network to generate proper answers to the queries related to medical images; and thirdly, we study the impact of QS in Medical-VQA by comparing the performance of the proposed model with QS and a model without QS. We evaluate the performance of our proposed model on two benchmark datasets, viz. RAD and CLEF18. Experimental results show that our proposed HQS-VQA technique outperforms the baseline models with significant margins. We also conduct a detailed quantitative and qualitative analysis of the obtained results and discover potential causes of errors and their solutions.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in ":" * " " * FUNCTION f...
-
[2]
author Abacha, A. B. , author Gayen, S. , author Lau, J. J. , author Rajaraman, S. , & author Demner-Fushman, D. ( year 2018 ). title Nlm at imageclef 2018 visual question answering in the medical domain. In booktitle CLEF (Working Notes) \/
2018
-
[3]
, author Agrawal, A
author Antol, S. , author Agrawal, A. , author Lu, J. , author Mitchell, M. , author Batra, D. , author Lawrence Zitnick, C. , & author Parikh, D. ( year 2015 ). title VQA: Visual Question Answering . In booktitle Proceedings of the IEEE international conference on computer vision \/ (pp. pages 2425--2433 )
2015
-
[4]
, & author Kapoor, S
author Arai, K. , & author Kapoor, S. ( year 2019 ). title Advances in Computer Vision: Proceedings of the 2019 Computer Vision Conference (CVC) \/ volume volume 2 . publisher Springer
2019
-
[5]
Enriching Word Vectors with Subword Information
author Bojanowski, P. , author Grave, E. , author Joulin, A. , & author Mikolov, T. ( year 2017 ). title Enriching Word Vectors with Subword Information . journal Transactions of the Association for Computational Linguistics \/ , volume 5 \/ , pages 135--146 . https://www.aclweb.org/anthology/Q17-1010. :10.1162/tacl_a_00051
-
[6]
author Bradley, E. , author Zeynettin, A. , author Jiri, S. , & author Panagiotis, K. ( year 2017 ). title Data From LGG-1p19qDeletion . DOI: https://doi.org/10.7937/K9/TCIA.2017.dwehtz9v
-
[7]
A Thorough Examination of the CNN/Daily Mail Reading Comprehension Task
author Chen, D. , author Bolton, J. , & author Manning, C. D. ( year 2016 ). title A thorough examination of the cnn/daily mail reading comprehension task . journal arXiv preprint arXiv:1606.02858 \/ ,
work page Pith review arXiv 2016
-
[8]
, author Mao, Y
author Chen, D. , author Mao, Y. , & author Zhou, J. ( year 2019 ). title Constructing medical image domain ontology with anatomical knowledge . In booktitle 2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) \/ (pp. pages 1750--1757 ). organization IEEE
2019
Show all 92 references
-
[9]
, author van Merrienboer, B
author Cho, K. , author van Merrienboer, B. , author Gulcehre, C. , author Bahdanau, D. , author Bougares, F. , author Schwenk, H. , & author Bengio, Y. ( year 2014 ). title Learning Phrase Representations using RNN Encoder -- Decoder for Statistical Machine Translation . In b...
2014
-
[10]
author Cid, Y. D. , author Liauchuk, V. , author Kovalev, V. , & author M \"u ller, H. ( year 2018 ). title Overview of ImageCLEFtuberculosis 2018-Detecting Multi-drug Resistance, Classifying Tuberculosis Type, and Assessing Severity Score . In booktitle CLEF2018 Working Notes...
2018
-
[11]
author Clark, A. T. , author Megerian, M. G. , author Petri, J. E. , & author Stevens, R. J. ( year 2018 ). title Question Classification and Feature Mapping in a Deep Question Answering System . note US Patent 9,911,082
2018
-
[12]
, & author Vapnik, V
author Cortes, C. , & author Vapnik, V. ( year 1995 ). title Support-Vector Networks . journal Machine learning \/ , volume 20 \/ , pages 273--297
1995
-
[13]
, author Shawe-Taylor, J
author Cristianini, N. , author Shawe-Taylor, J. et al. ( year 2000 ). title An Introduction to Support Vector Machines and Other Kernel-based Learning Methods \/ . publisher Cambridge university press
2000
-
[14]
, author Chang, M.-W
author Devlin, J. , author Chang, M.-W. , author Lee, K. , & author Toutanova, K. ( year 2018 ). title Bert: Pre-training of deep bidirectional transformers for language understanding . journal arXiv preprint arXiv:1810.04805 \/ ,
2018 arXiv
-
[15]
, author Yang, N
author Dong, L. , author Yang, N. , author Wang, W. , author Wei, F. , author Liu, X. , author Wang, Y. , author Gao, J. , author Zhou, M. , & author Hon, H.-W. ( year 2019 ). title Unified Language Model Pre-training for Natural Language Understanding and Generation . In book...
2019
-
[16]
, author Schwall, I
author Eickhoff, C. , author Schwall, I. , author de Herrera, A. G. S. , & author M \"u ller, H. ( year 2017 ). title Overview of ImageCLEFcaption 2017-Image Caption Prediction and Concept Detection for Biomedical Images . In booktitle CLEF (Working Notes) \/
2017
-
[17]
( year 2019 )
author Fu, Z. ( year 2019 ). title An Introduction of Deep Learning Based Word Representation Applied to Natural Language Processing . In booktitle 2019 International Conference on Machine Learning, Big Data and Business Intelligence (MLBDBI) \/ (pp. pages 92--104 ). organization IEEE
2019
-
[18]
, author Park, D
author Fukui, A. , author Park, D. H. , author Yang, D. , author Rohrbach, A. , author Darrell, T. , & author Rohrbach, M. ( year 2016 ). title Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding . In booktitle Proceedings of the 2016 Confere...
2016 doi
-
[19]
, author Mao, J
author Gao, H. , author Mao, J. , author Zhou, J. , author Huang, Z. , author Wang, L. , & author Xu, W. ( year 2015 ). title Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering . In booktitle Proceedings of the 28th International Confer...
2015
-
[20]
, author Jiang, Z
author Gao, P. , author Jiang, Z. , author You, H. , author Lu, P. , author Hoi, S. C. , author Wang, X. , & author Li, H. ( year 2019 ). title Dynamic Fusion with Intra-and Inter-modality Attention Flow for Visual Question Answering . In booktitle Proceedings of the IEEE Conf...
2019
-
[21]
, & author Wolf, M
author Gebhardt, E. , & author Wolf, M. ( year 2018 ). title Camel Dataset for Visual and Thermal Infrared Multiple Object Detection and Tracking . In booktitle 2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS) \/ (pp. pages 1--6 )....
2018
-
[22]
author Ger, R. B. , author Yang, J. , author Ding, Y. , author Jacobsen, M. C. , author Cardenas, C. E. , author Fuller, C. D. , & author Howell, R. M. ( year 2018 ). title Data from Synthetic and Phantom MR Images for Determining Deformable Image Registration Accuracy (MRI-DI...
2018 doi
-
[23]
, author Favre, B
author Ghannay, S. , author Favre, B. , author Esteve, Y. , & author Camelin, N. ( year 2016 ). title Word Embedding Evaluation and Combination . In booktitle LREC \/ (pp. pages 300--305 )
2016
-
[24]
, & author Schmidhuber, J
author Graves, A. , & author Schmidhuber, J. ( year 2005 ). title Framewise Phoneme Classification with Bidirectional LSTM and other Neural Network Architectures . journal Neural networks : the official journal of the International Neural Network Society \/ , volume 18 \/ , pa...
2005 doi
-
[25]
, author Srivastava, R
author Greff, K. , author Srivastava, R. K. , author Koutn \' k, J. , author Steunebrink, B. R. , & author Schmidhuber, J. ( year 2016 ). title LSTM: A Search Space Odyssey . journal IEEE transactions on neural networks and learning systems \/ , volume 28 \/ , pages 2222--2232
2016
-
[26]
, author He, H
author Guo, J. , author He, H. , author He, T. , author Lausen, L. , author Li, M. , author Lin, H. , author Shi, X. , author Wang, C. , author Xie, J. , author Zha, S. et al. ( year 2020 ). title Gluoncv and Gluonnlp: Deep Learning in Computer Vision and Natural Language Proc...
2020
-
[27]
, author Ekbal, A
author Gupta, D. , author Ekbal, A. , & author Bhattacharyya, P. ( year 2019 ). title A deep neural network framework for english hindi question answering . journal ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP) \/ , volume 19 \/ , pages 1--22
2019
-
[28]
, author Kumari, S
author Gupta, D. , author Kumari, S. , author Ekbal, A. , & author Bhattacharyya, P. ( year 2018 a ). title MMQA: A Multi-domain Multi-lingual Question-Answering Framework for English and Hindi . In editor N. C. C. chair) , editor K. Choukri , editor C. Cieri , editor T. Decle...
2018
-
[29]
, author Lenka, P
author Gupta, D. , author Lenka, P. , author Ekbal, A. , & author Bhattacharyya, P. ( year 2018 b ). title Uncovering code-mixed challenges: A framework for linguistically driven question generation and neural based question answering . In booktitle Proceedings of the 22nd Con...
2018
-
[30]
, author Pujari, R
author Gupta, D. , author Pujari, R. , author Ekbal, A. , author Bhattacharyya, P. , author Maitra, A. , author Jain, T. , & author Sengupta, S. ( year 2018 c ). title Can Taxonomy Help? Improving Semantic Question Matching using Question Taxonomy . In booktitle Proceedings of...
2018
-
[31]
author Hasan, S. A. , author Ling, Y. , author Farri, O. , author Liu, J. , author Lungren, M. , & author M \"u ller, H. ( year 2018 ). title Overview of the imageclef 2018 medical domain visual question answering task . In booktitle CLEF2018 Working Notes. CEUR Workshop Proce...
2018
-
[32]
, author Zhang, X
author He, K. , author Zhang, X. , author Ren, S. , & author Sun, J. ( year 2016 ). title Deep Residual Learning for Image Recognition . In booktitle Proceedings of the IEEE conference on computer vision and pattern recognition \/ (pp. pages 770--778 )
2016
-
[33]
author Hersh, W. R. , & author Bhupatiraju, R. T. ( year 2003 ). title TREC GENOMICS Track Overview . In booktitle TREC \/
2003
-
[34]
, & author Szegedy, C
author Ioffe, S. , & author Szegedy, C. ( year 2015 ). title Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift . In booktitle Proceedings of the 32Nd International Conference on International Conference on Machine Learning - Volume 37...
2015
-
[35]
, author M \"u ller, H
author Ionescu, B. , author M \"u ller, H. , author Villegas, M. , author de Herrera, A. G. S. , author Eickhoff, C. , author Andrearczyk, V. , author Cid, Y. D. , author Liauchuk, V. , author Kovalev, V. , author Hasan, S. A. et al. ( year 2018 ). title Overview of ImageCLEF ...
2018
-
[36]
, & author Kanan , C
author Kafle , K. , & author Kanan , C. ( year 2016 ). title Answer-Type Prediction for Visual Question Answering . In booktitle 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) \/ (pp. pages 4976--4984 ). :10.1109/CVPR.2016.538
2016 doi
-
[37]
, author Shrestha, R
author Kafle, K. , author Shrestha, R. , author Cohen, S. , author Price, B. , & author Kanan, C. ( year 2020 ). title Answering Questions about Data Visualizations using Efficient Bimodal Fusion . In booktitle The IEEE Winter Conference on Applications of Computer Vision \/ (...
2020
-
[38]
, & author Hamarneh, G
author Kawahara, J. , & author Hamarneh, G. ( year 2016 ). title Multi-Resolution-Tract CNN with Hybrid Pretrained and Skin-Lesion Trained Layers . In booktitle MLMI@MICCAI \/
2016
-
[39]
, & author Ba, J
author Kingma, D. , & author Ba, J. ( year 2014 ). title Adam: A Method for Stochastic Optimization . journal International Conference on Learning Representations \/ ,
2014
-
[40]
author Lau, J. J. , author Gayen, S. , author Abacha, A. B. , & author Demner-Fushman, D. ( year 2018 ). title A Dataset of Clinically Generated Visual Questions and Answers about Radiology Images . journal Scientific data \/ , volume 5 \/ , pages 180251
2018
-
[41]
, author Grandvalet, Y
author Li, X. , author Grandvalet, Y. , author Davoine, F. , author Cheng, J. , author Cui, Y. , author Zhang, H. , author Belongie, S. , author Tsai, Y.-H. , & author Yang, M.-H. ( year 2020 ). title Transfer Learning in Computer Vision Tasks: Remember where you come from . j...
2020
-
[42]
, author Maire, M
author Lin, T.-Y. , author Maire, M. , author Belongie, S. , author Hays, J. , author Perona, P. , author Ramanan, D. , author Doll \'a r, P. , & author Zitnick, C. L. ( year 2014 ). title Microsoft coco: Common objects in context . In booktitle European conference on computer...
2014
-
[43]
, author RoyChowdhury, A
author Lin, T.-Y. , author RoyChowdhury, A. , & author Maji, S. ( year 2017 ). title Bilinear Convolutional Neural Networks for Fine-grained Visual Recognition . journal IEEE transactions on pattern analysis and machine intelligence \/ , volume 40 \/ , pages 1309--1322
2017
-
[44]
, author Chen, L.-C
author Liu, C. , author Chen, L.-C. , author Schroff, F. , author Adam, H. , author Hua, W. , author Yuille, A. L. , & author Fei-Fei, L. ( year 2019 ). title Auto-deeplab: Hierarchical Neural Architecture Search for Semantic Image Segmentation . In booktitle Proceedings of th...
2019
-
[45]
, author Ouyang, W
author Liu, L. , author Ouyang, W. , author Wang, X. , author Fieguth, P. , author Chen, J. , author Liu, X. , & author Pietik \"a inen, M. ( year 2020 ). title Deep Learning for Generic Object Detection: A survey . journal International journal of computer vision \/ , volume ...
2020
-
[46]
, author Zhu, H
author Long, M. , author Zhu, H. , author Wang, J. , & author Jordan, M. I. ( year 2017 ). title Deep Transfer Learning with Joint Adaptation Networks . In booktitle Proceedings of the 34th International Conference on Machine Learning-Volume 70 \/ (pp. pages 2208--2217 ). orga...
2017
-
[47]
, author Yang, J
author Lu, J. , author Yang, J. , author Batra, D. , & author Parikh, D. ( year 2016 ). title Hierarchical Question-image Co-attention for Visual Question Answering . In booktitle Advances In Neural Information Processing Systems \/ (pp. pages 289--297 )
2016
-
[48]
author Magnuson, J. S. , author You, H. , author Luthra, S. , author Li, M. , author Nam, H. , author Escabi, M. , author Brown, K. , author Allopenna, P. D. , author Theodore, R. M. , author Monto, N. et al. ( year 2020 ). title EARSHOT: A Minimal Neural Network Model of Incr...
2020
-
[49]
, author Rohrbach, M
author Malinowski, M. , author Rohrbach, M. , & author Fritz, M. ( year 2015 ). title Ask Your Neurons: A Neural-based Approach to Answering Questions about Images . In booktitle Proceedings of the IEEE international conference on computer vision \/ (pp. pages 1--9 )
2015
-
[50]
, author Karafi \'a t, M
author Mikolov, T. , author Karafi \'a t, M. , author Burget, L. , author C ernock \`y , J. , & author Khudanpur, S. ( year 2010 ). title Recurrent Neural Network based Language Model . In booktitle Eleventh annual conference of the international speech communication association \/
2010
-
[51]
, author Krallinger, M
author Morante, R. , author Krallinger, M. , author Valencia, A. , & author Daelemans, W. ( year 2013 ). title Machine Reading of Biomedical Texts about Alzheimer's Disease . journal CEUR Workshop Proceedings \/ , volume 1179 \/
2013
-
[52]
, author Rohrbach, A
author Mukuze, N. , author Rohrbach, A. , author Demberg, V. , & author Schiele, B. ( year 2018 ). title A Vision-grounded Dataset for Predicting Typical Locations for Verbs . In booktitle Proceedings of the Eleventh International Conference on Language Resources and Evaluatio...
2018
-
[53]
, author Yadav, S
author Ningthoujam, D. , author Yadav, S. , author Bhattacharyya, P. , & author Ekbal, A. ( year 2019 ). title Relation extraction between the clinical entities based on the shortest dependency path based lstm . journal arXiv preprint arXiv:1903.09941 \/ ,
2019 arXiv
-
[54]
, author Roukos, S
author Papineni, K. , author Roukos, S. , author Ward, T. , & author Zhu, W.-J. ( year 2002 ). title BLEU: A Method for Automatic Evaluation of Machine Translation . In booktitle Proceedings of the 40th annual meeting on association for computational linguistics \/ (pp. pages ...
2002
-
[55]
, author Liu, F
author Peng, Y. , author Liu, F. , & author Rosen, M. P. ( year 2018 ). title Umass at imageclef medical visual question answering (med-vqa) 2018 task. In booktitle CLEF (Working Notes) \/
2018
-
[56]
, author Socher, R
author Pennington, J. , author Socher, R. , & author Manning, C. ( year 2014 ). title Glove: Global Vectors for Word Representation . In booktitle 2014 conference on empirical methods in natural language processing (EMNLP) \/ (pp. pages 1532--1543 )
2014
-
[57]
( year 2019 )
author Ruder, S. ( year 2019 ). title Neural Transfer Learning for Natural Language Processing \/ . Ph.D. thesis NUI Galway
2019
-
[58]
, author Peters, M
author Ruder, S. , author Peters, M. E. , author Swayamdipta, S. , & author Wolf, T. ( year 2019 ). title Transfer Learning in Natural Language Processing . In booktitle Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Lingu...
2019
-
[59]
, author Deng, J
author Russakovsky, O. , author Deng, J. , author Su, H. , author Krause, J. , author Satheesh, S. , author Ma, S. , author Huang, Z. , author Karpathy, A. , author Khosla, A. , author Bernstein, M. et al. ( year 2015 ). title ImageNet Large Scale Visual Recognition Challenge ...
2015
-
[60]
, author Olivier, G
author Shaimaa, B. , author Olivier, G. , author Sebastian, E. , author Kelsey, A. , author Mu, Z. , author Majid, S. , author Hong, Z. , author Weiruo, Z. , author Ann, L. , author Michael, K. , author Joseph, S. , author Andrew, Q. , author Daniel, R. , author Sylvia, P. , &...
2017 doi
-
[61]
, & author Zisserman, A
author Simonyan, K. , & author Zisserman, A. ( year 2015 ). title Very Deep Convolutional Networks for Large-Scale Image Recognition . In booktitle International Conference on Learning Representations \/
2015
-
[62]
, author Xv, H
author Sun, X. , author Xv, H. , author Dong, J. , author Zhou, H. , author Chen, C. , & author Li, Q. ( year 2020 a ). title Few-shot Learning for Domain-specific Fine-grained Image Classification . journal IEEE Transactions on Industrial Electronics \/ ,
2020
-
[63]
, author Xue, B
author Sun, Y. , author Xue, B. , author Zhang, M. , author Yen, G. G. , & author Lv, J. ( year 2020 b ). title Automatically Designing CNN Architectures Using the Genetic Algorithm for Image Classification . journal IEEE Transactions on Cybernetics \/ ,
2020
-
[64]
, author Ioffe, S
author Szegedy, C. , author Ioffe, S. , author Vanhoucke, V. , & author Alemi, A. A. ( year 2017 ). title Inception-v4, Inception-Resnet and the Impact of Residual Connections on Learning . In booktitle Thirty-First AAAI Conference on Artificial Intelligence \/
2017
-
[65]
, author Liu, W
author Szegedy, C. , author Liu, W. , author Jia, Y. , author Sermanet, P. , author Reed, S. E. , author Anguelov, D. , author Erhan, D. , author Vanhoucke, V. , & author Rabinovich, A. ( year 2014 ). title Going Deeper with Convolutions . journal 2015 IEEE Conference on Compu...
2014
-
[66]
, author Shin, J
author Tajbakhsh, N. , author Shin, J. Y. , author Gurudu, S. R. , author Hurst, R. T. , author Kendall, C. B. , author Gotway, M. B. , & author Liang, J. ( year 2016 ). title Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning? journal IEEE ...
2016
-
[67]
author Tarando, S. R. , author Fetita, C. , author Faccinetto, A. , & author Brillet, P.-Y. ( year 2016 ). title Increasing CAD System Efficacy for Lung Texture Analysis using a Convolutional Network . In booktitle Medical Imaging 2016: Computer-Aided Diagnosis \/ (p. pages 97...
2016
-
[68]
author Traore, B. B. , author Kamsu-Foguem, B. , & author Tangara, F. ( year 2018 ). title Deep Convolution Neural Network for Image Recognition . journal Ecological Informatics \/ , volume 48 \/ , pages 257--268
2018
-
[69]
, author Balikas, G
author Tsatsaronis, G. , author Balikas, G. , author Malakasiotis, P. , author Partalas, I. , author Zschunke, M. , author Alvers, M. R. , author Weissenborn, D. , author Krithara, A. , author Petridis, S. , author Polychronopoulos, D. , author Almirantis, Y. , author Pavlopou...
2015
-
[70]
, author Kay-Rivest, E
author Vallieres, M. , author Kay-Rivest, E. , author Perrin, L. J. , author Liem, X. , author Furstoss, C. , author Khaouam, N. , author Nguyen-Tan, P. F. , author Wang, C.-S. , & author Sultanem, K. ( year 2017 ). title Data from Head-Neck-PET-CT . The Cancer Imaging Archive...
2017 doi
-
[71]
, & author de Bruijne , M
author van Tulder , G. , & author de Bruijne , M. ( year 2016 ). title Combining Generative and Discriminative Representation Learning for Lung CT Analysis With Convolutional Restricted Boltzmann Machines . journal IEEE Transactions on Medical Imaging \/ , volume 35 \/ , pages...
2016
-
[72]
, author Zheng, B
author Wang, H. , author Zheng, B. , author Yoon, S. W. , & author Ko, H. S. ( year 2018 ). title A Support Vector Machine-based Ensemble Algorithm for Breast Cancer Diagnosis . journal European Journal of Operational Research \/ , volume 267 \/ , pages 687--699
2018
-
[73]
, & author Palmer, M
author Wu, Z. , & author Palmer, M. ( year 1994 ). title Verb semantics and lexical selection . In booktitle 32nd Annual Meeting of the Association for Computational Linguistics \/ (pp. pages 133--138 ). address Las Cruces, New Mexico, USA : publisher Association for Computati...
1994
-
[74]
, author Merity, S
author Xiong, C. , author Merity, S. , & author Socher, R. ( year 2016 ). title Dynamic Memory Networks for Visual and Textual Question Answering . In booktitle Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 \/ ICML...
2016
-
[75]
, author Ekbal, A
author Yadav, S. , author Ekbal, A. , author Saha, S. , & author Bhattacharyya, P. ( year 2019 ). title A unified multi-task adversarial learning framework for pharmacovigilance mining . In booktitle Proceedings of the 57th Annual Meeting of the Association for Computational L...
2019
-
[76]
, author Ekbal, A
author Yadav, S. , author Ekbal, A. , author Saha, S. , author Bhattacharyya, P. , & author Sheth, A. ( year 2018 ). title Multi-task learning framework for mining crowd intelligence towards clinical treatment . In booktitle Proceedings of the 2018 Conference of the North Amer...
2018
-
[77]
, author Ramteke, P
author Yadav, S. , author Ramteke, P. , author Ekbal, A. , author Saha, S. , & author Bhattacharyya, P. ( year 2020 ). title Exploring disorder-aware attention for clinical event extraction . journal ACM Transactions on Multimedia Computing, Communications, and Applications (T...
2020
-
[78]
, author Li, L
author Yan, X. , author Li, L. , author Xie, C. , author Xiao, J. , & author Gu, L. ( year 2019 ). title Zhejiang university at imageclef 2019 visual question answering in the medical domain. In booktitle CLEF (Working Notes) \/
2019
-
[79]
, author Liu, S
author Yang, M. , author Liu, S. , author Chen, K. , author Zhang, H. , author Zhao, E. , & author Zhao, T. ( year 2020 ). title A Hierarchical Clustering Approach to Fuzzy Semantic Representation of Rare Words in Neural Machine Translation . journal IEEE Transactions on Fuzzy...
2020
-
[80]
, author He, X
author Yang, Z. , author He, X. , author Gao, J. , author Deng, L. , & author Smola, A. ( year 2016 ). title Stacked Attention Networks for Image Question Answering . In booktitle Proceedings of the IEEE conference on computer vision and pattern recognition \/ (pp. pages 21--29 )
2016
-
[81]
, author Yu, J
author Yu, Z. , author Yu, J. , author Fan, J. , & author Tao, D. ( year 2017 ). title Multi-modal factorized bilinear pooling with co-attention learning for visual question answering . In booktitle Proceedings of the IEEE International Conference on Computer Vision (ICCV) \/
2017
-
[82]
, author Yu, J
author Yu, Z. , author Yu, J. , author Xiang, C. , author Fan, J. , & author Tao, D. ( year 2018 ). title Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering . journal IEEE transactions on neural networks and learning systems \/ ...
2018
-
[83]
, author Knoll, F
author Zbontar, J. , author Knoll, F. , author Sriram, A. , author Muckley, M. , author Bruno, M. , author Defazio, A. , author Parente, M. , author Geras, K. , author Katsnelson, J. , author Chandarana, H. , author Zhang, Z. , author Drozdzal, M. , author Romero, A. , author ...
2018
-
[84]
, author Sun, J
author Zhi, J. , author Sun, J. , author Wang, Z. , & author Ding, W. ( year 2018 ). title Support Vector Machine Classifier for Prediction of the Metastasis of Colorectal Cancer . journal International journal of molecular medicine \/ , volume 41 \/ , pages 1419--1426
2018
-
[85]
, author Kang, X
author Zhou, Y. , author Kang, X. , & author Ren, F. ( year 2018 ). title Employing inception-resnet-v2 and bi-lstm for medical domain visual question answering. In booktitle CLEF (Working Notes) \/
2018
-
[86]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
-
[87]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
-
[88]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
-
[89]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
-
[90]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
-
[91]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
-
[92]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 27, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.