Pith. sign in

REVIEW 4 major objections 5 minor 92 references

Hierarchical Deep Multi-modal Network for Medical Visual Question Answering

T0 review · 4 major / 5 minor · reviewed 2026-08-27 · deepseek-v4-flash

Pith's one-line read Medical VQA improves when questions are routed by answer type.

desk verdict Useful routing idea and a clean ablation, but the headline baseline comparison falls apart against the paper's own Table 6. read the letter →

arxiv 2009.12770 v1 pith:3MTXNWO6 submitted 2020-09-27 cs.CL cs.CVcs.LG

classification cs.CLcs.CVcs.LG
keywords VisualQuestionAnsweringMedicalVQAsegregationHierarchicalnetworkMulti-modalfusionSupportVectorMachineRadiologyAnswer-typerouting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that medical visual question answering is best done in two stages: first decide what kind of answer the question needs, then answer with a model built for that kind. The proposed system, HQS-VQA, puts a question-segregation module at the top of a hierarchy, splitting every question into a Yes/No type or an Others (descriptive) type; the two branches use different answer predictors. On the RAD and CLEF18 benchmark datasets the routed model reports higher BLEU and word-semantic-similarity scores than four prior medical VQA systems, and the same model without the routing step scores lower. The practical point is that patients and clinicians asking about medical images would receive answers of the right shape rather than a descriptive phrase where a Yes or No was expected, or a bare Yes/No where a description was required.

What carries the argument

The load-bearing mechanism is the two-level question-routing hierarchy. At the root, a linear SVM classifies each question using a binary presence vector for ten question-identifier words ('is', 'was', 'are', 'how', 'can', 'does', 'which', 'what', 'type', 'there') concatenated with a tf-idf vector over the top 500 vocabulary terms. The output class chooses the learning path: Yes/No questions go to a two-class softmax, while Others questions go to a multi-label decoder that emits answer words from a vocabulary. Images are encoded with Inception-Resnet-v2 into 1000 features, questions with a Bi-LSTM over concatenated word and subword embeddings, and the modalities are fused by concatenation followed by batch normalization. The routing works by reducing the output space for each question to the only plausible answer type.

What would settle it

Run the original released checkpoints or official challenge submissions of the four prior systems on the RAD and CLEF18 test sets with the paper's evaluation script. If they reproduce their originally reported scores, for example CLEF18 BLEU 0.161 and 0.134 for the two best-scoring prior systems, the paper's significant margins over those baselines would not hold.

Watch

Extended reading notes

Core claim

The central claim is that question segregation is itself a performance lever, not just a convenience. A linear SVM that marks a small set of question-identifier words and a tf-idf vector separates Yes/No from Others questions with F1 near 0.99 on RAD and 0.91 on CLEF18; routing on that decision into two specialist answer models — a two-class softmax over {Yes, No} and a word-by-word sequence predictor — improves BLEU and WBBS scores on RAD, CLEF18, and their combination, compared with the same multimodal network left unsegmented. The paper reports that the router raises Yes/No precision, recall, and F1 by as much as 0.5 points and keeps descriptive questions from being collapsed into Yes/No answers.

Load-bearing premise

The claimed margins over prior systems rest on the assumption that the authors' re-implementations of those systems are as strong as the original published systems; if the originals score higher on the same test sets, the reported margins shrink or vanish.

Editorial extensions

If this is right

  • Yes/No questions can be answered with a two-word output vocabulary, removing the dominant error mode in which a monolithic model answers them with descriptive phrases.
  • A question-segregation layer can be placed on top of an existing multimodal encoder and answer generator, and the paper shows that adding it raises scores on both RAD and CLEF18.
  • The segregation benefit survives training on RAD and CLEF18 combined, so it is not tied to one dataset's particular mix of question types.
  • On small medical datasets, simpler fusion without elaborate attention can beat more complex fusion mechanisms, because dedicated branches are easier to train and less prone to overfitting.
  • The error categories identified in the paper — semantic mismatch, modality or plane confusion, specification gaps, boundary loss, and miscellaneous reasoning errors — give concrete targets for the next generation of medical VQA systems.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the routing gain comes from shrinking the output space, then a finer taxonomy than the binary Yes/No vs Others split, such as RAD's eleven native question categories, should yield further gains; the paper does not test that extension.
  • The router is cheap and model-agnostic, so a clean test of its value would freeze the answer generators and toggle only the routing module across several general-domain VQA datasets.
  • The 'Others' branch still mixes one-word answers, short phrases, and long descriptions; separating those by expected answer length may improve sequence metrics further.
  • Evaluation with BLEU under-rewards synonymous medical terms, so a medical semantic-similarity metric could change the measured size of the routing benefit.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Request a human review

A listed scientist reviews the paper for a fee and the review publishes here regardless of verdict. See the reviewers or get listed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper proposes HQS-VQA, a hierarchical model for medical visual question answering that first uses an SVM with hand-crafted features to segregate questions into Yes/No and Others, then routes each question to a dedicated answer-prediction subnetwork. The Yes/No branch is a two-class softmax classifier; the Others branch is a sequence-generation model; image features come from Inception-ResNet-v2, question features from Bi-LSTM, and modalities are fused by concatenation. The model is evaluated on RAD and CLEF18, with and without the question segregation module. The authors report that adding QS improves performance and claim that the full model outperforms existing baselines by significant margins.

Significance. The question-routing idea is simple and potentially transferable, and the internal with/without-QS comparison is a clean ablation that shows consistent gains across both datasets. The paper also provides reproducible code links and a detailed error analysis, which are strengths. However, the headline claim of outperforming baselines is not established: the baseline comparison in Table 6 is uncalibrated, and the paper's own reported numbers contradict the abstract. If the authors reframe the contribution as a QS ablation study, the core result would be of interest to the medical VQA community.

major comments (4)
  1. [Abstract; Section 5.2, Table 6] The abstract's claim that the proposed HQS-VQA technique 'outperforms the baseline models with significant margins' is contradicted by the paper's own Table 6. On CLEF18, the proposed model obtains BLEU 0.132, which is below the officially reported Peng et al. (2018) score of 0.161 and essentially tied with the official Zhou et al. (2018) score of 0.134. The text after Table 6 explicitly states that the authors cannot directly compare their re-implementations with the proposed approach because the participants did not use the same evaluation setup. Therefore the headline result is not supported by the evidence presented.
  2. [Section 5.2] The baseline re-implementations are not validated as faithful reproductions. For Zhou et al. (2018), even the authors' run of the official code yields BLEU 0.072, far below the reported 0.134. For Peng et al. (2018) and Abacha et al. (2018), the re-implementations drop to 0.023 and 0.051, respectively, from reported scores of 0.161 and 0.121. These gaps suggest substantial differences in preprocessing, postprocessing, or evaluation protocol. The paper's conclusion that the proposed model outperforms these baselines is therefore based on an uncalibrated comparison, and the 'significant margins' claim is not justified.
  3. [Section 5.2.1, Tables 5 and 10] The internal with/without-QS comparison is the most credible result in the paper, but its interpretation is limited by the QS classifier's low recall for Yes/No questions on CLEF18 (0.28, F1-score 0.44 in Table 5). To separate routing accuracy from answer-generation quality, the authors should report an oracle-routing condition in which questions are assigned to the correct branch, and should quantify how much of the gain in Table 8 is lost to routing errors. This would strengthen the otherwise plausible claim that QS itself is responsible for the observed improvements.
  4. [Conclusion] The conclusion repeats the unsupported claim that the proposed model 'outperformed all the stated baseline models.' This sentence should either be removed or replaced with a statement that clearly distinguishes the QS ablation result from the uncalibrated baseline comparison, consistent with the caveat already acknowledged in Section 5.2.
minor comments (5)
  1. [Section 3.2.4] The metric name is misspelled as 'BiLingual' instead of 'Bilingual', and the paragraph following Eq. (12) refers to 'BLUE score' instead of 'BLEU score'.
  2. [Abstract] The sentence 'The existing techniques in VQA-Med fail to distinguish between the different question types sometimes complicates the simpler problems, or over-simplifies the complicated ones' is grammatically awkward and should be rewritten for clarity.
  3. [Section 3.2.3] The description of selecting the 10 question identifier words should clarify whether the selection was made using the training set, validation set, or external knowledge, to avoid the appearance of training-set peeking.
  4. [Table 11] The caption states 'question-type Yes/No' but the table shows Others-type questions; the caption should be corrected to avoid confusion.
  5. [Section 4.1] The list of 11 RAD question categories enumerates only 10 items; the category list should be completed or corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the QS benefit and baseline comparisons are empirical measurements, not reductions to the paper's own inputs.

full rationale

The paper's central claims are empirical: the QS module is an SVM trained on question features, and its effect is measured by comparing the same answer-prediction architecture with and without the segregation module (Table 8). No equation defines the improvement into existence; the with-QS and without-QS models are genuinely different training setups and are evaluated on held-out test data. The abstract's 'significant margins' claim rests on comparisons against re-implemented baselines, and the paper explicitly concedes in Section 5.2 that direct comparison with the official ImageCLEF2018 scores is not possible because 'the participants did not use their own evaluation setup/script.' That is an evaluation-validity problem, not a circularity problem. The hand-selected question identifier words and the tf-idf top-500 feature choice are tuned on training data, but they do not predetermine the downstream answer-generation results, and the QS accuracy itself is reported on test splits. The self-citations to Gupta et al. appear only in the introduction and related work as background on question answering and are not load-bearing for the proposed architecture or its evaluation. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no known result is repackaged as new under a new coordinate system. The derivation chain is therefore self-contained with respect to circularity, and the appropriate finding is no significant circularity.

Assumptions & free parameters 8 free parameters · 5 assumptions · 0 invented entities

This paper introduces no new theoretical entities. The central empirical claim rests on the hand-designed question segregation features and a set of architectural hyperparameters, all fitted to the development data. The most important free choice is the 10-word question identifier set, which defines the SVM feature space. The binary question-type split is an unvalidated domain assumption. The fairness of the baseline comparison depends on the fidelity of the authors' re-implementations, which is questionable given the large gap to official scores.

free parameters (8)
  • Question identifier word set = {is, was, are, how, can, does, which, what, type, there}
    Hand-selected after studying the training set (Section 3.2.3); directly defines the QS feature vector.
  • tf-idf top-k feature count = 500 (from vocabulary of 2000)
    Section 3.2.3: top 500 words with highest tf-idf are kept; this choice determines the QS feature dimension.
  • Question dictionary size = 1050
    Section 3.2.3: most frequent words in questions; controls embedding lookup vocabulary.
  • Max question length = 21
    Section 3.2.3: negligible number of questions exceed 21 words; used for padding/truncation.
  • Max answer length (Others) = 11
    Section 3.2.3: prunes long answers for the Others leaf model.
  • Bi-LSTM hidden units per direction = 128
    Section 3.2.3: hidden state size for question representation.
  • Training hyperparameters = dropout 0.5, batch size 256, epochs 251, Adam
    Section 3.2.3: tuned on validation; affects both leaf models.
  • Answer dictionary size = number of unique answer words in training data
    Section 3.2.3: defined by training answers; a data-derived size, not an external constant.
assumptions (5)
  • domain assumption Medical VQA questions can be partitioned into two useful types: Yes/No and Others.
    The whole hierarchy depends on this binary split (Section 3.1, Section 3.2.1). The paper does not compare against finer or alternative taxonomies, though RAD has 11 native categories.
  • domain assumption A dedicated per-type model outperforms a single joint model for VQA-Med.
    This is the hypothesis the paper tests, but it is assumed as the right design prior in Section 3.2 and used to justify the hierarchy.
  • domain assumption ImageNet-pretrained Inception-ResNet-v2 features transfer usefully to radiology images.
    Invoked in Section 3.2.2 with a citation to Tajbakhsh et al.; no in-paper evidence beyond final results.
  • domain assumption SVM with linear kernel and the chosen hand-crafted features is an adequate question segregator.
    Justified by small dataset size in Section 3.2.1, but the adequacy of this specific feature set is not independently validated.
  • ad hoc to paper The re-implementations of baseline systems in Table 6 faithfully reproduce the original systems.
    The fairness of the baseline comparison rests on this. The large gap between official and re-implemented scores makes this assumption questionable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hierarchical Deep Multi-modal Network for Medical Visual Question Answering." pith.science (2026). https://pith.science/paper/3MTXNWO6

@misc{pith2026200912770,
  author       = {Pith},
  title        = {Pith review of: Hierarchical Deep Multi-modal Network for Medical Visual Question Answering},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3MTXNWO6}},
  note         = {Machine review of arXiv:2009.12770}
}
read the original abstract

Visual Question Answering in Medical domain (VQA-Med) plays an important role in providing medical assistance to the end-users. These users are expected to raise either a straightforward question with a Yes/No answer or a challenging question that requires a detailed and descriptive answer. The existing techniques in VQA-Med fail to distinguish between the different question types sometimes complicates the simpler problems, or over-simplifies the complicated ones. It is certainly true that for different question types, several distinct systems can lead to confusion and discomfort for the end-users. To address this issue, we propose a hierarchical deep multi-modal network that analyzes and classifies end-user questions/queries and then incorporates a query-specific approach for answer prediction. We refer our proposed approach as Hierarchical Question Segregation based Visual Question Answering, in short HQS-VQA. Our contributions are three-fold, viz. firstly, we propose a question segregation (QS) technique for VQAMed; secondly, we integrate the QS model to the hierarchical deep multi-modal neural network to generate proper answers to the queries related to medical images; and thirdly, we study the impact of QS in Medical-VQA by comparing the performance of the proposed model with QS and a model without QS. We evaluate the performance of our proposed model on two benchmark datasets, viz. RAD and CLEF18. Experimental results show that our proposed HQS-VQA technique outperforms the baseline models with significant margins. We also conduct a detailed quantitative and qualitative analysis of the obtained results and discover potential causes of errors and their solutions.

Figures

Figures reproduced from arXiv: 2009.12770 by the authors.

Figure 1
Figure 1. Framework for VQA where, question and image are taken as input to generate or predict the answer. VQA systems differ from each other in the way they fuse multi-modal information. Although most open-ended VQA algorithms used the classification mechanism, this strategy can only produce answers seen during training. The multi-word response is generated one word at a time using LSTM (Gao et al., 2015; Malinowski et al.,… view at source ↗
Figure 2
Figure 2. Abstract representation of the proposed hierarchical HQS-VQA model. The first level segregates the questions while the second level generates the answer using the leaf node. Answer prediction strategy is decided based on the question type. 3.2.1. Question Segregation Question Segregation, in general, segregates the questions from a question list Q = [q1, q2 . . . qn] based on the q type, where n denotes the total nu… view at source ↗
Figure 3
Figure 3. Proposed question segregation module with linear SVM learner as base classifier. The extracted feature vectors are fed to the SVM for question segregation. The classifier, and input to QS module i.e. question feature vectors generated from the questions in the dataset are explained as follows: Question Feature Vectors: From each question, we extract the following two vectors: • Question Identifier Vector: We form a … view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Flowchart of generating question embedding. The word embedding is the concatenation of GloVe and Custom (sub-word) embedding, which are used along with the integer sequence representation of the questions by the embedding layer. (Bojanowski et al., 2017) work on FastTe…
Figure 5
Figure 5. Figure 5: Canonical form of a 2-layer ResNet block. Layer-2 is skipped over activation from layer-1 using residual link. 17 [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 6
Figure 6. Figure 6: The processing of a question by Bi-LSTM to get the question representation. The word embedding of each word in the question is fed to a Bi-LSTM network, and the forward and backward hidden state outputs are concatenated at each time-step to get the final representation…
Figure 7
Figure 7. Figure 7: The architecture represents the fusion of the two modalities’ feature vectors using concatenate layer, the output of which is normalized to generate the fused feature vector. Answer Prediction - Yes/No: We treat model m1 (c.f. Section 3.2.2) as a two-class classificati…
Figure 8
Figure 8. Figure 8: Sample images in the Medical-VQA dataset. The images in this dataset can be of different organs (a to d) and/or modalities (e to h). 25 [PITH_FULL_IMAGE:figures/full_fig_p025_8.png]
Figure 9
Figure 9. Figure 9: Word-Frequency distribution in the RAD dataset. The graph demonstrates that for both the questions and the answers, this distribution is almost similar for train and test data splits. radiology images and their respective captions extracted from the PubMed Central arti…
Figure 10
Figure 10. Figure 10: Sample images in the CLEF18 dataset. Some of the images in this dataset are blurred (hazy/not clear), and/or contains short information in the form of radiology markings, and/or contains stack of sub-images. Question categorization is not present in the dataset and on…
Figure 11
Figure 11. Figure 11: Word-Frequency distribution in the CLEF18 dataset. The graph shows that for the data splits, this distribution is not comparable for both the questions and the answers. Towards this, we use the following baseline models. 1. ResNet152 + LSTM + MFH (Peng et al., 2018): …
Figure 12
Figure 12. Figure 12: RAD CLEF18 CLEF18+RAD without with without with without with Yes/No 0.606 0.634 0.400 0.620 0.534 0.581 Others 0.099 0.129 0.015 0.080 0.064 0.102 Overall 0.392 0.411 0.053 0.132 0.213 0.257 [PITH_FULL_IMAGE:figures/full_fig_p035_12.png]
Figure 12
Figure 12. Figure 12: Impact of QS on the model performance. it shows that with QS the model performs better regardless of the type of question or dataset. Impact of QS on questions with type ‘Yes/No’: For this type of question, the main advantage of QS is that it prevents the model from p…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

92 extracted references · 72 canonical work pages

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in ":" * " " * FUNCTION f...

  2. [2]

    author Abacha, A. B. , author Gayen, S. , author Lau, J. J. , author Rajaraman, S. , & author Demner-Fushman, D. ( year 2018 ). title Nlm at imageclef 2018 visual question answering in the medical domain. In booktitle CLEF (Working Notes) \/

  3. [3]

    , author Agrawal, A

    author Antol, S. , author Agrawal, A. , author Lu, J. , author Mitchell, M. , author Batra, D. , author Lawrence Zitnick, C. , & author Parikh, D. ( year 2015 ). title VQA: Visual Question Answering . In booktitle Proceedings of the IEEE international conference on computer vision \/ (pp. pages 2425--2433 )

  4. [4]

    , & author Kapoor, S

    author Arai, K. , & author Kapoor, S. ( year 2019 ). title Advances in Computer Vision: Proceedings of the 2019 Computer Vision Conference (CVC) \/ volume volume 2 . publisher Springer

  5. [5]

    Enriching Word Vectors with Subword Information

    author Bojanowski, P. , author Grave, E. , author Joulin, A. , & author Mikolov, T. ( year 2017 ). title Enriching Word Vectors with Subword Information . journal Transactions of the Association for Computational Linguistics \/ , volume 5 \/ , pages 135--146 . https://www.aclweb.org/anthology/Q17-1010. :10.1162/tacl_a_00051

  6. [6]

    , author Zeynettin, A

    author Bradley, E. , author Zeynettin, A. , author Jiri, S. , & author Panagiotis, K. ( year 2017 ). title Data From LGG-1p19qDeletion . DOI: https://doi.org/10.7937/K9/TCIA.2017.dwehtz9v

  7. [7]

    A Thorough Examination of the CNN/Daily Mail Reading Comprehension Task

    author Chen, D. , author Bolton, J. , & author Manning, C. D. ( year 2016 ). title A thorough examination of the cnn/daily mail reading comprehension task . journal arXiv preprint arXiv:1606.02858 \/ ,

  8. [8]

    , author Mao, Y

    author Chen, D. , author Mao, Y. , & author Zhou, J. ( year 2019 ). title Constructing medical image domain ontology with anatomical knowledge . In booktitle 2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) \/ (pp. pages 1750--1757 ). organization IEEE

Show all 92 references
  1. [9]

    , author van Merrienboer, B

    author Cho, K. , author van Merrienboer, B. , author Gulcehre, C. , author Bahdanau, D. , author Bougares, F. , author Schwenk, H. , & author Bengio, Y. ( year 2014 ). title Learning Phrase Representations using RNN Encoder -- Decoder for Statistical Machine Translation . In b...

  2. [10]

    author Cid, Y. D. , author Liauchuk, V. , author Kovalev, V. , & author M \"u ller, H. ( year 2018 ). title Overview of ImageCLEFtuberculosis 2018-Detecting Multi-drug Resistance, Classifying Tuberculosis Type, and Assessing Severity Score . In booktitle CLEF2018 Working Notes...

  3. [11]

    author Clark, A. T. , author Megerian, M. G. , author Petri, J. E. , & author Stevens, R. J. ( year 2018 ). title Question Classification and Feature Mapping in a Deep Question Answering System . note US Patent 9,911,082

  4. [12]

    , & author Vapnik, V

    author Cortes, C. , & author Vapnik, V. ( year 1995 ). title Support-Vector Networks . journal Machine learning \/ , volume 20 \/ , pages 273--297

  5. [13]

    , author Shawe-Taylor, J

    author Cristianini, N. , author Shawe-Taylor, J. et al. ( year 2000 ). title An Introduction to Support Vector Machines and Other Kernel-based Learning Methods \/ . publisher Cambridge university press

  6. [14]

    , author Chang, M.-W

    author Devlin, J. , author Chang, M.-W. , author Lee, K. , & author Toutanova, K. ( year 2018 ). title Bert: Pre-training of deep bidirectional transformers for language understanding . journal arXiv preprint arXiv:1810.04805 \/ ,

  7. [15]

    , author Yang, N

    author Dong, L. , author Yang, N. , author Wang, W. , author Wei, F. , author Liu, X. , author Wang, Y. , author Gao, J. , author Zhou, M. , & author Hon, H.-W. ( year 2019 ). title Unified Language Model Pre-training for Natural Language Understanding and Generation . In book...

  8. [16]

    , author Schwall, I

    author Eickhoff, C. , author Schwall, I. , author de Herrera, A. G. S. , & author M \"u ller, H. ( year 2017 ). title Overview of ImageCLEFcaption 2017-Image Caption Prediction and Concept Detection for Biomedical Images . In booktitle CLEF (Working Notes) \/

  9. [17]

    ( year 2019 )

    author Fu, Z. ( year 2019 ). title An Introduction of Deep Learning Based Word Representation Applied to Natural Language Processing . In booktitle 2019 International Conference on Machine Learning, Big Data and Business Intelligence (MLBDBI) \/ (pp. pages 92--104 ). organization IEEE

  10. [18]

    , author Park, D

    author Fukui, A. , author Park, D. H. , author Yang, D. , author Rohrbach, A. , author Darrell, T. , & author Rohrbach, M. ( year 2016 ). title Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding . In booktitle Proceedings of the 2016 Confere...

  11. [19]

    , author Mao, J

    author Gao, H. , author Mao, J. , author Zhou, J. , author Huang, Z. , author Wang, L. , & author Xu, W. ( year 2015 ). title Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering . In booktitle Proceedings of the 28th International Confer...

  12. [20]

    , author Jiang, Z

    author Gao, P. , author Jiang, Z. , author You, H. , author Lu, P. , author Hoi, S. C. , author Wang, X. , & author Li, H. ( year 2019 ). title Dynamic Fusion with Intra-and Inter-modality Attention Flow for Visual Question Answering . In booktitle Proceedings of the IEEE Conf...

  13. [21]

    , & author Wolf, M

    author Gebhardt, E. , & author Wolf, M. ( year 2018 ). title Camel Dataset for Visual and Thermal Infrared Multiple Object Detection and Tracking . In booktitle 2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS) \/ (pp. pages 1--6 )....

  14. [22]

    author Ger, R. B. , author Yang, J. , author Ding, Y. , author Jacobsen, M. C. , author Cardenas, C. E. , author Fuller, C. D. , & author Howell, R. M. ( year 2018 ). title Data from Synthetic and Phantom MR Images for Determining Deformable Image Registration Accuracy (MRI-DI...

  15. [23]

    , author Favre, B

    author Ghannay, S. , author Favre, B. , author Esteve, Y. , & author Camelin, N. ( year 2016 ). title Word Embedding Evaluation and Combination . In booktitle LREC \/ (pp. pages 300--305 )

  16. [24]

    , & author Schmidhuber, J

    author Graves, A. , & author Schmidhuber, J. ( year 2005 ). title Framewise Phoneme Classification with Bidirectional LSTM and other Neural Network Architectures . journal Neural networks : the official journal of the International Neural Network Society \/ , volume 18 \/ , pa...

  17. [25]

    , author Srivastava, R

    author Greff, K. , author Srivastava, R. K. , author Koutn \' k, J. , author Steunebrink, B. R. , & author Schmidhuber, J. ( year 2016 ). title LSTM: A Search Space Odyssey . journal IEEE transactions on neural networks and learning systems \/ , volume 28 \/ , pages 2222--2232

  18. [26]

    , author He, H

    author Guo, J. , author He, H. , author He, T. , author Lausen, L. , author Li, M. , author Lin, H. , author Shi, X. , author Wang, C. , author Xie, J. , author Zha, S. et al. ( year 2020 ). title Gluoncv and Gluonnlp: Deep Learning in Computer Vision and Natural Language Proc...

  19. [27]

    , author Ekbal, A

    author Gupta, D. , author Ekbal, A. , & author Bhattacharyya, P. ( year 2019 ). title A deep neural network framework for english hindi question answering . journal ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP) \/ , volume 19 \/ , pages 1--22

  20. [28]

    , author Kumari, S

    author Gupta, D. , author Kumari, S. , author Ekbal, A. , & author Bhattacharyya, P. ( year 2018 a ). title MMQA: A Multi-domain Multi-lingual Question-Answering Framework for English and Hindi . In editor N. C. C. chair) , editor K. Choukri , editor C. Cieri , editor T. Decle...

  21. [29]

    , author Lenka, P

    author Gupta, D. , author Lenka, P. , author Ekbal, A. , & author Bhattacharyya, P. ( year 2018 b ). title Uncovering code-mixed challenges: A framework for linguistically driven question generation and neural based question answering . In booktitle Proceedings of the 22nd Con...

  22. [30]

    , author Pujari, R

    author Gupta, D. , author Pujari, R. , author Ekbal, A. , author Bhattacharyya, P. , author Maitra, A. , author Jain, T. , & author Sengupta, S. ( year 2018 c ). title Can Taxonomy Help? Improving Semantic Question Matching using Question Taxonomy . In booktitle Proceedings of...

  23. [31]

    author Hasan, S. A. , author Ling, Y. , author Farri, O. , author Liu, J. , author Lungren, M. , & author M \"u ller, H. ( year 2018 ). title Overview of the imageclef 2018 medical domain visual question answering task . In booktitle CLEF2018 Working Notes. CEUR Workshop Proce...

  24. [32]

    , author Zhang, X

    author He, K. , author Zhang, X. , author Ren, S. , & author Sun, J. ( year 2016 ). title Deep Residual Learning for Image Recognition . In booktitle Proceedings of the IEEE conference on computer vision and pattern recognition \/ (pp. pages 770--778 )

  25. [33]

    author Hersh, W. R. , & author Bhupatiraju, R. T. ( year 2003 ). title TREC GENOMICS Track Overview . In booktitle TREC \/

  26. [34]

    , & author Szegedy, C

    author Ioffe, S. , & author Szegedy, C. ( year 2015 ). title Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift . In booktitle Proceedings of the 32Nd International Conference on International Conference on Machine Learning - Volume 37...

  27. [35]

    , author M \"u ller, H

    author Ionescu, B. , author M \"u ller, H. , author Villegas, M. , author de Herrera, A. G. S. , author Eickhoff, C. , author Andrearczyk, V. , author Cid, Y. D. , author Liauchuk, V. , author Kovalev, V. , author Hasan, S. A. et al. ( year 2018 ). title Overview of ImageCLEF ...

  28. [36]

    , & author Kanan , C

    author Kafle , K. , & author Kanan , C. ( year 2016 ). title Answer-Type Prediction for Visual Question Answering . In booktitle 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) \/ (pp. pages 4976--4984 ). :10.1109/CVPR.2016.538

  29. [37]

    , author Shrestha, R

    author Kafle, K. , author Shrestha, R. , author Cohen, S. , author Price, B. , & author Kanan, C. ( year 2020 ). title Answering Questions about Data Visualizations using Efficient Bimodal Fusion . In booktitle The IEEE Winter Conference on Applications of Computer Vision \/ (...

  30. [38]

    , & author Hamarneh, G

    author Kawahara, J. , & author Hamarneh, G. ( year 2016 ). title Multi-Resolution-Tract CNN with Hybrid Pretrained and Skin-Lesion Trained Layers . In booktitle MLMI@MICCAI \/

  31. [39]

    , & author Ba, J

    author Kingma, D. , & author Ba, J. ( year 2014 ). title Adam: A Method for Stochastic Optimization . journal International Conference on Learning Representations \/ ,

  32. [40]

    author Lau, J. J. , author Gayen, S. , author Abacha, A. B. , & author Demner-Fushman, D. ( year 2018 ). title A Dataset of Clinically Generated Visual Questions and Answers about Radiology Images . journal Scientific data \/ , volume 5 \/ , pages 180251

  33. [41]

    , author Grandvalet, Y

    author Li, X. , author Grandvalet, Y. , author Davoine, F. , author Cheng, J. , author Cui, Y. , author Zhang, H. , author Belongie, S. , author Tsai, Y.-H. , & author Yang, M.-H. ( year 2020 ). title Transfer Learning in Computer Vision Tasks: Remember where you come from . j...

  34. [42]

    , author Maire, M

    author Lin, T.-Y. , author Maire, M. , author Belongie, S. , author Hays, J. , author Perona, P. , author Ramanan, D. , author Doll \'a r, P. , & author Zitnick, C. L. ( year 2014 ). title Microsoft coco: Common objects in context . In booktitle European conference on computer...

  35. [43]

    , author RoyChowdhury, A

    author Lin, T.-Y. , author RoyChowdhury, A. , & author Maji, S. ( year 2017 ). title Bilinear Convolutional Neural Networks for Fine-grained Visual Recognition . journal IEEE transactions on pattern analysis and machine intelligence \/ , volume 40 \/ , pages 1309--1322

  36. [44]

    , author Chen, L.-C

    author Liu, C. , author Chen, L.-C. , author Schroff, F. , author Adam, H. , author Hua, W. , author Yuille, A. L. , & author Fei-Fei, L. ( year 2019 ). title Auto-deeplab: Hierarchical Neural Architecture Search for Semantic Image Segmentation . In booktitle Proceedings of th...

  37. [45]

    , author Ouyang, W

    author Liu, L. , author Ouyang, W. , author Wang, X. , author Fieguth, P. , author Chen, J. , author Liu, X. , & author Pietik \"a inen, M. ( year 2020 ). title Deep Learning for Generic Object Detection: A survey . journal International journal of computer vision \/ , volume ...

  38. [46]

    , author Zhu, H

    author Long, M. , author Zhu, H. , author Wang, J. , & author Jordan, M. I. ( year 2017 ). title Deep Transfer Learning with Joint Adaptation Networks . In booktitle Proceedings of the 34th International Conference on Machine Learning-Volume 70 \/ (pp. pages 2208--2217 ). orga...

  39. [47]

    , author Yang, J

    author Lu, J. , author Yang, J. , author Batra, D. , & author Parikh, D. ( year 2016 ). title Hierarchical Question-image Co-attention for Visual Question Answering . In booktitle Advances In Neural Information Processing Systems \/ (pp. pages 289--297 )

  40. [48]

    author Magnuson, J. S. , author You, H. , author Luthra, S. , author Li, M. , author Nam, H. , author Escabi, M. , author Brown, K. , author Allopenna, P. D. , author Theodore, R. M. , author Monto, N. et al. ( year 2020 ). title EARSHOT: A Minimal Neural Network Model of Incr...

  41. [49]

    , author Rohrbach, M

    author Malinowski, M. , author Rohrbach, M. , & author Fritz, M. ( year 2015 ). title Ask Your Neurons: A Neural-based Approach to Answering Questions about Images . In booktitle Proceedings of the IEEE international conference on computer vision \/ (pp. pages 1--9 )

  42. [50]

    , author Karafi \'a t, M

    author Mikolov, T. , author Karafi \'a t, M. , author Burget, L. , author C ernock \`y , J. , & author Khudanpur, S. ( year 2010 ). title Recurrent Neural Network based Language Model . In booktitle Eleventh annual conference of the international speech communication association \/

  43. [51]

    , author Krallinger, M

    author Morante, R. , author Krallinger, M. , author Valencia, A. , & author Daelemans, W. ( year 2013 ). title Machine Reading of Biomedical Texts about Alzheimer's Disease . journal CEUR Workshop Proceedings \/ , volume 1179 \/

  44. [52]

    , author Rohrbach, A

    author Mukuze, N. , author Rohrbach, A. , author Demberg, V. , & author Schiele, B. ( year 2018 ). title A Vision-grounded Dataset for Predicting Typical Locations for Verbs . In booktitle Proceedings of the Eleventh International Conference on Language Resources and Evaluatio...

  45. [53]

    , author Yadav, S

    author Ningthoujam, D. , author Yadav, S. , author Bhattacharyya, P. , & author Ekbal, A. ( year 2019 ). title Relation extraction between the clinical entities based on the shortest dependency path based lstm . journal arXiv preprint arXiv:1903.09941 \/ ,

  46. [54]

    , author Roukos, S

    author Papineni, K. , author Roukos, S. , author Ward, T. , & author Zhu, W.-J. ( year 2002 ). title BLEU: A Method for Automatic Evaluation of Machine Translation . In booktitle Proceedings of the 40th annual meeting on association for computational linguistics \/ (pp. pages ...

  47. [55]

    , author Liu, F

    author Peng, Y. , author Liu, F. , & author Rosen, M. P. ( year 2018 ). title Umass at imageclef medical visual question answering (med-vqa) 2018 task. In booktitle CLEF (Working Notes) \/

  48. [56]

    , author Socher, R

    author Pennington, J. , author Socher, R. , & author Manning, C. ( year 2014 ). title Glove: Global Vectors for Word Representation . In booktitle 2014 conference on empirical methods in natural language processing (EMNLP) \/ (pp. pages 1532--1543 )

  49. [57]

    ( year 2019 )

    author Ruder, S. ( year 2019 ). title Neural Transfer Learning for Natural Language Processing \/ . Ph.D. thesis NUI Galway

  50. [58]

    , author Peters, M

    author Ruder, S. , author Peters, M. E. , author Swayamdipta, S. , & author Wolf, T. ( year 2019 ). title Transfer Learning in Natural Language Processing . In booktitle Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Lingu...

  51. [59]

    , author Deng, J

    author Russakovsky, O. , author Deng, J. , author Su, H. , author Krause, J. , author Satheesh, S. , author Ma, S. , author Huang, Z. , author Karpathy, A. , author Khosla, A. , author Bernstein, M. et al. ( year 2015 ). title ImageNet Large Scale Visual Recognition Challenge ...

  52. [60]

    , author Olivier, G

    author Shaimaa, B. , author Olivier, G. , author Sebastian, E. , author Kelsey, A. , author Mu, Z. , author Majid, S. , author Hong, Z. , author Weiruo, Z. , author Ann, L. , author Michael, K. , author Joseph, S. , author Andrew, Q. , author Daniel, R. , author Sylvia, P. , &...

  53. [61]

    , & author Zisserman, A

    author Simonyan, K. , & author Zisserman, A. ( year 2015 ). title Very Deep Convolutional Networks for Large-Scale Image Recognition . In booktitle International Conference on Learning Representations \/

  54. [62]

    , author Xv, H

    author Sun, X. , author Xv, H. , author Dong, J. , author Zhou, H. , author Chen, C. , & author Li, Q. ( year 2020 a ). title Few-shot Learning for Domain-specific Fine-grained Image Classification . journal IEEE Transactions on Industrial Electronics \/ ,

  55. [63]

    , author Xue, B

    author Sun, Y. , author Xue, B. , author Zhang, M. , author Yen, G. G. , & author Lv, J. ( year 2020 b ). title Automatically Designing CNN Architectures Using the Genetic Algorithm for Image Classification . journal IEEE Transactions on Cybernetics \/ ,

  56. [64]

    , author Ioffe, S

    author Szegedy, C. , author Ioffe, S. , author Vanhoucke, V. , & author Alemi, A. A. ( year 2017 ). title Inception-v4, Inception-Resnet and the Impact of Residual Connections on Learning . In booktitle Thirty-First AAAI Conference on Artificial Intelligence \/

  57. [65]

    , author Liu, W

    author Szegedy, C. , author Liu, W. , author Jia, Y. , author Sermanet, P. , author Reed, S. E. , author Anguelov, D. , author Erhan, D. , author Vanhoucke, V. , & author Rabinovich, A. ( year 2014 ). title Going Deeper with Convolutions . journal 2015 IEEE Conference on Compu...

  58. [66]

    , author Shin, J

    author Tajbakhsh, N. , author Shin, J. Y. , author Gurudu, S. R. , author Hurst, R. T. , author Kendall, C. B. , author Gotway, M. B. , & author Liang, J. ( year 2016 ). title Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning? journal IEEE ...

  59. [67]

    author Tarando, S. R. , author Fetita, C. , author Faccinetto, A. , & author Brillet, P.-Y. ( year 2016 ). title Increasing CAD System Efficacy for Lung Texture Analysis using a Convolutional Network . In booktitle Medical Imaging 2016: Computer-Aided Diagnosis \/ (p. pages 97...

  60. [68]

    author Traore, B. B. , author Kamsu-Foguem, B. , & author Tangara, F. ( year 2018 ). title Deep Convolution Neural Network for Image Recognition . journal Ecological Informatics \/ , volume 48 \/ , pages 257--268

  61. [69]

    , author Balikas, G

    author Tsatsaronis, G. , author Balikas, G. , author Malakasiotis, P. , author Partalas, I. , author Zschunke, M. , author Alvers, M. R. , author Weissenborn, D. , author Krithara, A. , author Petridis, S. , author Polychronopoulos, D. , author Almirantis, Y. , author Pavlopou...

  62. [70]

    , author Kay-Rivest, E

    author Vallieres, M. , author Kay-Rivest, E. , author Perrin, L. J. , author Liem, X. , author Furstoss, C. , author Khaouam, N. , author Nguyen-Tan, P. F. , author Wang, C.-S. , & author Sultanem, K. ( year 2017 ). title Data from Head-Neck-PET-CT . The Cancer Imaging Archive...

  63. [71]

    , & author de Bruijne , M

    author van Tulder , G. , & author de Bruijne , M. ( year 2016 ). title Combining Generative and Discriminative Representation Learning for Lung CT Analysis With Convolutional Restricted Boltzmann Machines . journal IEEE Transactions on Medical Imaging \/ , volume 35 \/ , pages...

  64. [72]

    , author Zheng, B

    author Wang, H. , author Zheng, B. , author Yoon, S. W. , & author Ko, H. S. ( year 2018 ). title A Support Vector Machine-based Ensemble Algorithm for Breast Cancer Diagnosis . journal European Journal of Operational Research \/ , volume 267 \/ , pages 687--699

  65. [73]

    , & author Palmer, M

    author Wu, Z. , & author Palmer, M. ( year 1994 ). title Verb semantics and lexical selection . In booktitle 32nd Annual Meeting of the Association for Computational Linguistics \/ (pp. pages 133--138 ). address Las Cruces, New Mexico, USA : publisher Association for Computati...

  66. [74]

    , author Merity, S

    author Xiong, C. , author Merity, S. , & author Socher, R. ( year 2016 ). title Dynamic Memory Networks for Visual and Textual Question Answering . In booktitle Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 \/ ICML...

  67. [75]

    , author Ekbal, A

    author Yadav, S. , author Ekbal, A. , author Saha, S. , & author Bhattacharyya, P. ( year 2019 ). title A unified multi-task adversarial learning framework for pharmacovigilance mining . In booktitle Proceedings of the 57th Annual Meeting of the Association for Computational L...

  68. [76]

    , author Ekbal, A

    author Yadav, S. , author Ekbal, A. , author Saha, S. , author Bhattacharyya, P. , & author Sheth, A. ( year 2018 ). title Multi-task learning framework for mining crowd intelligence towards clinical treatment . In booktitle Proceedings of the 2018 Conference of the North Amer...

  69. [77]

    , author Ramteke, P

    author Yadav, S. , author Ramteke, P. , author Ekbal, A. , author Saha, S. , & author Bhattacharyya, P. ( year 2020 ). title Exploring disorder-aware attention for clinical event extraction . journal ACM Transactions on Multimedia Computing, Communications, and Applications (T...

  70. [78]

    , author Li, L

    author Yan, X. , author Li, L. , author Xie, C. , author Xiao, J. , & author Gu, L. ( year 2019 ). title Zhejiang university at imageclef 2019 visual question answering in the medical domain. In booktitle CLEF (Working Notes) \/

  71. [79]

    , author Liu, S

    author Yang, M. , author Liu, S. , author Chen, K. , author Zhang, H. , author Zhao, E. , & author Zhao, T. ( year 2020 ). title A Hierarchical Clustering Approach to Fuzzy Semantic Representation of Rare Words in Neural Machine Translation . journal IEEE Transactions on Fuzzy...

  72. [80]

    , author He, X

    author Yang, Z. , author He, X. , author Gao, J. , author Deng, L. , & author Smola, A. ( year 2016 ). title Stacked Attention Networks for Image Question Answering . In booktitle Proceedings of the IEEE conference on computer vision and pattern recognition \/ (pp. pages 21--29 )

  73. [81]

    , author Yu, J

    author Yu, Z. , author Yu, J. , author Fan, J. , & author Tao, D. ( year 2017 ). title Multi-modal factorized bilinear pooling with co-attention learning for visual question answering . In booktitle Proceedings of the IEEE International Conference on Computer Vision (ICCV) \/

  74. [82]

    , author Yu, J

    author Yu, Z. , author Yu, J. , author Xiang, C. , author Fan, J. , & author Tao, D. ( year 2018 ). title Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering . journal IEEE transactions on neural networks and learning systems \/ ...

  75. [83]

    , author Knoll, F

    author Zbontar, J. , author Knoll, F. , author Sriram, A. , author Muckley, M. , author Bruno, M. , author Defazio, A. , author Parente, M. , author Geras, K. , author Katsnelson, J. , author Chandarana, H. , author Zhang, Z. , author Drozdzal, M. , author Romero, A. , author ...

  76. [84]

    , author Sun, J

    author Zhi, J. , author Sun, J. , author Wang, Z. , & author Ding, W. ( year 2018 ). title Support Vector Machine Classifier for Prediction of the Metastasis of Colorectal Cancer . journal International journal of molecular medicine \/ , volume 41 \/ , pages 1419--1426

  77. [85]

    , author Kang, X

    author Zhou, Y. , author Kang, X. , & author Ren, F. ( year 2018 ). title Employing inception-resnet-v2 and bi-lstm for medical domain visual question answering. In booktitle CLEF (Working Notes) \/

  78. [86]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  79. [87]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  80. [88]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  81. [89]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  82. [90]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  83. [91]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  84. [92]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 27, 2026 · model on record in the stance chip above.