REVIEW 4 major objections 4 minor 130 references
This survey claims that the state of the art in conversational question answering can be organized into three interacting components—history selection, question understanding, and answer prediction—plus five machine-learning techniques, fiv
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A review that categorizes ConvQA components, techniques, models, and datasets, with no new experimental result.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection A usable orientation survey of ConvQA with a coherent taxonomy, but the LLM comparison table contains concrete factual errors that need fixing before the paper can serve as a reference. the 4 major comments →
A Survey of the State-of-the-Art in Conversational Question Answering Systems
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that the current state of the art in ConvQA can be usefully organized as a taxonomy: four strategies for history selection (K-turn, immediate turn, entire history, dynamic), five for question understanding (rewriting, reformulation, named entity recognition, semantic parsing, attention), and four for answer prediction (retrieval-based, generative, retrieval-augmented generation, knowledge graph-based). On top of that architecture, the survey sorts recent work by machine-learning technique—reinforcement learning, knowledge distillation, contrastive learning, active learning, and transfer learning—and profiles the LLMs and datasets that anchor the field. It does no
What carries the argument
The carrying structure is the survey's three-component taxonomy—history selection, question understanding, answer prediction—shown as the system architecture and instantiated in Table 1. The taxonomy is what lets the paper place roughly thirty recent systems into comparable slots, and the same tripartite frame organizes the later sections on machine learning, LLMs, datasets, and open problems.
Load-bearing premise
The survey's claim to cover the state of the art depends on its selection of papers being representative and on its taxonomy categories matching what the cited systems actually do.
What would settle it
Re-run the paper's stated literature-selection protocol—keywords, venues, citation counts—for the same period and check whether Tables 1 and 2 capture every qualifying work with accurate category assignments. Finding a substantial cluster of qualifying ConvQA work that fits no category, or that is systematically misclassified, would settle the comprehensiveness claim against the survey.
If this is right
- A new ConvQA system can be located in the taxonomy by stating its history-selection, question-understanding, and answer-prediction strategies, which makes comparisons across papers more direct.
- No single history-selection strategy dominates: static K-turn selection works well on CoQA and QuAC, while dynamic and reinforcement-learning-based selection is the current frontier for long or noisy conversations.
- The LLM comparison table (parameter count, context length, training tokens, open-source status) gives practitioners a quick filter for choosing a model for ConvQA work.
- The six open research directions act as a checklist for future work, with dynamic conversational history management and ambiguity handling identified as the main barriers to real-world deployment.
- Because datasets differ in whether they require question rewriting, topic switching, or open-domain retrieval, benchmark choice changes what ability a ConvQA system must demonstrate.
Where Pith is reading between the lines
- A testable extension the survey leaves implicit: hold the dataset fixed and sweep the history-selection strategy across LLMs to quantify which component matters most as models grow.
- The taxonomy's boundary between question rewriting, reformulation, and structured representations is likely to blur in practice, since several cited systems combine them; a method-level re-coding would test the taxonomy's stability.
- The paper's multimodal direction could be operationalized by extending a topic-switching dataset like TopiOCQA with image- or audio-bearing turns, creating a benchmark that does not exist in Table 4.
- Given how fast the LLM landscape changes, the model tables will age quickly even if the three-component taxonomy persists; the survey's durable contribution is the architecture, not the model list.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of conversational question answering (ConvQA). It proposes a three-part taxonomy—history selection (k-turn, immediate-turn, entire-history, dynamic), question understanding (rewriting, reformulation, NER, semantic parsing, attention), and answer prediction (retrieval-based, generative, RAG, knowledge-graph)—and uses this taxonomy to organize a review of machine learning techniques (RL, KD, contrastive learning, active learning, transfer learning), large language models (RoBERTa, GPT-4, Gemini 2.0 Flash, Mistral 7B, LLaMA 3), and datasets (CoQA, QuAC, SQuAD 2.0, CANARD, QReCC, TopiOCQA). The paper claims to provide a comprehensive, reference-quality map of the field and identifies open research directions.
Significance. If accurate, this survey would be a useful structured entry point to ConvQA, with a coherent taxonomy, helpful figures, and recent coverage through 2025. The paper is an expository contribution, not a derivation, so its value rests on the accuracy and representativeness of its classifications and comparison tables. The taxonomy itself is defensible, and the manuscript makes its scope and contributions explicit. However, the reliability of the LLM and dataset comparisons is load-bearing for the 'comprehensive analysis' claim, and the factual inconsistencies identified below must be resolved before the survey can serve as a dependable reference.
major comments (4)
- [Table 3, GPT-4 row] The GPT-4 row reports 1.76T parameters and 13T training tokens, citing [97], the GPT-4 Technical Report. That report explicitly does not disclose GPT-4's parameter count or training-token count; these figures are unofficial public estimates, not facts from the cited source. Since Contribution 3 (§1.2) promises a 'dedicated analysis of LLMs' in terms of parameter size, context length, training tokens, and training data, Table 3 is the empirical backbone of that contribution. Please replace such entries with 'not disclosed' or explicitly mark them as unofficial estimates with a separate citation, and likewise verify the 'Training Dataset' column for proprietary models.
- [Table 3, other rows] Other rows in Table 3 need the same precision pass. 'RoBERTa [95] 355M' conflates RoBERTa-base (125M) and RoBERTa-large (355M). 'BERT [114] 137B' training tokens is not a figure reported in the BERT paper. 'PaLM-2 [113] Open Source = Yes' is inaccurate: while PaLM-2 was made available via APIs, its checkpoints were not open-sourced. Also, the column headed 'Token Size' is actually the context-window length; rename it. These issues matter because practitioners may extract Table 3 as a standalone quick reference.
- [Section 1.1] The paper-selection strategy is not reproducible. The text lists databases and journal/conference venues, but gives no search strings, date ranges, numbers of papers retrieved/screened/included, or explicit inclusion/exclusion criteria. 'Citation count' and 'significant impact' are subjective ranking criteria. For a survey whose central claim is comprehensiveness, this omission weakens confidence that the taxonomy and tables are representative. Add a systematic protocol or temper the 'comprehensive' claim accordingly.
- [Section 5.3 and Table 4] The dataset comparison is internally inconsistent. Section 5.3 states that SQuAD 2.0 combines '100,000+ questions from the original dataset with 50,000+ unanswerable questions,' which implies at least 150K questions, but Table 4 lists 129K. If 129K refers only to the train split, that should be stated explicitly. Since the survey promises a 'comprehensive analysis of key ConvQA datasets,' factual consistency in Table 4 is essential.
minor comments (4)
- [Table 1] The row/column alignment in Table 1 is hard to verify from the typeset version, and some classifications appear ambiguous. Please ensure each reference's check marks match the descriptions in Sections 2.1–2.3, and add a legend or notes to clarify overlapping categories such as question rewriting vs. question reformulation.
- [Section 4.3] Claims about Gemini 2.0 Flash's 'unified tokenization and cross-modal attention' should cite the primary technical/model documentation rather than only a comparative preprint [101].
- [References] Several web citations ([106], [110], [112]) lack consistent access dates or version identifiers. A uniform citation style would improve verifiability.
- [Section 5.4] CANARD is described as having 40K question-answer pairs and 5K conversations; this is fine, but add a sentence clarifying that it is derived from QuAC and is a rewriting benchmark, since that context is useful for readers.
Circularity Check
No circular derivation; survey is an independent literature exposition with only non-load-bearing self-citations.
full rationale
This paper is a survey, so its 'derivation chain' is a literature organization and summarization task rather than a mathematical derivation. The central claims ('comprehensive exploration', taxonomy of history selection/question understanding/answer prediction, ML techniques, LLMs, datasets) are supported by a cited corpus and the survey's own descriptive text, not by fitting parameters or deriving predictions from assumptions. The several self-citations ([23], [29], [41], [44], [46], [47]) are used as examples of prior approaches or as sources for descriptive distinctions such as hard vs. soft history selection; however, no benchmark result, equation, or prediction is forced by these citations, and the survey invokes no self-cited uniqueness theorem and does not forbid alternatives. The skeptical concerns about Table 3 — e.g., GPT-4's 1.76T parameters and 13T training tokens being unofficial estimates not disclosed in the cited GPT-4 technical report, and BERT's '137B' training-token figure not appearing in the BERT paper — are factual/correctness risks for the survey's reliability, not circularity, because they do not reduce the survey's claims to its own inputs. No specific circular reduction can be exhibited, so the appropriate circularity score is 0.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption The selected literature is representative of the state of the art in ConvQA.
- domain assumption The taxonomy (history selection, question understanding, answer prediction; plus the five ML techniques) is a faithful way to describe all surveyed systems.
- domain assumption The factual specifications in Table 3 are accurate.
Cite this review
Pith. "Pith review of A Survey of the State-of-the-Art in Conversational Question Answering Systems." pith.science (2026). https://pith.science/paper/I3KRGWX4
@misc{pith2026250905716,
author = {Pith},
title = {Pith review of: A Survey of the State-of-the-Art in Conversational Question Answering Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/I3KRGWX4}},
note = {Machine review of arXiv:2509.05716}
}
read the original abstract
Conversational Question Answering (ConvQA) systems have emerged as a pivotal area within Natural Language Processing (NLP) by driving advancements that enable machines to engage in dynamic and context-aware conversations. These capabilities are increasingly being applied across various domains, i.e., customer support, education, legal, and healthcare where maintaining a coherent and relevant conversation is essential. Building on recent advancements, this survey provides a comprehensive analysis of the state-of-the-art in ConvQA. This survey begins by examining the core components of ConvQA systems, i.e., history selection, question understanding, and answer prediction, highlighting their interplay in ensuring coherence and relevance in multi-turn conversations. It further investigates the use of advanced machine learning techniques, including but not limited to, reinforcement learning, contrastive learning, and transfer learning to improve ConvQA accuracy and efficiency. The pivotal role of large language models, i.e., RoBERTa, GPT-4, Gemini 2.0 Flash, Mistral 7B, and LLaMA 3, is also explored, thereby showcasing their impact through data scalability and architectural advancements. Additionally, this survey presents a comprehensive analysis of key ConvQA datasets and concludes by outlining open research directions. Overall, this work offers a comprehensive overview of the ConvQA landscape and provides valuable insights to guide future advancements in the field.
Reference graph
Works this paper leans on
-
[1]
Frontiers in Computer Science6, 1486581 (2025)
Orynbay, L., Bekmanova, G., Yergesh, B., Omarbekova, A., Sairanbekova, A., Sharipbay, A.: The Role of Cognitive Computing in NLP. Frontiers in Computer Science6, 1486581 (2025)
2025
-
[2]
Language Resources and Evaluation, 1–28 (2024)
Frenda, S., Abercrombie, G., Basile, V., Pedrani, A., Panizzon, R., Cignarella, A.T., Marco, C., Bernardi, D.: Perspectivist Approaches To Natural Language Processing: A Survey. Language Resources and Evaluation, 1–28 (2024)
2024
-
[3]
In: Proceedings of the 3rd International Confer- ence on Intelligent Communication Technologies and Virtual Mobile Networks, ICICV 2021, pp
Nagarhalli, T.P., Vaze, V., Rana, N.K.: Impact of Machine Learning in Natural Language Processing: A Review. In: Proceedings of the 3rd International Confer- ence on Intelligent Communication Technologies and Virtual Mobile Networks, ICICV 2021, pp. 1529–1534. Institute of Electrical and Electronics Engineers Inc., Online (2021) 32
2021
-
[4]
The Routledge Handbook of Translation and Technology, 419–436 (2019)
Melby, A.K.: Future of Machine Translation: Musings on Weaver’s Memo. The Routledge Handbook of Translation and Technology, 419–436 (2019)
2019
-
[5]
In: Proceedings of 2020 IEEE 3rd Interna- tional Conference of Safe Production and Informatization, IICSPI 2020, pp
Jiang, K., Lu, X.: Natural Language Processing and Its Applications in Machine Translation: A Diachronic Review. In: Proceedings of 2020 IEEE 3rd Interna- tional Conference of Safe Production and Informatization, IICSPI 2020, pp. 210–214. Institute of Electrical and Electronics Engineers Inc., Online (2020)
2020
-
[6]
(eds.) Prolog: Past, Present, and Future, pp
Gupta, G., Salazar, E., Shakerin, F., Arias, J., Varanasi, S.C., Basu, K., Wang, H., Li, F., Erbatur, S., Padalkar, P., Rajasekharan, A., Zeng, Y., Carro, M.: In: Warren, D.S., Dahl, V., Eiter, T., Hermenegildo, M.V., Kowalski, R., Rossi, F. (eds.) Prolog: Past, Present, and Future, pp. 48–61. Springer, Cham (2023)
2023
-
[7]
Johnson, M.L.: SHRDLU, Procedures, Mini-world, pp. 113–122. Palgrave Macmillan UK, London (1988)
1988
-
[8]
Philosoph- ical Transactions of the Royal Society A381(2251), 20220049 (2023)
Wahlster, W.: Understanding computational dialogue understanding. Philosoph- ical Transactions of the Royal Society A381(2251), 20220049 (2023)
2023
-
[9]
Expert Systems with Applications209, 118221 (2022)
Shao, Z., Zhao, R., Yuan, S., Ding, M., Wang, Y.: Tracing the Evolution of AI in the Past Decade and Forecasting the Emerging Trends. Expert Systems with Applications209, 118221 (2022)
2022
-
[10]
Engineering6(3), 275–290 (2020)
Zhou, M., Duan, N., Liu, S., Shum, H.-Y.: Progress in Neural NLP: Modeling, Learning, and Reasoning. Engineering6(3), 275–290 (2020)
2020
-
[11]
ACM Transactions on Asian and Low-Resource Language Information Processing
Rai, A., Borah, S.: Tokenization and Stemming of Limbu Language. ACM Transactions on Asian and Low-Resource Language Information Processing. (2025)
2025
-
[12]
Natural Language Processing Journal10, 100132 (2025)
Azmi, N.S.A.B.N., Ptaszynski, M., Masui, F., Eronen, J., Nowakowski, K.: Token and Part-of-speech Fusion for Pretraining of Transformers with Application in Automatic Cyberbullying Detection. Natural Language Processing Journal10, 100132 (2025)
2025
-
[13]
SN Computer Science6(2), 106 (2025)
Bharathi Mohan, G., Prasanna Kumar, R., Krishna Jayanth, K., Doss, S.: Telugu Language Analysis with XLM-RoBERTa: Enhancing Parts of Speech Tagging for Effective Natural Language Processing. SN Computer Science6(2), 106 (2025)
2025
-
[14]
IEEE Access13, 30444–30468 (2025)
Albladi, A., Islam, M., Seals, C.: Sentiment Analysis of Twitter Data Using NLP Models: A Comprehensive Review. IEEE Access13, 30444–30468 (2025)
2025
-
[15]
SN Computer Science6(1), 1–18 (2025)
Chabukswar, A., Shenoy, P.D., Venugopal, K.: Detecting Misinformation in COVID-19 Content: A Machine Learning and Deep Learning Approach with Word Embeddings. SN Computer Science6(1), 1–18 (2025)
2025
-
[16]
Human Trans- lation: A Comparative Study of Translation Quality
Haseeb, M., Akbar, M., Abbasi, W.S.: Machine Translation vs. Human Trans- lation: A Comparative Study of Translation Quality. Social Science Review Archives3(1), 885–894 (2025)
2025
-
[17]
Artificial Intelligence Review58(5), 127 (2025)
Jolfaei, S.A., Mohebi, A.: A Review on Persian Question Answering Systems: From Traditional to Modern Approaches. Artificial Intelligence Review58(5), 127 (2025)
2025
-
[18]
Rayhan, A., Kinzler, R., Rayhan, R.: Natural Language Processing: Transform- ing How Machines Understand Human Language (2023). https://doi.org/10. 13140/RG.2.2.34900.99200
-
[19]
International Journal of Computational Intelligence Systems18(1), 21 (2025)
Rana, M.R.R., Nawaz, A., Rehman, S.U., Abid, M.A., Garayevi, M., Kajanov´ a, 33 J.: BERT-BiGRU-Senti-GCN: An Advanced NLP Framework for Analyzing Customer Sentiments in e-Commerce. International Journal of Computational Intelligence Systems18(1), 21 (2025)
2025
-
[20]
RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question Answering
Bai, Y., Grant, C.E., Wang, D.Z.: RAMQA: A Unified Framework for Retrieval- augmented Multi-modal Question Answering (2025). https://arxiv.org/abs/ 2501.13297
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[21]
Knowledge and Information Systems65(4), 1399–1485 (2023)
Abdel-Nabi, H., Awajan, A., Ali, M.Z.: Deep Learning-based Question Answer- ing: A Survey. Knowledge and Information Systems65(4), 1399–1485 (2023)
2023
-
[22]
SDNet: Contextualized Attention-based Deep Network for Conversational Question Answering
Zhu, C., Zeng, M., Huang, X.: SDNet: Contextualized Attention-based Deep Network for Conversational Question Answering (2019). https://arxiv.org/abs/ 1812.03593
work page internal anchor Pith review Pith/arXiv arXiv 2019
-
[23]
Knowledge and Information Systems64(12), 3151–3195 (2022)
Zaib, M., Zhang, W., Sheng, Q., Mahmood, A., Zhang, Y.: Conversational question answering. Knowledge and Information Systems64(12), 3151–3195 (2022)
2022
-
[24]
In: Proceedings of the 17th ACM International Conference on Web Search and Data Mining
Abbasiantaeb, Z., Yuan, Y., Kanoulas, E., Aliannejadi, M.: Let the LLMs Talk: Simulating Human-to-human Conversational QA via Zero-Shot LLM-to-LLM Interactions. In: Proceedings of the 17th ACM International Conference on Web Search and Data Mining. WSDM ’24, pp. 8–17. Association for Computing Machinery, Merida, Mexico (2024)
2024
-
[25]
In: The 41st International ACM SIGIR Conference on Research & Development in Informa- tion Retrieval
Gao, J., Galley, M., Li, L.: Neural Approaches to Conversational AI. In: The 41st International ACM SIGIR Conference on Research & Development in Informa- tion Retrieval. SIGIR ’18, pp. 1371–1374. Association for Computing Machinery, Melbourne, Australia (2018)
2018
-
[26]
In: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval
Qu, C., Yang, L., Qiu, M., Croft, W.B., Zhang, Y., Iyyer, M.: Bert with his- tory answer embedding for conversational question answering. In: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval. SIGIR’19, pp. 1133–1136. Association for Computing Machinery, Paris, France (2019)
2019
-
[27]
ACM Comput
Biancofiore, G.M., Deldjoo, Y., Noia, T.D., Di Sciascio, E., Narducci, F.: Inter- active Question Answering Systems: Literature Review. ACM Comput. Surv. 56(9) (2024)
2024
-
[28]
IEEE Computational Intelligence Magazine13, 55–75 (2018)
Young, T., Hazarika, D., Poria, S., Cambria, E.: Recent trends in deep learn- ing based natural language processing [review article]. IEEE Computational Intelligence Magazine13, 55–75 (2018)
2018
-
[29]
Springer, Singapore (2021)
Zaib, M., Tran, D.H., Sagar, S., Mahmood, A., Zhang, W.E., Sheng, Q.Z.: BERT-CoQAC: BERT-based Conversational Question Answering in Context. Springer, Singapore (2021)
2021
-
[30]
In: Barzilay, R., Kan, M.-Y
Tian, Z., Yan, R., Mou, L., Song, Y., Feng, Y., Zhao, D.: How to Make Con- text More Useful? An Empirical Study on Context-Aware Neural Conversational Models. In: Barzilay, R., Kan, M.-Y. (eds.) Proceedings of the 55th Annual Meet- ing of the Association for Computational Linguistics, pp. 231–236. Association for Computational Linguistics, Vancouver, Cana...
2017
-
[31]
In: European Conference on Information Retrieval, pp
Raposo, G., Ribeiro, R., Martins, B., Coheur, L.: Question Rewriting? Assess- ing Its Importance For Conversational Question Answering. In: European Conference on Information Retrieval, pp. 199–206 (2022). Springer
2022
-
[32]
https://arxiv.org/ abs/2401.10225
Liu, Z., Ping, W., Roy, R., Xu, P., Lee, C., Shoeybi, M., Catanzaro, B.: ChatQA: 34 Surpassing GPT-4 on Conversational QA and RAG (2024). https://arxiv.org/ abs/2401.10225
Pith/arXiv arXiv 2024
-
[33]
In: Proceedings of the 14th ACM International Conference on Web Search and Data Mining, pp
Vakulenko, S., Longpre, S., Tu, Z., Anantha, R.: Question rewriting for con- versational question answering. In: Proceedings of the 14th ACM International Conference on Web Search and Data Mining, pp. 355–363 (2021)
2021
-
[34]
In: Proceedings of the First Workshop on NLP for Conversational AI
Ohsugi, Y., Saito, I., Nishida, K., Asano, H., Tomita, J.: A Simple But Effective Method To Incorporate Multi-turn Context With BERT For Conversational Machine comprehension. In: Proceedings of the First Workshop on NLP for Conversational AI. Association for Computational Linguistics, Florence, Italy (2019)
2019
-
[35]
In: Merlo, P., Tiedemann, J., Tsarfaty, R
Kacupaj, E., Plepi, J., Singh, K., Thakkar, H., Lehmann, J., Maleshkova, M.: Conversational Question Answering Over Knowledge Graphs with Transformer and Graph Attention Networks. In: Merlo, P., Tiedemann, J., Tsarfaty, R. (eds.) Proceedings of the 16th Conference of the European Chapter of the Associa- tion for Computational Linguistics, pp. 850–862. Ass...
2021
-
[36]
In: Vlachos, A., Augenstein, I
Perez-Beltrachini, L., Jain, P., Monti, E., Lapata, M.: Semantic Parsing for Conversational Question Answering over Knowledge Graphs. In: Vlachos, A., Augenstein, I. (eds.) Proceedings of the 17th Conference of the European Chap- ter of the Association for Computational Linguistics, pp. 2507–2522. Association for Computational Linguistics, Dubrovnik, Croa...
2023
-
[37]
In: Pro- ceedings of the 17th ACM International Conference on Web Search and Data Mining
Kaiser, M., Saha Roy, R., Weikum, G.: Robust Training for Conversational Question Answering Models with Reinforced Reformulation Generation. In: Pro- ceedings of the 17th ACM International Conference on Web Search and Data Mining. WSDM ’24, pp. 322–331. Association for Computing Machinery, Merida, Mexico (2024)
2024
-
[38]
Few-shot Policy (de)composition in Conversational Question Answering
Erwin, K., Axelrod, G., Chang, M., Fokoue, A., Crouse, M., Dan, S., Gao, T., Uceda-Sosa, R., Makondo, N., Khan, N., et al.: Few-shot Policy (de) composition in Conversational Question Answering. arXiv preprint arXiv:2501.11335 (2025)
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[39]
(eds.) Proceedings of the Third Work- shop on Insights from Negative Results in NLP, pp
Ishii, E., Xu, Y., Cahyawijaya, S., Wilie, B.: Can Question Rewriting Help Conversational Question Answering? In: Tafreshi, S., Sedoc, J., Rogers, A., Drozd, A., Rumshisky, A., Akula, A. (eds.) Proceedings of the Third Work- shop on Insights from Negative Results in NLP, pp. 94–99. Association for Computational Linguistics, Dublin, Ireland (2022)
2022
-
[40]
In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval
Ye, L., Lei, Z., Yin, J., Chen, Q., Zhou, J., He, L.: Boosting Conversational Ques- tion Answering with Fine-grained Retrieval-augmentation and Self-check. In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. SIGIR ’24, pp. 2301–2305. Association for Computing Machinery, Washington DC, USA (2024)
2024
-
[41]
In: Advanced Data Mining and Applications, pp
Perera, M.M., Mahmood, A., Wijethilake, K.E., Sheng, Q.Z.: Towards Adap- tive Context Management for Intelligent Conversational Question Answering. In: Advanced Data Mining and Applications, pp. 360–375. Springer, Singapore (2025)
2025
-
[42]
In: Zong, C., Xia, F., Li, W., Navigli, R
Kim, G., Kim, H., Park, J., Kang, J.: Learn to Resolve Conversational Dependency: A Consistency Training Framework for Conversational Question Answering. In: Zong, C., Xia, F., Li, W., Navigli, R. (eds.) Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language P...
2021
-
[43]
Preference-based Learning with Retrieval Augmented Generation for Conversational Question Answering
Kaiser, M., Weikum, G.: Preference-based Learning with Retrieval Aug- mented Generation for Conversational Question Answering. arXiv preprint arXiv:2503.22303 (2025)
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[44]
In: 2023 International Joint Conference on Neural Net- works (IJCNN)
Zaib, M., Sheng, Q., Zhang, W., Mahmood, A.: Keeping the questions conversa- tional: Using structured representations to resolve dependency in conversational question answering. In: 2023 International Joint Conference on Neural Net- works (IJCNN). Institute of Electrical and Electronics Engineers (IEEE), Online (2023). 2023 International Joint Conference ...
2023
-
[45]
Pro- ceedings of the AAAI Conference on Artificial Intelligence35(15), 13718–13726 (2021)
Qiu, M., Huang, X., Chen, C., Ji, F., Qu, C., Wei, W., Huang, J., Zhang, Y.: Reinforced History Backtracking for Conversational Question Answering. Pro- ceedings of the AAAI Conference on Artificial Intelligence35(15), 13718–13726 (2021)
2021
-
[46]
In: Lecture Notes in Computer Science, vol
Zaib, M., Zhang, W.E., Sheng, Q.Z., Sagar, S., Mahmood, A., Zhang, Y.: Learn- ing to Select the Relevant History Turns in Conversational Question Answering. In: Lecture Notes in Computer Science, vol. 14306 LNCS, pp. 334–348. Springer, Berlin, Heidelberg (2023)
2023
-
[47]
In: International Conference on Web Information Systems Engineering, pp
Zaib, M., Sheng, Q.Z., Zhang, W.E., Alhazmi, E., Mahmood, A.: Learning contrastive representations for dense passage retrieval in open-domain conver- sational question answering. In: International Conference on Web Information Systems Engineering, pp. 3–13 (2024). Springer
2024
-
[48]
https://arxiv.org/abs/2407.21712
Wang, X., Sen, P., Li, R., Yilmaz, E.: Adaptive Retrieval-augmented Generation for Conversational Systems (2024). https://arxiv.org/abs/2407.21712
Pith/arXiv arXiv 2024
-
[49]
SIGIR ’24, pp
Kostric, I., Balog, K.: A Surprisingly Simple yet Effective Multi-query Rewrit- ing Method for Conversational Passage Retrieval. SIGIR ’24, pp. 2271–2275. Association for Computing Machinery, New York, NY, USA (2024)
2024
-
[50]
In: 2024 IEEE 18th International Conference on Semantic Computing (ICSC), pp
Rashid, M.S., Meem, J.A., Hristidis, V.: NORMY: Non-Uniform History Model- ing for Open Retrieval Conversational Question Answering. In: 2024 IEEE 18th International Conference on Semantic Computing (ICSC), pp. 101–109 (2024). IEEE
2024
-
[51]
In: Ku, L.-W., Martins, A., Srikumar, V
Liu, L., Hill, B., Du, B., Wang, F., Tong, H.: Conversational Question Answering with Language Models Generated Reformulations over Knowledge Graph. In: Ku, L.-W., Martins, A., Srikumar, V. (eds.) Findings of the Association for Com- putational Linguistics ACL 2024, pp. 839–850. Association for Computational Linguistics, Bangkok, Thailand (2024)
2024
-
[52]
Neural Computing and Applications36(16), 8995–9022 (2024)
Hu, Z., Hou, W., Liu, X.: Deep Learning For Named Entity Recognition: A Survey. Neural Computing and Applications36(16), 8995–9022 (2024)
2024
-
[53]
IEEE Transactions on Emerging Topics in Computational Intelligence8(3), 2640–2653 (2024)
Liu, Z., He, J., Gong, T., Weng, H., Wang, F.L., Liu, H., Hao, T.: Improv- ing Topic Tracing with a Textual Reader for Conversational Knowledge Based Question Answering. IEEE Transactions on Emerging Topics in Computational Intelligence8(3), 2640–2653 (2024)
2024
-
[54]
Symmetry16(9) (2024)
Jiang, P., Cai, X.: A Survey of Semantic Parsing Techniques. Symmetry16(9) (2024)
2024
-
[55]
Journal of Computer and Communications12(10), 1–13 (2024) 36
Wu, Z.: Large Language Model Based Semantic Parsing for Intelligent Database Query Engine. Journal of Computer and Communications12(10), 1–13 (2024) 36
2024
-
[56]
Schneider, P., Klettner, M., Jokinen, K., Simperl, E., Matthes, F.: Evaluat- ing Large Language Models in Semantic Parsing for Conversational Question Answering over Knowledge Graphs (2024). https://arxiv.org/abs/2401.01711
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[57]
In: Proceedings of the 10th SIGHAN Workshop on Chinese Language Processing (SIGHAN- 10), pp
Bai, Z., Wang, B., Liang, B., Xu, R.: Auto-ACE: An Automatic Answer Correct- ness Evaluation Method for Conversational Question Answering. In: Proceedings of the 10th SIGHAN Workshop on Chinese Language Processing (SIGHAN- 10), pp. 80–87. Association for Computational Linguistics, Bangkok, Thailand (2024)
2024
-
[58]
IEEE Transactions on Computational Social Systems11(2), 1888–1906 (2024)
Ahmed, M., Khan, H.U., Munir, E.U.: Conversational ai: An explication of few-shot learning problem in transformers-based chatbot systems. IEEE Transactions on Computational Social Systems11(2), 1888–1906 (2024)
1906
-
[59]
https://arxiv.org/abs/2405.17822
Pan, Z., Luo, H., Li, M., Liu, H.: Conv-CoA: Improving Open-domain Question Answering in Large Language Models via Conversational Chain-of-action (2024). https://arxiv.org/abs/2405.17822
Pith/arXiv arXiv 2024
-
[60]
Soft Computing25(11), 7341–7378 (2021)
Ghahramani, F., Tahayori, H., Visconti, A.: Effects of central tendency mea- sures on term weighting in textual information retrieval. Soft Computing25(11), 7341–7378 (2021)
work page 2021
-
[61]
In: Rogers, A., Boyd-Graber, J., Okazaki, N
Jeong, S., Baek, J., Hwang, S.J., Park, J.: Phrase Retrieval for Open Domain Conversational Question Answering with Conversational Dependency Modeling via Contrastive Learning. In: Rogers, A., Boyd-Graber, J., Okazaki, N. (eds.) Findings of the Association for Computational Linguistics: ACL 2023, pp. 6019–
work page 2023
-
[62]
Xia, F., Li, B., Weng, Y., He, S., Liu, K., Sun, B., Li, S., Zhao, J.: MedConQA: Medical conversational question answering system based on knowledge graphs. In: Che, W., Shutova, E. (eds.) Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 148–158. Association for Computational Linguistics, A...
work page 2022
-
[63]
Transactions of the Association for Computational Linguistics11, 1–17 (2023)
Siriwardhana, S., Weerasekera, R., Wen, E., Kaluarachchi, T., Rana, R., Nanayakkara, S.: Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for Open Domain Question Answering. Transactions of the Association for Computational Linguistics11, 1–17 (2023)
work page 2023
-
[64]
https://arxiv.org/abs/2501.12789
Filice, S., Horowitz, G., Carmel, D., Karnin, Z., Lewin-Eytan, L., Maarek, Y.: Generating Diverse Q&A Benchmarks for RAG Evaluation with DataMorgana (2025). https://arxiv.org/abs/2501.12789
Pith/arXiv arXiv 2025
-
[65]
In: Proceedings of the 31st ACM International Conference on Information & Knowledge Management
Kacupaj, E., Singh, K., Maleshkova, M., Lehmann, J.: Contrastive Representa- tion Learning for Conversational Question Answering over Knowledge Graphs. In: Proceedings of the 31st ACM International Conference on Information & Knowledge Management. CIKM ’22, pp. 925–934. Association for Computing Machinery, Atlanta, GA, USA (2022)
work page 2022
-
[66]
Computers & Indus- trial Engineering, 110856 (2025)
Khadivi, M., Charter, T., Yaghoubi, M., Jalayer, M., Ahang, M., Shojaeinasab, A., Najjaran, H.: Deep Reinforcement Learning For Machine Scheduling: Methodology, The State-of-the-art, and Future Directions. Computers & Indus- trial Engineering, 110856 (2025)
work page 2025
-
[67]
IEEE Transactions on Neural Networks and Learning Systems, 1–21 (2024) 37
Cao, Y., Zhao, H., Cheng, Y., Shu, T., Chen, Y., Liu, G., Liang, G., Zhao, J., Yan, J., Li, Y.: Survey on Large Language Model-enhanced Reinforcement Learning: Concept, Taxonomy, and Methods. IEEE Transactions on Neural Networks and Learning Systems, 1–21 (2024) 37
work page 2024
-
[68]
https://arxiv.org/abs/2405.11106
Sun, C., Huang, S., Pompili, D.: LLM-based Multi-agent Reinforcement Learn- ing: Current and Future Directions (2024). https://arxiv.org/abs/2405.11106
Pith/arXiv arXiv 2024
-
[69]
Chen, Z., Zhao, J., Fang, A., Fetahu, B., Rokhlenko, O., Malmasi, S.: Rein- forced Question Rewriting for Conversational Question Answering. In: Li, Y., Lazaridou, A. (eds.) Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 357–370. Association for Computational Linguistics, Abu Dhabi, UAE (2022)
work page 2022
-
[70]
Reinforcement Learning for Conversational Question Answering over Knowledge Graph
Wu, M.: Reinforcement Learning for Conversational Question Answering over Knowledge Graph (2024). https://arxiv.org/abs/2401.08460
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[71]
Pattern Recognition159, 111095 (2025)
Sun, T., Chen, H., Hu, G., Zhao, C.: Explainability-based knowledge distillation. Pattern Recognition159, 111095 (2025)
work page 2025
-
[72]
DDK: Distilling Domain Knowledge for Efficient Large Language Models
Liu, J., Zhang, C., Guo, J., Zhang, Y., Que, H., Deng, K., Bai, Z., Liu, J., Zhang, G., Wang, J., Wu, Y., Liu, C., Su, W., Wang, J., Qu, L., Zheng, B.: DDK: Distilling Domain Knowledge for Efficient Large Language Models (2024). https://arxiv.org/abs/2407.16154
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[73]
In: Carpuat, M., Marneffe, M.-C., Meza Ruiz, I.V
You, C., Chen, N., Liu, F., Ge, S., Wu, X., Zou, Y.: End-to-end spoken conversa- tional question answering: Task, dataset and model. In: Carpuat, M., Marneffe, M.-C., Meza Ruiz, I.V. (eds.) Findings of the Association for Computational Linguistics: NAACL 2022. Association for Computational Linguistics, Seattle, United States (2022)
work page 2022
-
[74]
https://arxiv.org/abs/1909.10772
Ju, Y., Zhao, F., Chen, S., Zheng, B., Yang, X., Liu, Y.: Technical Report on Conversational Question Answering (2019). https://arxiv.org/abs/1909.10772
Pith/arXiv arXiv 2019
-
[75]
Engineering Applications of Artificial Intelligence123, 106137 (2023)
Boreshban, Y., Mirbostani, S.M., Ghassem-Sani, G., Mirroshandel, S.A., Amiri- parian, S.: Improving question answering performance using knowledge distilla- tion and active learning. Engineering Applications of Artificial Intelligence123, 106137 (2023)
work page 2023
-
[76]
Computer Modeling in Engineering & Sciences
Yin, L., Wang, L., Cai, Z., Lu, S., Wang, R., AlSanad, A., AlQahtani, S.A., Chen, X., Yin, Z., Li, X., et al.: Dpal-bert: A faster and lighter question answering model. Computer Modeling in Engineering & Sciences
-
[77]
In: Carpuat, M., Marneffe, M.-C., Meza Ruiz, I.V
Caciularu, A., Dagan, I., Goldberger, J., Cohan, A.: Long context question answering via supervised contrastive learning. In: Carpuat, M., Marneffe, M.-C., Meza Ruiz, I.V. (eds.) Proceedings of the 2022 Conference of the North Ameri- can Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2872–2879. Association for C...
work page 2022
-
[78]
Computerized Medical Imaging and Graphics121, 102500 (2025)
Xu, X., Wong, S.T.C.: Contrastive Learning in Brain Imaging. Computerized Medical Imaging and Graphics121, 102500 (2025)
work page 2025
-
[79]
In: Proceedings of the AAAI Conference on Artificial Intelligence, vol
Gao, X., Das, K.: Customizing Language Model Responses with Contrastive In-Context Learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, pp. 18039–18046 (2024)
work page 2024
-
[80]
In: Goldberg, Y., Kozareva, Z., Zhang, Y
Zhang, Z., Strubell, E., Hovy, E.: A Survey of Active Learning for Natural Lan- guage Processing. In: Goldberg, Y., Kozareva, Z., Zhang, Y. (eds.) Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 6166–6190. Association for Computational Linguistics, Online (2022)
work page 2022
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.