REVIEW 4 major objections 5 minor 59 references
Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Chatty-Gen is a fully automated RAG platform that produces dialogue benchmarks from knowledge graphs, cutting DBpedia generation from about 30 hours to roughly 10 minutes while matching human-built quality.
desk verdict A genuinely useful system for KG-based dialogue benchmark generation, but the paper overclaims quality and the evaluation contradicts that claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the multi-stage pipeline with assertion-based validation. Each stage is given a simple zero-shot prompt rather than one complex prompt; after each stage, a validator checks the output against explicit conditions: questions must explicitly name the entity, triples must belong to the provided subgraph, SPARQL queries must be syntactically correct and consistent with the subgraph, and the dialogue must start with an independent question and use pronouns appropriately in later questions. Invalid outputs are retried up to three times before the seed entity is discarded. A second mechanism is subgraph summarization, which collapses repeated predicates to a single modified triple with the object removed, shrinking the prompt while preserving the facts needed for question and query generation.
What would settle it
Take any dialogue that passed Chatty-Gen's validation, run its SPARQL queries directly against the full KG endpoint, and check each returned answer against the source triples in the original entity subgraph; if a nontrivial fraction of answers name entities or values that are not in the subgraph, or if the queries return empty results for claimed questions, the hallucination-mitigation claim is disproved.
Extended reading notes
Core claim
The central discovery is that decomposing KG-to-dialogue generation into small validated stages makes the task tractable for LLMs of widely varying capability, including open-source models, without sacrificing quality. The paper shows that a single-prompt approach fails for most models, while a pipeline of (1) subgraph summarization, (2) independent question generation, (3) SPARQL query generation, and (4) dialogue assembly, with assertion rules checked after each stage, raises success rates from near zero to over 90 percent for GPT-3.5 and to 100 percent for GPT-4o and a combination of open-source models. The platform also eliminates the expensive preprocessing of the whole KG by querying the SPARQL endpoint on demand to find seed entities and their subgraphs. The authors claim this makes Chatty-Gen the first fully automated RAG-based dialogue benchmark generator for KGs and that it significantly outperforms Maestro in both question quality and time.
Load-bearing premise
The approach assumes that the stage-level assertion rules—checking that triples appear in the subgraph, questions are self-contained, SPARQL syntax parses, and queries agree with the summarized subgraph—are strong enough to guarantee the generated dialogues are factually correct against the full knowledge graph.
Editorial extensions
If this is right
- Benchmark creation stops being a manual bottleneck: domain-specific dialogue benchmarks can be generated on demand for any SPARQL-enabled KG without writing templates or labeling data.
- Open-source LLMs can replace commercial APIs for this task, so high-quality benchmarks no longer require expensive proprietary models.
- The stage-wise checkpoints catch hallucinations early, meaning the system rarely has to restart from scratch on a large KG.
- The same four-stage recipe can be reused for other KG-to-text tasks such as QA-pair generation or entity summarization, since the assertion rules are KG-agnostic.
Reading between the lines
- Because correctness is defined by the assertion rules against the summarized subgraph, not by independent verification against the full KG, a released Chatty-Gen benchmark would be stronger if it also published the supporting triples and the executed SPARQL answers for external checks.
- The entity-skipping policy trades coverage for cost; a natural extension the paper leaves open is a correction module that repairs bad questions or queries instead of abandoning an entity, which would matter for users who need specific head or tail entities.
- Sampling proportionally to node-type distribution means benchmark content follows the KG's own biases; a testable variant would compare uniform sampling against distribution-based sampling to see which yields more useful evaluation coverage.
- The quality comparison to ConvQuestions rested on an LLM judge; a human-annotator replication would test whether the parity claim survives outside automated evaluation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Chatty-Gen, a multi-stage retrieval-augmented generation platform that automatically produces dialogue benchmarks from a knowledge graph. The pipeline selects representative node types and seed entities, predicts human-readable entity labels, extracts subgraphs as dialogue context, and then generates independent questions, SPARQL queries, and a final coherent dialogue in separate stages, with assertion-based validation between stages. The system is evaluated on DBpedia, YAGO, DBLP, and MAG using several commercial and open-source LLMs. The authors report that Chatty-Gen reduces benchmark generation time dramatically compared with Maestro and produces dialogues whose quality is comparable to the human-generated ConvQuestions benchmark, while claiming in the abstract and contributions that Chatty-Gen 'significantly outperforms state-of-the-art systems in both quality and time performance.'
Significance. If fully substantiated, the paper would describe a useful, cost-effective, and largely automatic way to construct KG-grounded dialogue benchmarks, and it would be the first fully automated RAG-based platform for this task. Strengths of the work include public code, a pipeline that decomposes generation into independently validatable stages, experiments across multiple real KGs and many LLMs, and a subgraph summarization method that reduces token consumption while improving SPARQL correctness in the reported experiments. However, the current evidence does not support the headline quality-superiority claim: the LLM-judge comparison in Section 7 shows mostly ties and a preference for the human benchmark when forced, no human evaluation is reported, and the sample sizes are too small for statistical significance. The time comparison also rests on a partially completed Maestro run for DBpedia.
major comments (4)
- [Abstract and Section 7, Table 5] The claim that Chatty-Gen 'significantly outperforms state-of-the-art systems in both quality and time performance' is not supported by the paper's own quality evaluation. In Table 5, the Gemini judge marks 70% (GPT-4o) and 80% (Multi-LLM-1) of dialogue pairs as ties; when a preference is forced, ConvQuestions is favored in 20-25% of cases while Chatty-Gen is favored in only 0-5%. This is evidence of rough equivalence at best, and of inferiority on forced choices, not significant superiority. Moreover, ConvQuestions is a human-generated benchmark rather than a system-generated dialogue benchmark, so the comparison does not directly support the stated claim. I recommend reframing the abstract and contributions to claim comparable quality at lower cost, and supplementing the LLM-judge result with a human evaluation or at least significance testing on repeated runs.
- [Section 6.4, Table 4] All headline comparisons across LLMs are based on only 20 generated dialogues per configuration, with no repeated trials, confidence intervals, or significance tests. The reported success rates vary widely across models (e.g., 22% for Gemini-1.5-pro, 41% for LLAMA-3-8b-inst, 100% for GPT-4o on YAGO), so the claim of 'consistent model and system performance across multiple LLMs' is not established. The reader cannot tell whether the differences between the multi-stage and single-prompt approaches, or between different LLMs, are real or within run-to-run noise. Please report variance over repeated runs or otherwise quantify the stability of the success-rate metric.
- [Section 6.3, Table 3, and Section 6.2] The DBpedia time comparison appears to count an incomplete Maestro run. The text states that for DBpedia 'the process was halted at 18,211 out of 994,592 predicates due to memory limitations,' yet Table 3 reports a DBpedia Maestro time of 30.77 hours. If this number is projected or partial, the 99% time improvement claim is not a fair end-to-end comparison. Please clarify whether the Maestro entry reflects a completed run, a projection, or a partial run, and otherwise report a like-for-like comparison on a KG where Maestro can complete.
- [Section 5.3 and Algorithm 3] The assertion-based validation is weaker than the paper's hallucination-mitigation claims. Algorithm 3 summarizes triples by removing object values, yielding modified triples of the form ⟨e, p, None⟩ or ⟨None, p, e⟩. The query validator checks syntactic validity and consistency with the summarized subgraph information, but it cannot check whether the SPARQL query answers the natural-language question: a question about a birth date paired with a query over birthPlace could pass if both predicates occur in the subgraph. Thus the validation does not guarantee factual correctness of the generated dialogue or semantic equivalence between question and query. Please either add a stronger validation step (e.g., execute the query against the full KG and compare answers with the original triple objects) or soften the correctness claims accordingly.
minor comments (5)
- [Section 1] The introduction says evaluation uses 'four diverse real-world KGs: DBpedia, Yago, and DBLP,' but lists only three; MAG is described later. Please correct the enumeration.
- [Table 4] The column headers 'Dialogue-E' and 'Parsing-E' are not defined in the table or text; please define these error categories and clarify whether they are counts per seed entity or per generated dialogue.
- [Section 5.3] When an output fails validation, the retry mechanism sends 'the same prompt and inputs' without providing feedback about the validation failure; this may cause repeated failures and could partly explain the low success rates of some open-source models. Consider reporting whether retries use corrective feedback.
- [Section 7] The use-case comparison reports that Chatty-Gen produces 'dialogues of comparable quality in about 15 minutes at a cost of just $0.27 USD,' but the time and cost numbers are not clearly tied to a reproducible configuration (number of entities, LLM, retries). Please specify the configuration used for this statement.
- [Section 8] There is a typo in 'we propsoed a multi-stage approach' (should be 'proposed'). A careful proofread would also catch other minor errors, e.g., 'an entity' vs. 'a entity' in Section 4.3.
Circularity Check
No significant circularity: the pipeline's outputs are checked against external KG facts and external benchmarks, with only a non-load-bearing self-citation.
full rationale
Chatty-Gen's claimed derivation chain is not circular. The generation pipeline consumes KG subgraphs retrieved by SPARQL as external context; the summarized subgraph (Algorithm 3) is a deterministic projection of KG triples, not an LLM output. SPARQL queries are generated from questions plus triples and then executed against the KG, so answers (the benchmark ground truth) come from an external source. Assertion rules in Section 5.3 are self-consistency checks (membership of triples in the provided subgraph, syntactic validity of SPARQL, dialogue-form constraints); they are engineering gates, not fitted parameters renamed as predictions. The self-citation [26] (KGQAn) appears only in a limitations sentence about existing QAS lacking dialogue support and is not load-bearing. The quality comparison against ConvQuestions via Gemini 1.5 is weak evidence, and Table 5 actually shows mostly ties, contradicting the abstract's 'significantly outperforms' wording; however, that is an evidential/correctness problem, not a definitional equivalence. Time comparisons against Maestro are external. No equation in the paper reduces a claimed result to its own inputs.
Assumptions & free parameters
free parameters (7)
- Rare types threshold R =
1% of KG entities
- Shadowed parent threshold S =
99%
- Entity batch size BZ =
10,000
- Hop count h =
1
- Predicate direction =
both outgoing and incoming
- Questions per dialogue nq =
5 in evaluation
- Retry limit =
3
assumptions (5)
- domain assumption KG node-type distribution is a reliable proxy for domain topics and benchmark coverage.
- domain assumption An LLM can select a correct entity-label predicate from the list of string-literal predicates under zero-shot prompting.
- domain assumption A subgraph with enough unique predicates is a sufficient context for generating a meaningful dialogue.
- domain assumption Assertion-based validators catch hallucinations and guarantee output correctness.
- domain assumption SPARQL endpoints with built-in indices support efficient count and subgraph queries on large KGs.
Cite this review
Pith. "Pith review of Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs." pith.science (2026). https://pith.science/paper/ZH4W2F56
@misc{pith2026250109928,
author = {Pith},
title = {Pith review of: Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZH4W2F56}},
note = {Machine review of arXiv:2501.09928}
}
read the original abstract
Dialogue benchmarks are crucial in training and evaluating chatbots engaging in domain-specific conversations. Knowledge graphs (KGs) represent semantically rich and well-organized data spanning various domains, such as DBLP, DBpedia, and YAGO. Traditionally, dialogue benchmarks have been manually created from documents, neglecting the potential of KGs in automating this process. Some question-answering benchmarks are automatically generated using extensive preprocessing from KGs, but they do not support dialogue generation. This paper introduces Chatty-Gen, a novel multi-stage retrieval-augmented generation platform for automatically generating high-quality dialogue benchmarks tailored to a specific domain using a KG. Chatty-Gen decomposes the generation process into manageable stages and uses assertion rules for automatic validation between stages. Our approach enables control over intermediate results to prevent time-consuming restarts due to hallucinations. It also reduces reliance on costly and more powerful commercial LLMs. Chatty-Gen eliminates upfront processing of the entire KG using efficient query-based retrieval to find representative subgraphs based on the dialogue context. Our experiments with several real and large KGs demonstrate that Chatty-Gen significantly outperforms state-of-the-art systems and ensures consistent model and system performance across multiple LLMs of diverse capabilities, such as GPT-4o, Gemini 1.5, Llama 3, and Mistral.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
AI@Meta. 2024. Llama 3 Model Card. (2024). https://github.com/meta-llama/ llama3/blob/main/MODEL_CARD.md
2024
-
[2]
Debayan Banerjee, Sushil Awale, Ricardo Usbeck, and Chris Biemann. 2023. DBLP- QuAD: A Question Answering Dataset over the DBLP Scholarly Knowledge Graph. In Proceedings of the International Workshop on Bibliometric-enhanced Information Retrieval, Vol. 3617. 37–51. https://ceur-ws.org/Vol-3617/paper- 05.pdf
work page 2023
-
[3]
Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, and et al
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, and et al. 2020. Language Models are Few-Shot Learners. In Advances in Neu- ral Information Processing Systems: Annual Conference on Neural Information Processing Systems (NeurIPS). https://proceedings.neurips.cc/paper/2020/hash/ 1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html
work page 2020
-
[4]
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking Large Language Models in Retrieval-Augmented Generation. InProceedings of the AAAI Conference on Artificial Intelligence. 17754–17762. https://doi.org/10.1609/AAAI. V38I16.29728
doi:10.1609/aaai 2024
-
[5]
Chi, Xuezhi Wang, and Denny Zhou
Xinyun Chen, Ryan A. Chi, Xuezhi Wang, and Denny Zhou. 2024. Premise Order Matters in Reasoning with Large Language Models. In International Conference on Machine Learning, ICML . https://openreview.net/forum?id=4zAHgkiCQg
work page 2024
-
[6]
Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, and el al. 2018. QuAC: Question Answering in Context. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP) . 2174–2184. https: //doi.org/10.18653/V1/D18-1241
-
[7]
Philipp Christmann, Rishiraj Saha Roy, Abdalghani Abujabal, Jyotsna Singh, and Gerhard Weikum. 2019. Look before you Hop: Conversational Question Answer- ing over Knowledge Graphs Using Judicious Context Expansion. In Proceedings of the ACM International Conference on Information and Knowledge Management (CIKM). 729–738. https://doi.org/10.1145/3357384.3358016
arXiv 2019
-
[8]
Philipp Christmann, Rishiraj Saha Roy, and Gerhard Weikum. 2022. Conversa- tional Question Answering on Heterogeneous Sources. InSIGIR: The International ACM SIGIR Conference on Research and Development in Information Retrieval. 144–
work page 2022
Show all 59 references
- [9]
-
[10]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Lang...
2019 doi
-
[11]
Mohnish Dubey, Debayan Banerjee, Abdelrahman Abdelkawi, and Jens Lehmann
-
[12]
Bahare Fatemi, Jonathan Halcrow, and Bryan Perozzi. 2024. Talk like a Graph: Encoding Graphs for Large Language Models. In The International Conference on Learning Representations, ICLR. https://openreview.net/forum?id=IuXR1CCrSi
2024
- [13]
-
[14]
Google. 2023. https://gemini.google.com
2023
-
[15]
Xixin Hu, Yiheng Shu, Xiang Huang, and Yuzhong Qu. 2021. EDG-Based Question Decomposition for Complex Question Answering over Knowledge Bases. In Proceedings of the International Semantic Web Conference, (ISWC) , Vol. 12922. 128–145. https://doi.org/10.1007/978-3-030-88361-4_8
2021 doi
-
[16]
Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Deven- dra Singh Chaplot, and et al
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Deven- dra Singh Chaplot, and et al. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023). https://doi.org/10.48550/ARXIV.2310.06825
-
[17]
Yohan Jo, Xinyan Zhao, Arijit Biswas, Nikoletta Basiou, Vincent Auvray, and et al. 2023. Multi-User MultiWOZ: Task-Oriented Dialogues among Multiple Users. In Findings of the Association for Computational Linguistics: (EMNLP) . 3237–3269. https://doi.org/10.18653/V1/2023.FINDI...
2023 doi
-
[18]
Jaehun Jung, Bokyung Son, and Sungwon Lyu. 2020. AttnIO: Knowledge Graph Exploration with In-and-Out Attention Flow for Knowledge-Grounded Dialogue. In Proceedings of the Conference on Empirical Methods in Natural Language Process- ing (EMNLP). 3484–3497. https://doi.org/10.18...
2020 doi
-
[19]
Adam Tauman Kalai and Santosh S. Vempala. 2024. Calibrated Language Models Must Hallucinate. In Proceedings of the Annual ACM Symposium on Theory of Computing, STOC. 160–171. https://doi.org/10.1145/3618260.3649777
2024
-
[20]
Gray et al
Pavan Kapanipathi, Ibrahim Abdelaziz, Srinivas Ravishankar, Salim Roukos, and Alexander G. Gray et al. 2021. Leveraging Abstract Meaning Representation for Knowledge Base Question Answering. In Findings of the Association for Compu- tational Linguistics: (ACL/IJCNLP). 3884–389...
2021 doi
-
[21]
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large Language Models are Zero-Shot Reasoners. In Advances in Neural Information Processing Systems : Annual Conference on Neural Information Processing Systems (NeurIPS). http://papers.ni...
2022
-
[22]
Seungjun Lee, Yoonna Jang, Chanjun Park, Jungseob Lee, Jaehyung Seo, and et al
-
[23]
Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, and et al. 2015. DBpedia - A large-scale, multilingual knowledge base extracted from Wikipedia. Semantic Web 6 (2015), 167–195. https://doi.org/10.3233/SW-140134
2015 doi
-
[24]
Zekun Li, Zhiyu Chen, Mike Ross, Patrick Huber, Seungwhan Moon, and et al. 2024. Large Language Models as Zero-shot Dialogue State Tracker through Function Calling. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A...
2024 doi
-
[25]
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. In Advances in Neural In- formation Processing Systems: Annual Conference on Neural Informa...
2023
-
[26]
Reham Omar, Ishika Dhall, Panos Kalnis, and Essam Mansour. 2023. A Universal Question-Answering Platform for Knowledge Graphs. Proc. ACM Manag. Data 1, 1 (2023), 57:1–57:25. https://doi.org/10.1145/3588911
2023 doi
-
[27]
OpenAI. 2022. https://openai.com/blog/chatgpt
2022
- [28]
-
[29]
OpenAI. 2024. GPT-4o. (2024). https://openai.com/index/hello-gpt-4o/
2024
-
[30]
Abdelghny Orogat and Ahmed El-Roby. 2023. Maestro: Automatic Generation of Comprehensive Benchmarks for Question Answering Over Knowledge Graphs. Proceedings of the ACM on Management of Data 1, 2 (2023), 177:1–177:24. https: //doi.org/10.1145/3589322
2023 doi
-
[31]
Wain- wright, and et al
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wain- wright, and et al. 2022. Training language models to follow in- structions with human feedback. In Advances in Neural Information Processing Systems: Annual Conference on Neural Information Processing Systems, ...
2022
-
[32]
Aaron Pham, Chaoyu Yang, Sean Sheng, Shenyang Zhao, Sauyon Lee, Bo Jiang, Fog Dong, Xipeng Guan, and Frost Ming. 2023. OpenLLM: Operating LLMs in production. https://github.com/bentoml/OpenLLM
2023
-
[33]
Mohammadreza Pourreza and Davood Rafiei. 2023. DIN-SQL: Decomposed In- Context Learning of Text-to-SQL with Self-Correction. In Advances in Neural Information Processing Systems: Annual Conference on Neural Information Pro- cessing Systems (NeurIPS) . http://papers.nips.cc/pap...
2023
-
[34]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.Journal of Machine Learning Research 21 (2020), 140:1–140:67. h...
2020
-
[35]
Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019. CoQA: A Conversa- tional Question Answering Challenge. Transactions of the Association for Compu- tational Linguistics 7 (2019), 249–266. https://doi.org/10.1162/TACL_A_00266
2019 doi
- [36]
-
[37]
Khapra, Karthik Sankaranarayanan, and Sarath Chandar
Amrita Saha, Vardaan Pahuja, Mitesh M. Khapra, Karthik Sankaranarayanan, and Sarath Chandar. 2018. Complex Sequential Question Answering: Towards Learning to Converse Over Linked Question Answer Pairs with a Knowledge Graph. In Proceedings of AAAI Conference on Artificial Inte...
2018 doi
-
[38]
Kai Sun, Yifan Ethan Xu, Hanwen Zha, Yue Liu, and Xin Luna Dong. 2024. Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?. In Proceedings of the Conference of the North American Chapter of the Association for Computatio...
2024 doi
-
[39]
Kai Sun, Dian Yu, Jianshu Chen, Dong Yu, Yejin Choi, and et al. 2019. DREAM: A Challenge Dataset and Models for Dialogue-Based Reading Comprehension. Trans. Assoc. Comput. Linguistics 7 (2019), 217–231. https://doi.org/10.1162/ TACL_A_00264
2019
-
[40]
Chang-Yu Tai, Ziru Chen, Tianshu Zhang, Xiang Deng, and Huan Sun. 2023. Exploring Chain of Thought Style Prompting for Text-to-SQL. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, EMNLP . 5376–5393. https://doi.org/10.18653/V1/2023.EMNLP-MAIN.327
2023 doi
- [41]
- [42]
- [43]
- [44]
-
[45]
Priyansh Trivedi, Gaurav Maheshwari, Mohnish Dubey, and Jens Lehmann. 2017. LC-QuAD: A Corpus for Complex Question Answering over Knowledge Graphs. In Proceedings of the International Semantic Web Conference (ISWC) , Vol. 10588. 210–218. https://doi.org/10.1007/978-3-319-68204-4_22
2017 doi
-
[46]
Ricardo Usbeck, Ria Gusmita, Axel-Cyrille Ngomo, and Muhammad Saleem
-
[47]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, and et al. 2022. Chain-of-Thought Prompting Elicits Reason- ing in Large Language Models. In Advances in Neural Information Pro- cessing Systems: Annual Conference on Neural Information Processing Systems (N...
2022
- [48]
-
[49]
Wen-tau Yih, Matthew Richardson, Christopher Meek, Ming-Wei Chang, and Jina Suh. 2016. The Value of Semantic Parse Labeling for Knowledge Base Question Answering. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, (ACL). https://doi.org/10.1...
2016 doi
-
[50]
Diliara Zharikova, Daniel Kornev, Fedor Ignatov, Maxim Talimanchuk, Dmitry Evseev, and et al. 2023. DeepPavlov Dream: Platform for Building Gener- ative AI Assistants. In Proceedings of the Annual Meeting of the Association for Computational Linguistics: System Demonstrations ...
2023 doi
-
[51]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. InProceedings of the Annual Confere...
2023
-
[52]
Li Zhong and Zilong Wang. 2024. Can LLM Replace Stack Overflow? A Study on Robustness and Reliability of Large Language Model Code Generation. In Proceedings of the AAAI Conference on Artificial Intelligence . 21841–21849. https: //doi.org/10.1609/AAAI.V38I19.30185
2024 doi
-
[53]
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, and et al
-
[54]
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, and et al. 2023. Large Language Models are Human-Level Prompt Engineers. In The International Conference on Learning Representations (ICLR) . https:// openreview.net/pdf?id=92gvk82DE-
2023
-
[58]
In The International Conference on Learning Representations, (ICLR)
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models. In The International Conference on Learning Representations, (ICLR). https: //openreview.net/pdf?id=WZH7099tgfM
-
[154]
https://doi.org/10.1145/3477495.3531815
-
[2018]
9th Challenge on Question Answering over Linked Data (QALD-9). In Joint proceedings of the Workshop on Semantic Deep Learning (SemDeep-4) and NLIWoD4: Natural Language Interfaces for the Web of Data (NLIWOD-4) and 9th Question Answering over Linked Data challenge (QALD-9) co-l...
-
[2019]
In Proceedings of The Semantic Web - (ISWC) , Vol
LC-QuAD 2.0: A Large Dataset for Complex Question Answering over Wikidata and DBpedia. In Proceedings of The Semantic Web - (ISWC) , Vol. 11779. 69–78. https://doi.org/10.1007/978-3-030-30796-7_5
-
[2023]
In Proceedings of the Annual Meeting of the Association for Computational Linguistics: System Demonstrations, (ACL)
PEEP-Talk: A Situational Dialogue-based Chatbot for English Education. In Proceedings of the Annual Meeting of the Association for Computational Linguistics: System Demonstrations, (ACL). 190–207. https://doi.org/10.18653/V1/2023.ACL- DEMO.18
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.