Pith. sign in

REVIEW 3 major objections 41 references

Expert and editor agents improve long-document summaries by stepwise questioning that forces targeted revision of an initial draft.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 12:05 UTC pith:7VRZCWHP

load-bearing objection Incremental multi-agent recipe that sometimes lifts auto metrics on long scientific summaries via expert/editor questions, but mixed results and a circular LLM evaluator leave the claim under-supported. the 3 major comments →

arxiv 2607.10390 v1 pith:7VRZCWHP submitted 2026-07-11 cs.CL cs.AI

A Stepwise Questioning Expert-Editor Multi-Agent Framework for Long-Document Summarization

classification cs.CL cs.AI
keywords long-document summarizationmulti-agent systemsstepwise questioningprompt engineeringscientific summarizationArXivPubMedLLM agents
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Long scientific papers often exceed what a large language model can take in at once, so summaries lose facts or invent content. This paper claims that a multi-agent setup modeled on academic review can fix that. An author agent first produces a draft (directly or by section-and-aggregate). An expert agent then asks content questions about completeness, coherence, relevance and factual consistency; an editor agent asks surface yes/no questions about grammar and fidelity to the source. After each answer the author revises, and an evaluator keeps the better version. On ArXiv and PubMed samples the method raises ROUGE, BERTScore and FactCC scores relative to direct generation and a recent packaging baseline, showing that stepwise questions can surface the clues needed for faithful revision without reading the whole document again.

Core claim

The Expert-Editor stepwise questioning multi-agent method (SQ2E) improves long-document scientific summarization over both direct generation and the HERA segment-and-aggregate baseline. Expert questions supply missing content clues while editor questions catch surface inconsistencies; the author revises after each answer and an evaluator retains the superior draft, yielding higher automatic scores on ArXiv and PubMed.

What carries the argument

The Expert-Editor stepwise questioning pipeline: expert and editor agents formulate a short question plan from a fixed bank, pose the questions one by one to the author agent, trigger targeted revision, and let an evaluator keep the better of the previous and revised summaries under the same four quality criteria.

Load-bearing premise

That an LLM evaluator using the same completeness-coherence-relevance-consistency criteria already built into the questions can reliably pick the better summary after each revision, and that a short hand-designed question list (capped at five expert plus four editor questions) generalizes without new errors.

What would settle it

Run the identical pipeline on a held-out sample of ArXiv or PubMed papers and check whether the final ROUGE-L, BERTScore and FactCC numbers still exceed both direct generation and HERA; if they do not, or if human raters reverse the ranking, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Long-document summarizers can replace full re-reading with a short, role-specialized question plan that surfaces missing content clues.
  • Separating content questions (expert) from surface fidelity questions (editor) reduces irrelevant or unfaithful generations that single-agent methods produce.
  • The same question-revision-evaluate loop can be applied to other long-form generation tasks that currently rely on simple segment-and-aggregate.
  • Because valid questions are finite, systems can safely cap the number of revision rounds without losing most of the gain.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The method implicitly treats academic peer-review roles as a reusable prompt template; the same template may transfer to legal or medical long documents with only a change of question bank.
  • If the evaluator is itself an LLM, the pipeline risks circular preference for fluent but still incomplete text; an external non-LLM judge would be a direct next test.
  • The drop observed on the 70B model under limited compute suggests the gains may be fragile to quantization or memory constraints that smaller models avoid.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper proposes SQ2E, an expert–editor stepwise-questioning multi-agent framework for long-document summarization. An author agent first produces an initial summary (via direct generation or section-wise aggregation). An expert agent then poses content-oriented questions (completeness, coherence, relevance, factual consistency) and an editor agent poses surface-level yes/no questions; after each answer the author may revise, and an evaluator agent selects the better of the previous and revised versions. Experiments on 300-sample subsets of ArXiv and PubMed with LLaMA-3.1-8B/70B and DeepSeek-R1 compare SQ2E against direct generation (DG) and HERA, reporting ROUGE-1/2/L, BERTScore and FactCC. The authors claim that the questioning loop yields more accurate and faithful summaries than the baselines.

Significance. If the gains were robust, the work would supply a practical, training-free multi-agent recipe that exploits role-play and iterative questioning to mitigate context-length and faithfulness problems in scientific long-document summarization. The explicit expert/editor division of labor and the publicly described question-bank design are concrete engineering contributions that other multi-agent pipelines could reuse. The manuscript does not, however, ship code, human evaluations, or statistical tests, so the claimed advance remains provisional.

major comments (3)
  1. Table 3 (main results) does not uniformly support the effectiveness claim. On LLaMA-3.1-70B PubMed, SQ2E under-performs DG on ROUGE-1 (36.82 vs 39.58), ROUGE-2 (12.28 vs 14.25) and ROUGE-L (20.60 vs 20.7); FactCC also drops relative to DG or HERA in several cells. The authors attribute the 70B degradation to “loss of precision” from resource constraints (§4.5) without quantification or a corrected run, leaving the central claim only partially evidenced.
  2. §3.4 and §4.4: the evaluator that decides whether each revision is retained is itself an LLM that re-uses the identical completeness/coherence/relevance/consistency criteria already encoded in the expert/editor questions. No human agreement study, inter-annotator reliability, or ablation of the evaluator is reported. Consequently it is unclear whether the positive metric cells reflect genuine quality gains or self-reinforcement / metric gaming.
  3. §4.1–4.5: evaluation uses only 300 randomly sampled documents, reports no error bars or significance tests, and contains no human evaluation. With free parameters (max 5 expert + 4 editor questions, hand-designed question bank) fixed after a small pilot, the statistical support for the blanket effectiveness claim is insufficient.

Circularity Check

0 steps flagged

Empirical multi-agent pipeline with external automatic metrics; no derivation collapses into its inputs by construction.

full rationale

This is an empirical systems paper proposing a multi-agent refinement loop (expert/editor questions + author revision + LLM evaluator) for long-document summarization, then measuring final outputs with standard external automatic metrics (ROUGE-1/2/L, BERTScore, FactCC) against DG and HERA baselines on ArXiv/PubMed. There is no mathematical derivation, no fitted parameter renamed as a prediction, no uniqueness theorem, and no load-bearing self-citation that forces the central claim. The only mild internal alignment is that the hand-designed question bank and the evaluator both reference the same qualitative criteria (completeness, coherence, relevance, consistency); this is by design of the critic loop and does not make the reported ROUGE/FactCC gains tautological, because those metrics are independent of the internal evaluator. The paper is therefore self-contained against external benchmarks; score remains near zero.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 2 invented entities

The central empirical claim rests on a small set of design choices (question caps, qualitative criteria, role prompts) that are not derived from theory and on the domain assumption that LLM role-play plus short questions reliably improves faithfulness. No free physical constants appear; the free parameters are purely engineering knobs chosen after a pilot. Invented entities are the named roles and the SQ2E pipeline itself, which have no independent existence outside the prompting setup.

free parameters (3)
  • max_expert_questions = 5
    Hard cap of 5 expert questions per document, chosen after a small pilot to avoid infinite loops (Section 4.4); directly controls how much refinement occurs.
  • max_editor_questions = 4
    Hard cap of 4 editor yes/no questions, likewise set by pilot (Section 4.4).
  • question_bank_and_ordering
    The specific list and sequence of questions around completeness/coherence/relevance/consistency is hand-designed and selected per document by the expert/editor agents; not fixed by theory.
axioms (3)
  • domain assumption LLM agents assigned expert/editor/writer/evaluator roles via prompting can reliably improve summary quality through stepwise questioning without external supervision.
    Stated throughout Sections 1 and 3 as the motivating premise drawn from prior multi-agent literature; never independently validated outside the automatic metrics of this paper.
  • domain assumption Automatic metrics ROUGE, BERTScore and FactCC are sufficient proxies for the qualitative criteria (completeness, coherence, relevance, factual consistency) used by the agents.
    Evaluation section (4.3) equates these metrics with the desired summary properties; no human correlation study is supplied.
  • domain assumption Segment-by-section local summaries followed by aggregation produce a usable initial draft for subsequent refinement.
    Adopted from prior work (Li et al. 2025, Zhong & Litman 2025) in Section 3.2 without re-derivation.
invented entities (2)
  • SQ2E Expert-Editor stepwise questioning multi-agent pipeline no independent evidence
    purpose: Orchestrates four LLM roles and a fixed question-driven revision loop to produce the final long-document summary.
    The named framework and its four roles are introduced by the paper; they exist only as prompt configurations and have no independent empirical handle outside the reported automatic scores.
  • Evaluator agent that selects better summary after each revision no independent evidence
    purpose: Gates whether a revision is kept, using the same qualitative criteria as the questions.
    Introduced in Section 3.1/3.4; its judgments are internal to the same LLM family and criteria, providing no external falsifiable signal.

pith-pipeline@v1.1.0-grok45 · 14351 in / 3197 out tokens · 37846 ms · 2026-07-14T12:05:12.051777+00:00 · methodology

0 comments
read the original abstract

Although large language models (LLMs) have shown promising potential in news summarization tasks, their performance on long-document summarization remains challenging as their length often exceeds the input limits. As the agent investment, which provide possibility to improve the inherent capabilities of LLMs. To enhance the effectiveness of long-document summarization based on LLMs, this paper proposes an expert-editor stepwise questioning multi-agent method, in which the expert and the editor guide another agent to refine the summary by posing questions on different aspects of the content and providing targeted clues for revision. We conducted experiments on two representative long-document scientific datasets and evaluated the results through widely recognized automatic metrics. The results demonstrated the effectiveness of our method.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

41 extracted references · 1 canonical work pages

  1. [1]

    Aneesha Bakharia. 2025. Iterative proof-driven development LLM prompt. In Companion proceedings of the ACM on web conference 2025, pages 1596 – 1597, Sydney NSW, Australia and New York, NY, USA. Association for Computing Machinery. Citation Key: 10.1145/3701716.3717811

  2. [2]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, et al. 2020. Language models are few-shot learners. Citat...

  3. [3]

    Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shan Zhang, Jie Fu, and Zhiyuan Liu. 2023. ChatEval: Towards better LLM-based evaluators through multi-agent debate. ArXiv, abs/2308.07201. Citation Key: Chan2023ChatEvalTB

  4. [4]

    Sangwoo Cho, Kaiqiang Song, Xiaoyang Wang, Fei Liu, and Dong Yu. 2022. Toward Unifying Text Segmentation and Long Document Summarization. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 106–118, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics. Citation Key: cho-etal-2022-toward

  5. [5]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sashank Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, et al. 2023. PaLM: scaling language modeling with pathways...

  6. [6]

    Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018. A Discourse-Aware Attention Model for Abstractive Summarization of Long Documents. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short...

  7. [7]

    Tianyi and Li Dong Zican and Tang. 2024. BAMBOO: a comprehensive benchmark for evaluating long text modeling capacities of large language models. In Min-Yen and Hoste Calzolari Nicoletta and Kan, editor, Proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (LREC-COLING 2024), pages 2086–209...

  8. [8]

    Collaborative Document Simplification Using Multi-Agent Systems

    Dengzhao Fang, Jipeng Qiang, Xiaoye Ouyang, Yi Zhu, Yunhao Yuan, and Yun Li. Collaborative Document Simplification Using Multi-Agent Systems

  9. [9]

    JiaLe and Wang Gu WenYuan and Han. 2025. Explain-analyze-generate: a sequential multi-agent collaboration method for complex reasoning. In Leo and Apidianaki Rambow Owen and Wanner, editor, Proceedings of the 31st international conference on computational linguistics, pages 7127 – 7140, Abu Dhabi, UAE. Association for Computational Linguistics. Citation K...

  10. [10]

    Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. 2024. RULER: What’s the real context size of your long-context language models? arXiv e-prints:arXiv:2404.06654. Citation Key: 10 2024arXiv240406654HarXiv: 2404.06654 [cs.CL]number: arXiv:2404.06654tex.adsnote: Provided by the SAO/NASA Astr...

  11. [11]

    Hou Pong and Li Hu Zhe and Chan. 2025. Debate-to-write: a persona-driven multi-agent framework for diverse argument generation. In Leo and Apidianaki Rambow Owen and Wanner, editor, Proceedings of the 31st international conference on computational linguistics, pages 4689 – 4703, Abu Dhabi, UAE. Association for Computational Linguistics. Citation Key: hu-e...

  12. [12]

    Bastian, Alvaro Velasquez, and Sandeep Neema

    Susmit Jha, Sumit Kumar Jha, Patrick Lincoln, Nathaniel D. Bastian, Alvaro Velasquez, and Sandeep Neema. 2023. Dehallucinating large language models using formal methods guided iterative prompting. In 2023 IEEE international conference on assured autonomy (ICAA), pages 149–152. Citation Key: 10207581

  13. [13]

    Wojciech Kryści ń ski, Bryan McCann, Caiming Xiong, and Richard Socher. 2019. Evaluating the factual consistency of abstractive text summarization. Citation Key: kryściń ski2019evaluatingfactualconsistencyabstractivearXiv: 1910.12840 [cs.CL]

  14. [15]

    Taiji Li, Hao Chen, Fei Yu, and Yin Zhang. 2025b. HERA: Improving Long Document Summarization using Large Language Models with Context Packaging and Reordering. arXiv:2502.00448 [cs]

  15. [16]

    Chin-Yew Lin. 2004. ROUGE: a package for automatic evaluation of summaries. In Text summarization branches out, pages 74 – 81, Barcelona, Spain. Association for Computational Linguistics. Citation Key: lin-2004-rouge

  16. [17]

    Alexander and Chen Liu Yixin and Fabbri. 2024. Benchmarking generation and evaluation capabilities of large language models for instruction controllable summarization. In Helena and Bethard Duh Kevin and Gomez, editor, Findings of the association for computational linguistics: NAACL 2024, pages 4481 – 4501, Mexico City, Mexico. Association for Computation...

  17. [18]

    Ran Liu, Xian-Ling Mao, and Heyan Huang. 2025. DSciSum: Detailed summarization of long scientific documents. Knowledge-Based Systems, 317:113409

  18. [19]

    Kaijie Mo and Renfen Hu. 2024. ExpertEase: A Multi-Agent Framework for Grade- Specific Document Simplification with Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 9080 – 9099, Miami, Florida, USA. Association for Computational Linguistics. Citation Key: mo-hu-2024- expertease

  19. [20]

    Lingyun Shen and Xiaoqiu Le. 2023. An Enhanced Method on Transformer-Based Model for ONE2SEQ Keyphrase Generation. Electronics, 12(13)

  20. [21]

    Nathalia Nascimento, Paulo Alencar, and Donald Cowan. 2023. GPT-in-the-loop: Adaptive decision-making for multiagent systems. Citation Key: nascimento2023gptintheloopadaptivedecisionmakingmultiagentarXiv: 2308.10435 [cs.MA]

  21. [22]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, et al. 2024. GPT-4 technical report. Citation Key: openai2024gpt4technic...

  22. [23]

    Muru and Min Press Ofir and Zhang. 2023. Measuring and narrowing the compositionality gap in language models. In Juan and Bali Bouamor Houda and Pino, editor, Findings of 11 the association for computational linguistics: EMNLP 2023, pages 5687–5711, Singapore. Association for Computational Linguistics. Citation Key: press-etal-2023-measuring

  23. [25]

    Gaurav Sahu, Olga Vechtomova, and Issam H. Laradji. 2025b. A guide to effectively leveraging llms for low-resource text summarization: Data augmentation and semi- supervised approaches. Citation Key: sahu2025guideeffectivelyleveragingllmsarXiv: 2407.07341 [cs.CL]

  24. [26]

    Schmidt, Jesse Spencer-Smith, Quchen Fu, and Jules White

    Douglas C. Schmidt, Jesse Spencer-Smith, Quchen Fu, and Jules White. 2023. Cataloging prompt patterns to enhance the discipline of prompt engineering. In Citation Key: Schmidt2023CatalogingPP

  25. [27]

    Sander Schulhoff, Michael Ilie, Nishant Balepur, Konstantine Kahadze, Amanda Liu, Chenglei Si, Yinheng Li, Aayush Gupta, HyoJung Han, Sevien Schulhoff, Pranav Sandeep Dulepet, Saurav Vidyadhara, Dayeon Ki, Sweta Agrawal, Chau Pham, Gerson Kroiz, Feileen Li, Hudson Tao, Ashay Srivastava, et al. 2025. The prompt report: a systematic survey of prompt enginee...

  26. [28]

    Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Jiaxiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, Weixin Liu, Zhihua Wu, Weibao Gong, Jianzhong Liang, Zhizhou Shang, Peng Sun, Wei Liu, Xuan Ouyang, Dianhai Yu, et al

  27. [29]

    Citation Key: sun2021ernie30largescaleknowledgearXiv: 2107.02137 [cs.CL]

    ERNIE 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation. Citation Key: sun2021ernie30largescaleknowledgearXiv: 2107.02137 [cs.CL]

  28. [30]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. LLaMA: Open and efficient foundation language models. Citation Key: touvron2023llamaopenefficientfoundationarXiv: 2302...

  29. [31]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. Cit...

  30. [32]

    David Wan, Justin Chih-Yao Chen, Elias Stengel-Eskin, and Mohit Bansal. 2025. MAMM-refine: a recipe for improving faithfulness in generation with multi-agent collaboration. Citation Key: wan2025mammrefinerecipeimprovingfaithfulnessarXiv: 2503.15272 [cs.CL]

  31. [33]

    MAMM-refine: a recipe for improving faithfulness in generation with multi-agent collaboration

    David Wan, Justin Chen, Elias Stengel-Eskin, and Mohit Bansal. MAMM-refine: a recipe for improving faithfulness in generation with multi-agent collaboration. Citation Key: osti_10610633

  32. [34]

    Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen. 2024a. A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6). Citation Key: Wang_2024

  33. [35]

    Ting Wang, Chuan Yang, Maoyang Zou, Jiaying Liang, Dong Xiang, Wenjie Yang, Hongyang Wang, and Jia Li. 2024b. A study of extractive summarization of long documents incorporating local topic and hierarchical information. Scientific Reports, 14(1):10140. 12

  34. [36]

    Huang, Jie Fu, and Junran Peng

    Zekun Moore Wang, Zhongyuan Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang, Wanli Ouyang, Ke Xu, Stephen W. Huang, Jie Fu, and Junran Peng. 2024c. RoleLLM: Benchmarking, eliciting, and enhancing role-playing abilities of large language models. Citation Key: wang2024rol...

  35. [37]

    Weinberger, and Yoav Artzi

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating text generation with BERT. Citation Key: zhang2020bertscoreevaluatingtextgenerationarXiv: 1904.09675 [cs.CL]

  36. [38]

    Hashimoto

    Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B. Hashimoto. 2023. Benchmarking large language models for news summarization. Citation Key: zhang2023benchmarkinglargelanguagemodelsarXiv: 2301.13848 [cs.CL]

  37. [39]

    Hashimoto

    Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B. Hashimoto. 2024. Benchmarking Large Language Models for News Summarization. Transactions of the Association for Computational Linguistics, 12:39–57

  38. [40]

    Mingqian Zheng, Jiaxin Pei, Lajanugen Logeswaran, Moontae Lee, and David Jurgens

  39. [41]

    a helpful assistant

    When “a helpful assistant” is not really helpful: Personas in system prompts do not improve performances of large language models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the association for computational linguistics: EMNLP 2024, pages 15126–15154, Miami, Florida, USA. Association for Computational Linguistics. Citation ...

  40. [42]

    Yang Zhong and Diane Litman. 2025. Discourse-Driven Evaluation: Unveiling Factual Inconsistency in Long Document Summarization. arXiv:2502.06185 [cs]

  41. [43]

    Factual Dialogue Summarization via Learning from Large Language Models

    Rongxin Zhu, Jey Han Lau, and Jianzhong Qi. Factual Dialogue Summarization via Learning from Large Language Models