REVIEW 3 major objections 5 minor 70 references
A Unified Retrieval Framework with Document Ranking and EDU Filtering for Multi-document Summarization
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read ReREF replaces naive truncation with document ranking plus EDU filtering and reports consistent ROUGE gains across seven summarizers on four MDS datasets.
desk verdict A coherent, model-agnostic retrieval plug-in for MDS with a genuinely new EDU-level ranking/filtering combination, but the paper's internal retrieval analysis is circular and the end-to-end comparison does not isolate the mechanism. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the elementary discourse unit (EDU), a minimal coherent segment of text, smaller than a sentence, obtained with the DMRST discourse parser. The machinery has three moving parts: a filtering model that scores every EDU by salience and takes the top-scoring ones as latent queries; a ranking model that scores each document by averaging its dot-product similarity to all latent queries; and an EM loop that alternates between choosing the latent queries (E-step) and updating both models through two Bayesian Personalized Ranking losses (M-step). Training labels come from cosine similarity between EDUs or documents and the reference summary, computed with a pre-trained semantic-search embedding model. This design lets query selection, ranking, and filtering share one representation space and one optimization procedure.
What would settle it
Build a test set whose reference summaries paraphrase the source in vocabulary that the embedding model scores as low-similarity to the informative EDUs; if ReREF's selected inputs then fail to beat truncation while an oracle that selects EDUs by their marginal ROUGE contribution succeeds, the embedding-similarity training signal is the weak link.
Extended reading notes
Core claim
ReREF claims that the retrieve-then-summarize bottleneck in multi-document summarization can be broken without human-written queries or coarse passage/sentence retrieval. The framework picks the top-scoring elementary discourse units (EDUs), minimal coherent text segments smaller than sentences, from the input documents and uses them as latent queries to rank the documents; it then replaces naive token truncation with filtering that removes the lowest-scoring EDUs until the input fits the summarizer's context window. An expectation-maximization loop alternates between selecting these latent queries and updating the filtering and ranking models. Trained this way and applied as a model-agnostic front end, ReREF reports consistent ROUGE gains over the original truncation for seven summarizers on Multi-News, Multi-XScience, Wikisum, and WCEP-10, in both fully supervised and 1% few-shot settings, and human evaluation on 100 Multi-News samples shows improved informativeness and fluency with unchanged succinctness.
Load-bearing premise
The whole framework is trained and judged against one proxy: how similar each piece of text or whole document is to the reference summary under a pre-trained semantic-similarity model; if that measure does not track what a summarizer actually needs, the reported gains may not transfer.
Editorial extensions
If this is right
- Because ReREF is model-agnostic and applied before the summarizer, a trained ReREF front end could upgrade any fixed-context summarizer without changing its parameters.
- The ablation results imply that document ranking and EDU filtering contribute separately: dropping either lowers ROUGE, and dropping both lowers it further, so both mechanisms are needed for the full gain.
- In few-shot settings with only 1% of training data, ReREF still improves most model-dataset pairs, with the largest reported jump being PEGASUS gaining +7.47 ROUGE-2 on Wikisum.
- In retrieval evaluation, ReREF beats BM25 for both query selection and filtering even when BM25 is given a gold query derived from the reference summary, so the learned latent queries carry information beyond keyword overlap.
Reading between the lines
- A natural next test the paper does not run is cross-domain transfer: since training labels come from reference summaries, a retriever trained on news data may need retraining or substitution before it works on legal or medical MDS.
- The same EDU-filtering mechanism could be used as a denoiser for models with very long context windows, not just as a way to fit a length limit; the paper's core mechanism targets irrelevant content, not merely token budget.
- Because the latent queries are just embeddings of selected EDUs, an explicit user query could be injected as an additional query embedding, which would convert ReREF into a query-focused multi-document summarizer without changing the EM loop.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes ReREF, a model-agnostic retrieval framework for multi-document summarization (MDS) that unifies document ranking and elementary discourse unit (EDU) filtering. The framework automatically selects latent queries from the most salient EDUs, uses those queries to rank documents, and then replaces naive last-token truncation with EDU-level filtering to fit the summarizer's context limit. The training signal for both EDU scoring and document ranking is cosine similarity to the reference summary using the pre-trained embedding model multi-qa-mpnet-base-cos-v1, and the whole pipeline is optimized by an EM-style alternation between query selection and parameter updates. The authors evaluate ReREF with seven summarizers on four MDS datasets in fully supervised and few-shot settings, reporting ROUGE improvements over the original truncation baselines, along with retrieval-side precision/NDCG analyses and a small human evaluation. The core claim is that ReREF consistently improves summarization quality across models and datasets while avoiding manual query construction and coarse-grained retrieval.
Significance. If the central claim holds, ReREF offers a practical, model-agnostic way to use limited context windows more effectively in MDS, with the notable advantages of eliminating handcrafted queries and filtering at a finer granularity than sentences or passages. The paper's strengths include a broad experimental matrix (seven summarizers, four datasets, both fully supervised and few-shot settings), the use of paired bootstrap significance testing, a released code repository, and a human evaluation component. The ROUGE point estimates in Tables 2 and 3 are mostly positive, and the framework's plug-and-play nature is genuinely appealing. However, the paper's mechanism-level validation is substantially weakened by the circular retrieval evaluation in Section 5.4, and the claim of 'significant performance gain in all datasets' in Section 5.2 is contradicted by the manuscript's own significance markers. The lack of an end-to-end comparison against a standard retriever under the same context budget means the observed gains could plausibly stem from 'any selection beats naive truncation' rather than from ReREF's specific latent-query and EDU-filtering mechanism.
major comments (3)
- [§5.4, §4.3] The retrieval-side evaluation in Section 5.4 is circular. The ground-truth labels for EDU selection (top-k and bottom-k EDUs) and document ranking are constructed by the same cosine-similarity-to-reference-summary scoring described in Section 4.3 using multi-qa-mpnet-base-cos-v1. The retrieval model is trained to reproduce exactly these labels, so the high Precision@K and NDCG/MRR values in Figures 4 and 5 only demonstrate that the model fits the training proxy; they do not validate that the proxy captures what downstream summarizers actually need. This circularity undermines the paper's claim to have validated the latent-query and EDU-filtering mechanism through 'in-depth analysis.' I recommend replacing or supplementing this evaluation with a non-circular test, such as an end-to-end comparison in which the same summarizer is fed, under an identical token budget, by (a) ReREF, (b) a standard retriever like BM25+RAKE or DYLE+RAKE, and (c) an inference-time oracle that directly uses the Section 4.3 similarity scores without any learned retrieval module. Alternatively, human annotation of whether the selected/filtered EDUs are actually summary-relevant would break the circularity.
- [§5.2, Tables 2 and 3] The text in Section 5.2 states that 'Our retrieval framework achieves a significant performance gain in all datasets compared to the original base models,' but this is contradicted by the minus signs in Tables 2 and 3, which denote p>=0.05. For example, in Table 2, all three ROUGE improvements for StableLM-Zephyr-3B on Multi-XScience and WCEP-10 are marked not significant, as are all three for BartGraphSum on Multi-XScience. In Table 3, BartGraphSum, BART (partial), PEGASUS (partial), and StableLM show multiple non-significant cells, especially on Multi-XScience and WCEP-10. Across both tables, roughly 26 of 168 model-dataset-metric cells are not significant. The claim of consistent significant improvement is therefore overstated. The authors should temper the language to 'improvements in most settings, with significance varying by dataset and model,' or provide additional evidence (e.g., multiple random seeds, larger sample sizes, or meta-analytic aggregation) to support the significance claim.
- [Table 5 and footnote 11] The ablation study in Table 5 does not isolate the mechanism responsible for the end-to-end gains. Footnote 11 explicitly states that the 'w/o both' variant (random document order plus last-token drop) is different from and worse than PRIMERA's original even truncation, and that its Table 5 results do not align with the 'original' row in Table 2. Consequently, the ablation demonstrates only that ReREF improves over a deliberately degraded baseline (random order plus last-token drop), not that ranking and EDU filtering are the cause of the improvement over the actual baseline used in Table 2. This is a strawman comparison. To support the claim that the learned ranking and filtering are what produce the gains, the authors should compare ReREF end-to-end against a standard retriever (e.g., BM25+RAKE or DYLE+RAKE) with the same context budget feeding the same summarizer, or against a non-learned selection based directly on the Section 4.3 similarity proxy. Without such a comparison, the observed ROUGE gains could be explained by any content-selection strategy that avoids the information loss of naive truncation.
minor comments (5)
- [Figure 5] The caption contains a typo: 'MRR_2rd' should be 'MRR_2nd'.
- [§5.5] The human evaluation in Section 5.5 is based on only 100 Multi-News samples with a single base model (PRIMERA), and no inter-annotator agreement statistic (e.g., Krippendorff's alpha or Cohen's kappa) is reported. Adding agreement metrics and ideally a second dataset would strengthen the reliability of the human evaluation.
- [Eq. (11)] The notation in Equation (11) is confusing: the two loss terms use P_q and P_f for the positive sets and what appear to be complement sets, but the complement notation is not visible in the typeset equation. Please clarify the sampling distributions and the definition of the negative sets in the BPR loss.
- [§5.4] Figures 4 and 5 report retrieval-side precision and ranking metrics without confidence intervals or significance tests. Given the circularity issue, at minimum the variance across test instances should be reported so that the reader can assess the stability of the comparisons against BM25 and DYLE.
- [§5.1] The sentence 'For fair comparison, we used the same input context as PRIMERA when evaluating LLaMA variants' is unclear: LLaMA-3.1 supports 128k tokens, so it is not obvious why a 4096-token limit is 'fair' to LLaMA. Please clarify the rationale for constraining all models to the same budget.
Circularity Check
Section 5.4 evaluates the trained retrieval module against the same reference-summary cosine-similarity labels used to create its training supervision; the ROUGE headline is independent but this mechanism claim is circular.
-
fitted input called prediction
[Section 4.3 (training labels) vs Section 5.4 (evaluation ground truth), Figures 4-5]
"The relevance of each EDU is calculated as the cosine similarity between its embedding and the embedding of the reference summary. These scores reflect the salience of the EDUs within the overall context and are leveraged to supervise the EDU ranking process, including both latent queries selection and EDU filtering. ... For the EDU selection task, we used the ranked EDUs based on the EDU scoring from Section 4.3 as the ground truth. For the document ranking task, we used the document ranking labels from Section 4.3 as the ground truth."
The retrieval model is trained with the BPR losses in Eqs. 11-13, whose positive/negative pairs are derived from exactly these cosine-similarity-to-reference labels: top EDUs become P_q, bottom EDUs become P_f, and document pairs P_g come from the same reference-summary scores. Section 5.4 then measures Precision@K, NDCG@K, and MRR against 'the ranked EDUs based on the EDU scoring from Section 4.3' and 'the document ranking labels from Section 4.3' as ground truth. Thus the high retrieval/filtering/ranking accuracy reported in Figures 4-5 only shows that the trained module reproduces its own supervisory signal; it does not independently establish that the cosine-similarity proxy captures summary-relevant content, nor does it isolate the mechanism responsible for the ROUGE gains.
full rationale
The paper's headline ROUGE results compare ReREF against original truncation for seven summarizers on four datasets (Tables 2-3); those are external, non-circular benchmarks, and the human evaluation in Section 5.5 is also independent. The circularity is confined to the in-depth analysis of Section 5.4 and the corresponding claims in the abstract/conclusion about 'dynamically select appropriate queries and accurately rank documents based on their relevance scores.' The DMRST parser, multi-qa-mpnet-base-cos-v1, SpanExt, and the BM25/RAKE/DYLE baselines are all external resources, and there is no load-bearing self-citation chain. One experimental-design caveat in footnote 11, that the 'w/o both' ablation is not the same as the original truncation, weakens the attribution of ROUGE gains to ranking and filtering, but that is a baseline-comparison issue rather than circularity in the derivation. Overall score 5 reflects partial circularity of the mechanism-validation analysis, while the main ROUGE claim remains independently supported.
Assumptions & free parameters
free parameters (3)
- Query number k =
10
- Loss balance weight lambda =
1.0
- Chunk size c =
1024 tokens
assumptions (5)
- domain assumption EDUs are the correct granularity for relevance assessment and filtering.
- domain assumption Cosine similarity between an EDU and the reference summary in the embedding space of multi-qa-mpnet-base-cos-v1 is a valid training signal for salience.
- domain assumption Cross-attention between EDU embeddings and document embeddings highlights EDUs shared across documents, which are more salient.
- standard math BPR loss is an appropriate objective for EDU and document ranking.
- domain assumption Reference summaries in the datasets are accurate and sufficient for supervision.
Cite this review
Pith. "Pith review of A Unified Retrieval Framework with Document Ranking and EDU Filtering for Multi-document Summarization." pith.science (2026). https://pith.science/paper/KAPRMZQ6
@misc{pith2026250416711,
author = {Pith},
title = {Pith review of: A Unified Retrieval Framework with Document Ranking and EDU Filtering for Multi-document Summarization},
year = {2026},
howpublished = {\url{https://pith.science/paper/KAPRMZQ6}},
note = {Machine review of arXiv:2504.16711}
}
read the original abstract
In the field of multi-document summarization (MDS), transformer-based models have demonstrated remarkable success, yet they suffer an input length limitation. Current methods apply truncation after the retrieval process to fit the context length; however, they heavily depend on manually well-crafted queries, which are impractical to create for each document set for MDS. Additionally, these methods retrieve information at a coarse granularity, leading to the inclusion of irrelevant content. To address these issues, we propose a novel retrieval-based framework that integrates query selection and document ranking and shortening into a unified process. Our approach identifies the most salient elementary discourse units (EDUs) from input documents and utilizes them as latent queries. These queries guide the document ranking by calculating relevance scores. Instead of traditional truncation, our approach filters out irrelevant EDUs to fit the context length, ensuring that only critical information is preserved for summarization. We evaluate our framework on multiple MDS datasets, demonstrating consistent improvements in ROUGE metrics while confirming its scalability and flexibility across diverse model architectures. Additionally, we validate its effectiveness through an in-depth analysis, emphasizing its ability to dynamically select appropriate queries and accurately rank documents based on their relevance scores. These results demonstrate that our framework effectively addresses context-length constraints, establishing it as a robust and reliable solution for MDS.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Chenxin An, Ming Zhong, Zhichao Geng, Jianqiang Yang, and Xipeng Qiu. 2021. RetrievalSum: A Retrieval Enhanced Framework for Abstractive Summarization. CoRR abs/2109.07943 (2021). https://arxiv.org/abs/2109.07943
arXiv 2021
-
[2]
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The Long- Document Transformer. CoRR abs/2004.05150 (2020). https://arxiv.org/abs/2004. 05150
arXiv 2020
-
[3]
Ziqiang Cao, Wenjie Li, Sujian Li, and Furu Wei. 2018. Retrieve, Rerank and Rewrite: Soft Template Based Neural Summarization. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL) . 152–161
work page 2018
-
[4]
Lynn Carlson, Daniel Marcu, and Mary Ellen Okurowski. 2003. Building a discourse-tagged corpus in the framework of rhetorical structure theory. Current and new directions in discourse and dialogue (2003), 85–112
work page 2003
-
[5]
Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda. 2021. CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification . In2021 IEEE/CVF International Conference on Computer Vision (ICCV) . 347–356
work page 2021
-
[6]
Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu, Kaixin Ma, Xinran Zhao, Hongming Zhang, and Dong Yu. 2024. Dense X Retrieval: What Retrieval Granu- larity Should We Use?. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 15159–15177
work page 2024
-
[7]
Xiuying Chen, Mingzhe Li, Shen Gao, Xin Cheng, Qingqing Zhu, Rui Yan, Xin Gao, and Xiangliang Zhang. 2024. Flexible and Adaptable Summarization via Expertise Separation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) . ACM, 2018–2027
work page 2024
-
[8]
Nachshon Cohen, Oren Kalinsky, Yftah Ziser, and Alessandro Moschitti. 2021. WikiSum: Coherent Summarization Dataset for Efficient Human-Evaluation. In Proceedings of the 59th Annual Meeting of the Association for Computational Lin- guistics and the 11th International Joint Conference on Natural Language Processing (ACL: Short Papers). 212–219
work page 2021
Show all 70 references
-
[9]
Peng Cui and Le Hu. 2021. Topic-Guided Abstractive Multi-Document Summa- rization. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021 , Marie- Francine Moens, Xuanjing Huang, Lucia Spec...
2021
-
[10]
A. P. Dempster, N. M. Laird, and D. B. Rubin. 1977. Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society: Series B 39 (1977), 1–38
1977
-
[11]
Jay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl, and Lucy Lu Wang
-
[12]
Masafumi Enomoto, Kunihiro Takeoka, Kosuke Akimoto, Kiril Gashteovski, and Masafumi Oyamada. 2024. LightPAL: Lightweight Passage Retrieval for Open Domain Multi-Document Summarization. CoRR abs/2406.12494 (2024). https://arXiv.2406.12494
2024 arXiv
-
[13]
Alexander Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir Radev. 2019. Multi- News: A Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL). 1074–1084
2019
-
[14]
Demian Gholipour Ghalandari, Chris Hokamp, Nghia The Pham, John Glover, and Georgiana Ifrim. 2020. A Large-Scale Multi-Document Summarization Dataset from the Wikipedia Current Events Portal. In Proceedings of the 58th Annual Meeting of the Association for Computational Lingui...
2020
-
[15]
Giorgi, Luca Soldaini, Bo Wang, Gary D
John M. Giorgi, Luca Soldaini, Bo Wang, Gary D. Bader, Kyle Lo, Lucy Lu Wang, and Arman Cohan. 2023. Open Domain Multi-document Summarization: A Comprehensive Study of Model Brittleness under Retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2023 ....
2023
-
[16]
Jade Goldstein, Vibhu O Mittal, Jaime G Carbonell, and Mark Kantrowitz. 2000. Multi-document summarization by sentence extraction. In NAACL-ANLP 2000 workshop: automatic summarization
2000
-
[17]
Ho, Christopher Ré, and et al
Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher Ré, and et al. 2023. Legal- Bench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Pro...
2023
-
[18]
Itay Harel, Hagai Taitelbaum, Idan Szpektor, and Oren Kurland. 2022. A Dataset for Sentence Retrieval for Open-Ended Dialogues. In SIGIR ’22: The 45th Inter- national ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR). ACM, 2960–2969
2022
-
[19]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In The Tenth International Conference on Learning Representa- tions (ICLR)
2022
-
[20]
Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated gain-based evaluation of IR techniques. ACM Trans. Inf. Syst. 20, 4 (2002), 422–446
2002
-
[21]
Hanqi Jin, Tianming Wang, and Xiaojun Wan. 2020. Multi-Granularity Interaction Network for Extractive and Abstractive Multi-Document Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL). 6244–6254
2020
-
[22]
Jordan and R.A
M.I. Jordan and R.A. Jacobs. 1993. Hierarchical mixtures of experts and the EM algorithm. In Proceedings of 1993 International Conference on Neural Networks (IJCNN). 1339–1344
1993
-
[23]
Subhendu Khatuya, Koushiki Sinha, Niloy Ganguly, Saptarshi Ghosh, and Pawan Goyal. 2024. Instruction-Guided Bullet Point Summarization of Long Financial Earnings Call Transcripts. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Info...
2024
-
[24]
Philipp Koehn. 2004. Statistical Significance Tests for Machine Translation Evaluation. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing (EMNLP). 388–395
2004
-
[25]
Jingun Kwon, Naoki Kobayashi, Hidetaka Kamigaito, and Manabu Okumura
-
[26]
Kenton Lee, Luheng He, Mike Lewis, and Luke Zettlemoyer. 2017. End-to-end Neural Coreference Resolution. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 188–197
2017
-
[27]
In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP)
Considering Nested Tree Structure in Sentence Extractive Summarization with Pre-trained Transformer. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 4039–4044
2021
-
[28]
Miao Li, Jianzhong Qi, and Jey Han Lau. 2023. Compressed Heterogeneous Graph for Abstractive Multi-Document Summarization. In Thirty-Seventh AAAI Conference on Artificial Intelligence (AAAI). 13085–13093
2023
-
[29]
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58t...
2020
-
[30]
Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, 74–81
2004
-
[31]
Wei Li, Xinyan Xiao, Jiachen Liu, Hua Wu, Haifeng Wang, and Junping Du
-
[32]
Shengjie Liu, Jing Wu, Jingyuan Bao, and et al. 2024. Towards a Robust Retrieval- Based Summarization System. CoRR abs/2403.19889 (2024). https://arxiv.org/ abs/2403.19889
2024 arXiv
-
[33]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR abs/1907.11692 (2019). http://arxiv.org/abs/1907.11692
2019 arXiv
-
[34]
Chin-Yew Lin and Eduard Hovy. 2002. From single to multi-document summariza- tion. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL). 457–464
2002
-
[35]
Yao Lu, Yue Dong, and Laurent Charlin. 2020. Multi-XScience: A Large-scale Dataset for Extreme Multi-document Summarization of Scientific Articles. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 8068–8074
2020
-
[36]
Congbo Ma. 2021. Improving Deep Learning based Multi-document Summa- rization through Linguistic Knowledge. In SIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, Canada, July 11-15, 2021 , Fernando Diaz, ...
2021
-
[37]
Zhengyuan Liu, Ke Shi, and Nancy F. Chen. 2021. DMRST: A Joint Framework for Document-Level Multilingual RST Discourse Segmentation and Parsing. CoRR abs/2110.04518. https://arxiv.org/abs/2110.04518
2021 arXiv
-
[38]
Manuj Malik, Zheng Zhao, Marcio Fonseca, Shrisha Rao, and Shay B. Cohen. 2024. CivilSum: A Dataset for Abstractive Summarization of Indian Court Decisions. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR...
2024
-
[39]
William C Mann and Sandra A Thompson. 1988. Rhetorical structure theory: Toward a functional theory of text organization. Text-interdisciplinary Journal for the Study of Discourse 8, 3 (1988), 243–281
1988
-
[40]
Congbo Ma, Wei Emma Zhang, Mingyu Guo, Hu Wang, and Quan Z. Sheng. 2023. Multi-document Summarization via Deep Learning Techniques: A Survey. ACM Comput. Surv. 55, 5 (2023), 102:1–102:37
2023
-
[41]
Ziming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang, Rui Zhang, Tao Yu, Budhaditya Deb, Chenguang Zhu, Ahmed Hassan Awadallah, and Dragomir R. Radev. 2022. DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization. In Proceedings of the 60th Annual Meeting of the ...
2022
-
[42]
Geoffrey J McLachlan and Thriyambakam Krishnan. 2008. The EM algorithm and extensions. John Wiley & Sons
2008
-
[43]
Yuning Mao, Yanru Qu, Yiqing Xie, Xiang Ren, and Jiawei Han. 2020. Multi- document Summarization with Maximal Marginal Relevance-guided Reinforce- ment Learning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 1737–1751
2020
-
[44]
Tran, and Cong Yu
Richard Yuanzhe Pang, Ádám Dániel Lelkes, Vinh Q. Tran, and Cong Yu. 2021. AgreeSum: Agreement-Oriented Multi-Document Summarization. In Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021 (Findings of ACL, Vol. ACL/IJCNLP...
2021 doi
-
[45]
Ramakanth Pasunuru, Mengwen Liu, Mohit Bansal, Sujith Ravi, and Markus Dreyer. 2021. Efficiently Summarizing Text and Graph Encodings of Multi- Document Clusters. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguisti...
2021
-
[46]
Şaziye Betül Özateş, Arzucan Özgür, and Dragomir Radev. 2016. Sentence sim- ilarity based on dependency tree kernels for multi-document summarization. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16). 2833–2838
2016
-
[47]
Yutong Qu. 2024. Leveraging Knowledge-aware Methodologies for Multi- document Summarization. In Companion Proceedings of the ACM on Web Con- ference 2024, WWW 2024, Singapore, Singapore, May 13-17, 2024 , Tat-Seng Chua, Chong-Wah Ngo, Roy Ka-Wei Lee, Ravi Kumar, and Hady W. La...
2024
-
[48]
Radev, Hong Qi, Harris Wu, and Weiguo Fan
Dragomir R. Radev, Hong Qi, Harris Wu, and Weiguo Fan. 2002. Evaluating Web-based Question Answering Systems. InProceedings of the Third International Conference on Language Resources and Evaluation (LREC)
2002
-
[49]
Ratish Surendran Puduppully, Parag Jain, Nancy Chen, and Mark Steedman. 2023. Multi-Document Summarization with Centroid-Based Pretraining. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, (ACL: Short Papers). 128–138
2023
-
[50]
Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Em- pirical Methods in Natural Language Processing . Association for Computational Linguistics. https://arxiv.org/abs/1908.10084
2019 arXiv
-
[51]
Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme
-
[52]
Thilina Chaturanga Rajapakse, Andrew Yates, and Maarten de Rijke. 2024. Nega- tive Sampling Techniques for Dense Passage Retrieval in a Multilingual Setting. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIG...
2024
-
[53]
Stuart Rose, Dave Engel, Nick Cramer, and Wendy Cowley. 2010. Automatic keyword extraction from individual documents. Text mining: applications and theory (2010), 1–20
2010
-
[54]
Le, Geoffrey E
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In 5th International Conference on Learning Representations (ICLR)
2017
-
[55]
Zejiang Shen, Kyle Lo, Lauren Yu, Nathan Dahlberg, Margo Schlanger, and Doug Downey. 2022. Multi-LexSum: Real-World Summaries of Civil Rights Lawsuits at Multiple Granularities. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Pr...
2022
-
[56]
Stephen Robertson and Hugo Zaragoza. 2009. The Probabilistic Relevance Frame- work: BM25 and Beyond. Found. Trends Inf. Retr. 3, 4 (2009), 333–389
2009
-
[57]
Zhengliang Shi, Shen Gao, Zhen Zhang, Xiuying Chen, Zhumin Chen, Pengjie Ren, and Zhaochun Ren. 2023. Towards a Unified Framework for Reference Retrieval and Related Work Generation. In Findings of the Association for Compu- tational Linguistics: EMNLP. 5785–5799
2023
-
[58]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurPS). 6000–6010
2017
-
[59]
Pancheng Wang, Shasha Li, Dong Li, Kehan Long, Jintao Tang, and Ting Wang
-
[60]
Chi, Nathanael Schärli, and Denny Zhou
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, and Denny Zhou. 2023. Large Language Models Can Be Easily Distracted by Irrelevant Context. In International Conference on Machine Learning (ICML), Vol. 202. PMLR, 31210–31227
2023
-
[61]
Wen Xiao, Iz Beltagy, Giuseppe Carenini, and Arman Cohan. 2022. PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document Summariza- tion. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL). 5245–5263
2022
-
[62]
Wenhao Yu, Hongming Zhang, Xiaoman Pan, Peixin Cao, Kaixin Ma, Jian Li, Hongwei Wang, and Dong Yu. 2024. Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP...
2024
-
[63]
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization. In Proceedings of the 37th International Conference on Machine Learning (ICML) , Vol. 119. PMLR, 11328–11339
2020
-
[64]
Weijia Zhang, Svitlana Vakulenko, Thilina Rajapakse, Yumo Xu, and Evangelos Kanoulas. 2021. Tackling query-focused summarization as a knowledge-intensive task: A pilot study. arXiv preprint arXiv:2112.07536 (2021)
2021 arXiv
-
[65]
Ruben Wolhandler, Arie Cattan, Ori Ernst, and Ido Dagan. 2022. How “Multi” is Multi-Document Summarization?. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 5761–5769
2022
-
[70]
Zixuan Zhang, Heba Elfardy, Markus Dreyer, Kevin Small, Heng Ji, and Mohit Bansal. 2023. Enhancing Multi-Document Summarization with Cross-Document Graph-based Information Extraction. In Proceedings of the 17th Conference of the European Chapter of the Association for Computat...
2023
-
[2009]
In Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence (UAI)
BPR: Bayesian personalized ranking from implicit feedback. In Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence (UAI). AUAI Press, 452–461
-
[2020]
In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL)
Leveraging Graph to Improve Abstractive Multi-Document Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL). 6232–6243
-
[2021]
InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP)
MSˆ2: Multi-Document Summarization of Medical Studies. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP). 7494–7513
2021
-
[2024]
InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR)
Disentangling Instructive Information from Ranked Multiple Candidates for Multi-Document Scientific Summarization. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR). ACM, 2028–2037
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.