REVIEW 5 major objections 4 minor 56 references
Generating an Overview Report over Many Documents
T0 review · 5 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A five-module pipeline called NDORGS generates structured overview reports over thousands of documents, and its 20% summary-length setting ranks best overall.
desk verdict A coherent pipeline for a real task, but the headline claim about lambda=0.2 is not reproducible until the AHP weight matrix is reported; raw human scores on BBC favor lambda=0.3. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the $\lambda$-summary: each original document is first compressed by a single-document summarizer to a fixed length fraction $\lambda$ of itself, and that compressed set is what the rest of the pipeline clusters and re-summarizes. $\lambda$ acts as the information budget for the whole report, because all later stages see only these summaries. The two-level topic hierarchy is built by LDA, with second-level sub-clusters created whenever a top-level cluster exceeds a preset size; each node's cluster score orders sections by salience. GFLOW then produces a coherent multi-document summary for each cluster, DTATG generates the section and subsection titles, and the evaluation uses a 9-point pairwise-comparison weighting of the four criteria feeding into TOPSIS, which ranks the three $\lambda$ alternatives by closeness to an ideal solution.
What would settle it
Recompute the TOPSIS ranking for the classified corpus with the pairwise-comparison weight vector made explicit: because the 0.3 report has the highest human-evaluation score (4.03 versus 3.71), there is a computable threshold on the human-evaluation weight at which the 0.3 report overtakes the 0.2 report, and the claim that the decision is stable fails if that threshold lies inside the weight range covered by the sensitivity analysis.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that the NDORGS configuration — filtering plus Semantic WordRank single-document summarization at $\lambda = 0.2$, two-level hierarchical clustering with LDA, GFLOW cluster summarization, DTATG title generation, and salience-ordered report assembly — is the best overall choice on both test corpora. On BBC News the 0.2 report has the best information-coverage score (0.54), the best topic-diversity score (0.1444), and the middle human-evaluation score; on Factiva-Marx it has the best human-evaluation score (4.03) and the best diversity score (0.1167). The paper therefore claims that $\lambda = 0.2$ is the best overall setting under its four-criteria ordering, and that the TOPSIS ranking is stable when criterion weights are adjusted one at a time.
Load-bearing premise
The claim that the 20% setting is best overall rests on the authors' four evaluation criteria and the weights assigned through pairwise comparisons; if those weights shift toward human readability, the 20% report no longer wins on at least one dataset.
Editorial extensions
If this is right
- A corpus of several thousand documents can be reduced to a bounded two-level overview report in hours of CPU time on a commonplace desktop, with sections ordered by salience and each section carrying an automatically generated title.
- For teams that prioritize speed, $\lambda = 0.1$ is the fastest configuration, while for readability alone the BBC News human scores favor $\lambda = 0.3$; the $\lambda = 0.2$ recommendation is specifically the best balance of the four weighted criteria.
- The observation that the much larger unclassified corpus preferred a smaller $\lambda$ suggests that optimal compression should scale down as corpus size and document length grow, so deployments on even larger corpora should start near $\lambda = 0.2$ or below.
- The evaluation protocol itself — standard multi-document summarization quality scoring plus text-mining proxies for coverage and diversity, aggregated by TOPSIS — can be reused to compare other summarization pipelines when reading the whole corpus is infeasible.
Reading between the lines
- The 'best overall' label is tied to the authors' criterion ordering: because human annotators scored the 0.3 report above the 0.2 report on the classified corpus (4.03 vs 3.71), a stakeholder who weights readability heavily would likely prefer $\lambda = 0.3$, so the result should be read as best under the stated weights, not best in an absolute sense.
- Treating $\lambda$ as an information budget suggests a principled extension: set $\lambda$ as a function of corpus size, average document length, and desired report length rather than grid-searching three fixed values.
- Because coverage and diversity are measured as top-word overlap and cluster-overlap scores, both metrics reward lexical and topical resemblance to the corpus; substituting a semantic or human-rated novelty metric could reorder the three candidates and is a cheap robustness test.
- The pipeline is modular, so the $\lambda$ result is not a claim about any single component: swapping the single-document summarizer, clustering method, multi-document summarizer, or title generator while keeping $\lambda = 0.2$ could change the ranking, and jointly searching module choice and $\lambda$ is a natural next experiment.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents NDORGS, a five-module pipeline (preprocessing, hierarchical topic clustering, cluster summarizing, cluster titling, and report generating) for producing a two-level overview report from thousands of documents. The pipeline uses Semantic WordRank for single-document summaries at length ratios 0.1, 0.2, and 0.3, LDA-based hierarchical clustering, GFLOW for cluster summaries, and DTATG for title generation. The authors evaluate six generated reports (three length ratios on BBC News and three on Factiva-Marx) using DUC-style human ratings, running time, and self-defined information coverage and topic diversity scores, then combine these four criteria with AHP-derived weights and TOPSIS to claim that 0.2-summaries produce the best overall report on both datasets. The paper also includes a sensitivity analysis and trending graphs.
Significance. If the evaluation were fully specified and reproducible, the paper would contribute a useful modular architecture for large-corpus overview generation and a concrete practical finding about the SDS compression ratio in such pipelines. The provision of the six generated reports at ndorg.net is a strength, as is the explicit use of a multi-criteria decision framework rather than a single metric. However, the headline result currently rests on an unreported AHP weight vector and on self-defined metrics that have not been validated, so the practical recommendation is not yet reliably established.
major comments (5)
- [§4.5, Table 4] The headline claim that 0.2-summaries are best overall is not reproducible because the Saaty pairwise comparison matrix and the resulting AHP weight vector are never reported. The raw scores in Table 4 do not imply the stated ranking: on BBC News, the 0.3 report has the best human evaluation (4.03 vs. 3.71), the 0.1 report has the best running time (3310 s vs. 5060 s), and 0.2 wins only on the two self-defined text-mining metrics; on Factiva-Marx, 0.2 wins on human and diversity but loses on time and coverage. Consequently, which alternative TOPSIS selects is entirely determined by the four undisclosed weights. The sensitivity analysis in Fig. 9 starts from an unstated base weight vector and perturbs weights in 0.02 increments, so it demonstrates only local stability; as the weight on human evaluation approaches 1, TOPSIS must converge to the human-only ranking, which on BBC News places 0.3 first. Please report the full pairwise comparison matrix, the derived weight vector, and the region of the weight simplex in which 0.2 is ranked first.
- [§4.5, Human evaluations] The human evaluation component has four annotators per report and no inter-annotator agreement measure or significance test. The raw scores in Appendix Table 5 show large disagreement, for example BBC-0.1 receives scores of 1, 2, 3, and 4 on the UselessText criterion, so the reported means may not be stable. Since human evaluation is the highest-weighted criterion in the AHP/TOPSIS procedure, the absence of agreement statistics and inferential tests undermines the comparative claim between λ=0.2 and λ=0.3. Please report per-criterion annotator agreement, such as Krippendorff's alpha, and at minimum a paired or bootstrap test for the human-score differences.
- [§4.5, Information coverage and Topic diversity] The information coverage score S_k(A,B) and the topic diversity score based on CSD F1 are author-defined measures, and the value of k is never stated in Section 4.5 even though the appendix lists top 50 words. These two metrics are the only criteria on which the 0.2 report beats the 0.3 report on BBC News, so the central ranking depends on metrics that have not been validated against reading quality or shown to be stable under the choice of k. Please report the exact k used, justify the choice, and provide evidence that these text-mining proxies correlate with human judgments or at least are robust to k in a reasonable range.
- [§3, Step 2] The cluster score formula is garbled: it reads SC = 1/n^2 ∑_{j=i} 2p_i, with an undefined summation index and an unexplained factor of 2. Because cluster scores determine the ordering of sections and subsections in the Report Generating module, this is part of the algorithm's specification and must be corrected, presumably to a sum over i of 2p_i or similar, with p_i's role clarified. Without a correct formula, the structure of the generated reports cannot be reproduced.
- [§4, Evaluations] The evaluation compares only three configurations of NDORGS (λ=0.1, 0.2, and 0.3) and does not include any baseline system, such as T-DMCA [25], a flat MDS pipeline, or a random/lead-based summary. Consequently, the experiments cannot support the contribution that NDORGS is an effective approach for large-corpus overview generation; they only rank the system's own compression choices. Please add at least one competitive baseline or a meaningful ablation, such as a non-hierarchical MDS report or a pipeline without title generation, and evaluate it with the same criteria.
minor comments (4)
- [§1 and §2.2.2] The system name is written as T-CMDA in the Introduction but as T-DMCA in Section 2.2.2; please use a single consistent acronym for reference [25].
- [§4.2] The text says 'Parameters α, β, and λ are learned using the DUC-03 dataset,' but λ is elsewhere the SDS length ratio that the authors set to 0.1, 0.2, and 0.3, and the displayed ILP formulation does not contain λ. Please clarify what λ means in the GFLOW parameters and how this relates to the three fixed values used in the experiments.
- [Fig. 9 caption] The caption states that the x-axis shows increments/decrements of 0.02 each time, but the plotted x-axis values appear to be absolute weight values starting near 0. Please mark the base weight vector and the direction of perturbation on the figure.
- [Appendix Table 5] Table 5 lists individual annotator scores, but the main text reports a single human-evaluation mean in Table 4; please state explicitly how the seven criterion scores are aggregated into that single number and whether all seven criteria are weighted equally.
Circularity Check
No circularity: empirical pipeline comparison; unreported AHP weights are a reproducibility gap, not a circular derivation.
full rationale
The paper makes an empirical, not derivational, claim: it constructs three NDORGS variants that differ only in the SDS length ratio lambda in {0.1, 0.2, 0.3}, measures them on four criteria, and applies AHP/TOPSIS to rank the alternatives. No equation in the paper defines the superiority of lambda = 0.2 in terms of the input data; the lambda values are genuinely varied, and the raw scores in Table 4 do not by construction force the TOPSIS ranking. The main limitation is that the Saaty pairwise-comparison matrix and the resulting AHP weight vector are never reported, so the TOPSIS ranking is not fully reproducible and the 'best overall' label is conditional on an unstated weight choice. However, an unreported parameter is a transparency or validity issue, not a circular reduction: the conclusion does not reduce to the inputs by definition, and no fitted parameter is renamed as a prediction. The self-citations (Semantic WordRank [47] and DTATG [40]) are used as component algorithms and are supported by external evaluations on DUC-02, SummBank, and human-judged title adequacy, so they are not load-bearing self-referential support. The coverage and diversity metrics are author-defined but are measured against the corpus and clusters rather than being defined as 'whatever makes lambda = 0.2 win.' No circular step meeting the evidentiary bar was found.
Assumptions & free parameters
free parameters (6)
- K (number of top-level clusters) =
9
- N (second-level split threshold) =
200
- MDS cluster summary length coefficients =
150/200 and 300 formula
- TOPSIS/AHP weights =
not fully reported
- k for information coverage score =
not stated (likely 50)
- lambda (SDS length ratio) =
0.1, 0.2, 0.3
assumptions (5)
- domain assumption Only a small part of the most important content of an article will ultimately contribute to the overview report, so SDS summaries are sufficient (Section 3, Step 1).
- domain assumption A higher cluster score S_C represents a more significant topic (Section 3, Step 2 and RG module).
- domain assumption The number of top-level sections in an overview report should not exceed 10 (Section 4.4).
- domain assumption Clustering can be performed on either original documents or SDS summaries, and the input choice for the final reports is not explicitly stated (Section 3, Step 2 and Section 4.4).
- domain assumption GFLOW's ILP formulation with parameters learned on DUC-03 remains valid for the corpora and cluster sizes in this paper (Section 4.2).
Cite this review
Pith. "Pith review of Generating an Overview Report over Many Documents." pith.science (2026). https://pith.science/paper/7KWGNCX2
@misc{pith2026190806216,
author = {Pith},
title = {Pith review of: Generating an Overview Report over Many Documents},
year = {2026},
howpublished = {\url{https://pith.science/paper/7KWGNCX2}},
note = {Machine review of arXiv:1908.06216}
}
read the original abstract
How to efficiently generate an accurate, well-structured overview report (ORPT) over thousands of related documents is challenging. A well-structured ORPT consists of sections of multiple levels (e.g., sections and subsections). None of the existing multi-document summarization (MDS) algorithms is directed toward this task. To overcome this obstacle, we present NDORGS (Numerous Documents' Overview Report Generation Scheme) that integrates text filtering, keyword scoring, single-document summarization (SDS), topic modeling, MDS, and title generation to generate a coherent, well-structured ORPT. We then devise a multi-criteria evaluation method using techniques of text mining and multi-attribute decision making on a combination of human judgments, running time, information coverage, and topic diversity. We evaluate ORPTs generated by NDORGS on two large corpora of documents, where one is classified and the other unclassified. We show that, using Saaty's pairwise comparison 9-point scale and under TOPSIS, the ORPTs generated on SDS's with the length of 20% of the original documents are the best overall on both datasets.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[25]
P. J. Liu, M. Saleh, E. Pot, B. Goodrich, R. Sepassi, L. Kaiser, and N. Shazeer. Generating wikipedia by summarizing long sequences. CoRR, abs/1801.10198, 2018
arXiv 2018
-
[1]
D. Bahdanau, K. Cho, and Y . Bengio. Neural machine translation by jointly learning to align and translate. CoRR, abs/1409.0473, 2014
arXiv 2014
-
[2]
D. M. Blei, A. Y . Ng, and M. I. Jordan. Latent dirichlet allocation. In Journal of Machine Learning Research, 2003
work page 2003
-
[3]
O. Buyukkokten, H. Garcia-Molina, and A. Paepcke. Seeing the whole in parts: Text summarization for web browsing on handheld devices. In Proceedings of the 10th International Conference on World Wide Web, WWW ’01, pages 652–662. ACM, 2001
work page 2001
-
[4]
Z. Cao, W. Li, S. Li, and F. Wei. Improving multi- document summarization via text classification. In AAAI, pages 3053–3059, 2017
work page 2017
-
[5]
Z. Cao, W. Li, S. Li, F. Wei, and Y . Li. Attsum: Joint learning of focusing and summarization with neural attention. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguis- tics: Technical Papers, pages 547–556, 2016
work page 2016
-
[6]
J. Christensen, Mausam, S. Soderland, and O. Etzioni. Towards coherent multi-document summarization. In HLT-NAACL, 2013
work page 2013
-
[7]
J. Christensen, S. Soderland, G. Bansal, et al. Hier- archical summarization: Scaling up multi-document summarization. In Proceedings of the 52nd annual meeting of the association for computational linguis- tics, volume 1, pages 902–912, 2014
work page 2014
Show all 56 references
-
[8]
Duc2004 quality questions
DUC. Duc2004 quality questions. http: //duc.nist.gov/duc2004/quality. questions.txt, 2004
2004
-
[9]
Document understanding confer- ence
DUC. Document understanding confer- ence. https://www-nlpir.nist.gov/ projects/duc/intro.html, 2014
2014
-
[10]
Erkan and D
G. Erkan and D. R. Radev. Lexrank: Graph-based lex- ical centrality as salience in text summarization. Jour- nal of Artificial Intelligence Research , 22:457–479, 2004
2004
-
[11]
Marx dataset
Factiva-Marx. Marx dataset. http://www. ndorg.net, 2018
2018
-
[12]
W. Gao, P. Li, and K. Darwish. Joint topic model- ing for event summarization across news and social media streams. In Proceedings of the 21st ACM in- ternational conference on Information and knowledge management, pages 1173–1182. ACM, 2012
2012
-
[13]
Gillick and B
D. Gillick and B. Favre. A scalable global model for summarization. In Proceedings of the Workshop on In- teger Linear Programming for Natural Langauge Pro- cessing, pages 10–18. Association for Computational Linguistics, 2009
2009
-
[14]
Greene and P
D. Greene and P. Cunningham. Practical solutions to the problem of diagonal dominance in kernel docu- ment clustering. In Proc. 23rd International Confer- ence on Machine learning (ICML’06), pages 377–384. ACM Press, 2006
2006
-
[15]
J. A. Hartigan and M. A. Wong. Algorithm as 136: A k-means clustering algorithm. Journal of the Royal Statistical Society. Series C (Applied Statistics), 28(1):100–108, 1979
1979
-
[16]
K. Hong, J. M. Conroy, B. Favre, A. Kulesza, H. Lin, and A. Nenkova. A repository of state of the art and competitive baseline summaries for generic news summarization. In LREC, pages 1608–1616, 2014
2014
-
[17]
Hwang and K
C.-L. Hwang and K. Yoon. Methods for multiple at- tribute decision making. In Multiple attribute decision making, pages 58–191. Springer, 1981
1981
-
[18]
D. Jones. Factiva global news database. https:// www.dowjones.com/products/factiva/, 2018
2018
-
[19]
T. N. Kipf and M. Welling. Semi-supervised clas- sification with graph convolutional networks. arXiv preprint arXiv:1609.02907, 2016
2016 arXiv
-
[20]
M. Lapata. Probabilistic text structuring: Experiments with sentence ordering. In Proceedings of the 41st Annual Meeting on Association for Computational Linguistics-Volume 1, pages 545–552. Association for Computational Linguistics, 2003. 15
2003
-
[21]
C. Li, Y . Lu, J. Wu, Y . Zhang, Z. Xia, T. Wang, D. Yu, X. Chen, P. Liu, and J. Guo. Lda meets word2vec: A novel model for academic abstract clus- tering. In Companion of the The Web Conference 2018 on The Web Conference 2018, pages 1699–1706. International World Wide Web Con...
2018
-
[22]
C. Li, X. Qian, and Y . Liu. Using supervised bigram- based ilp for extractive summarization. In Proceed- ings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , volume 1, pages 1004–1013, 2013
2013
-
[23]
P. Li, J. Jiang, and Y . Wang. Generating templates of entity summaries with an entity-aspect model and pat- tern mining. In Proceedings of the 48th Annual Meet- ing of the Association for Computational Linguistics , ACL ’10, pages 640–649. Association for Computa- tional Ling...
2010
-
[24]
S. Li, Y . Ouyang, W. Wang, and B. Sun. Multi- document summarization using support vector regres- sion. In Proceedings of DUC. Citeseer, 2007
2007
-
[26]
Lulli, T
A. Lulli, T. Debatty, M. Dell’Amico, P. Michiardi, and L. Ricci. Scalable k-nn based text clustering. In Big Data (Big Data), 2015 IEEE International Conference on, pages 958–963. IEEE, 2015
2015
-
[27]
Mihalcea and P
R. Mihalcea and P. Tarau. Textrank: Bringing or- der into texts. In Proceedings of the 2004 conference on empirical methods in natural language processing, 2004
2004
-
[28]
Nallapati, B
R. Nallapati, B. Xiang, and B. Zhou. Sequence- to-sequence rnns for text summarization. CoRR, abs/1602.06023, 2016
2016 arXiv
-
[29]
Nallapati, F
R. Nallapati, F. Zhai, and B. Zhou. Summarunner: A recurrent neural network based sequence model for extractive summarization of documents. In AAAI, pages 3075–3081, 2017
2017
-
[30]
M. T. Nayeem and Y . Chali. Extract with order for coherent multi-document summarization. In Proceed- ings of TextGraphs-11: the Workshop on Graph-based Methods for Natural Language Processing, pages 51–
-
[31]
A. Y . Ng, M. I. Jordan, and Y . Weiss. On spectral clustering: Analysis and an algorithm. In Advances in neural information processing systems, pages 849– 856, 2002
2002
-
[32]
Otterbacher, D
J. Otterbacher, D. Radev, and O. Kareem. News to go: hierarchical text summarization for mobile devices. In Proc. of ACM SIGIR, pages 589–596, 2006
2006
-
[33]
L. Page, S. Brin, R. Motwani, and T. Winograd. The pagerank citation ranking: Bringing order to the web. Technical report, Stanford InfoLab, 1999
1999
-
[34]
D. R. Radev, H. Jing, M. Sty, and D. Tam. Centroid- based summarization of multiple documents. Inf. Pro- cess. Manage., 40:919–938, 2004
2004
-
[35]
S. J. Rose, D. Engel, N. Cramer, and W. Cowley. Auto- matic keyword extraction from individual documents. In In book: Text Mining: Applications and Theory, pp.1 - 20, 2010
2010
-
[36]
A. M. Rush, S. Chopra, and J. Weston. A neural at- tention model for abstractive sentence summarization. CoRR, abs/1509.00685, 2015
2015 arXiv
-
[37]
Salton and C
G. Salton and C. Buckley. Term-weighting approaches in automatic text retrieval. Inf. Process. Manage. , 24:513–523, 1988
1988
-
[38]
Sauper and R
C. Sauper and R. Barzilay. Automatically generating wikipedia articles: A structure-aware approach. In Proceedings of the Joint Conference of the 47th An- nual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP: Volume 1, ...
2009
-
[39]
Shang, J
J. Shang, J. Liu, M. Jiang, X. Ren, C. R. V oss, and J. Han. Automated phrase mining from massive text corpora. IEEE Transactions on Knowledge and Data Engineering, 30(10):1825–1837, 2018
2018
-
[40]
Shao and J
L. Shao and J. Wang. Dtatg: An automatic title gen- erator based on dependency trees. In Proceedings of the International Joint Conference on Knowledge Dis- covery, Knowledge Engineering and Knowledge Man- agement, IC3K 2016, pages 166–173, Portugal, 2016. SCITEPRESS - Science...
2016
-
[41]
X. Wang, M. Nishino, T. Hirao, K. Sudoh, and M. Na- gata. Exploring text links for coherent multi-document summarization. In Proceedings of COLING 2016, the 26th International Conference on Computational Lin- guistics: Technical Papers, pages 213–223, 2016
2016
-
[42]
J. Xu, P. Wang, G. Tian, B. Xu, J. Zhao, F. Wang, and H. Hao. Short text clustering via convolutional neural 16 Table 5. Human evaluation Scores, where C1, C2,..., C7 represent, respectively, Coher- ence, UselessText, Redundancy, Referents, OverlyExplicit, Grammatical, and For...
2015
-
[43]
C. Yao, X. Jia, S. Shou, S. Feng, F. Zhou, and H. Liu. Autopedia: Automatic domain-independent wikipedia article generation. In Proceedings of the 20th Inter- national Conference Companion on World Wide Web, WWW ’11, pages 161–162. ACM, 2011
2011
-
[44]
Yasunaga, R
M. Yasunaga, R. Zhang, K. Meelu, A. Pareek, K. Srinivasan, and D. R. Radev. Graph-based neural multi-document summarization. In CoNLL, 2017
2017
-
[45]
Yin and J
J. Yin and J. Wang. A dirichlet multinomial mixture model-based approach for short text clustering. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 233–242. ACM, 2014
2014
-
[46]
Yogatama, F
D. Yogatama, F. Liu, and N. A. Smith. Extractive sum- marization by maximizing semantic volume. In Pro- ceedings of the 2015 Conference on Empirical Meth- ods in Natural Language Processing , pages 1961– 1966, 2015
2015
-
[47]
Zhang and J
H. Zhang and J. Wang. Semantic WordRank: Gener- ating Finer Single-Document Summarizations. ArXiv e-prints, Sept. 2018. A Appendix A.1 Human evaluation scores A.2 Comparisons of top words Top words are listed below in the corpus of BBC News and the corpus of Factiva-Marx, resp...
2018
-
[49]
people, told, best, government, time, year, num- ber, three, film, music, bbc, set, game, going, years, labour, good, well, top, british, european, win, mar- ket, won, company, public, second, play,mobile, work, firm, blair, games, minister, expected, england, chief,technology, ...
-
[50]
people, best, number, government, film, year, three, game, howard, music, london, british, face, biggest, net, action, firm, deal, rise, national, foreign, singer, michael, leader, oil, blair , dollar, stock, star, cup, on- line, future, games, 2004, work, won, list, interna- ti...
2004
-
[51]
year, people, number, three, best, british, film, company, won, labour, music, net, bbc, government, leader, shares, european, earlier, chart, third, games, state, win , coach, expected , second , months, politi- cal, house, economic, game, years, team, start, manch- ester, eng...
-
[52]
people, england , year, film, labour, boss, firm, de- spite, number, three, wales, british, nations, best, company, music, blair , set, record, oil, time, years , won, prices, plans, net, online, including, films, bbc , court, games, game, brown, david, government, expected, club...
-
[53]
party, chinese, china, political, people, communist, economic, national, state, government, years, so- cial, great, time , rights, development , international , president, central, war, north, university , power, united, work, country , foreign, global, military, history, sout...
-
[54]
party, china, chinese, communist, political, years, economic, rights, human, president, people, na- tional, year, state, leaders, government, central, countries, news, social, country, leader, time, foreign, power, north, nuclear, top,marxism, ideological, led, media, war, bei...
-
[55]
china, party, communist, chinese, political, eco- nomic, years, rights, president, people, central, state, social, united, north, beijing, western, news, media, mao, cpc, war, human, anniversary, public , members, country , jinping, leader, states , govern- ment, south, marxis...
-
[56]
Association for Computational Linguistics, 2017
2017
-
[57]
china, party, communist, chinese, economic, years, people, political, human, news, state, social, gov- ernment, central, national, leader, president, me- dia, cultural, rights, mao, power, development, year, international, university, leaders, history, united, bei- jing, copyr...
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.