REVIEW 3 major objections 5 minor 54 references
Domain Regeneration: How well do LLMs match syntactic properties of text domains?
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read LLMs regenerate text domains as compressed, less diverse versions of the human originals, with reduced variance and a reduced syntactic long tail.
desk verdict A useful observational benchmark, but the headline claim about syntactic simplification is confounded by sentence length; needs a length-matched control before the diversity finding is believable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the regeneration paradigm itself: feed the first 256 words of a Wikipedia article, the first 180 words of a news article, or the title of an ELI5 thread into an instruction-tuned LLM, ask it to complete the text in that domain's style, and compare the resulting corpus with the human original. The comparison runs through a set of syntactic measures—Flesch-Kincaid grade level, sentence length, number of unique dependency tags, constituency parse depth, number of unique constituency labels, and Yngve score, which is the average number of left branches from the root to a leaf and so measures departure from right-branching structure. Three distributional signatures carry the argument: mean shift, narrowing of the standard deviation, and reduction of the long right tail. Their joint pattern is what separates a compressed domain representation from a neutral fallback.
What would settle it
Compare the parse-failure rate per sentence-length bin for human and model corpora, and re-run the analysis with failed or zero-depth parses included rather than dropped. If the human long tail shrinks toward the model tail once those sentences are counted, the reported long-tail reduction is partly a measurement artifact; if it survives, the compression is a property of generation.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that LLM regeneration does not resample a human text domain faithfully: for the majority of metrics and datasets, the model distribution is narrower and its long tail is reduced, while its mean shifts in the direction that makes the domain more extreme. ELI5 regenerations come out simpler, while Wikipedia and CCNews regenerations come out more complex, which the authors take as evidence that models encode a genuine notion of domain rather than collapsing to a neutral middle ground. The reduced variance and shortened tail mean that the models underproduce rare syntactic configurations, such as strongly left-branching structures reflected in high Yngve scores, even when average complexity is roughly matched. The paper concludes that LLMs hold a compressed, not fully humanlike, representation of domain syntax.
Load-bearing premise
The paper's result depends on the assumption that sentences dropped by the parser—very long or very deeply nested ones—fail at comparable rates for human and model text; if human text loses more of its extreme tail to parse failures, the reduced long tail is partly an artifact of measurement.
Editorial extensions
If this is right
- Synthetic-text detection can exploit the three signatures—mean shift, lower variance, and reduced long tail—as domain-conditional evidence rather than relying only on a single global style measure.
- Prompted domain regeneration will not simply interpolate between domains: for complex domains like Wikipedia and news it will overshoot the human complexity level, so downstream systems that rely on model-written domain-style text should expect harder text than the original.
- Rare recursive structures such as strongly left-branching sentences will be systematically underrepresented in regenerated corpora, which matters for corpus-based syntax research that uses LLM text as a stand-in for human data.
- The domain-conditional nature of the mean shift implies that models preserve register distinctions through post-training, so instruction tuning and alignment have not collapsed all domains into one neutral style.
Reading between the lines
- An immediate extension would hold the prompt fixed and vary only the decoding temperature; the paper uses a single temperature, so whether the compression is partly a decoding artifact remains untested.
- Cross-lingual replication with a predominantly left-branching language would separate a general compression bias from an English-specific syntactic prior: if high-Yngve tails also vanish where left-branching is canonical, the effect concerns low-probability syntax rather than English alone.
- A parser that never fails on deep trees, applied to the same corpora, would settle the main measurement caveat and could be run without any new data collection.
- The ELI5 depth exception suggests a floor effect: once a domain is already simple, variance cannot shrink much further; sampling even simpler domains would test whether the compression asymptotes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces an 'LLM-regeneration' protocol: the authors take the first 256 words of Wikipedia articles, the first 180 words of CCNews articles, or the title of an ELI5 thread, prompt open-weight instruction-tuned Llama and Mistral models to continue the article in the target style, and then compare the regenerated corpus to the original human corpus on five distributional metrics: Flesch-Kincaid grade level, unique dependency tags per sentence, constituency parse depth, Yngve score, and unique constituency labels per sentence. The main descriptive finding is that regenerated distributions are generally narrower, have shifted means, and have a reduced right tail compared with the human originals; for mean complexity, the models appear to overshoot the domain, simplifying ELI5 while making Wikipedia and CCNews more complex. The paper interprets these patterns as evidence that models encode a real but compressed, not fully humanlike, notion of text domain.
Significance. If the central claim is established, this is a useful and timely empirical contribution: it connects the study of syntactic diversity and long-tail phenomena to practical questions of model fidelity, synthetic-text detection, and model collapse. The paper has notable strengths: it uses several open-weight model families, large-scale corpora, a publicly reproducible toolchain (Stanza, vLLM), and an explicit cleaning ablation in Appendix A showing that the reported trends persist under less aggressive filtering. The qualitative examples in Appendices G and H are also valuable for grounding the distributional results. However, the main interpretive claim—that the reduced variance and lost long tail reflect syntactic properties of LLM generation—is not yet supported because the analysis does not control for sentence length, which is correlated with or bounds every metric studied.
major comments (3)
- [§3.1, Table 2, Figure 14, and Appendix B] The paper never conditions on sentence length, yet the headline signatures—reduced variance and a reduced right tail—are already present in the sentence-length distributions themselves, and every syntactic metric is either a function of length or bounded by it: Flesch-Kincaid is computed from words per sentence, unique dependency tags and unique constituency labels are bounded by token count, and parse depth and Yngve score grow with sentence length. The regeneration protocol in Appendix B also imposes target article lengths and fixed seed lengths, so the length difference is partly an experimental choice. To support the central claim that LLMs produce syntactically less diverse text, the authors should compare human and model syntax within matched length strata or adjust for length quantitatively; without this, the reduced diversity and lost long tail are not independently established for any metric.
- [§2.3 and §7] The filtering of sentences that fail Stanza parsing and of zero depth/Yngve scores is applied to both human and model corpora, but the manuscript does not report failure rates by source or by sentence length. If long or deeply nested human sentences fail parsing more often than model-generated sentences, then the observed reduction of the human long tail in the depth and Yngve metrics is partly a measurement artifact. The authors should quantify exclusion rates by source and length, and rerun the depth and Yngve analyses on successfully parsed sentences within matched length strata.
- [§3.7 and Table 1] The summary that 'the majority of metrics and datasets' show reduced variability and reduced long tails is based on visual inspection of normal-fit curves; no effect sizes, confidence intervals, or tail-mass statistics are reported. Given the extremely large sample sizes, even negligible differences would be statistically significant with a conventional test, so the meaningful quantity is the size of the distributional shift. Reporting standardized effect sizes, distribution-overlap measures, or explicit tail-mass differences would make the cross-domain and cross-model claims falsifiable and would align the paper with best practice for observational corpus comparisons.
minor comments (5)
- [Abstract and §1] The abstract says the work investigates 'three domains' while the introduction at the start of §1 refers to 'two domains' before listing three; please make the count consistent.
- [Throughout] There are frequent typos such as 'Flesh-Kincaid' for 'Flesch-Kincaid', 'Intruct' for 'Instruct', and 'on Wikipedia writing style' / 'on news writing style'; a careful proofreading pass is needed.
- [§2.5 footnote 5] The sentence describing why readability scores were calculated without data cleaning is grammatically unclear; please revise to state the threshold and rationale explicitly.
- [§3.1 and Appendix C] Figure 14 is the first place where sentence-length distributions are described, but it appears in an appendix; consider moving a version to the main text since it is central to the interpretation.
- [§3.7] The claim that the models do not land at a 'neutral' middle ground would be stronger with an explicit neutral-domain regeneration baseline; as written, the conclusion relies on comparing domain-specific regenerations to each other rather than to a control condition.
Circularity Check
No significant circularity: the paper is an observational benchmark comparing LLM regenerations against external human corpora, with no fitted parameter, self-imported uniqueness theorem, or definitionally forced prediction.
full rationale
This paper does not derive any quantity from itself. The central comparison is between distributions of standard corpus-linguistic metrics computed identically on human-written text and on LLM-regenerated text: Flesch-Kincaid grade level, unique dependency tags, constituency parse depth, Yngve scores, and unique constituency labels. These metrics are established external instruments, not quantities defined in terms of the paper's conclusions, and no parameter is fitted to the human data and then renamed as a prediction. The regeneration protocol, including prompt seeds and minimum article lengths, is an experimental design choice; it constrains the generation setup but does not logically force any of the reported signatures (mean shift, reduced variance, or truncated long tail) to appear. The paper explicitly acknowledges that complexity metrics can correlate with sentence length (Section 3.1, citing Salkar et al. 2022) and that parser failures are possible (Section 7), but these are stated as caveats about measurement and interpretation, not as premises from which the results are derived. The possible length confound raised by an external reader is a threat to causal interpretation, not a circularity: observing a narrower length distribution does not make the reduced syntactic variance true by construction. Self-citations to Ju et al. (2024) and Williams et al. (2021) are methodological borrowings for the regeneration approach and parsing pipeline, respectively, and the paper's empirical claims do not depend on the truth of those prior papers' conclusions. The findings are benchmarked against external human corpora, and the qualitative analyses and ablations provide independent checks. Accordingly, there is no load-bearing step that reduces to its own inputs, and the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Prompt prefix length =
Wikipedia 256 words, CCNews 180 words
- Target article length in prompt =
700 words (Wikipedia), 500 words (CCNews), 100 words (ELI5)
- Decoding temperature =
1.0
assumptions (3)
- domain assumption Wikipedia, CCNews, and ELI5 are internally consistent, well-circumscribed domains and are plausibly contained in LLM training data.
- domain assumption Stanza parsing errors and exclusions affect human and regenerated corpora comparably.
- domain assumption Prompting with the opening paragraph plus style instructions yields semantically-controlled regenerations.
Cite this review
Pith. "Pith review of Domain Regeneration: How well do LLMs match syntactic properties of text domains?." pith.science (2026). https://pith.science/paper/N2L5ADO3
@misc{pith2026250507784,
author = {Pith},
title = {Pith review of: Domain Regeneration: How well do LLMs match syntactic properties of text domains?},
year = {2026},
howpublished = {\url{https://pith.science/paper/N2L5ADO3}},
note = {Machine review of arXiv:2505.07784}
}
read the original abstract
Recent improvement in large language model performance have, in all likelihood, been accompanied by improvement in how well they can approximate the distribution of their training data. In this work, we explore the following question: which properties of text domains do LLMs faithfully approximate, and how well do they do so? Applying observational approaches familiar from corpus linguistics, we prompt a commonly used, opensource LLM to regenerate text from two domains of permissively licensed English text which are often contained in LLM training data -- Wikipedia and news text. This regeneration paradigm allows us to investigate whether LLMs can faithfully match the original human text domains in a fairly semantically-controlled setting. We investigate varying levels of syntactic abstraction, from more simple properties like sentence length, and article readability, to more complex and higher order properties such as dependency tag distribution, parse depth, and parse complexity. We find that the majority of the regenerated distributions show a shifted mean, a lower standard deviation, and a reduction of the long tail, as compared to the human originals.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Danial Alihosseini, Ehsan Montahaei, and Mahdieh Soleymani Baghshah. 2019. https://doi.org/10.18653/v1/W19-2311 Jointly measuring diversity and quality in text generation models . In Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation, pages 90--98, Minneapolis, Minnesota. Association for Computational Linguistics
-
[4]
Kevin Bache, David Newman, and Padhraic Smyth. 2013. Text-based measures of document diversity. In Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 23--31
work page 2013
-
[5]
Douglas Biber. 1991. Variation across speech and writing. Cambridge University Press
1991
-
[6]
BigScience Workshop , Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili\' c , Daniel Hesslow, Roman Castagn\' e , Alexandra Sasha Luccioni, Fran c ois Yvon, Matthias Gall' e , Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova de...
arXiv 2023
-
[7]
Steven Bird, Ewan Klein, and Edward Loper. 2009. Natural language processing with Python: analyzing text with the natural language toolkit. O'Reilly Media, Incorporated
work page 2009
-
[8]
Iacer Calixto, Alessandro Raganato, and Tommaso Pasini. 2021. https://doi.org/10.18653/v1/2021.naacl-main.286 W ikipedia entities as rendezvous across languages: Grounding multilingual language models by predicting W ikipedia hyperlinks . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: ...
Show all 54 references
-
[9]
Danqi Chen and Christopher Manning. 2014. https://doi.org/10.3115/v1/D14-1082 A fast and accurate dependency parser using neural networks . In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 740--750, Doha, Qatar. Associ...
2014 doi
-
[10]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[11]
Nigel Dewdney, Carol VanEss-Dykema, and Richard MacMillan. 2001. https://aclanthology.org/W01-1007/ The form is the substance: Classification of genres in text . In Proceedings of the ACL 2001 Workshop on Human Language Technology and Knowledge Management
2001
-
[12]
Chrysanne DiMarco and Graeme Hirst. 1993. https://aclanthology.org/J93-3002/ A computational theory of goal-directed style in syntax . Computational Linguistics, 19(3):451--500
1993
-
[13]
Emily Dinan, Angela Fan, Ledell Wu, Jason Weston, Douwe Kiela, and Adina Williams. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.23 Multi-dimensional gender bias classification . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMN...
2020 doi
-
[14]
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019. W izard of W ikipedia: Knowledge-powered conversational agents. In Proceedings of the International Conference on Learning Representations (ICLR)
2019
-
[15]
Elvis Dohmatob, Yunzhen Feng, Arjun Subramonian, and Julia Kempe. 2024. https://arxiv.org/abs/2410.04840 Strong model collapse
2024 arXiv
-
[16]
Liat Ein-Dor, Ariel Gera, Orith Toledo-Ronen, Alon Halfon, Benjamin Sznajder, Lena Dankin, Yonatan Bilu, Yoav Katz, and Noam Slonim. 2019. https://doi.org/10.18653/v1/D19-5102 Financial event extraction using W ikipedia-based weak supervision . In Proceedings of the Second Wor...
2019 doi
-
[17]
Julian Eisenschlos, Bhuwan Dhingra, Jannis Bulian, Benjamin B \"o rschinger, and Jordan Boyd-Graber. 2021. https://doi.org/10.18653/v1/2021.naacl-main.32 Fool me twice: Entailment from W ikipedia gamification . In Proceedings of the 2021 Conference of the North American Chapte...
2021 doi
-
[18]
Angela Fan and Claire Gardent. 2022. https://doi.org/10.18653/v1/2022.acl-long.586 Generating biographies on W ikipedia: The impact of gender bias on the retrieval-based generation of women biographies . In Proceedings of the 60th Annual Meeting of the Association for Computat...
2022 doi
-
[19]
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. https://doi.org/10.18653/v1/P19-1346 ELI 5: Long form question answering . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3558--356...
2019 doi
-
[20]
Rudolph Flesch. 1948. A new readability yardstick. Journal of applied psychology, 32(3):221
1948
-
[21]
Robert Gunning. 1952. The technique of clear writing
1952
-
[22]
Sil Hamilton. 2024. https://aclanthology.org/2024.scalellm-1.5/ Detecting mode collapse in language models via narration . In Proceedings of the First edition of the Workshop on the Scaling Behavior of Large Language Models (SCALE-LLM 2024), pages 65--72, St. Julian ' s, Malta...
2024
-
[23]
Colby Horn, Cathryn Manduca, and David Kauchak. 2014. https://doi.org/10.3115/v1/P14-2075 Learning a lexical simplifier using W ikipedia . In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 458--463, Balti...
2014 doi
-
[24]
J. D. Hunter. 2007. https://doi.org/10.1109/MCSE.2007.55 Matplotlib: A 2d graphics environment . Computing in Science & Engineering, 9(3):90--95
2007 doi
-
[25]
Dieuwke Hupkes, Mario Giulianelli, Verna Dankers, Mikel Artetxe, Yanai Elazar, Tiago Pimentel, Christos Christodoulopoulos, Karim Lasri, Naomi Saphra, Arabella Sinclair, et al. 2023. A taxonomy and review of generalization research in nlp. Nature Machine Intelligence, 5(10):1161--1174
2023
-
[26]
Da Ju, Karen Ullrich, and Adina Williams. 2024. https://doi.org/10.18653/v1/2024.findings-acl.253 Are female carpenters like blue bananas? a corpus investigation of occupation gender typicality . In Findings of the Association for Computational Linguistics: ACL 2024, pages 425...
2024 doi
-
[27]
Juzek and Zina B
Tom S. Juzek and Zina B. Ward. 2024. https://arxiv.org/abs/2412.11385 Why does chatgpt "delve" so much? exploring the sources of lexical overrepresentation in large language models
2024 arXiv
-
[28]
Marcus Klang and Pierre Nugues. 2019. https://aclanthology.org/W19-6148/ D ocria: Processing and storing linguistic data with W ikipedia . In Proceedings of the 22nd Nordic Conference on Computational Linguistics, pages 400--405, Turku, Finland. Link \"o ping University Electr...
2019
-
[29]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating S...
2023
-
[30]
Jack Lanchantin, Angelica Chen, Shehzaad Dhuliawala, Ping Yu, Jason Weston, Sainbayar Sukhbaatar, and Ilia Kulikov. 2025. https://arxiv.org/abs/2501.18101 Diverse preference optimization
2025 arXiv
-
[31]
David Lee. 2002. Genres, registers, text types, domains and styles: clarifying the concepts and navigating a path through the bnc jungle. In Teaching and learning by doing corpus analysis, pages 245--292. Brill
2002
-
[32]
Dianqi Li, Yizhe Zhang, Zhe Gan, Yu Cheng, Chris Brockett, Bill Dolan, and Ming-Ting Sun. 2019. https://doi.org/10.18653/v1/D19-1325 Domain adaptive text style transfer . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Inte...
2019 doi
-
[33]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. https://arxiv.org/abs/1907.11692 Roberta: A robustly optimized bert pretraining approach
2019 arXiv
-
[34]
AI@Meta Llama Team. 2024. The L lama 3 H erd of M odels
2024
-
[35]
Jian Ni and Radu Florian. 2016. https://doi.org/10.18653/v1/D16-1135 Improving multilingual named entity recognition with W ikipedia entity type mapping . In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1275--1284, Austin, Texas...
2016 doi
-
[36]
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. https://doi.org/10.18653/v1/2020.acl-main.441 Adversarial NLI : A new benchmark for natural language understanding . In Proceedings of the 58th Annual Meeting of the Association for Comp...
2020 doi
-
[37]
Vishakh Padmakumar and He He. 2023. Does writing with language models reduce content diversity? arXiv preprint arXiv:2309.05196
2023 arXiv
-
[38]
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rockt \"a schel, and Sebastian Riedel. 2021. https://doi.org/10.18653/v1/2021.naacl-main.200 KIL...
2021 doi
-
[39]
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020. Stanza: A python natural language processing toolkit for many human languages
2020
-
[40]
Brian Roark, Margaret Mitchell, and Kristy Hollingshead. 2007. https://aclanthology.org/W07-1001/ Syntactic complexity measures for detecting mild cognitive impairment . In Biological, translational, and clinical language processing, pages 1--8, Prague, Czech Republic. Associa...
2007
-
[41]
Peters, Swabha Swayamdipta, and Thomas Wolf
Sebastian Ruder, Matthew E. Peters, Swabha Swayamdipta, and Thomas Wolf. 2019. https://doi.org/10.18653/v1/N19-5004 Transfer learning in natural language processing . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Ling...
2019 doi
-
[42]
Jenna Russell, Marzena Karpinska, and Mohit Iyyer. 2025. https://arxiv.org/abs/2501.15654 People who frequently use chatgpt for writing tasks are accurate and robust detectors of ai-generated text
2025 arXiv
-
[43]
Nikita Salkar, Thomas Trikalinos, Byron Wallace, and Ani Nenkova. 2022. https://doi.org/10.18653/v1/2022.aacl-short.42 Self-repetition in abstractive neural summarizers . In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Ling...
2022 doi
-
[44]
Sina Semnani, Violet Yao, Heidi Zhang, and Monica Lam. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.157 W iki C hat: Stopping the hallucination of large language model chatbots by few-shot grounding on W ikipedia . In Findings of the Association for Computational Ling...
2023 doi
-
[45]
Chantal Shaib, Yanai Elazar, Junyi Jessy Li, and Byron C Wallace. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.368 Detection and measurement of syntactic templates in generated text . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processin...
2024 doi
-
[46]
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennigh...
2024
-
[47]
George Spache. 1953. A new readability formula for primary-grade reading materials. The Elementary School Journal, 53(7):410--413
1953
-
[48]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...
2023
-
[49]
Michael L. Waskom. 2021. https://doi.org/10.21105/joss.03021 seaborn: statistical data visualization . Journal of Open Source Software, 6(60):3021
2021 doi
-
[50]
Adina Williams, Ryan Cotterell, Lawrence Wolf-Sonkin, Dami \'a n Blasi, and Hanna Wallach. 2021. https://doi.org/10.1162/tacl_a_00355 On the relationships between the grammatical genders of inanimate nouns and their co-occurring adjectives and verbs . Transactions of the Assoc...
2021 doi
-
[51]
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. https://doi.org/10.18653/v1/N18-1101 A broad-coverage challenge corpus for sentence understanding through inference . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computation...
2018 doi
-
[52]
Fei Wu and Daniel S. Weld. 2010. https://aclanthology.org/P10-1013/ Open information extraction using W ikipedia . In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, pages 118--127, Uppsala, Sweden. Association for Computational Linguistics
2010
-
[53]
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. https://doi.org/10.18653/v1/D18-1259 H otpot QA : A dataset for diverse, explainable multi-hop question answering . In Proceedings of the 2018 Conference...
2018 doi
-
[54]
Victor H Yngve. 1960. A model and an hypothesis for language structure. Proceedings of the American philosophical society, 104(5):444--466
1960
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.