REVIEW 3 major objections 5 minor 227 references
Computational Approaches to Understanding Large Language Model Impact on Writing and Information Ecosystems
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This dissertation claims that the fraction of text substantially modified by large language models can be estimated at the population level by fitting a two-component word-frequency mixture model, and reports that 6.5% to 16.9% of…
desk verdict A well-structured compilation of published work; the method is sound and the validation is extensive, but the headline LLM-adoption numbers rest on an untested stationarity assumption and should be read as upper bounds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the two-component mixture model with maximum-likelihood estimation over word-occurrence statistics. The paper defines a vocabulary of adjectives, chosen for stability over adverbs, verbs, and nouns, and models each sentence's probability as the product of per-token occurrence probabilities under the human distribution $P$ and under the LLM distribution $Q$; the corpus log-likelihood is $\sum_i \log((1-\alpha)P(x_i)+\alpha Q(x_i))$, maximized to estimate $\alpha$. What makes it work is the empirical observation that words such as 'commendable', 'meticulous', and 'intricate' rose sharply in frequency in post-ChatGPT review corpora while staying flat for years beforehand, so the estimator detects a distributional shift toward LLM-flavored adjectives rather than attempting to label any one sentence.
What would settle it
Collect a matched corpus of post-ChatGPT, human-only peer reviews from a community that demonstrably avoided LLM assistance, with topics and author demographics matched to the machine-learning conferences studied; if the estimation method still reports a share of LLM-modified text well above its measured error when the true share is zero, the pre-ChatGPT human baseline is shifting for reasons other than LLM modification.
Extended reading notes
Core claim
On its own terms, the central discovery is that the fraction of a corpus that has been substantially modified by an LLM can be recovered from the relative frequencies of a small vocabulary of words without classifying a single document. The model assumes every sentence is drawn from $(1-\alpha)P + \alpha Q$, where $P$ is the human-written reference distribution and $Q$ the LLM distribution, both estimated from labeled reference corpora; maximum likelihood over the occurrence probabilities yields $\alpha$. On semi-synthetic blends of official human reviews and LLM-generated reviews, the estimator recovers the true $\alpha$ within 0 to 2.4 percentage points across in-distribution and out-of-distribution venues. On real post-ChatGPT corpora it finds a sharp rise in $\alpha$ for machine-learning conference reviews (ICLR 2024 at 10.6%, EMNLP 2023 at 16.9%, NeurIPS 2023 at 9.1%, CoRL 2023 at 6.5%) but no significant rise in Nature-portfolio reviews, and it shows that proofreading alone cannot explain the increase while expanding a bulleted outline into full review text can. The dissertation also reports that reviews with higher estimated LLM modification are more likely to be submitted near the deadline, to carry low self-rated confidence, to omit citations, and to sit close to the centroid of all reviews of the same paper.
Load-bearing premise
The load-bearing assumption is that the human writing patterns estimated from pre-ChatGPT reviews still describe human writing after ChatGPT's launch; if reviewers shifted toward AI-flavored adjectives because of exposure, style change, or topic drift rather than direct LLM use, the estimated share of LLM-modified text would overstate true LLM modification.
Editorial extensions
If this is right
- A substantial minority of peer-review sentences at top machine-learning venues, between 6.5% and 16.9%, were substantially modified by LLMs in the first post-ChatGPT cycle, with the share varying by venue and review behavior.
- LLM-assisted text is not evenly distributed: it concentrates near deadlines, in low-confidence reviews, in reviews without citations, and in reviews that converge toward the average review, implying measurable homogenization of feedback.
- The same estimator finds double-digit shares of LLM-modified sentences in consumer complaints, corporate press releases, job postings, and UN press releases after ChatGPT's launch, so AI-assisted writing is not confined to academia.
- Individual-level GPT detectors flag non-native English writing as AI-generated at high rates and can be bypassed with a single self-edit prompt, so institutions that rely on them will both penalize non-native writers and miss AI text.
- LLM-generated feedback on scientific manuscripts overlaps with human reviewer comments at rates comparable to reviewer-reviewer agreement, about 30-39% versus 28-35%, and is paper-specific, suggesting a viable complement to human review.
Reading between the lines
- A consequence the dissertation leaves implicit: the telltale adjective set it identifies could be used as a diachronic tracer, allowing later corpora to be dated or audited for LLM influence even if the originating model changes.
- A testable extension is to apply the estimator to any domain that has a stable pre-2022 human-written archive, such as news op-eds, student essays, policy documents, or clinical notes, and compare adoption rates across them; the method's requirement is only a matched human baseline and a plausible LLM reference corpus.
- Because the estimator reads word frequencies, prompt-engineering to avoid AI-flavored adjectives would push the measured share down; if that became common practice, the current figures would be a lower bound on true LLM modification rather than an upper bound.
- The homogenization correlation suggests a direct test: within venues where the estimated LLM share rises, the average pairwise similarity of reviews of the same paper should increase over time; if it does not, the convergence signal may reflect reviewer demographics or topic shift rather than LLM use.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The dissertation develops computational methods for measuring and characterizing the impact of large language models on writing and information ecosystems. Chapter 2 shows that widely used GPT detectors systematically misclassify non-native English writing as AI-generated and that simple prompting can bypass detectors. Chapter 3 introduces a maximum-likelihood 'distributional GPT quantification' framework: using pre-ChatGPT human-written reviews and ChatGPT-generated reference reviews to estimate the token-occurrence distributions P and Q, the method estimates α, the fraction of sentences in a target corpus substantially modified by an LLM. Applying this to ML conference peer reviews, the dissertation reports α estimates of 6.5%–16.9% after ChatGPT's release, with higher estimates near deadlines, for low-confidence reviewers, and for reviews without citations. Chapter 4 extends the framework to scientific abstracts and introductions across arXiv, bioRxiv, and Nature portfolio journals, showing rising LLM-modified content after 2022. Chapter 5 applies the same method to consumer complaints, corporate press releases, job postings, and UN press releases, reporting double-digit percentages of LLM-modified sentences. Chapter 6 evaluates GPT-4's ability to provide scientific feedback, finding substantial overlap with human reviewer comments and generally positive assessments in a prospective user study.
Significance. If the identification assumptions hold, this is an important and timely contribution. The population-level estimation framework is genuinely novel: it avoids instance-level detection, is many orders of magnitude cheaper than classifier-based detectors, and is validated extensively on semi-synthetic mixtures, under prompt shift, across parts of speech, for proofreading and outline-expansion use cases, and with alternative LLMs. The temporal patterns across multiple independent corpora—peer reviews, scientific abstracts, press releases, and complaints—are internally consistent and make a strong circumstantial case that LLM-assisted writing increased sharply after late 2022. The dissertation also provides a valuable cautionary result on detector bias against non-native writers. However, the central quantitative claims rest on an untested identification assumption: that the pre-ChatGPT human distribution P remains the correct counterfactual for post-ChatGPT human writing. The paper's own limitations section acknowledges temporal distribution shift and changing non-native-speaker populations as potential error sources without bounding them.
major comments (3)
- [§3.3.1–3.3.5, Eq. (3.3), §3.4.4] The central identification assumption is that the human token-occurrence distribution P, estimated from pre-ChatGPT reviews, remains the correct distribution for human-written text after November 30, 2022. The post-ChatGPT increase in α is then attributed to LLM adoption, but the counterfactual human distribution is unobserved. If human style drifted toward AI-flavored adjectives (e.g., 'commendable', 'meticulous', 'intricate') through exposure to LLM output, changing review norms, or reviewer-pool shifts, the estimated α would overstate true LLM modification. The semi-synthetic validation in Section 3.3.6 mixes pre-ChatGPT human reviews with the same ChatGPT-generated Q used for estimation, so it verifies the estimator only under the model's own generative assumptions. This is the load-bearing point for the headline 6.5–16.9% claim, and it needs a direct sensitivity analysis or an external post-ChatGPT human-only reference corpus.
- [Chapter 4, Fig. 4.3] The temporal-split validation in Section 4.2.2 is the strongest validation in the dissertation because it separates training data (up to 2020) from validation data (January–November 2022) by more than a year and still achieves estimation error below 3.5%. However, this split stops before ChatGPT's launch, so it does not test the regime where human style drift from LLM exposure is most plausible. To support the population-level adoption estimates in Chapters 4 and 5, the author needs either a post-ChatGPT human-only validation corpus (e.g., writing from communities or venues with verified zero LLM use, or pre-registered human-written text) or a formal bound on the bias in α under a plausible range of human style drift.
- [§3.5 Limitations] The Limitations section concedes that 'temporal distribution shift in token frequencies due to, e.g., changes in topics, reviewers, etc.' introduces error and that substantial shifts in the non-native speaker population could affect accuracy, but it provides no quantification or bound. Since the magnitude of the headline estimates (6.5–16.9%) is comparable to the estimated baseline false-positive rate of 1.6–2.4% plus the post-ChatGPT increase of roughly 5–15 percentage points, an unquantified drift of even a few percentage points could materially change the conclusions. The manuscript should report a sensitivity analysis that perturbs P in the direction of observed post-ChatGPT token-frequency shifts and shows how α changes.
minor comments (5)
- [§2.2.3] There is a typo: 'non-native English witters' should be 'non-native English writers', and the sentence 'A lot of Room of improvement, it is crucial to develop more robust detection methods' is grammatically incomplete and should be revised.
- [§2.3] The word 'discrepency' should be 'discrepancy'.
- [Chapter 3, Figure 3.4] The caption states 'EMNLP '23' as an orange dot, but the figure legend is small and the point is easy to miss; consider adding a label directly or a table reference in the caption.
- [§3.4.4] The sentence 'Among the conferences with pre- and post-ChatGPT data, ICLR experienced the most significant increase' is slightly confusing because NeurIPS also has pre- and post-ChaGPT data; clarifying that ICLR showed the largest absolute increase would improve readability.
- [Chapter 5, Fig. 5.1] The figure reports adoption 'plateauing at 17.7% through August 2024' for consumer complaints, but the discussion in §5.2.2 describes geographic and demographic disparities; a brief note on how the national estimate relates to the state-level estimates in Fig. 5.3 would help the reader connect the two analyses.
Circularity Check
No by-construction circularity: alpha is an MLE output on new corpora, validated on held-out OOD and temporally split data, and anchored by raw word-frequency discontinuities and a Nature negative control; the style-drift concern is an acknowledged identification limitation, not a circular reduction.
full rationale
Derivation chain: Eqs. (3.1)-(3.3) define the mixture model; P and Q are estimated from pre-ChatGPT human reviews and LLM-generated reference reviews (Sec. 3.3.4-3.3.5); Sec. 3.3.6 validates the MLE on held-out corpora (ICLR 2023, NeurIPS 2022, CoRL 2022, Nature 2022) at known alpha with <2.4% error; Sec. 3.4.4 applies the estimator to post-ChatGPT corpora disjoint from training and validation. Alpha is an output, not a fitted input renamed as a prediction: the central finding is visible in raw adjective-frequency discontinuities at the ChatGPT launch ('commendable', 'meticulous', 'intricate' showing 9.8, 34.7, and 11.2-fold increases, Fig. 3.1), absent in a contemporaneous negative control (Nature portfolio alpha stays within the alpha=0 margin of error, Fig. 3.4), and robust under prompt shift (Table 3.9), model shift (Tables 3.28-3.29, 3.32-3.33), and a temporal split with training up to 2020 and validation in 2022 (Fig. 4.3). The reader's concern (post-launch human style drift inflating alpha) is a genuine identification assumption about the counterfactual P, but nothing in Eqs. (3.1)-(3.3) forces the post-ChatGPT shift to be attributed to LLM use; the paper's own Limitations (Sec. 3.5) concede 'the temporal distribution shift in token frequencies due to, e.g., changes in topics, reviewers, etc.' and likely 'shifts in the non-native speaker population' as error sources. An untested counterfactual with frank acknowledgement is a correctness risk, not a by-construction reduction. Chapter 6's GPT-4-as-evaluator design is likewise not circular: the extraction/matching pipeline is human-verified (Table 6.3: 639 feedbacks, 12,035 pairs) and the shuffle control drops the hit rate from 57.55% to 1.13%, excluding generic-text confounds. Self-citations (Chapters 2-6 adapt the author's own ICML/Patterns papers) are pervasive but not load-bearing, since every method is re-derived in-chapter. Verdict: no significant circularity; score 2 reflects only the minor, non-load-bearing self-citation pattern.
Assumptions & free parameters
free parameters (3)
- Alpha (MLE estimate of LLM-modified sentence fraction) =
0-25% depending on corpus; e.g., 10.6% ICLR 2024, 16.9% EMNLP 2023
- Vocabulary selection (adjectives, excluding technical terms) =
Adjectives; technical terms removed
- Pre-ChatGPT baseline alpha (false positive rate) =
1-2% for ML conferences
assumptions (5)
- domain assumption Documents/sentences in the target corpus are generated by the mixture model (1-alpha)P + alpha Q (Eq. 3.1)
- domain assumption P is stationary over time; pre-ChatGPT human writing is representative of post-ChatGPT human writing absent LLM use
- domain assumption Token occurrences are independent within a document (naive Bayes factorization in Eq. 3.3)
- domain assumption GPT-4's semantic matching in Chapter 6 reliably identifies overlapping comments between LLM and human feedback
- standard math Standard MLE consistency and bootstrap inference
Cite this review
Pith. "Pith review of Computational Approaches to Understanding Large Language Model Impact on Writing and Information Ecosystems." pith.science (2026). https://pith.science/paper/6MHK4U4A
@misc{pith2026250617467,
author = {Pith},
title = {Pith review of: Computational Approaches to Understanding Large Language Model Impact on Writing and Information Ecosystems},
year = {2026},
howpublished = {\url{https://pith.science/paper/6MHK4U4A}},
note = {Machine review of arXiv:2506.17467}
}
read the original abstract
Large language models (LLMs) have shown significant potential to change how we write, communicate, and create, leading to rapid adoption across society. This dissertation examines how individuals and institutions are adapting to and engaging with this emerging technology through three research directions. First, I demonstrate how the institutional adoption of AI detectors introduces systematic biases, particularly disadvantaging writers of non-dominant language varieties, highlighting critical equity concerns in AI governance. Second, I present novel population-level algorithmic approaches that measure the increasing adoption of LLMs across writing domains, revealing consistent patterns of AI-assisted content in academic peer reviews, scientific publications, consumer complaints, corporate communications, job postings, and international organization press releases. Finally, I investigate LLMs' capability to provide feedback on research manuscripts through a large-scale empirical analysis, offering insights into their potential to support researchers who face barriers in accessing timely manuscript feedback, particularly early-career researchers and those from under-resourced settings.
Figures
Figures from the paper (106 more)
Reference graph
Works this paper leans on
-
[1]
Simons Institute Talk on Watermarking of Large Language Models, 2023
Scott Aaronson. Simons Institute Talk on Watermarking of Large Language Models, 2023
2023
-
[2]
An online platform for interactive feedback in biomedical machine learning.Nature Machine Intelligence, 2(2):86–88, 2020
Abubakar Abid, Ali Abdalla, Ali Abid, Dawood Khan, Abdulrahman Alfozan, and James Zou. An online platform for interactive feedback in biomedical machine learning.Nature Machine Intelligence, 2(2):86–88, 2020
2020
-
[3]
A billion-dollar donation: estimating the cost of researchers’ time spent on peer review.Research Integrity and Peer Review, 6(1):1–8, 2021
Balazs Aczel, Barnabas Szaszi, and Alex O Holcombe. A billion-dollar donation: estimating the cost of researchers’ time spent on peer review.Research Integrity and Peer Review, 6(1):1–8, 2021
2021
-
[4]
Reviewing peer review, 2008
Bruce Alberts, Brooks Hanson, and Katrina L Kelner. Reviewing peer review, 2008
2008
-
[5]
Claude 2
Anthropic. Claude 2. Anthropic News, 2023. Accessed: 2023-04-25
2023
-
[6]
Arnav Arora, Lucie-Aimée Kaffee, and Isabelle Augenstein. Probing pre-trained language models for cross-cultural differences in values.arXiv preprint arXiv:2203.13722, 2022
arXiv 2022
-
[7]
ACL-IJCNLP 2021 Instructions for Reviewers
Association for Computational Linguistics. ACL-IJCNLP 2021 Instructions for Reviewers. https://2021.aclweb.org/blog/instructions-for-reviewers/, 2021
2021
-
[8]
ACL’23 Peer Review Policies.https://2023
Association for Computational Linguistics. ACL’23 Peer Review Policies.https://2023. aclweb.org/blog/review-acl23/, 2023
2023
Show all 227 references
-
[9]
Atallah, Victor Raskin, Michael Crogan, Christian F
Mikhail J. Atallah, Victor Raskin, Michael Crogan, Christian F. Hempelmann, Florian Ker- schbaum, Dina Mohamed, and Sanket Naik. Natural Language Watermarking: Design, Analysis, and a Proof-of-Concept Implementation. InInformation Hiding, 2001
2001
-
[10]
Triezenberg
MikhailJ.Atallah, VictorRaskin, ChristianF.Hempelmann, MercanTopkara, RaduSion, Umut Topkara, and Katrina E. Triezenberg. Natural Language Watermarking and Tamperproofing. In Information Hiding, 2002
2002
-
[11]
Comparing 176 BIBLIOGRAPHY 177 physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum.JAMA internal medicine, 2023
John W Ayers, Adam Poliak, Mark Dredze, Eric C Leas, Zechariah Zhu, Jessica B Kelley, DennisJFaix, AaronMGoodman, ChristopherALonghurst, MichaelHogarth, etal. Comparing 176 BIBLIOGRAPHY 177 physician and artificial intelligence chatbot responses to patient questions posted to ...
2023
-
[12]
Artificial intelligence, firm growth, and product innovation.Journal of Financial Economics, 151:103745, January 2024
Tania Babina, Anastassia Fedyk, Alex He, and James Hodson. Artificial intelligence, firm growth, and product innovation.Journal of Financial Economics, 151:103745, January 2024
2024
-
[13]
Identifying Real or Fake Articles: Towards betterLanguageModeling
Sameer Badaskar, Sachin Agarwal, and Shilpa Arora. Identifying Real or Fake Articles: Towards betterLanguageModeling. In International Joint Conference on Natural Language Processing, 2008
2008
-
[14]
Real or Fake? Learning to Discriminate Machine from Human Generated Text.ArXiv, abs/1906.03351, 2019
Anton Bakhtin, Sam Gross, Myle Ott, Yuntian Deng, Marc’Aurelio Ranzato, and Arthur Szlam. Real or Fake? Learning to Discriminate Machine from Human Generated Text.ArXiv, abs/1906.03351, 2019
1906 arXiv
-
[15]
Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, and Yue Zhang. Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature. ArXiv, abs/2310.05130, 2023
2023 arXiv
-
[16]
Discourses of artificial intelligence in higher education: A critical literature review.Higher Education, 86(2):369–385, 2023
Margaret Bearman, Juliana Ryan, and Rola Ajjawi. Discourses of artificial intelligence in higher education: A critical literature review.Higher Education, 86(2):369–385, 2023
2023
-
[17]
Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Smitchell
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Smitchell. On the dangers of stochastic parrots: Can language models be too big? InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 610–623. ACM, March 2021
2021
-
[18]
Computer-Generated Text Detection Using Machine Learning: A System- atic Review
Daria Beresneva. Computer-Generated Text Detection Using Machine Learning: A System- atic Review. InInternational Conference on Applications of Natural Language to Data Bases, 2016
2016
-
[19]
Selection bias in web surveys.International statistical review, 78(2):161–188, 2010
Jelke Bethlehem. Selection bias in web surveys.International statistical review, 78(2):161–188, 2010
2010
-
[20]
Rahul Bhagat and Eduard H. Hovy. Squibs: What Is a Paraphrase?Computational Linguistics, 39:463–472, 2013
2013
-
[21]
ConDA:Contrastive Domain Adaptation for AI-generated Text Detection.ArXiv, abs/2309.03992, 2023
AmritaBhattacharjee, TharinduKumarage, RahaMoraffah, andHuanLiu. ConDA:Contrastive Domain Adaptation for AI-generated Text Detection.ArXiv, abs/2309.03992, 2023
2023 arXiv
-
[22]
Drivers and barriers of ai adoption and use in scientific research.arXiv preprint arXiv:2312.09843, 2023
Stefano Bianchini, Moritz Müller, and Pierre Pelletier. Drivers and barriers of ai adoption and use in scientific research.arXiv preprint arXiv:2312.09843, 2023. BIBLIOGRAPHY 178
2023 arXiv
-
[23]
Should we use characteristics of conversation to measure grammatical complexity in L2 writing development?Tesol Quarterly, 45(1):5–35, 2011
Douglas Biber, Bethany Gray, and Kornwipa Poonpon. Should we use characteristics of conversation to measure grammatical complexity in L2 writing development?Tesol Quarterly, 45(1):5–35, 2011
2011
-
[24]
Alexander Bick, Adam Blandin, and David J. Deming. The rapid adoption of generative ai. Working Paper 32966, National Bureau of Economic Research, September 2024. Revised February 2025
2024
-
[25]
The values encoded in machine learning research
Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. The values encoded in machine learning research. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’22, page 173–184, New York, NY, USA, 202...
2022
-
[26]
The publishing delay in scholarly peer-reviewed journals
Bo-Christer Björk and David Solomon. The publishing delay in scholarly peer-reviewed journals. Journal of informetrics, 7(4):914–923, 2013
2013
-
[27]
Are ideas getting harder to find?American Economic Review, 110(4):1104–1144, 2020
Nicholas Bloom, Charles I Jones, John Van Reenen, and Michael Webb. Are ideas getting harder to find?American Economic Review, 110(4):1104–1144, 2020
2020
-
[28]
Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization? Advances in Neural Information Processing Systems, 35:3663–3678, 2022
Rishi Bommasani, Kathleen A Creel, Ananya Kumar, Dan Jurafsky, and Percy S Liang. Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization? Advances in Neural Information Processing Systems, 35:3663–3678, 2022
2022
-
[29]
Bernstein, et al
Rishi Bommasani, Daniel Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, et al. On the opportunities and risks of foundation models.arXiv preprint arXiv:2108.07258, 2022
2022 arXiv
-
[30]
Cultural reproduction and social reproduction
Pierre Bourdieu. Cultural reproduction and social reproduction. InKnowledge, education, and cultural change, pages 71–112. Routledge, 2018
2018
-
[31]
A large annotated corpus for learning natural language inference.arXiv preprint arXiv:1508.05326, 2015
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. A large annotated corpus for learning natural language inference.arXiv preprint arXiv:1508.05326, 2015
2015 arXiv
-
[32]
The rise of ai-generated content in wikipedia, 2024
Creston Brooks, Samuel Eggert, and Denis Peskoff. The rise of ai-generated content in wikipedia, 2024
2024
-
[33]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020
1901
-
[34]
Generative ai at work
Erik Brynjolfsson, Danielle Li, and Lindsey Raymond. Generative ai at work. Working Paper w31161, National Bureau of Economic Research, April 2023. NBER Working Paper No. 31161. BIBLIOGRAPHY 179
2023
-
[35]
Nearly 50 news websites are ‘AI-generated’, a study says
Matthew Cantor. Nearly 50 news websites are ‘AI-generated’, a study says. Would I be able to tell?, 2023. Accessed: 2024-02-24
2023
-
[36]
Assessing Cross-Cultural Alignment between ChatGPT and Human Societies: An Empirical Study, 2023
Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. Assessing Cross-Cultural Alignment between ChatGPT and Human Societies: An Empirical Study, 2023
2023
-
[37]
On the possibilities of ai-generated text detection.arXiv preprint arXiv:2304.04736, 2023
Souradip Chakraborty, Amrit Singh Bedi, Sicheng Zhu, Bang An, Dinesh Manocha, and Furong Huang. On the possibilities of ai-generated text detection.arXiv preprint arXiv:2304.04736, 2023
2023 arXiv
-
[38]
GPT- Sentinel: Distinguishing Human and ChatGPT Generated Content.ArXiv, abs/2305.07969, 2023
Yutian Chen, Hao Kang, Vivian Zhai, Liang Li, Rita Singh, and Bhiksha Ramakrishnan. GPT- Sentinel: Distinguishing Human and ChatGPT Generated Content.ArXiv, abs/2305.07969, 2023
2023 arXiv
-
[39]
Natu- ral Language Watermarking Using Semantic Substitution for Chinese Text
Yuei-Lin Chiang, Lu-Ping Chang, Wen-Tai Hsieh, and Wen-Chih Chen. Natu- ral Language Watermarking Using Semantic Substitution for Chinese Text. In International Workshop on Digital Watermarking, 2003
2003
-
[40]
What data can do: A typology of mechanisms
Angèle Christin. What data can do: A typology of mechanisms. International Journal of Communication, 14:20, 2020
2020
-
[41]
Slowed canonical progress in large fields of science
Johan SG Chu and James A Evans. Slowed canonical progress in large fields of science. Proceedings of the National Academy of Sciences, 118(41):e2021636118, 2021
2021
-
[42]
All That’s ‘Human’Is Not Gold: Evaluating Human Evaluation of Generated Text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith. All That’s ‘Human’Is Not Gold: Evaluating Human Evaluation of Generated Text. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11t...
2021
-
[43]
How ChatGPT and other AI tools could disrupt scientific publishing.Nature, October 2023
Gemma Conroy. How ChatGPT and other AI tools could disrupt scientific publishing.Nature, October 2023
2023
-
[44]
Scientific sleuths spot dishonest ChatGPT use in papers.Nature, September 2023
Gemma Conroy. Scientific sleuths spot dishonest ChatGPT use in papers.Nature, September 2023
2023
-
[45]
Does writing development equal writing quality? A computational investigation of syntactic complexity in L2 learners.Journal of Second Language Writing, 26:66–79, 2014
Scott A Crossley and Danielle S McNamara. Does writing development equal writing quality? A computational investigation of syntactic complexity in L2 learners.Journal of Second Language Writing, 26:66–79, 2014
2014
-
[46]
Machine Generated Text: A Compre- hensive Survey of Threat Models and Detection Methods.arXiv preprint arXiv:2210.07321, 2022
Evan Crothers, Nathalie Japkowicz, and Herna Viktor. Machine Generated Text: A Compre- hensive Survey of Threat Models and Detection Methods.arXiv preprint arXiv:2210.07321, 2022. BIBLIOGRAPHY 180
2022 arXiv
-
[47]
Lexical richness in the spontaneous speech of bilinguals.Applied linguistics, 24(2):197–222, 2003
Helmut Daller, Roeland Van Hout, and Jeanine Treffers-Daller. Lexical richness in the spontaneous speech of bilinguals.Applied linguistics, 24(2):197–222, 2003
2003
-
[48]
Fred D. Davis. Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Quarterly, 13(3):319–340, 1989
1989
-
[49]
Evaluating web lectures: A case study from hci
Jason Day and Jim Foley. Evaluating web lectures: A case study from hci. InCHI’06 Extended Abstracts on Human Factors in Computing Systems, pages 195–200, 2006
2006
-
[50]
Indexing by latent semantic analysis.Journal of the American society for information science, 41(6):391–407, 1990
Scott Deerwester, Susan T Dumais, George W Furnas, Thomas K Landauer, and Richard Harsh- man. Indexing by latent semantic analysis.Journal of the American society for information science, 41(6):391–407, 1990
1990
-
[51]
AI-generated nonsense is leaking into scientific journals .Popular Science, March 2024
Mack Deguerin. AI-generated nonsense is leaking into scientific journals .Popular Science, March 2024
2024
-
[52]
Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality, 2023
Fabrizio Dell’Acqua, Edward McFowland, Ethan R Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R Lakhani. Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge wor...
2023
-
[53]
so what if chatgpt wrote it?
Yogesh K. Dwivedi, Nir Kshetri, Laurie Hughes, Emma Louise Slade, Anand Jeyaraj, Arpan Ku- mar Kar, Abdullah M. Baabdullah, Alex Koohang, Vishnupriya Raghavan, Manju Ahuja, Hanaa Albanna, Mousa Ahmad Albashrawi, Adil S. Al-Busaidi, Janarthanan Balakrishnan, Yves Barlette, Srip...
2023
-
[54]
Yes-yes-yes: Proactive data collection for ACL rolling review and beyond
Nils Dycke, Ilia Kuznetsov, and Iryna Gurevych. Yes-yes-yes: Proactive data collection for ACL rolling review and beyond. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, BIBLIOGRAPHY 181 Findings of the Association for Computational Linguistics: EMNLP 2022, pages ...
2022
-
[55]
NLPeer: A unified resource for the computa- tional study of peer review
Nils Dycke, Ilia Kuznetsov, and Iryna Gurevych. NLPeer: A unified resource for the computa- tional study of peer review. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Vo...
2023
-
[56]
Sciencebeam—using computer vision to extract pdf data
Daniel Ecer and Giuliano Maciocci. Sciencebeam—using computer vision to extract pdf data. Elife Blog Post, 8 2017. [Online; accessed 2023-Sep-8]
2017
-
[57]
Tools such as ChatGPT threaten transparent science; here are our ground rules for their use.Nature, 613(7945):612–612, 2023
Nature Editorial. Tools such as ChatGPT threaten transparent science; here are our ground rules for their use.Nature, 613(7945):612–612, 2023
2023
-
[58]
New methods in automatic extracting.Journal of the ACM (JACM), 16(2):264–285, 1969
Harold P Edmundson. New methods in automatic extracting.Journal of the ACM (JACM), 16(2):264–285, 1969
1969
-
[59]
What’s In My Big Data? InThe Twelfth International Conference on Learning Representations, 2023
Yanai Elazar, Akshita Bhagia, Ian Helgi Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Evan Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, et al. What’s In My Big Data? InThe Twelfth International Conference on Learning Representations, 2023
2023
-
[60]
Abstracts written by ChatGPT fool scientists.Nature, Jan 2023
Holly Else. Abstracts written by ChatGPT fool scientists.Nature, Jan 2023
2023
-
[61]
Art and the science of generative ai.Science, 380(6650):1110–1111, 2023
Ziv Epstein, Aaron Hertzmann, Investigators of Human Creativity, Memo Akten, Hany Farid, Jessica Fjeld, Morgan R Frank, Matthew Groh, Laura Herman, Neil Leach, et al. Art and the science of generative ai.Science, 380(6650):1110–1111, 2023
2023
-
[62]
Lexrank: Graph-based lexical centrality as salience in text summarization
Günes Erkan and Dragomir R Radev. Lexrank: Graph-based lexical centrality as salience in text summarization. Journal of artificial intelligence research, 22:457–479, 2004
2004
-
[63]
TweepFake: About detecting deepfake tweets.Plos one, 16(5):e0251415, 2021
Tiziano Fagni, Fabrizio Falchi, Margherita Gambini, Antonio Martella, and Maurizio Tesconi. TweepFake: About detecting deepfake tweets.Plos one, 16(5):e0251415, 2021
2021
-
[64]
Three Bricks to Consolidate Watermarks for Large Language Models
Pierre Fernandez, Antoine Chaffin, Karim Tit, Vivien Chappelier, and Teddy Furon. Three Bricks to Consolidate Watermarks for Large Language Models. 2023 IEEE International Workshop on Information Forensics and Security (WIFS), pages 1–6, 2023
2023
-
[65]
Microeconomics of technology adoption.Annu
Andrew D Foster and Mark R Rosenzweig. Microeconomics of technology adoption.Annu. Rev. Econ., 2(1):395–424, 2010
2010
-
[66]
Tradition and innovation in scientists’ research strategies
Jacob G Foster, Andrey Rzhetsky, and James A Evans. Tradition and innovation in scientists’ research strategies. American sociological review, 80(5):875–908, 2015. BIBLIOGRAPHY 182
2015
-
[67]
Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic-Aware Watermark Remedy.ArXiv, abs/2307.13808, 2023
Yu Fu, Deyi Xiong, and Yue Dong. Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic-Aware Watermark Remedy.ArXiv, abs/2307.13808, 2023
2023 arXiv
-
[68]
Catherine A Gao, Frederick M Howard, Nikolay S Markov, Emma C Dyer, Siddhi Ramesh, Yuan Luo, and Alexander T Pearson. Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded hu...
2022
-
[69]
The pile: An 800gb dataset of diverse text for language modeling.arXiv preprint arXiv:2101.00027, 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling.arXiv preprint arXiv:2101.00027, 2020
2020 arXiv
-
[70]
GLTR: Statistical Detection and Visualization of Generated Text
Sebastian Gehrmann, Hendrik Strobelt, and Alexander M Rush. GLTR: Statistical Detection and Visualization of Generated Text. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 111–116, 2019
2019
-
[71]
Soumya Suvra Ghosal, Souradip Chakraborty, Jonas Geiping, Furong Huang, Dinesh Manocha, and A. S. Bedi. Towards Possibilities & Impossibilities of AI-generated Text Detection: A Survey. ArXiv, abs/2310.15264, 2023
2023 arXiv
-
[72]
’Person’== Light-skinned, Western Man, and Sexualization of Women of Color: Stereotypes in Stable Diffusion.arXiv preprint arXiv:2310.19981, 2023
Sourojit Ghosh and Aylin Caliskan. ’Person’== Light-skinned, Western Man, and Sexualization of Women of Color: Stereotypes in Stable Diffusion.arXiv preprint arXiv:2310.19981, 2023
2023 arXiv
-
[73]
Manuscript quality before and after peer review and editing at annals of internal medicine.Annals of internal medicine, 121(1):11–21, 1994
Steven N Goodman, Jesse Berlin, Suzanne W Fletcher, and Robert H Fletcher. Manuscript quality before and after peer review and editing at annals of internal medicine.Annals of internal medicine, 121(1):11–21, 1994
1994
-
[74]
Water- marking Pre-trained Language Models with Backdooring.arXiv preprint arXiv:2210.07543, 2022
Chenxi Gu, Chengsong Huang, Xiaoqing Zheng, Kai-Wei Chang, and Cho-Jui Hsieh. Water- marking Pre-trained Language Models with Backdooring.arXiv preprint arXiv:2210.07543, 2022
2022 arXiv
-
[75]
How to spot AI-generated text.MIT TechnologyReview, Dec 2022
Melissa Heikkilä. How to spot AI-generated text.MIT TechnologyReview, Dec 2022
2022
-
[76]
Dialect prejudice predicts AI decisions about people’s character, employability, and criminality.arXiv preprint arXiv:2403.00742, 2024
Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. Dialect prejudice predicts AI decisions about people’s character, employability, and criminality.arXiv preprint arXiv:2403.00742, 2024
2024 arXiv
-
[77]
Biasinperceptionofartproducedbyartificialintelligence
Joo-WhaHong. Biasinperceptionofartproducedbyartificialintelligence. In Human-Computer Interaction. Interaction in Context: 20th International Conference, HCI International 2018, Las Vegas, NV, USA, July 15–20, 2018, Proceedings, Part II 20, pages 290–303. Springer, 2018. BIBLI...
2018
-
[78]
The changing forms and expectations of peer review
Serge PJM Horbach and Willem Halffman. The changing forms and expectations of peer review. Research integrity and peer review, 3(1):1–15, 2018
2018
-
[79]
Adopting ai: how familiarity breeds both trust and contempt.AI & society, 39(4):1721–1735, 2024
Michael C Horowitz, Lauren Kahn, Julia Macdonald, and Jacquelyn Schneider. Adopting ai: how familiarity breeds both trust and contempt.AI & society, 39(4):1721–1735, 2024
2024
-
[80]
SemStamp: A SemanticWatermarkwithParaphrasticRobustnessforTextGeneration
Abe Bohan Hou, Jingyu Zhang, Tianxing He, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, and Yulia Tsvetkov. SemStamp: A SemanticWatermarkwithParaphrasticRobustnessforTextGeneration. ArXiv, abs/2310.03991, 2023
-
[81]
Chatgpt sets record for fastest-growing user base - analyst note.Reuters, February 2023
Krystal Hu. Chatgpt sets record for fastest-growing user base - analyst note.Reuters, February 2023
2023
-
[82]
RADAR: Robust AI-Text Detection via Adversarial Learning.ArXiv, abs/2307.03838, 2023
Xiaobing Hu, Pin-Yu Chen, and Tsung-Yi Ho. RADAR: Robust AI-Text Detection via Adversarial Learning.ArXiv, abs/2307.03838, 2023
2023 arXiv
-
[83]
Unbiased Watermark for Large Language Models.ArXiv, abs/2310.10669, 2023
Zhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu, Hongyang Zhang, and Heng Huang. Unbiased Watermark for Large Language Models.ArXiv, abs/2310.10669, 2023
2023 arXiv
-
[84]
The adoption of chatgpt.IZA Discussion Paper No
Anders Humlum and Emilie Vestergaard. The adoption of chatgpt.IZA Discussion Paper No. 16992, 2024
2024
-
[85]
Clarification on large language model policy LLM
ICML. Clarification on large language model policy LLM. https://icml.cc/ Conferences/2023/llm-policy, 2023
2023
-
[86]
ICML 2023 Reviewer Tutorial
International Conference on Machine Learning. ICML 2023 Reviewer Tutorial. https: //icml.cc/Conferences/2023/ReviewerTutorial, 2023
2023
-
[87]
Automatic detection of generated text is easiest when humans are fooled.arXiv preprint arXiv:1911.00650, 2019
Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. Automatic detection of generated text is easiest when humans are fooled.arXiv preprint arXiv:1911.00650, 2019
1911 arXiv
-
[88]
Ai-mediated communication: How the perception that profile text was written by ai affects trustworthiness
Maurice Jakesch, Megan French, Xiao Ma, Jeffrey T Hancock, and Mor Naaman. Ai-mediated communication: How the perception that profile text was written by ai affects trustworthiness. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–13, 2019
2019
-
[89]
Short texts, best-fitting curves and new measures of lexical diversity.Language Testing, 19(1):57–84, 2002
Scott Jarvis. Short texts, best-fitting curves and new measures of lexical diversity.Language Testing, 19(1):57–84, 2002
2002
-
[90]
Automatic detection of machine generated text: A critical survey.arXiv preprint arXiv:2011.01314, 2020
Ganesh Jawahar, Muhammad Abdul-Mageed, and Laks VS Lakshmanan. Automatic detection of machine generated text: A critical survey.arXiv preprint arXiv:2011.01314, 2020. BIBLIOGRAPHY 184
2011 arXiv
-
[91]
death of the renaissance man
Benjamin F Jones. The burden of knowledge and the “death of the renaissance man”: Is innovation getting harder?The Review of Economic Studies, 76(1):283–317, 2009
2009
-
[92]
Martin.Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models
Daniel Jurafsky and James H. Martin.Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. Prentice Hall, 3rd edition, 2025. Online manuscript released January 12, 2025
2025
-
[93]
Generative ai and perceptual harms: Who’s suspected of using llms?arXiv preprint arXiv:2410.00906, 2024
Kowe Kadoma, Danaë Metaxa, and Mor Naaman. Generative ai and perceptual harms: Who’s suspected of using llms?arXiv preprint arXiv:2410.00906, 2024
2024 arXiv
-
[94]
The hewlett foundation: Automated essay scoring.https://www.kaggle.com/c/ asap-aes, 2012
Kaggle. The hewlett foundation: Automated essay scoring.https://www.kaggle.com/c/ asap-aes, 2012. Accessed: 2023-03-15
2012
-
[95]
The diffusion of new technologies*.The Quarterly Journal of Economics, page qjaf002, 01 2025
Aakash Kalyani, Nicholas Bloom, Marcela Carvalho, Tarek Hassan, Josh Lerner, and Ahmed Tahoun. The diffusion of new technologies*.The Quarterly Journal of Economics, page qjaf002, 01 2025
2025
-
[96]
The diffusion of new technologies
Aakash Kalyani, Nicholas Bloom, Marcela Carvalho, Tarek Alexander Hassan, Josh Lerner, and Ahmed Tahoun. The diffusion of new technologies. Working Paper 28999, National Bureau of Economic Research, July 2021
2021
-
[97]
Chatgpt for good? on opportunities and challenges of large language models for education.Learning and Individual Differences, 103:102274, April 2023
Enkelejda Kasneci, Kathrin Sessler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Jürgen Pfeffer, Oleksandra Poquet, Michael Sail...
2023
-
[98]
ChatGPT creator pulls AI detection tool due to ‘low rate of accuracy’
Samantha Murphy Kelly. ChatGPT creator pulls AI detection tool due to ‘low rate of accuracy’. CNN Business, Jul 2023
2023
-
[99]
A watermark for large language models.International Conference on Machine Learning, 2023
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. A watermark for large language models.International Conference on Machine Learning, 2023
2023
-
[100]
New AI classifier for indicating AI-written text, 2023
Jan Hendrik Kirchner, Lama Ahmad, Scott Aaronson, and Jan Leike. New AI classifier for indicating AI-written text, 2023
2023
-
[101]
A review of mobile hci research methods
Jesper Kjeldskov and Connor Graham. A review of mobile hci research methods. InInternational Conference on Mobile Human-Computer Interaction, pages 317–335. Springer, 2003
2003
-
[102]
Algorithmic monoculture and social welfare.Proceedings of the National Academy of Sciences, 118(22):e2018340118, 2021
Jon Kleinberg and Manish Raghavan. Algorithmic monoculture and social welfare.Proceedings of the National Academy of Sciences, 118(22):e2018340118, 2021. BIBLIOGRAPHY 185
2021
-
[103]
Reduced, reused and recycled: The life of a dataset in machine learning research
Bernard Koch, Emily Denton, Alex Hanna, and Jacob Gates Foster. Reduced, reused and recycled: The life of a dataset in machine learning research. InThirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track(Round 2), 2021
2021
-
[104]
The global burden of journal peer review in the biomedical literature: Strong imbalance in the collective enterprise
Michail Kovanis, Raphaël Porcher, Philippe Ravaud, and Ludovic Trinquart. The global burden of journal peer review in the biomedical literature: Strong imbalance in the collective enterprise. PloS one, 11(11):e0166387, 2016
2016
-
[105]
All the News That’s Fit to Fabricate: AI- Generated Text as a Tool of Media Misinformation.Journal of Experimental Political Science, 9(1):104–117, 2022
Sarah Kreps, R McCain, and Miles Brundage. All the News That’s Fit to Fabricate: AI- Generated Text as a Tool of Media Misinformation.Journal of Experimental Political Science, 9(1):104–117, 2022
2022
-
[106]
Paraphras- ing evades detectors of AI-generated text, but retrieval is an effective defense.arXiv preprint arXiv:2303.13408, 2023
Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, and Mohit Iyyer. Paraphras- ing evades detectors of AI-generated text, but retrieval is an effective defense.arXiv preprint arXiv:2303.13408, 2023
2023 arXiv
-
[107]
Robust Distortion- free Watermarks for Language Models.ArXiv, abs/2307.15593, 2023
Rohith Kuditipudi, John Thickstun, Tatsunori Hashimoto, and Percy Liang. Robust Distortion- free Watermarks for Language Models.ArXiv, abs/2307.15593, 2023
2023 arXiv
-
[108]
The structure of scientifi revolutions.The Un of Chicago Press, 2:90, 1962
Thomas S Kuhn. The structure of scientifi revolutions.The Un of Chicago Press, 2:90, 1962
1962
-
[109]
Per- formance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models
Tiffany H Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepaño, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, et al. Per- formance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language mod...
2023
-
[110]
Har- vard University Press, 2009
Michèle Lamont.How professors think: Inside the curious world of academic judgment. Har- vard University Press, 2009
2009
-
[111]
Toward a comparative sociology of valuation and evaluation.Annual review of sociology, 38:201–221, 2012
Michèle Lamont. Toward a comparative sociology of valuation and evaluation.Annual review of sociology, 38:201–221, 2012
2012
-
[112]
Vocabulary size and use: Lexical richness in L2 written production
Batia Laufer and Paul Nation. Vocabulary size and use: Lexical richness in L2 written production. Applied linguistics, 16(3):307–322, 1995
1995
-
[113]
Ethical and social risks of harm from language models.arXiv preprint arXiv:2112.04359, 2021
Maribeth Rauh Laura Weidinger, John Mellor and other authors. Ethical and social risks of harm from language models.arXiv preprint arXiv:2112.04359, 2021
2021 arXiv
-
[114]
Detecting Fake Content with Relative Entropy Scoring.Pan, 2008
Thomas Lavergne, Tanguy Urvoy, and François Yvon. Detecting Fake Content with Relative Entropy Scoring.Pan, 2008
2008
-
[115]
Bias in peer review.Journal of the American Society for information Science and Technology, 64(1):2–17, 2013
Carole J Lee, Cassidy R Sugimoto, Guo Zhang, and Blaise Cronin. Bias in peer review.Journal of the American Society for information Science and Technology, 64(1):2–17, 2013. BIBLIOGRAPHY 186
2013
-
[116]
Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L.C
Mina Lee, Katy Ilonka Gero, John Joon Young Chung, Simon Buckingham Shum, Vipul Raheja, Hua Shen, Subhashini Venugopalan, Thiemo Wambsganss, David Zhou, Emad A. Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L.C. Guo, Md Naimul Hoque, Yewon Kim, S...
2024
-
[117]
Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities
Mina Lee, Percy Liang, and Qian Yang. Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities. InProceedings of the 2022 CHI conference on human factors in computing systems, pages 1–19, 2022
2022
-
[118]
Evaluating Human- Language Model Interaction.arXiv preprint arXiv:2212.09746, 2022
Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, et al. Evaluating Human- Language Model Interaction.arXiv preprint arXiv:2212.09746, 2022
2022 arXiv
-
[119]
Deepfake Text Detection in the Wild.ArXiv, abs/2305.13242, 2023
Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Longyue Wang, Linyi Yang, Shuming Shi, and Yue Zhang. Deepfake Text Detection in the Wild.ArXiv, abs/2305.13242, 2023
2023 arXiv
-
[120]
McFarland, and James Y
Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, Daniel A. McFarland, and James Y. Zou. Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Revie...
2024 arXiv
-
[121]
How random is the review outcome? a systematic study of the impact of external factors on elife peer review.bioRxiv, pages 2023–01, 2023
Weixin Liang, Kyle Mahowald, Jennifer Raymond, Vamshi Krishna, Daniel Smith, Daniel Jurafsky, Daniel McFarland, and James Zou. How random is the review outcome? a systematic study of the impact of external factors on elife peer review.bioRxiv, pages 2023–01, 2023
2023
-
[122]
Systematic analysis of 32,111 ai model cards characterizes documentation practice in ai.Nature Machine Intelligence, 6(7):744–753, 2024
Weixin Liang, Nazneen Rajani, Xinyu Yang, Ezinwanne Ozoani, Eric Wu, Yiqun Chen, Daniel Scott Smith, and James Zou. Systematic analysis of 32,111 ai model cards characterizes documentation practice in ai.Nature Machine Intelligence, 6(7):744–753, 2024
2024
-
[123]
Mixture-of-mamba: Enhancing multi-modal state-space models with modality-aware sparsity
Weixin Liang, Junhong Shen, Genghan Zhang, Ning Dong, Luke Zettlemoyer, and Lili Yu. Mixture-of-mamba: Enhancing multi-modal state-space models with modality-aware sparsity. arXiv preprint arXiv:2501.16295, 2025
2025 arXiv
-
[124]
Mixture-of-transformers: A BIBLIOGRAPHY 187 sparse and scalable architecture for multi-modal foundation models.Transactions on Machine Learning Research, 2025
Weixin Liang, LILI YU, Liang Luo, Srini Iyer, Ning Dong, Chunting Zhou, Gargi Ghosh, Mike Lewis, Wen tau Yih, Luke Zettlemoyer, and Xi Victoria Lin. Mixture-of-transformers: A BIBLIOGRAPHY 187 sparse and scalable architecture for multi-modal foundation models.Transactions on M...
2025
-
[125]
Code and Data for: GPT Detectors Are Biased Against Non-Native English Writers, May 2023
Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, and James Zou. Code and Data for: GPT Detectors Are Biased Against Non-Native English Writers, May 2023
2023
-
[126]
Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, and James Y. Zou. GPT detectors are biased against non-native English writers.ArXiv, abs/2304.02819, 2023
2023 arXiv
-
[127]
The widespread adoption of large language model-assisted writing across society.arXiv preprint arXiv:2502.09747, 2025
Weixin Liang, Yaohui Zhang, Mihai Codreanu, Jiayu Wang, Hancheng Cao, and James Zou. The widespread adoption of large language model-assisted writing across society.arXiv preprint arXiv:2502.09747, 2025
2025 arXiv
-
[128]
Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, Diyi Yang, Christopher Potts, Christopher D Manning, and James Y. Zou. Mapping the increasing use of LLMs in scientific papers. InFirst Conference on L...
2024
-
[129]
Can large language models provide useful feedback on research papers? a large-scale empirical analysis
Weixin Liang, Yuhui Zhang, Hancheng Cao, Binglu Wang, Daisy Yi Ding, Xinyu Yang, Kailas Vodrahalli, Siyu He, Daniel Scott Smith, Yian Yin, et al. Can large language models provide useful feedback on research papers? a large-scale empirical analysis. NEJM AI, 1(8):AIoa2400196, 2024
2024
-
[130]
The unlocking spell on base llms: Rethinking alignment via in-context learning.arXiv preprint arXiv:2312.01552, 2023
Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. The unlocking spell on base llms: Rethinking alignment via in-context learning.arXiv preprint arXiv:2312.01552, 2023
2023 arXiv
-
[131]
A Semantic Invariant Robust Watermark for Large Language Models.ArXiv, abs/2310.06356, 2023
Aiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng, and Lijie Wen. A Semantic Invariant Robust Watermark for Large Language Models.ArXiv, abs/2310.06356, 2023
2023 arXiv
-
[132]
When ChatGPT is gone: Creativity reverts and homogeneity persists, 2024
Qinghan Liu, Yiyong Zhou, Jihao Huang, and Guiquan Li. When ChatGPT is gone: Creativity reverts and homogeneity persists, 2024
2024
-
[133]
Reviewergpt? an exploratory study on using large language models for paper reviewing.arXiv preprint arXiv:2306.00622, 2023
Ryan Liu and Nihar B Shah. Reviewergpt? an exploratory study on using large language models for paper reviewing.arXiv preprint arXiv:2306.00622, 2023
2023 arXiv
-
[134]
CoCo: Coherence- Enhanced Machine-Generated Text Detection Under Data Limitation With Contrastive Learn- ing
Xiaoming Liu, Zhaohan Zhang, Yichen Wang, Yu Lan, and Chao Shen. CoCo: Coherence- Enhanced Machine-Generated Text Detection Under Data Limitation With Contrastive Learn- ing. ArXiv, abs/2212.10341, 2022
2022 arXiv
-
[135]
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A Robustly Optimized BERT Pretraining Approach. ArXiv, abs/1907.11692, 2019. BIBLIOGRAPHY 188
1907 arXiv
-
[136]
S2ORC: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. S2ORC: The semantic scholar open research corpus. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4969–4983, Online, July 2020. Association for Computational L...
2020
-
[137]
Science as social knowledge: Values and objectivity in scientific inquiry
Helen E Longino. Science as social knowledge: Values and objectivity in scientific inquiry. Princeton university press, 1990
1990
-
[138]
A corpus-based evaluation of syntactic complexity measures as indices of college-level ESL writers’ language development.TESOL quarterly, 45(1):36–62, 2011
Xiaofei Lu. A corpus-based evaluation of syntactic complexity measures as indices of college-level ESL writers’ language development.TESOL quarterly, 45(1):36–62, 2011
2011
-
[139]
The automatic creation of literature abstracts.IBM Journal of research and development, 2(2):159–165, 1958
Hans Peter Luhn. The automatic creation of literature abstracts.IBM Journal of research and development, 2(2):159–165, 1958
1958
-
[140]
The Global AI Talent Tracker, 2024
MacroPolo. The Global AI Talent Tracker, 2024
2024
-
[141]
The matthew effect in science: The reward and communication systems of science are considered.Science, 159(3810):56–63, 1968
Robert K Merton. The matthew effect in science: The reward and communication systems of science are considered.Science, 159(3810):56–63, 1968
1968
-
[142]
Lisa Messeri and M. J. Crockett. Artificial intelligence and illusions ofunderstanding in scientific research. Nature, 627:49–58, 2024
2024
-
[143]
Household surveys in crisis.Journal of Economic Perspectives, 29(4):199–226, 2015
Bruce D Meyer, Wallace KC Mok, and James X Sullivan. Household surveys in crisis.Journal of Economic Perspectives, 29(4):199–226, 2015
2015
-
[144]
Textrank: Bringing order into text
Rada Mihalcea and Paul Tarau. Textrank: Bringing order into text. InProceedings of the 2004 conference on empirical methods in natural language processing, pages 404–411, 2004
2004
-
[145]
DetectGPT: Zero-shot machine-generated text detection using probability curvature
Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn. DetectGPT: Zero-shot machine-generated text detection using probability curvature. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett...
2023
-
[146]
Ryan, Alan Ritter, and Wei Xu
Tarek Naous, Michael J. Ryan, Alan Ritter, and Wei Xu. Having Beer after Prayer? Measuring Cultural Bias in Large Language Models, 2024
2024
-
[147]
How to Write a Report.https://www.nature.com/nature/for-referees/ how-to-write-a-report
Nature. How to Write a Report.https://www.nature.com/nature/for-referees/ how-to-write-a-report. Accessed: 21 September 2023
2023
-
[148]
Nature will publish peer review reports as a trial.Nature, 578(7793):8, 2020
Nature. Nature will publish peer review reports as a trial.Nature, 578(7793):8, 2020
2020
-
[149]
Nature is trialling transparent peer review - the early results are encouraging.Nature, 603(1):8, 2022
Nature. Nature is trialling transparent peer review - the early results are encouraging.Nature, 603(1):8, 2022. BIBLIOGRAPHY 189
2022
-
[150]
Writing Your Report
Nature Communications. Writing Your Report. https://www.nature.com/ncomms/ for-reviewers/writing-your-report. Accessed: 21 September 2023
2023
-
[151]
The double dixie cup problem.The American Mathematical Monthly, 67(1):58–61, 1960
Donald J Newman. The double dixie cup problem.The American Mathematical Monthly, 67(1):58–61, 1960
1960
-
[152]
TrackingAI-enabledMisinformation: 713‘UnreliableAI-GeneratedNews’Websites (and Counting), Plus the Top False Narratives Generated by Artificial Intelligence Tools, 2023
NewsGuard. TrackingAI-enabledMisinformation: 713‘UnreliableAI-GeneratedNews’Websites (and Counting), Plus the Top False Narratives Generated by Artificial Intelligence Tools, 2023. Accessed: 2024-02-24
2023
-
[153]
A quick guide to writing a solid peer review.Eos, Transactions American Geophysical Union, 92(28):233–234, 2011
Kimberly A Nicholas and Wendy S Gordon. A quick guide to writing a solid peer review.Eos, Transactions American Geophysical Union, 92(28):233–234, 2011
2011
-
[154]
Experimental evidence on the productivity effects of generative artificial intelligence.Availableat SSRN 4375283, 2023
Shakked Noy and Whitney Zhang. Experimental evidence on the productivity effects of generative artificial intelligence.Availableat SSRN 4375283, 2023
2023
-
[155]
Google search exposes academics using ChatGPT in research papers
Paulina Okunyt˙ e. Google search exposes academics using ChatGPT in research papers. Cybernews, November 2023
2023
-
[156]
GPT-2: 1.5B release
OpenAI. GPT-2: 1.5B release. https://openai.com/research/ gpt-2-1-5b-release, 2019. Accessed: 2019-11-05
2019
-
[157]
OpenAI. ChatGPT. https://chat.openai.com/, 2022. Accessed: 2022-12-31
2022
-
[158]
GPT-4 Technical Report.ArXiv, abs/2303.08774, 2023
OpenAI. GPT-4 Technical Report.ArXiv, abs/2303.08774, 2023
2023 arXiv
-
[159]
Papers and peer reviews with evidence of ChatGPT writing
Ivan Oransky and Adam Marcus. Papers and peer reviews with evidence of ChatGPT writing . Retraction Watch, 2024
2024
-
[160]
Syntactic complexity measures and their relationship to L2 proficiency: A research synthesis of college-level L2 writing.Applied linguistics, 24(4):492–518, 2003
Lourdes Ortega. Syntactic complexity measures and their relationship to L2 proficiency: A research synthesis of college-level L2 writing.Applied linguistics, 24(4):492–518, 2003
2003
-
[161]
’helpfulness’ in online communities: a measure of message quality
Jahna Otterbacher. ’helpfulness’ in online communities: a measure of message quality. In Proceedings of the SIGCHI conference on human factors in computing systems, pages 955–964, 2009
2009
-
[162]
Training language models to follow instructions with human feedback.Advancesin Neural Information Processing Systems, 35:27730–27744, 2022
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback.Advancesin Neural Information Processing Systems, 35:27730–...
2022
-
[163]
Survey design and implementation in hci
A Ant Ozok. Survey design and implementation in hci. InThe human-computer interaction handbook, pages 1177–1196. CRC Press, 2007. BIBLIOGRAPHY 190
2007
-
[164]
Multilingual BERT has an accent: Evaluating English influences on fluency in multilingual models, 2023
Isabel Papadimitriou, Kezia Lopez, and Dan Jurafsky. Multilingual BERT has an accent: Evaluating English influences on fluency in multilingual models, 2023
2023
-
[165]
ChatGPT Hits 100 Million Users, Google invests in AI Bot And ChatGPT Goes Viral
Martine Paris. ChatGPT Hits 100 Million Users, Google invests in AI Bot And ChatGPT Goes Viral. Forbes, February 2023
2023
-
[166]
The impact of ai on developer productivity: Evidence from github copilot.arXiv preprint arXiv:2302.06590, 2023
Sida Peng, Eirini Kalliamvakou, Peter Cihon, and Mert Demirer. The impact of ai on developer productivity: Evidence from github copilot.arXiv preprint arXiv:2302.06590, 2023
2023 arXiv
-
[167]
Columbia University Press, 1963
Derek J De Solla Price.Little science, big science. Columbia University Press, 1963
1963
-
[168]
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020
2020
-
[169]
Gpt4 is slightly helpful for peer-review assistance: A pilot study.arXiv preprint arXiv:2307.05492, 2023
Zachary Robertson. Gpt4 is slightly helpful for peer-review assistance: A pilot study.arXiv preprint arXiv:2307.05492, 2023
2023 arXiv
-
[170]
How to review for ACL Rolling Review
Anna Rogers and Isabelle Augenstein. How to review for ACL Rolling Review. https: //aclrollingreview.org/reviewertutorial, 11 2021
2021
-
[171]
Diffusion of innovations
Everett M Rogers, Arvind Singhal, and Margaret M Quinlan. Diffusion of innovations. In Michael B. Salwen and Don W. Stacks, editors,An Integrated Approach to Communication Theory and Research, pages 432–448. Routledge, 2 edition, 2008
2008
-
[172]
Demo- graphics, attitudes, and technology readiness: A cross-cultural analysis and model validation
Jose I Rojas-Mendez, Ananthanarayanan Parasuraman, and Nicolas Papadopoulos. Demo- graphics, attitudes, and technology readiness: A cross-cultural analysis and model validation. Marketing Intelligence & Planning, 35(1):18–39, 2017
2017
-
[173]
ChatGPT banned from New York City public schools’ devices and networks
Kalhan Rosenblatt. ChatGPT banned from New York City public schools’ devices and networks. NBC News, Jan 2023. Accessed: 22.01.2023
2023
-
[174]
Women arecreditedless in sciencethan men.Nature, 608(7921):135– 145, 2022
Matthew B Ross, Britta M Glennon, Raviv Murciano-Goroff, Enrico G Berkes, Bruce A Weinberg, and JuliaI Lane. Women arecreditedless in sciencethan men.Nature, 608(7921):135– 145, 2022
2022
-
[175]
Balasubramanian, Wenxiao Wang, and Soheil Feizi
Vinu Sankar Sadasivan, Aounon Kumar, S. Balasubramanian, Wenxiao Wang, and Soheil Feizi. Can AI-Generated Text be Reliably Detected?ArXiv, abs/2303.11156, 2023
2023 arXiv
-
[176]
Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971–30004
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971–30004. PMLR, 2023. BIBLIOGRAPHY 191
2023
-
[177]
Do datasets have politics? disciplinary values in computer vision dataset development
Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton. Do datasets have politics? disciplinary values in computer vision dataset development. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW2):1–37, 2021
2021
-
[178]
Four years in review: Statistical practices of likert scales in human-robot interaction studies
Mariah L Schrum, Michael Johnson, Muyleng Ghuy, and Matthew C Gombolay. Four years in review: Statistical practices of likert scales in human-robot interaction studies. InCompanion of the 2020 ACM/IEEE International Conference on Human-Robot Interaction, pages 43–52, 2020
2020
-
[179]
Is the future of peer review automated?BMC Research Notes, 15(1):1–5, 2022
Robert Schulz, Adrian Barnett, René Bernard, Nicholas JL Brown, Jennifer A Byrne, Peter Eckmann, Malgorzata A Gazda, Halil Kilicoglu, Eric M Prager, Maia Salholz-Hillel, et al. Is the future of peer review automated?BMC Research Notes, 15(1):1–5, 2022
2022
-
[180]
Challenges, experiments, and computational solutions in peer review
Nihar B Shah. Challenges, experiments, and computational solutions in peer review. Communications of the ACM, 65(6):76–87, 2022
2022
-
[181]
A mathematical theory of communication.The Bell system technical journal, 27(3):379–423, 1948
Claude Elwood Shannon. A mathematical theory of communication.The Bell system technical journal, 27(3):379–423, 1948
1948
-
[182]
Red Teaming Language Model Detectors with Language Models.ArXiv, abs/2305.19713, 2023
Zhouxing Shi, Yihan Wang, Fan Yin, Xiangning Chen, Kai-Wei Chang, and Cho-Jui Hsieh. Red Teaming Language Model Detectors with Language Models.ArXiv, abs/2305.19713, 2023
2023 arXiv
-
[183]
The adoption and efficacy of large language models: Evidence from consumer complaints in the financial industry.Availableat SSRN 5004194, 2024
Minkyu Shin, Jin Kim, and Jiwoong Shin. The adoption and efficacy of large language models: Evidence from consumer complaints in the financial industry.Availableat SSRN 5004194, 2024
2024
-
[184]
The curse of recursion: Training on generated data makes models forget.arXiv preprint arXiv:2305.17493, 2023
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson. The curse of recursion: Training on generated data makes models forget.arXiv preprint arXiv:2305.17493, 2023
2023 arXiv
-
[185]
New Evaluation Metrics Capture Quality Degradation due to LLM Watermarking.arXiv preprint arXiv:2312.02382, 2023
Karanpartap Singh and James Zou. New Evaluation Metrics Capture Quality Degradation due to LLM Watermarking.arXiv preprint arXiv:2312.02382, 2023
2023 arXiv
-
[186]
Real ml: Recognizing, exploring, and articulating limitations of machine learning research
Jessie J Smith, Saleema Amershi, Solon Barocas, Hanna Wallach, and Jennifer Wort- man Vaughan. Real ml: Recognizing, exploring, and articulating limitations of machine learning research. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, p...
2022
-
[187]
Dynamic pooling and unfolding recursive autoencoders for paraphrase detection.Advances in neural information processing systems, 24, 2011
Richard Socher, Eric Huang, Jeffrey Pennin, Christopher D Manning, and Andrew Ng. Dynamic pooling and unfolding recursive autoencoders for paraphrase detection.Advances in neural information processing systems, 24, 2011. BIBLIOGRAPHY 192
2011
-
[188]
Release strategies and the social impacts of language models.arXiv preprint arXiv:1908.09203, 2019
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al. Release strategies and the social impacts of language models.arXiv preprint arXiv:1908.09203, 2019
1908 arXiv
-
[189]
When the interface is a face.Human-computer interaction, 11(2):97–124, 1996
Lee Sproull, Mani Subramani, Sara Kiesler, Janet H Walker, and Keith Waters. When the interface is a face.Human-computer interaction, 11(2):97–124, 1996
1996
-
[190]
Artificial intelligence index report
Stanford Institute for Human-Centered Artificial Intelligence. Artificial intelligence index report
-
[191]
Why do scientists disagree? PsyArXiv, 2023
Justin Sulik, Nakwon Rim, Elizabeth Pontikes, James Evans, and Gary Lupyan. Why do scientists disagree? PsyArXiv, 2023. Preprint
2023
-
[192]
The sociology of scientific validity: How professional networks shape judgement in peer review
Misha Teplitskiy, Daniel Acuna, Aïda Elamrani-Raoult, Konrad Körding, and James Evans. The sociology of scientific validity: How professional networks shape judgement in peer review. Research Policy, 47(9):1825–1841, 2018
2018
-
[193]
Christian Terwiesch. Would chat GPT3 get a Wharton MBA? A prediction based on its performance in the operations management course.Mack Institute for InnovationManagement at the Wharton School, University of Pennsylvania, 2023
2023
-
[194]
Blueprint for an AI bill of rights,
The White House Office of Science and Technology Policy. Blueprint for an AI bill of rights,
-
[195]
Holden Thorp
H. Holden Thorp. ChatGPT is fun, but not an author.Science, 379(6630):313–313, 2023
2023
-
[196]
Hakkani-Tür, and Mikhail J
Mercan Topkara, Giuseppe Riccardi, Dilek Z. Hakkani-Tür, and Mikhail J. Atallah. Natural language watermarking: challenges in building a practical system. InElectronic imaging, 2006
2006
-
[197]
Umut Topkara, Mercan Topkara, and Mikhail J. Atallah. The hiding virtues of ambiguity: quantifiably resilient watermarking of natural language text through synonym substitutions. In Workshop on Multimedia & Security, 2006
2006
-
[198]
Yunus Topsakal. How familiarity, ease of use, usefulness, and trust influence the accep- tance of generative artificial intelligence (ai)-assisted travel planning.International Journal of Human–Computer Interaction, pages 1–14, 2024
2024
-
[199]
Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[200]
Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023. BIBLIOGRAPHY 193
2023 arXiv
-
[201]
Energyandinformation
MyronTribusandEdwardCMcIrvine. Energyandinformation. ScientificAmerican, 225(3):179– 190, 1971
1971
-
[202]
Barannikov, Irina Piontkovskaya, Sergey I
Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Daniil Cherniavskii, S. Barannikov, Irina Piontkovskaya, Sergey I. Nikolenko, and Evgeny Burnaev. Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts.ArXiv, abs/2306.04723, 2023
2023 arXiv
-
[203]
Authorship Attribution for Neural Text Generation
Adaku Uchendu, Thai Le, Kai Shu, and Dongwon Lee. Authorship Attribution for Neural Text Generation. In Conference on Empirical Methods in Natural Language Processing, 2020
2020
-
[204]
AI and science: what 1,600 researchers think
Richard Van Noorden and Jeffrey M Perkel. AI and science: what 1,600 researchers think. Nature, 621(7980):672–675, 2023
2023
-
[205]
Van Rossum
Dann. Van Rossum. Generative AI Top 150: The World’s Most Used AI Tools.https: //www.flexos.work/learn/generative-ai-top-150, February 2024
2024
-
[206]
User acceptance of information technology: Toward a unified view.MIS quarterly, pages 425–478, 2003
Viswanath Venkatesh, Michael G Morris, Gordon B Davis, and Fred D Davis. User acceptance of information technology: Toward a unified view.MIS quarterly, pages 425–478, 2003
2003
-
[207]
Artificial artificial artificial intelligence: Crowd workers widely use large language models for text production tasks.arXiv preprint arXiv:2306.07899, 2023
Veniamin Veselovsky, Manoel Horta Ribeiro, and Robert West. Artificial artificial artificial intelligence: Crowd workers widely use large language models for text production tasks.arXiv preprint arXiv:2306.07899, 2023
2023 arXiv
-
[208]
‘As an AI language model’: the phrase that shows how AI is pollulating the web
James Vincent. ‘As an AI language model’: the phrase that shows how AI is pollulating the web. The Verge, Apr 2023
2023
-
[209]
Scientific discovery in the age of artificial intelligence.Nature, 620(7972):47–60, 2023
Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, Ziming Liu, Payal Chandak, Shengchao Liu, Peter Van Katwyk, Andreea Deac, et al. Scientific discovery in the age of artificial intelligence.Nature, 620(7972):47–60, 2023
2023
-
[210]
Testing of detection tools for AI-generated text
Debora Weber-Wulff, Alla Anohina-Naumeca, Sonja Bjelobaba, Tomáš Foltýnek, Jean Guerrero- Dib, Olumide Popoola, Petr Šigut, and Lorna Waddington. Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19(1):26, 2023
2023
-
[211]
Emma Wiles and John J. Horton. More, but worse: The impact of ai writing assistance on the supply and quality of job posts.Massachusetts Institute of Technology(MIT) Sloan, March 2024
2024
-
[212]
ExplanAItions: an artificial intelligence study by Wiley.https://www.wiley.com/ en-us/ai-study, 2023
Wiley. ExplanAItions: an artificial intelligence study by Wiley.https://www.wiley.com/ en-us/ai-study, 2023
2023
-
[213]
Attacking Neural Text Detectors.ArXiv, abs/2002.11768, 2020
Max Wolff. Attacking Neural Text Detectors.ArXiv, abs/2002.11768, 2020
2002 arXiv
-
[214]
DiPmark: A Stealthy, Efficient and Resilient Watermark for Large Language Models.ArXiv, abs/2310.07710, 2023
Yihan Wu, Zhengmian Hu, Hongyang Zhang, and Heng Huang. DiPmark: A Stealthy, Efficient and Resilient Watermark for Large Language Models.ArXiv, abs/2310.07710, 2023. BIBLIOGRAPHY 194
2023 arXiv
-
[215]
DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text.ArXiv, abs/2305.17359, 2023
Xianjun Yang, Wei Cheng, Linda Petzold, William Yang Wang, and Haifeng Chen. DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text.ArXiv, abs/2305.17359, 2023
2023 arXiv
-
[216]
A Survey on Detection of LLMs-Generated Content
Xianjun Yang, Liangming Pan, Xuandong Zhao, Haifeng Chen, Linda Ruth Petzold, William Yang Wang, and Wei Cheng. A Survey on Detection of LLMs-Generated Content. ArXiv, abs/2310.15654, 2023
2023 arXiv
-
[217]
Robust Multi-bit Natural Language Watermarking through Invariant Features
Kiyoon Yoo, Wonhyuk Ahn, Jiho Jang, and No Jun Kwak. Robust Multi-bit Natural Language Watermarking through Invariant Features. In Annual Meeting of the Association for Computational Linguistics, 2023
2023
-
[218]
Is your paper being reviewed by an llm? investigating ai text detectability in peer review.arXiv preprint arXiv:2410.03019, 2024
Sungduk Yu, Man Luo, Avinash Madasu, Vasudev Lal, and Phillip Howard. Is your paper being reviewed by an llm? investigating ai text detectability in peer review.arXiv preprint arXiv:2410.03019, 2024
2024 arXiv
-
[219]
Xiao Yu, Yuang Qi, Kejiang Chen, Guoqiang Chen, Xi Yang, Pengyuan Zhu, Weiming Zhang, and Neng H. Yu. GPT Paternity Test: GPT Generated Text Detection with GPT Genetic Inheritance. ArXiv, abs/2305.12519, 2023
2023 arXiv
-
[220]
Defending Against Neural Fake News.ArXiv, abs/1905.12616, 2019
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. Defending Against Neural Fake News.ArXiv, abs/1905.12616, 2019
1905 arXiv
-
[221]
Adaptive self-improvement llm agentic system for ml library development.arXiv preprint arXiv:2502.02534, 2025
Genghan Zhang, Weixin Liang, Olivia Hsu, and Kunle Olukotun. Adaptive self-improvement llm agentic system for ml library development.arXiv preprint arXiv:2502.02534, 2025
2025
-
[222]
Assaying on the Robustness of Zero-Shot Machine-Generated Text Detectors.ArXiv, abs/2312.12918, 2023
Yi-Fan Zhang, Zhang Zhang, Liang Wang, Tien-Ping Tan, and Rong Jin. Assaying on the Robustness of Zero-Shot Machine-Generated Text Detectors.ArXiv, abs/2312.12918, 2023
2023 arXiv
-
[223]
Provable Robust Watermarking for AI-Generated Text
Xuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, and Yu-Xiang Wang. Provable Robust Watermarking for AI-Generated Text. InInternationalConference onLearning Representations (ICLR), 2024
2024
-
[224]
Permute-and-Flip: An optimally robust and watermarkable decoder for LLMs.arXiv preprint arXiv:2402.05864, 2024
Xuandong Zhao, Lei Li, and Yu-Xiang Wang. Permute-and-Flip: An optimally robust and watermarkable decoder for LLMs.arXiv preprint arXiv:2402.05864, 2024
2024 arXiv
-
[225]
Protecting Language Generation Models via Invisi- ble Watermarking
Xuandong Zhao, Yu-Xiang Wang, and Lei Li. Protecting Language Generation Models via Invisi- ble Watermarking. InProceedings of the 40th International Conference on Machine Learning, pages 42187–42199, 2023
2023
-
[2022]
Accessed: 2023-09-08
2023
-
[2024]
Technical report, Stanford University, 2024
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.