{"id":"57562e9a-8a86-49cd-8657-72fd32c2a4eb","arxiv_id":"2501.00598","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Analysis of tens of thousands of papers shows NLP/AI research mixes 'dialogue' and 'dialog' with no clear trend, author, or context explanation.","lead":"A corpus study of 87,000 papers finds NLP and AI research is split roughly 72-24-5 between 'dialogue', 'dialog', and both spellings in titles and abstracts, with no clear trend over 20 years. The paper tests and mostly rules out American-English, author, context, and time explanations, leaving the field's spelling mix as an unresolved quirk.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline statistics rest on a hand-built S2 venue list and field labels; no robustness check or release shows that the 72/24/5 split or the no-shift trend survives alternative corpus definitions.","rationale":"The reader's verdict is CONDITIONAL for good reasons. My independent read agrees: the paper is an honest descriptive study with explicit limitations, but the primary statistics rest on a corpus that is not independently checkable. The weakest load-bearing premise is not the statistics themselves but the definition of the population: S2 field labels and mean-citation venue ranking. I identify two concrete mechanisms by which this could fail: (i) venue selection by citations rather than by a canonical NLP/AI set admits non-NLP venues and excludes relevant high-output venues with lower citations; (ii) using a 2010+ citation-based venue list to plot a 24-year trend can distort the early part of the curve. The RoBERTa context null is also weak, but it is a secondary, carefully hedged claim ('suggests limited influence'), and the paper's own noun-phrase analysis shows some context effects, so it is less load-bearing than the corpus definition. The manuscript provides no data or code, so the recommended condition is to require a reproducibility artifact or a robustness check; this does not change the reader's verdict, which already conditions on that. My recommendation is UNCHANGED (remain CONDITIONAL).","tokens_in":12763,"tokens_out":7542,"duration_ms":75570,"concrete_test":"Recompute the headline statistics and Figure 4 using the ACL Anthology (or DBLP) directly: define a fixed, pre-registered venue list (ACL, EMNLP, NAACL, COLING, EACL, SIGDIAL, TACL, CL, etc.), extract titles and abstracts, count 'dialog' versus 'dialogue' with a word-boundary regex (excluding compounds), and compare the 72/24/5 split and the yearly dialogue fraction to the paper's values. Then repeat after adding or removing the five borderline venues (LAK, CHIIR, KDD, CVPR, CHI) to see whether the split moves outside the reported confidence intervals. If the alternative corpus gives a split outside the reported CIs, or the time trend shows a monotonic shift, the S2/venue-selection step is the cause.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims ('72% dialogue / 24% dialog / 5% both' and 'no clear evidence of a shift over time') are defined over a corpus assembled in Appendix A from Semantic Scholar metadata and a hand-chosen 'High Impact Dialog(ue) Venues' list. That list is not a canonical set of NLP/AI venues: it is the top 25 venues by mean citation count of dialog(ue) papers, subject to a >=10-paper threshold and a CS-field-label filter, and it includes non-NLP/AI venues (CHI, KDD, CVPR, LAK, CHIIR) while excluding any venue with lower average citations regardless of total dialog(ue) output. The venue table itself contains a likely S2 metadata error ('International Conference on Computational Logic' for COLING), illustrating that the underlying labels are imperfect. Because the same venue set is used for the year-by-year trend (Fig. 4), and because no data or code is released, the reader cannot tell whether the split or the null time trend is an artifact of this selection. The authors honestly list many limitations, but they do not provide a sensitivity analysis over venue inclusion or an alternative corpus construction, so the headline statistics remain conditional on S2 metadata that is known to be noisy (their own footnote on author misidentification, §5.1).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a descriptive corpus study of the spelling variation \"dialog\" vs \"dialogue\" in NLP/AI research, based on Semantic Scholar metadata for papers whose title or abstract contains the term. The headline finding is that among papers in a hand-selected list of top venues, 72% use \"dialogue\", 24% use \"dialog\", and 5% use both in the same title and abstract. The paper also reports no clear time trend toward \"dialog\" over roughly two decades, a weak association between author nationality and spelling, substantial within-author mixing of the two spellings, and limited influence of linguistic context on spelling choice. The study is explicitly framed as descriptive, and the authors are appropriately cautious in their negative conclusions.","tokens_in":13008,"tokens_out":5650,"duration_ms":55861,"significance":"If the headline statistics are robust, this is a useful, systematic quantitative contribution to a long-standing and practically relevant orthographic question in the NLP/AI community. The paper is honest about its limitations, reports confidence intervals for many comparisons, and uses multiple complementary probes (syntactic parses, RoBERTa embeddings, source-code corpora). It makes no claim to a parameter-free derivation and does not overstate the explanatory power of nationality or context. The main deficit is that the core corpus construction is not accompanied by a robustness analysis or a data/code release, which limits the verifiability of the central claims.","major_comments":[{"comment":"The definition of \"High Impact Dialog(ue) Venues\" is a hand-built list of the top 25 venues by mean citation count of Dialog(ue) Papers, and it includes non-NLP/AI venues such as CHI, KDD, CVPR, LAK, and CHIIR while excluding venues with lower average citations. Because the headline 72/24/5 split (§3) and the time-trend analysis (Fig. 4) are computed only over papers in this list, the central descriptive claims are conditional on this particular cutoff and on Semantic Scholar field-label accuracy, which the paper itself notes is imperfect (footnote in §5.1; the \"Computational Logic\" row in Table 2 is a clear metadata error). The paper should provide a sensitivity analysis over alternative venue sets—for example, a fixed canonical NLP/AI venue list, a different citation threshold, or the full set of CS venues—to show that the split and the null trend are not artifacts of this selection.","section":"Appendix A, Table 2"},{"comment":"The manuscript does not release the raw data, the exact Semantic Scholar query, the venue list, or the analysis code. Since all claims are built on a corpus that cannot be fully reconstructed from the description alone (e.g., the handling of \"Both\" across title and abstract, the search pattern details, and the sub-noun-phrase filtering in Appendix E), the reader cannot independently verify the headline percentages or the null shift. I would ask that the authors release a data/code package, or at minimum a table of paper IDs with venue, year, and spelling category, so that the central statistics can be checked.","section":"General (reproducibility)"}],"minor_comments":[{"comment":"Row 9 lists \"International Conference on Computational Logic\"; the usual name of the venue is the International Conference on Computational Linguistics (COLING). This should be corrected or explicitly flagged as a Semantic Scholar metadata label.","section":"Table 2"},{"comment":"The sentence \"The accuracy showed no improvement over the baseline of always predicting dialogue (0.725 vs 0.739 baseline)\" is confusing: if the baseline is always predicting the majority class, it should equal the dialogue prevalence (about 72%), not 73.9%. Please clarify these numbers, since the current wording suggests the model was actually slightly worse than the majority-class baseline.","section":"Appendix H"},{"comment":"The caption reads \"CS Dialogu(ue) Publications\" — a typo for \"Dialog(ue)\".","section":"Figure 2 caption"},{"comment":"The text says \"23.0 percent-points less than authors at British intuitions\"; this should be \"British institutions\".","section":"Section 5.2"},{"comment":"The abstract says \"over ~20 years of NLP/AI research,\" but Figure 4 shows that before 2017 there were fewer than 100 papers per year, so the effective power to detect a shift is concentrated in the recent period. The paper already acknowledges this in §4, but the abstract should be more explicit that the null time-trend conclusion is primarily based on the post-2017 data.","section":"Abstract and Section 4"},{"comment":"The noun-phrase analysis identifies \"visual dialog\" as the most significant phrase, but this is likely driven by the proper-noun task \"Visual Dialog\" from a highly cited paper. The paper discusses proper nouns in §6.3, but it would be helpful to explicitly separate named tasks/datasets from generic noun phrases in the logistic regression.","section":"Section 6.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-written, honest descriptive study of a niche but real community-level question. Its main weakness is that the headline statistics are inherently tied to a non-canonical corpus construction, and the absence of a sensitivity analysis or data release makes the central claims hard to verify. I would encourage the editor to request the sensitivity analysis and a data/code release as part of the revision. The scope fit is reasonable for a venue that accepts empirical meta-analyses of research communities; if the journal expects methodological novelty or generalizable theory, the fit is more marginal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is genuinely new descriptive work on the dialog/dialogue split in NLP/AI publishing, and the headline 72/24/5 split is probably right for the corpus the authors built. The soft spots are exactly where your reader's report puts them: hand-built venue list, no sensitivity analysis, no code or data release. That makes the central claims conditional, not wrong.\n\nWhat's actually new: the venue-level, author-level, and context-level statistics. I don't know of prior work that documents the 72/24/5 split, the author-level mixing, the closed-compound preference, or the null time trend for this community. The authors are careful. They explicitly say they do not conclude there is a shift. They flag S2 author misidentification, admit body text is unexamined, and give a clear appendix explaining the venue selection. The RoBERTa null result is reported with the caveat that they didn't fine-tune the model, which is honest. The compound finding (81.6% dialog, CI 68.4–92.1) is a concrete result that doesn't depend as much on the venue set.\n\nThe soft spot, and it's a real one, is the corpus definition. The top-25 venues by mean citations is a reasonable guess but not a canonical set of NLP/AI venues. It includes CHI, KDD, CVPR, LAK, and CHIIR, and it excludes lower-mean-citation venues regardless of total dialog(ue) output. The table even shows a likely S2 metadata error ('International Conference on Computational Logic' for COLING). Because the same venue set drives the year-by-year trend, the null 'no shift' conclusion could change with a different set. The authors don't provide a sensitivity analysis, and they don't release data or code. That is the difference between a solid descriptive study and an airtight one. I don't think it's load-bearing in the sense of making the claims false; it just caps how strongly they can be stated.\n\nFor a reader interested in bibliometric corpus construction or orthographic variation in a technical community, this is worth a look. For a general NLP reader, it's a curiosity. It deserves a serious referee: the question is well scoped, the methods are standard and reported honestly, and the limitations are stated. I'd send it out with a request for robustness checks over venue selection and a data/code release or detailed replication appendix. After those, accept.","headline":"Careful, honest description of a real orthographic split, probably right about the 72/24/5 numbers, but the hand-built corpus and missing robustness checks keep it conditional.","tokens_in":13587,"tokens_out":2715,"would_cite":false,"duration_ms":26186,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A study of tens of thousands of NLP/AI papers claims the field is stuck on a stable, unexplained mix of 'dialogue' and 'dialog': 72% use 'dialogue', 24% use 'dialog', and 5% use both in the same title or abstract.","keywords":["dialogue vs dialog","orthographic variation","spelling variation","NLP/AI research","corpus statistics","author-level variation","contextual embeddings","scientific writing"],"falsifier":"Recompute the dialogue/dialog/both percentages from the full text of every paper at the same venues, rather than from titles and abstracts alone; if the split changes by more than a few points or a clear time trend toward either spelling emerges, the stable-mix conclusion gives way.","tokens_in":12500,"feed_emoji":"💬","tokens_out":6270,"duration_ms":59422,"temperature":0.7,"pith_summary":"The paper sets out to establish that NLP/AI research is caught in a stable, large-scale spelling disagreement over 'dialogue' vs 'dialog', and that the disagreement is not explained by the usual suspects. Counting more than 87,000 bibliographic records and narrowing to high-impact computing venues, it reports that 72% of the sampled publications use 'dialogue', 24% use 'dialog', and 5% use both in the same title or abstract. Over roughly two decades the mix does not shift consistently, author nationality explains little, and surrounding text barely predicts which spelling appears. If the paper is right, the field is accepting two spellings at once for a word at the center of its subject matter.","feed_headline":"AI papers split 72/24/5 on 'dialogue' vs 'dialog'","feed_subtitle":"A 24-year corpus study finds the split is stable, and author nationality and context explain little.","key_machinery":"The load-bearing object is the 'Dialog(ue) Publication': any paper with 'dialog(s)' or 'dialogue(s)' in its title or abstract, classed into three mutually exclusive categories—dialogue-only, dialog-only, and both—after filtering to 25 high-impact NLP/AI venues chosen by mean citation count. The counts in these categories produce the 72/24/5 headline and the venue, year, and author breakdowns. Around that object the paper builds a supporting apparatus: noun-phrase extraction to test whether phrases like 'visual dialog' bias spelling, a logistic regression with false-discovery-rate correction to identify the few significant phrases, and a masked-language-model embedding classifier that asks whether surrounding context predicts the spelling. The negative result from that classifier is what supports the paper's reading that the spellings are largely interchangeable in context.","core_discovery":"On the paper's own terms, the discovery is that the orthographic split in NLP/AI is real, large, and stable: 72% of the sampled publications use 'dialogue', 24% use 'dialog', and 5% use both in the same title and abstract, and two decades of data give no clear trend toward 'dialog'. The paper reports that this split is more common in computing than in other academic disciplines, that author nationality has only a weak association, that prolific authors often publish under both spellings, and that context predicts spelling only in narrow corners: plural forms favor 'dialogue', proper nouns and closed compounds favor 'dialog', and phrases like 'visual dialog' are outliers. The paper concludes that none of the three common explanations—American English, a computing-specific norm, or full interchangeability—completely accounts for the observed facts.","pith_inferences":["A natural next check is full-text analysis: if 'dialog box' appears in the bodies of papers, the computing-specific explanation may be stronger than the title/abstract data suggest.","If the author-level mixing reflects indifference, then a controlled survey of authors could test whether the choice is effectively random given identical contexts.","The 5% of papers that use both spellings in the same title or abstract is itself a signal that editors and authors do not enforce consistency; a direct extension would be to trace whether the two spellings occupy different sections or refer to different senses within those papers.","For downstream systems, the persistent mix implies practical tooling: search, evaluation, and code repositories that treat 'dialog' and 'dialogue' as the same token would become more robust."],"forward_implications":["There is no empirical support in the title/abstract record for the claim that 'dialog' is taking over NLP/AI writing.","The 'dialog box' rule does not show up in the analyzed titles and abstracts, so the computing-specific explanation is at best incomplete for prose.","Prolific authors and coauthorship networks mix spellings, so individual-level accounts of the choice are unlikely to explain the aggregate pattern.","Context-sensitive spelling is rare: phrases such as 'visual dialog' are outliers, while plural forms and proper or compound forms carry small, systematic biases.","Source code already favors 'dialog' far more than paper prose, suggesting a code-to-prose carryover rather than the reverse."],"supporting_citations":[{"why":"Supplies the bibliographic search API and corpus that define the population of Dialog(ue) Papers.","marker":"Kinney et al., 2023"},{"why":"Supplies the dependency parser used to extract noun phrases for the context analysis.","marker":"Honnibal et al., 2020"},{"why":"Supplies the pretrained masked language model whose contextual embeddings fail to predict spelling.","marker":"Liu et al., 2019"},{"why":"Provides the multiple-comparison correction that keeps the few significant noun-phrase findings from being Type I errors.","marker":"Benjamini and Hochberg, 1995"},{"why":"Provides the Google Ngram book data used as a comparison showing a clearer time trend in books than in papers.","marker":"Lin et al., 2012"},{"why":"Supplies the broad source-code corpus used to measure the stronger preference for 'dialog' in code.","marker":"Kocetkov et al., 2022"}],"fun_headline_variants":["AI papers split 72/24/5 on 'dialogue' vs 'dialog'","NLP's spelling split: 72% dialogue, 24% dialog, no trend","Two spellings, one field: NLP's dialogue/dialog conundrum persists","Why AI can't decide between 'dialogue' and 'dialog' after 20 years","AI's orthographic mystery: 72% dialogue, 24% dialog, 5% both"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole analysis rests on the assumption that the bibliographic database used to find papers correctly labels the venues and fields of NLP/AI research, so that the filtered sample really is the population of papers the conclusions are about.","fun_headline_variants_meta":{"raw":{"variants":["AI papers split 72/24/5 on 'dialogue' vs 'dialog'","NLP's spelling split: 72% dialogue, 24% dialog, no trend","Two spellings, one field: NLP's dialogue/dialog conundrum persists","Why AI can't decide between 'dialogue' and 'dialog' after 20 years","AI's orthographic mystery: 72% dialogue, 24% dialog, 5% both"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000314,"raw_usage":{"total_tokens":1756,"prompt_tokens":891,"completion_tokens":865,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":507,"completion_tokens_details":{"reasoning_tokens":763}},"tokens_in":507,"tokens_out":865,"duration_ms":8799,"temperature":1.0,"reasoning_tokens":763,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:46:56.923379+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the dialogue/dialog/both percentages from the full text of every paper at the same venues, rather than from titles and abstracts alone; if the split changes by more than a few points or a clear time trend toward either spelling emerges, the stable-mix conclusion gives way.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the multiple-comparison correction that keeps the few significant noun-phrase findings from being Type I errors."},{"cited_title":"Brockman, and Slav Petrov","cited_arxiv_id":null,"evidence_quote":"Provides the Google Ngram book data used as a comparison showing a clearer time trend in books than in papers."}],"review_version":1}