{"id":"7849a80f-8912-456e-a609-6bb98bbe2a9d","arxiv_id":"2505.12970","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured literature review of 2023 ACM papers finds that traditional, non-neural NLP techniques are still used in classification, information extraction, relation extraction, text simplification, and text summarization, mainly as baselines or pipeline parts.","lead":"An analysis of 2023 papers in the ACM Digital Library shows that traditional NLP methods such as SVMs, rule-based systems, and extractive summarization still appear in all five surveyed tasks, usually as baselines or pipeline components. The review maps where these older, cheaper approaches remain useful, helping practitioners decide when a large language model is not necessary.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's all-five-scenarios claim is contradicted by the paper's own Table 2 for text simplification: the sole retrieved paper is a survey with no pipeline/comparison/core usage.","rationale":"The reader's weakest_assumption focused on the representativeness of the ACM-DL/title-only/2023 sample. That is a valid limitation, but it is explicitly acknowledged and does not undermine an existence claim, since even a narrow sample can demonstrate that traditional models are still used. A more direct problem is the internal inconsistency between the abstract's universal claim and the paper's own Table 2 for text simplification. The paper transparently labels [61] as a survey without any of the three usage roles, yet the abstract says all five scenarios exhibit traditional models in those roles. This is a correctness issue in the central claim, not a sampling issue. The paper could be accepted after a revision that either qualifies the claim for text simplification (e.g., 'referenced in a survey') or identifies a paper that actually uses traditional models in one of the three roles. Hence CONDITIONAL rather than ACCEPT, and the reader's representativeness concern, while reasonable, was not the most load-bearing issue.","tokens_in":16242,"tokens_out":4866,"duration_ms":50291,"concrete_test":"Inspect Tables 1 and 2 for the Text Simplification row. Verify whether [61] has any checkmark in the pipeline/comparison/core columns. If it has none, the abstract's claim that 'all five application scenarios still exhibit traditional models … as part of a processing pipeline, as a comparison/baseline, or as the main model(s)' is not supported. Additionally, re-run the Section 4 text-simplification search using the same title-only query applied to the other four scenarios; if zero papers are retrieved, the inclusion of text simplification depends entirely on an ad hoc abstract-query expansion. If the authors intend to count survey references as a fourth usage category, the abstract and Section 5 must be revised accordingly.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract claims that in all five application scenarios, traditional models are still used 'as part of a processing pipeline, as a comparison/baseline to the core model of the respective paper, or as the main model(s) of the paper.' The paper's evidence for text simplification is a single paper, North et al. [61] ('Lexical Complexity Prediction: An Overview'), which is a survey. In Table 2, entry [61] has no checkmarks in the Pipeline, Comparison, or Core columns. Section 5 states that only one paper could be retrieved for text simplification and describes it as a survey on lexical complexity prediction—it does not fit any of the three usage categories. Furthermore, the search protocol for this scenario had to be expanded from title-only to abstract search because the title query returned zero results. Thus the universal claim as stated is internally contradicted by the paper's own classification. This is not merely a representativeness concern; if Table 2 is accurate, the abstract's central claim is unsupported for one of the five scenarios.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a structured literature review of five NLP application scenarios (classification, information extraction, relation extraction, text simplification, and text summarization) based on publications retrieved from the ACM Digital Library, with a declared focus on 2023. The authors define \"traditional models\" pragmatically for each scenario, then classify retrieved papers according to three usage roles: part of a processing pipeline, comparison or baseline, or core method. They report that 28 of 119 retrieved papers incorporate traditional models and claim that all five application scenarios still exhibit traditional models in one of these roles. The paper also addresses why traditional approaches remain in use (RQ2) and what advantages they offer (RQ3), and includes a brief discussion of future perspectives.","tokens_in":16408,"tokens_out":6901,"duration_ms":70573,"significance":"If the universal claim were fully supported, the paper would be a useful, reproducible evidence point that traditional NLP methods remain relevant alongside neural and LLM-based approaches, with practical implications for reproducibility, efficiency, and domain-constrained applications. The study's strengths are its transparent search and classification protocol, public dataset on Zenodo, explicit terminology discussion, and an honest limitations section. The main caveat is that the evidence for one of the five scenarios (text simplification) is a single survey paper that is not classified as using traditional models in any of the three stated roles, and the sample includes papers outside the declared 2023 window. After aligning the claims with the actual evidence, the survey would be a worthwhile contribution; in its present form, the headline claim overstates the findings.","major_comments":[{"comment":"The abstract's central claim that all five application scenarios \"still exhibit traditional models\" as part of a pipeline, as a comparison/baseline, or as the main model is not supported by the paper's own evidence for text simplification. The only retrieved text-simplification paper, North et al. [61], is a survey; Table 2 assigns it no checkmarks in any of the three usage columns, and §5 states only that \"a survey on lexical complexity prediction is done.\" A survey that discusses classifiers is not an instance of a model being used in a pipeline, as a baseline, or as the core method. The authors should either revise the abstract and RQ1 answer to distinguish \"discussed in the literature\" from \"used in the paper,\" or explicitly report that no direct usage evidence was found for text simplification.","section":"Abstract, §4, §5, Table 2"},{"comment":"The methodology states that the pre-determined publication-year criterion is 2023, but the qualitative analysis includes several papers outside that window: [40], [50], [52], [65], [78], and [90] are from 2022, and [12] is from 2024. This contradicts the declared selection criterion and makes the claim about \"current\" NLP research ambiguous. If these papers were included intentionally (for example, to increase scenario coverage or because of publication-date conventions), the methodology must state this explicitly and report the inclusion criteria; otherwise the analysis should be restricted to 2023.","section":"§4, Tables 1–2"},{"comment":"The universal conclusion \"for all five application scenarios traditional models still exist\" rests on very small per-scenario samples: one retrieved paper for text simplification and three for information extraction. The abstract's phrasing \"in one way or another\" obscures this sparsity and suggests equal evidentiary strength across scenarios. The authors should report per-scenario counts and confidence levels in the abstract and conclusion, and qualify the strength of the claim accordingly.","section":"§4, §7, Figure 1"}],"minor_comments":[{"comment":"The abstract cites the complete statistics at https://zenodo.org/records/13683801, while the footnote in §4 cites https://doi.org/10.5281/zenodo.13683800; please unify the identifier so readers can locate the dataset.","section":"Abstract, §4 footnote"},{"comment":"The sentence claiming that \"nearly every retrieved publication has undergone a peer review process\" is stronger than the cited 80% peer-reviewed figure supports; please rephrase to \"the large majority\" or similar.","section":"§4"},{"comment":"The column headers \"Compe.\" and \"Tr. Be.\" are undefined; please spell them out as \"Competitive\" and \"Trailing Behind\" so the reader can interpret the table without referring back to the caption.","section":"Table 2"},{"comment":"The usage category \"comparison/baseline\" conflates two distinct functions. If the coding scheme is kept, please clarify whether a paper counts as comparison even when the traditional model is not presented as a baseline in the experimental design.","section":"§5, Table 2"},{"comment":"The answer to RQ3 is inferred by the authors from the selected models' properties rather than drawn from explicit discussion in the retrieved papers; this should be labeled as the authors' interpretation, not as a finding of the SLR.","section":"§6"},{"comment":"Please fix typographical errors such as \"Furtermore\" (§2.3), \"text simplificationas well astext summarization\" (Abstract), and \"applications scenarios\" (§7).","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you read it. First, the abstract overclaims: it says all five scenarios still use traditional models in one of three ways (pipeline, baseline, core), but for text simplification the only retrieved paper is a survey with no checkmarks in the paper's own Table 2 — the stress-test note is correct. Second, the '2023 corpus' is quietly salted with 2022 and 2024 papers in the qualitative tables, with no year markers. Both are real, and both are fixable.\n\nWhat's new: a systematic count, for 2023 ACM DL papers across classification, information/relation extraction, simplification, and summarization, of where traditional models still appear. That count — 28 of 119 papers, summarization the only task with over half using traditional methods — is original and reproducible, with full statistics on Zenodo. The pipeline/baseline/core categorization is a useful framing, and the paper is transparent about its search protocol, including the awkward fact that text simplification had to be searched by abstract because the title query returned nothing. The body is honest where it counts: it explicitly flags the two survey papers in Table 2 that fall in none of the three usage categories. The abstract just doesn't carry that nuance.\n\nSoft spots, in proportion. The biggest is the abstract vs. the evidence: with n=1 for text simplification, and that one paper a survey, the universal claim should have been softened to 'referenced in all five, actively used in four,' or reworded to match the body's looser RQ1 answer ('publications incorporating traditional models were retrieved'). The body's wording is fine; the abstract's enumeration is not. Second, the year-mixing in Tables 1-2 matters for a snapshot whose entire value is '2023' — at least seven entries are 2022 or 2024, and a reader can't tell without cross-checking the reference list. Third, the single-database, task-specific-definition, tiny-sample caveats are all acknowledged in Section 8 and are proportional to the paper's scoped claim.\n\nWho it's for: people writing related-work sentences about baselines persisting, and instructors choosing methods to teach. It's a citable dataset and a fair empirical snapshot, not a breakthrough.\n\nRecommendation: yes, send it to review. The reader's ACCEPT is defensible, but I'd require aligning the abstract with Table 2 and year-annotating the tables before publication.","headline":"A transparent, useful SLR snapshot whose abstract overclaims the text-simplification evidence and whose 2023 tables quietly include 2022/2024 papers — fixable, and worth refereeing.","tokens_in":16908,"tokens_out":7960,"would_cite":true,"duration_ms":68153,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that traditional NLP models—SVMs, rule-based extractors, extractive summarizers—still appear across all five surveyed tasks, in 28 of 119 papers.","keywords":["traditional NLP models","large language models","structured literature review","text classification","information extraction","relation extraction","text simplification","text summarization"],"falsifier":"Re-running the survey on a different database, a different year, or broader search terms to see whether traditional models still appear as core methods would settle it; for instance, if a comparable sample from 2025 proceedings shows zero or near-zero papers using rule-based, SVM, or extractive methods as the main approach outside a few low-resource domains, the paper's generalization holds only for its narrow snapshot. A sharper check is to take 50 recent relation-extraction papers from venues outside the original database and count how many use SVM, k-nearest neighbors, or hand-written rules as the main method; the paper's claim predicts several, and finding none would falsify it.","tokens_in":1466,"feed_emoji":"📚","tokens_out":2471,"duration_ms":87075,"temperature":0.7,"pith_summary":"Large language models dominate current NLP, but this paper asks whether the older, simpler techniques they displaced have actually disappeared. Surveying 119 papers published in one year across classification, information extraction, relation extraction, text simplification, and text summarization, it finds that 28 of them still use traditional models—support vector machines, naive Bayes, decision trees, rule-based extraction, or extractive summarization—either inside a processing pipeline, as comparison baselines, or as the main method. The paper's central claim is that traditional techniques remain live options in today's research, chosen deliberately for reproducibility, efficiency, and explainability rather than kept only as historical artifacts. That matters because it complicates the story of a wholesale neural takeover and points to concrete situations where simpler methods still earn their place.","feed_headline":"28 of 119 surveyed NLP papers still use traditional models","feed_subtitle":"SVMs, rule-based extraction, and extractive summarization still serve as baselines, pipeline steps, and main models.","key_machinery":"The survey works by fixing an operational definition of \"traditional\" and applying a three-way usage taxonomy to each paper. A model counts as traditional when it is deterministic and reproducible, cheap in time and hardware, and well documented or explainable; concretely this means SVMs, naive Bayes, decision trees, n-gram or TF-IDF features, rule-based extraction, or extractive summarization algorithms. The authors then classify each relevant paper by whether the traditional component sits in the pipeline, serves as comparison or baseline, or acts as the core model. This definition-and-taxonomy machinery is what turns scattered examples into a quantitative claim with a clear answer to each research question.","core_discovery":"Within the surveyed snapshot, every one of the five application scenarios contains papers using traditional approaches. Summarization shows the strongest presence, with extractive algorithms such as LexRank, TextRank, and lead-based selection appearing in over half of the retrieved papers, often as baselines or hybrid components; classification relies mainly on SVM and naive Bayes as comparisons; information extraction and relation extraction use rule-based systems or SVM/kNN as core methods in narrow domains; text simplification is the scarcest, surfacing only in lexical-complexity prediction. The authors interpret the pattern as evidence that traditional models persist in three clearly distinguishable roles—pipeline parts, baselines, and main models—and that their persistence is tied to properties LLMs do not reliably offer: deterministic repeatability, low compute cost, documentation, and freedom from hallucination in extractive settings.","pith_inferences":["Editorial extension: the paper's own data suggest the traditional-versus-modern boundary is not a timeline but a design trade-off, so rising energy and API costs could strengthen the case for deterministic low-cost methods in production.","Editorial extension: a testable follow-up would be to measure whether traditional methods cluster in low-resource languages or narrowly scoped domains; if so, they may be the default wherever annotated data or compute is scarce.","Editorial extension: because the paper reports usage rather than performance, one could extend it by collecting the relative scores of traditional baselines to identify when simplicity actually matches or beats modern models.","Editorial extension: since the search was limited to exact phrases in titles, the true prevalence of traditional methods may be higher than the reported counts, making the paper's figures a lower bound."],"forward_implications":["Evaluation practice should keep traditional baselines, because current papers still measure newer models against SVM, naive Bayes, and extractive systems.","Hybrid systems are a live design pattern: transformer models and large language models are combined with extractive selection or rule-based extraction in current work.","For hallucination-sensitive settings such as legal texts, news, or low-resource languages, extractive and rule-based methods remain a defensible choice that some surveyed papers explicitly motivate.","The uneven distribution—summarization rich in traditional methods, text simplification nearly absent—implies that continued relevance is task-dependent rather than uniform across NLP.","If traditional models still serve as reference points, then claims that a task is solved based only on LLM-versus-LLM comparisons are incomplete."],"supporting_citations":[{"why":"exemplifies a traditional SVM baseline that matched or beat modern models on some text-classification benchmarks.","marker":"[14]"},{"why":"shows a rule-based extraction model used as the main method for stock announcements.","marker":"[86]"},{"why":"shows SVM and k-nearest neighbors as the core classifiers in a hybrid Arabic relation-extraction approach.","marker":"[63]"},{"why":"shows traditional classifiers such as SVM, decision trees, and random forests in lexical-complexity prediction for text simplification.","marker":"[61]"},{"why":"shows TextRank used as the main model in a hybrid Chinese summarization system.","marker":"[52]"},{"why":"shows a legal-domain hybrid where BERT ranks sentences and a LexRank-based extractive model selects them.","marker":"[47]"},{"why":"shows graph-based extractive summarization with similarity measures and a Lead-300 baseline for a low-resource language.","marker":"[18]"},{"why":"shows extractive and traditional models used as baselines and comparisons in an abstractive summarization study.","marker":"[6]"}],"fun_headline_variants":["Traditional NLP methods appear in all five surveyed tasks","Survey: 28 of 119 papers still rely on traditional NLP","Classic NLP persists as baselines, pipelines, and core models","Extractive summarization keeps traditional NLP alive","Rule-based and SVM approaches endure in NLP survey"],"cache_read_input_tokens":19200,"weakest_assumption_plain":"The paper's conclusion rests on treating one year's papers from a single bibliographic database with exact-title searches as a fair sample of current NLP research; if that sample is unrepresentative, the claim that traditional models remain broadly relevant does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Traditional NLP methods appear in all five surveyed tasks","Survey: 28 of 119 papers still rely on traditional NLP","Classic NLP persists as baselines, pipelines, and core models","Extractive summarization keeps traditional NLP alive","Rule-based and SVM approaches endure in NLP survey"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000289,"raw_usage":{"total_tokens":1674,"prompt_tokens":906,"completion_tokens":768,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":522,"completion_tokens_details":{"reasoning_tokens":690}},"tokens_in":522,"tokens_out":768,"duration_ms":8315,"temperature":1.0,"reasoning_tokens":690,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:21:37.045640+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-running the survey on a different database, a different year, or broader search terms to see whether traditional models still appear as core methods would settle it; for instance, if a comparable sample from 2025 proceedings shows zero or near-zero papers using rule-based, SVM, or extractive methods as the main approach outside a few low-resource domains, the paper's generalization holds only for its narrow snapshot. A sharper check is to take 50 recent relation-extraction papers from venues outside the original database and count how many use SVM, k-nearest neighbors, or hand-written rules as the main method; the paper's claim predicts several, and finding none would falsify it.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"shows a rule-based extraction model used as the main method for stock announcements."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"shows SVM and k-nearest neighbors as the core classifiers in a hybrid Arabic relation-extraction approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"shows traditional classifiers such as SVM, decision trees, and random forests in lexical-complexity prediction for text simplification."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"shows TextRank used as the main model in a hybrid Chinese summarization system."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"shows a legal-domain hybrid where BERT ranks sentences and a LexRank-based extractive model selects them."}],"review_version":1}