{"id":"9e08d70e-cc65-4684-ad5a-25fa752dc828","arxiv_id":"2606.29859","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Applies pretrained deep learning models with data augmentation to classify algorithm mention motivations in NLP papers, reporting that direct use dominates and motivation diversity has increased over time.","lead":"The paper develops a sentence-level deep learning framework to classify motivations for mentioning algorithms in NLP research papers, such as direct use or description. The analysis of trends could support better tracking of how methods spread and gain value in scientific fields.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Annotation reliability and sampling strategy are the load-bearing assumptions for all distributional and temporal claims","rationale":"The reader's weakest_assumption matches the single point on which every quantitative result depends. Because the work is purely empirical and contains no parameter-free derivation or external validation set, annotation quality is the only place where the argument can fail internally. A concrete check on agreement and sampling directly tests whether that failure occurs.","tokens_in":1760,"tokens_out":340,"duration_ms":14063,"concrete_test":"Locate the methods subsection on corpus construction and annotation; extract the reported inter-annotator agreement metric (Cohen's kappa or equivalent) and the exact sampling frame (number of papers, conferences, year range, stratification). If kappa is unreported or <0.65, or if sampling is non-stratified, re-annotate a 200-sentence hold-out set with two new annotators and recompute the motivation distribution; a shift >8 percentage points in the 'use' or 'description' category falsifies the central empirical claims.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"All headline results (motivation frequencies, temporal replacement of description by use, category-specific patterns, and DL vs. traditional ML performance) are derived from a manually annotated corpus of algorithm-related sentences. The paper provides no reported inter-annotator agreement, no explicit definition or validation of the motivation taxonomy exhaustiveness, and no description of how papers were sampled across venues and years. If annotator consistency is low or the sample over-represents recent arXiv preprints, both the classification accuracy numbers and the evolutionary trends become unreliable.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes a sentence-level framework for identifying algorithm entities and related sentences in NLP papers via manual annotation and machine learning, then classifying mention motivations (e.g., use, description, improvement) with pretrained deep learning models augmented by data augmentation techniques. It reports that DL models outperform traditional ML, that over half of algorithm-related sentences express direct use (with improvement least frequent), that motivation diversity has increased over time while the number of types per algorithm has declined, and that use motivations have gradually replaced description motivations, with category-specific patterns (e.g., grammar-based algorithms more often described, ML algorithms more often used).","tokens_in":1847,"tokens_out":395,"duration_ms":22125,"significance":"If the annotated corpus proves reliable and representative, the work supplies a scalable method for tracing how algorithms are invoked in scientific writing and yields concrete distributional and diachronic findings that could support downstream tasks such as algorithm relationship extraction and impact assessment. The empirical model comparison and temporal analysis constitute the primary contributions.","major_comments":[{"comment":"Abstract and dataset description: The manuscript reports performance gains from deep learning models with augmentation and all distributional/temporal claims, yet supplies no dataset size, inter-annotator agreement, sampling protocol across venues/years, or error analysis. These omissions are load-bearing because every headline result (motivation frequencies, use replacing description, category patterns) derives directly from the manually labeled corpus.","section":"Abstract"},{"comment":"Annotation and taxonomy section: No validation is provided for the exhaustiveness or consistent application of the motivation categories, nor any measure of annotator reliability. Without these, the claims that improvement is least frequent, that use has replaced description over time, and that diversity has increased cannot be assessed for robustness.","section":"Methods / Annotation"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments. We address each major comment below and will make the indicated revisions to improve the transparency of the annotation process and corpus details.","responses":[{"response":"We agree that these details are essential for evaluating the reliability of the reported findings. Although the Methods section outlines the annotation workflow, the manuscript does not report dataset size, inter-annotator agreement, sampling protocol, or error analysis in the abstract or dataset description. We will add a dedicated subsection on corpus construction that includes these elements (number of sentences and papers annotated, Cohen's kappa for agreement, venue/year sampling strategy, and error analysis) so that the distributional and temporal claims can be properly assessed.","revision_made":"yes","referee_comment":"[Abstract] Abstract and dataset description: The manuscript reports performance gains from deep learning models with augmentation and all distributional/temporal claims, yet supplies no dataset size, inter-annotator agreement, sampling protocol across venues/years, or error analysis. These omissions are load-bearing because every headline result (motivation frequencies, use replacing description, category patterns) derives directly from the manually labeled corpus."},{"response":"We acknowledge that the current version does not include explicit validation of taxonomy exhaustiveness or annotator reliability metrics. The motivation categories were developed through iterative pilot annotation, but this process and any agreement statistics are not reported. In the revision we will add a description of the taxonomy development (including how coverage of observed motivations was verified) together with inter-annotator agreement figures for the motivation labels, thereby supporting the robustness of the frequency and diachronic claims.","revision_made":"yes","referee_comment":"[Methods / Annotation] Annotation and taxonomy section: No validation is provided for the exhaustiveness or consistent application of the motivation categories, nor any measure of annotator reliability. Without these, the claims that improvement is least frequent, that use has replaced description over time, and that diversity has increased cannot be assessed for robustness."}],"tokens_in":1407,"tokens_out":434,"duration_ms":25832,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The punchline is that this paper sets up a new classification task for the motivations behind algorithm mentions in NLP papers and applies deep learning to it, but the absence of basic validation on the annotations makes the distributional and temporal results hard to rely on.\n\nWhat stands out as new is the focus on sentence-level motivation identification for algorithms specifically in the NLP domain. They use pretrained models with data augmentation to classify things like direct use, description, comparison, and improvement. The abstract shows they get better results than traditional ML and pull out some patterns, such as use being the most common motivation and a shift away from description over time. That framing could be useful for people tracking how methods spread in the literature.\n\nThe paper does a decent job of laying out the task and running the models on it. Credit for identifying that grammar-based algorithms get mentioned differently than machine learning ones.\n\nThe soft spots are in the data side. As the stress-test note says, there's no mention of inter-annotator agreement, dataset size, or how the papers were sampled across time and venues. All the claims about frequencies, category differences, and the replacement of description by use depend on those annotations being solid and the sample being fair. If the annotators disagreed a lot or if recent papers dominate the set, the trends don't hold. That's a real issue, not a minor one, because the empirical findings are the main output.\n\nThis work is for scientometric researchers or NLP people interested in analyzing their own field's writing. A reader who wants to see how DL can be used for literature mapping might get some ideas, but without the validation details it's difficult to take the specific numbers seriously.\n\nI'd recommend engaging with it only after they provide the annotation process and sampling strategy in detail. It could go to peer review if those are addressed, as the task is novel enough to warrant a look, but right now the evidence is too thin.","headline":"This paper defines a fresh task for classifying algorithm mention motivations in NLP but its findings depend on unvalidated manual annotations.","tokens_in":2313,"tokens_out":458,"would_cite":false,"duration_ms":30370,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Deep learning classifiers on NLP papers find direct use as the most common motivation for algorithm mentions while improvement is the rarest.","keywords":["algorithm mention motivations","natural language processing","deep learning classification","academic writing analysis","temporal evolution","data augmentation"],"falsifier":"Independent re-annotation of a random sample of the same full-text papers yielding substantially different motivation distributions or lower inter-annotator agreement on the four main categories would undermine the reported proportions and temporal trends.","tokens_in":2646,"feed_emoji":"📊","tokens_out":693,"duration_ms":26298,"temperature":0.7,"pith_summary":"The paper builds a sentence-level system to detect algorithm entities in full-text NLP papers, extract related sentences, and classify the author's purpose for mentioning each algorithm. Pretrained deep learning models trained on manually annotated data plus augmentation achieve better accuracy than traditional machine learning for this classification task. Analysis of the resulting labels shows more than half of algorithm-related sentences express direct use, improvement is least frequent, grammar-based algorithms are mentioned more for description while machine learning algorithms are mentioned more for use, and use motivations have steadily replaced description motivations across the literature. The diversity of motivations tied to any single algorithm has also declined over time. A reader can see how these patterns trace the changing roles algorithms play in research writing.","feed_headline":"Direct use dominates algorithm mentions in NLP papers","feed_subtitle":"Full-text analysis shows improvement rarest, use replacing description over time, and machine learning algorithms cited differently than gra","key_machinery":"Sentence-level motivation classification model that first identifies algorithm entities and algorithm-related sentences via manual annotation and machine learning, then assigns one of several motivation labels using pretrained deep learning models trained with data augmentation.","core_discovery":"Deep learning models trained with augmented data outperform traditional machine learning models in motivation classification. In NLP papers, more than half of algorithm-related sentences express direct use, whereas improvement is the least frequent motivation. Grammar-based algorithms are more often mentioned for description, while machine learning algorithms are more often mentioned for use. Over time, use motivations have gradually replaced description motivations across different algorithms, and the number of motivation types associated with individual algorithms has declined significantly.","pith_inferences":["The same extraction and classification pipeline could be run on papers from other domains to test whether use-versus-description patterns are unique to NLP.","The drop in the number of motivation types per algorithm may reflect growing specialization; this could be checked by linking the motivation labels to citation or co-occurrence graphs.","Literature-search tools could surface papers according to the dominant motivation type attached to a given algorithm rather than simple keyword matches."],"forward_implications":["Motivation patterns can serve as input for identifying relationships among algorithms.","Mention motivations provide one measurable basis for evaluating the roles and value of algorithms in research.","Observed shifts from description to use motivations indicate how algorithms move from being introduced to being applied in practice.","Category-specific differences show that grammar-based and machine learning algorithms follow distinct mention trajectories."],"fun_headline_variants":["Use dominates NLP algorithm mentions","Improvement least frequent algorithm mention motivation","Grammar-based for description, ML for use in NLP","Use replaces description in algorithm mentions over time","Motivation types per algorithm decline significantly"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The manually annotated dataset of algorithm entities and sentences is accurate, representative of the full NLP literature, and the motivation categories are exhaustive and consistently applied by annotators.","fun_headline_variants_meta":{"raw":{"variants":["Use dominates NLP algorithm mentions","Improvement least frequent algorithm mention motivation","Grammar-based for description, ML for use in NLP","Use replaces description in algorithm mentions over time","Motivation types per algorithm decline significantly"]},"model":"grok-4.3","cost_usd":0.004598,"raw_usage":{"total_tokens":2218,"prompt_tokens":704,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":45978000,"prompt_tokens_details":{"text_tokens":704,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1454,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":704,"tokens_out":60,"duration_ms":11298,"temperature":1.0,"reasoning_tokens":1454,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T06:02:46.143054+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Independent re-annotation of a random sample of the same full-text papers yielding substantially different motivation distributions or lower inter-annotator agreement on the four main categories would undermine the reported proportions and temporal trends.","supporting_citations":[],"review_version":1}