{"id":"82f2461f-9fc7-4493-9396-321c6a3d061d","arxiv_id":"2501.03370","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A model-scan on an existing social support dataset whose claimed macro F1 improvements are contradicted by its own tables.","lead":"This paper benchmarks six transformer models, several zero-shot NLI models, and GPT-3.5/GPT-4 prompts on an existing YouTube dataset for detecting supportive versus non-supportive comments. It reports small gains that are internally inconsistent, and it does not release code or data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Internal contradiction in the headline result: RoBERTa-base Task 3 macro F1 is 0.7951 (Table 7, confirmed by Table 12) but 0.67 (Table 10) on the same normal dataset; the claimed across-the-board improvement rests on an unresolved inconsistency.","rationale":"The reader's verdict is REJECT with high confidence, and I agree the paper should not be accepted. My stress-test identifies a sharper, more decisive problem than the one named in the reader's weakest_assumption. The reader focused on possible differences in test splits and lack of confidence intervals; those concerns are real but secondary. The manuscript's own tables contradict each other for what is described as the same model, same task, and same normal dataset: Table 7 and Table 12 report Task 3 roberta-base macro F1 around 0.795, while Table 10 reports 0.67. Since Section 4.1 explicitly builds the balanced-versus-normal discussion on Table 10, and the conclusion builds the '7% and 8%' improvement claim on the same set of numbers, the central argument has an unresolved internal inconsistency. A single analytical check -- averaging the class F1-scores in Table 12 -- already shows which value is consistent with the per-class table, but the authors would need to rerun or clarify which table is authoritative. I also note the paper provides no public code, no data, no confidence intervals, and no multi-seed results, so there is no external way to adjudicate. These issues are not about novelty or framing; they concern the correctness of the headline empirical claim. The limitation section itself concedes over-reliance on a single architecture and limited generalizability, but that is not the main problem. The main problem is that the numbers supporting the claimed improvement are mutually inconsistent. I therefore keep the reader's REJECT verdict unchanged; no adjustment is needed.","tokens_in":12609,"tokens_out":5863,"duration_ms":50338,"concrete_test":"Analytical check: average the six Task-3 class F1-scores in Table 12: (0.7894 + 0.9173 + 0.9425 + 0.8519 + 0.8040 + 0.4776) / 6 = 0.7971. This matches Table 7's 0.7951, not Table 10's 0.67. Then rerun roberta-base on the same normal train/test split given in Table 6 with at least five seeds and report mean and standard deviation of macro F1 per task. If Task 3 is approximately 0.795, Table 10 is wrong and the balanced-data comparison must be recomputed; if Task 3 is approximately 0.67, Table 7 and Table 12 are wrong and the 'superior performance' claim loses Task 3 entirely.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that fine-tuned transformers achieve 'superior performance' and 'a notable increase in macro F1-scores across all three tasks' over the prior SVM baselines in Table 1. Section 4.0.1, Table 7 reports roberta-base Task 3 macro F1 as 0.7951, with accuracy 0.8725. Section 5, Table 12 gives the per-class F1-scores for the best Task 3 model, also identified as roberta-base: Other 0.7894, Nation 0.9173, LGBTQ 0.9425, Black Community 0.8519, Women 0.8040, Religion 0.4776. Averaging these six class F1-scores gives 0.7971, matching Table 7's 0.7951. Section 5 also reports Task 3 accuracy of 87.31%, again matching Table 7's 0.8725. Yet Section 4.1 and Table 10, also described as roberta-base on the normal dataset, report Task 3 macro F1 as 0.67 and accuracy 0.83. These two accounts of the same model and same normal split cannot both be correct. If Table 10 is the source for the headline comparison, then Task 3 actually decreases relative to the prior baseline (0.67 vs 0.7262), so the claim of improvement across all three tasks is false. If Table 7 and Table 12 are the correct results, then Table 10 and the subsequent balanced-versus-normal comparison (Task 3 dropping to 0.11) are misreported. Either way, the paper's central quantitative claim is not verifiable as written. The additional inconsistency between the abstract ('0.4% and 0.7% increase'), the conclusion ('7% and 8% increase'), and the actual deltas in Table 10 versus Table 1 (+6.3 points on Task 2, -5.6 points on Task 3) reinforces that the reported gains have not been pinned down.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates fine-tuned transformer models (BERT, DistilBERT, RoBERTa, etc.), zero-shot NLI models, and GPT-based zero-shot classifiers on a three-level social support detection dataset introduced by Ahani et al. (2024b). The authors report macro F1 scores for three tasks (support vs. non-support; individual vs. group; six group-support categories) and compare these against earlier SVM-based baselines. They also experiment with k-means based dataset balancing and claim improved macro F1 over prior work. The manuscript's central claim is that advanced models, especially roberta-base, achieve a 'notable increase' in macro F1 across all three tasks.","tokens_in":13011,"tokens_out":4561,"duration_ms":39918,"significance":"If the numerical results were internally consistent, the paper would provide a useful benchmark showing that pretrained transformers outperform traditional psycholinguistic/TF-IDF classifiers on a socially relevant task. The dataset itself, introduced in prior work, is a reasonable target for benchmarking. The authors are transparent about limitations, including class imbalance and potential bias in the dataset. However, the paper is currently not reliable as a benchmark contribution because the same model and dataset are reported with irreconcilable scores in different tables, and the claimed improvements over baselines are internally contradictory. The lack of statistical uncertainty, undocumented train/test split alignment with prior work, and the degenerate balanced-training experiment (15 examples per class) further undermine confidence in any conclusion.","major_comments":[{"comment":"The paper reports three mutually incompatible macro F1 values for the same model (roberta-base) on the same 'normal' Task 3 dataset: 0.7951 in Table 7, 0.67 in Table 10, and approximately 0.7971 (average of the per-class F1 scores in Table 12: 0.7894, 0.9173, 0.9425, 0.8519, 0.8040, 0.4776). These values cannot all be correct. If Table 10 is the source for the headline comparison, then Task 3 performance (0.67) is actually below the prior baseline (0.7262 in Table 1), contradicting the claim of improvement across all three tasks. If Tables 7 and 12 are instead correct, then Table 10 and the normal-versus-balanced discussion in §4.1 (which reports Task 3 normal F1 = 0.67) are misreported. The central quantitative claim is therefore unverifiable as written.","section":"§4.0.1, Table 7 vs §4.1, Table 10 vs §5, Table 12"},{"comment":"The claimed improvements over prior work are inconsistent across the paper. The abstract states a 0.4% increase for Task 2 and 0.7% for Task 3; the conclusion states a 7% increase for Task 2 and 8% for Task 3. Using Table 10 against Table 1, the absolute macro F1 deltas are +0.0631 (Task 2) and -0.0562 (Task 3). Using Table 7 against Table 1, the deltas are +0.0388 (Task 2) and +0.0689 (Task 3). None of these numbers match the abstract or the conclusion. This is not a presentation nuance; the paper does not identify which comparison is meant by 'percent increase,' and the two possible readings lead to opposite conclusions about Task 3.","section":"Abstract, §7 Conclusion, and §4.1/Table 10 vs Table 1"},{"comment":"The balanced dataset experiment for Task 3 is methodologically degenerate. Table 6 shows that the balanced training set contains exactly 15 examples per class for all six Task 3 categories, while the test set keeps its original tiny size (e.g., Religion: 4 test examples; Women: 7). Training a classifier on 15 examples per class and then comparing macro F1 against a model trained on the full data does not provide meaningful evidence about the effect of balancing. The reported collapse to 0.11 macro F1 is unsurprising and does not support the discussion in §4.1 about 'removal of valuable data'; it is an artifact of an underspecified and extreme undersampling procedure.","section":"§3.4, Table 6"},{"comment":"The comparison between this paper's transformer results and the prior baselines in Table 1 is only valid if the train/test splits are identical to those in Ahani et al. (2024b). The paper does not state how the 80/20-like split in Table 6 was generated, what random seed was used, or whether the prior work used the same split. Without this information, the reported deltas (positive or negative) may reflect different evaluation sets rather than model quality. This is a load-bearing omission for a benchmark paper whose main claim is comparative.","section":"§3.4 and §4.1 (baseline comparison protocol)"},{"comment":"All experimental results appear to be single-run point estimates with no confidence intervals, significance tests, or variance information. This is especially problematic for Task 3, where Table 6 lists classes with 4, 7, and 26 test examples. A change of one or two predictions in the Religion or Women class shifts macro F1 by several points. The paper's claims that one model 'outperforms' another are not supported without accounting for this noise; this applies to the transformer comparisons and to the zero-shot and GPT comparisons.","section":"§4.0.1, Table 7; §4.0.2, Table 8; §4.0.3, Table 9"}],"minor_comments":[{"comment":"Table 8 is captioned 'Results for different zero-shot models,' while Table 9 (the GPT models) is also captioned 'Results for different zero-shot models.' The captions should be distinguished so that readers know which table contains the Hugging Face NLI models and which contains the GPT results.","section":"Table 8 and Table 9 captions"},{"comment":"The text says 'A graphical representation of the confusion matrix is presented in Table 12,' but Table 12 is a numerical table of per-class precision, recall, and F1, not a confusion matrix. The confusion matrices in Figures 2-4 are referenced but not visible in the manuscript text; the figures should be included or the text should refer only to the table.","section":"§5, Table 12"},{"comment":"The sentence 'The parameter details can be found in Table 3' appears to refer to the prompt details that are actually given in Table 5. Please correct the cross-reference.","section":"§3.3.2 and Table 5"},{"comment":"The sentence 'also in this part you can find the Overview of Contributions' is an incomplete and informal placeholder; it should be removed or rewritten as a proper transition.","section":"§1"},{"comment":"The reference list contains incomplete entries, such as 'Ahani, Z., Tash, M., Zamir, M., Gelbukh, I., 2024a' with missing page numbers, and the Statista reference is missing its title. Please complete all bibliographic entries.","section":"References"}],"recommendation":"reject","confidential_remarks":"This manuscript appears to be a preliminary preprint that has not been checked for internal numerical consistency. The contradiction between Tables 7, 10, and 12 alone is sufficient to reject, and the additional mismatch between the abstract, conclusion, and tables makes the headline result unrecoverable without a full re-run of all experiments. The authors should be encouraged to redo the evaluation, document the split protocol, and report uncertainty before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Nothing in this paper's headline claim survives contact with its own tables. The abstract says +0.4% and +0.7% over prior work, the conclusion says +7% and +8%, and Table 10 versus Table 1 gives +6.3 points on Task 2 and -5.6 points on Task 3. Worse, Table 7 reports roberta-base Task 3 macro F1 as 0.7951 with accuracy 0.8725, and Table 12's per-class numbers average to 0.7971, so those two are consistent. But Table 10, also labeled roberta-base on the same normal dataset, reports 0.67 and 0.83. Both cannot be right. The central claim of improvement across all three tasks rests on which table you pick; one of them shows a drop on Task 3.\n\nWhat the paper does well is the breadth of the scan. Six transformers, five zero-shot NLI models, three GPT variants, and a k-means balancing experiment is a lot of off-the-shelf evaluation, and the per-class error analysis for the best model is genuinely informative. The zero-shot prompts are quoted in full, and the authors are honest that balancing hurts or collapses performance. That is credit where it is due. But the techniques are not new, the dataset is from the authors' own prior work, and there is no code or data release - only 'available upon request.' No seeds, no confidence intervals, no hyperparameters. Several Task 3 test classes contain four to seven examples, so single-run macro F1 differences are likely noise. The limitation section concedes over-reliance on roberta-base and generalizability concerns, which is honest but undercuts the practical claim further.\n\nI agree with the stress-test note: the internal contradiction is real and load-bearing. This looks like a paper assembled from different experimental runs without checking that the tables describe the same setup. The fix is not a patch; the authors would need to re-run and report a single consistent evaluation protocol. Even if that were done, the contribution is a model-scan on an existing dataset, so the extra insight would be modest. I would not send this to a referee in its current state. It should go back to the authors to reconcile the numbers, or be withdrawn. There is enough useful error analysis here to salvage a shorter, corrected version - but only if the numbers are made to agree.","headline":"The paper's own tables contradict each other on the headline result, so the claimed improvement over prior work is not verifiable as written.","tokens_in":13642,"tokens_out":3517,"would_cite":false,"duration_ms":29758,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuned RoBERTa-base outperforms earlier SVM and TF-IDF classifiers on detecting social support in YouTube comments, the paper reports.","keywords":["social support detection","transformers","RoBERTa","zero-shot learning","GPT-4","class imbalance","K-means clustering","YouTube comments"],"falsifier":"Re-run RoBERTa-base and the best prior classifier on the exact shared train/test split from the dataset paper across at least five random seeds, then compare confidence intervals for macro $F_1$; the across-the-board improvement claim would fail if the Task 3 score does not rise above the old baseline or if the intervals overlap on Task 1 or Task 2.","tokens_in":12366,"feed_emoji":"💬","tokens_out":8525,"duration_ms":71917,"temperature":0.7,"pith_summary":"This paper tries to establish that modern pretrained transformers are the best available method for detecting social support in social-media comments. The authors fine-tune several transformer encoders on a 10,000-comment YouTube dataset and report that RoBERTa-base reaches macro $F_1$ of 0.80 on supportive-versus-non-supportive, 0.86 on individual-versus-group, and 0.67 on the six-way type-of-support task, improving on earlier SVM and TF-IDF/LIWC results in the first two tasks. They also test zero-shot GPT-3/GPT-4 prompting and K-means dataset balancing, finding that zero-shot lags behind fine-tuning and that balancing hurts. The paper matters because social-support classification could give content moderation and mental-health research a positive signal rather than only hate-speech filtering, and the reported results suggest a task that an off-the-shelf transformer can handle.","feed_headline":"RoBERTa beats SVM and TF-IDF at spotting supportive comments","feed_subtitle":"Fine-tuned transformer reaches macro F1 of 0.80, 0.86, 0.67 across three social-support tasks.","key_machinery":"The load-bearing mechanism is fine-tuning a pretrained transformer encoder, with RoBERTa-base (a BERT-style English encoder pretrained on more text and for longer) as the best performer. Its self-attention layers replace hand-built psycholinguistic and TF-IDF features by learning contextual representations of supportiveness, addressee type, and support category directly from comments. For comparison, the paper evaluates zero-shot classification through NLI-style DeBERTa and BART models and through GPT-3, GPT-4, and GPT-4-o with prompt voting, and it uses K-means clustering to undersample majority classes and replicate minority classes in the balanced experiments.","core_discovery":"The paper's central claim is that fine-tuning a pretrained transformer, specifically RoBERTa-base, lifts social-support classification above the earlier SVM with TF-IDF and LIWC features. On the unbalanced test set, the paper reports RoBERTa-base macro $F_1$ scores of 0.80, 0.86, and 0.67 for Tasks 1, 2, and 3, against 0.7830, 0.7969, and 0.7262 for the prior best classifiers. The authors read the overall pattern as showing that transformer fine-tuning gives superior performance across the three tasks, and they also report that K-means balancing reduced Task 1 and Task 2 scores and collapsed Task 3 from 0.67 to 0.11. The paper's own tables show the largest gains on Task 2 and a Task 3 score below the listed prior baseline, and its Limitation section says the results may be tied to this specific dataset and architecture.","pith_inferences":["If the paper's numbers are taken at face value, the 'across all three tasks' phrasing is stronger than the evidence: Task 3 is below the prior baseline, so a fair reading is that transformers clearly help on binary and direction tasks but the six-way type task remains open.","A natural follow-up is to replace K-means balancing with class-weighted losses or label-preserving oversampling; the sharp Task-3 drop from 0.67 to 0.11 is a ready-made testbed.","The extreme scarcity of Religion and Women test samples, 4 and 7 examples respectively, means the reported Task-3 macro scores are highly sensitive to a handful of comments; collecting more annotations for those classes would likely change rankings.","The same protocol could be carried to other platforms or languages using multilingual checkpoints, which the paper does not test; that would show whether the gains are tied to YouTube comment style or transfer."],"forward_implications":["A standard-size pretrained transformer, rather than a massive model or elaborate prompt pipeline, is enough to make social-support detection practical.","Traditional feature engineering with LIWC and TF-IDF becomes a weaker default once fine-tuned transformers are available for this task.","K-means undersampling is not advisable for highly skewed support-type labels; it can destroy performance on minority classes such as Religion and Women.","Zero-shot LLM prompting is close to older supervised classifiers but still below fine-tuning, so labeled examples for this task continue to pay.","A reliable supportive-comment classifier would let platforms filter for positive engagement instead of only detecting toxic content."],"supporting_citations":[{"why":"Supplies the 10,000-comment YouTube dataset, the three task definitions, and the prior SVM/soft-voting baseline scores in Table 1.","marker":"Ahani et al. (2024b)"},{"why":"Introduces BERT and the fine-tuning recipe that the transformer experiments follow.","marker":"Devlin et al. (2018)"},{"why":"Introduces the Transformer attention architecture underlying all fine-tuned and zero-shot models tested.","marker":"Vaswani et al. (2017)"},{"why":"Establishes GPT-3 and few-shot/zero-shot prompting, the basis for the GPT-3-Turbo experiments.","marker":"Brown et al. (2020)"},{"why":"Describes GPT-4, the model family used for the GPT-4-Turbo and GPT-4-O zero-shot runs.","marker":"Achiam et al. (2023)"},{"why":"Supplies the comparison of LLMs for text classification from zero-shot to fine-tuning that motivates the experimental design.","marker":"Chae and Davidson (2023)"},{"why":"Provides the clustering-based undersampling method that the paper adapts for K-means dataset balancing.","marker":"Lin et al. (2017)"}],"fun_headline_variants":["RoBERTa beats SVM on social support, but not all tasks","Fine-tuned RoBERTa lifts social support detection, except one task","Transformer wins social support tasks, but balancing backfires","RoBERTa tops social support detection, K-means hurts","Social support AI: RoBERTa gains, but balancing fails"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the assumption that the earlier SVM and TF-IDF baselines were evaluated under the same train/test split and that single-run macro $F_1$ differences are not noise; the paper does not establish either.","fun_headline_variants_meta":{"raw":{"variants":["RoBERTa beats SVM on social support, but not all tasks","Fine-tuned RoBERTa lifts social support detection, except one task","Transformer wins social support tasks, but balancing backfires","RoBERTa tops social support detection, K-means hurts","Social support AI: RoBERTa gains, but balancing fails"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0006,"raw_usage":{"total_tokens":2830,"prompt_tokens":997,"completion_tokens":1833,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":1746}},"tokens_in":613,"tokens_out":1833,"duration_ms":11494,"temperature":1.0,"reasoning_tokens":1746,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:52:15.096310+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run RoBERTa-base and the best prior classifier on the exact shared train/test split from the dataset paper across at least five random seeds, then compare confidence intervals for macro $F_1$; the across-the-board improvement claim would fail if the Task 3 score does not rise above the old baseline or if the intervals overlap on Task 1 or Task 2.","supporting_citations":[],"review_version":1}