Pith. sign in

It's about Time: Rethinking Evaluation on Rumor Detection Benchmarks using Chronological Splits

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

New events emerge over time influencing the topics of rumors in social media. Current rumor detection benchmarks use random splits as training, development and test sets which typically results in topical overlaps. Consequently, models trained on random splits may not perform well on rumor classification on previously unseen topics due to the temporal concept drift. In this paper, we provide a re-evaluation of classification models on four popular rumor detection benchmarks considering chronological instead of random splits. Our experimental results show that the use of random splits can significantly overestimate predictive performance across all datasets and models. Therefore, we suggest that rumor detection models should always be evaluated using chronological splits for minimizing topical overlaps.

citation-role summary

background 1

citation-polarity summary

fields

cs.SI 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

unclear 1

representative citing papers

Framework of Voting Prediction of Parliament Members

cs.SI · 2025-05-18 · reject · novelty 4.0

A multi-country framework predicts individual parliamentary votes with up to 85% accuracy and bill outcomes with up to 84% accuracy, but the evaluation does not include trivial baselines.

citing papers explorer

Showing 1 of 1 citing paper.

  • Framework of Voting Prediction of Parliament Members cs.SI · 2025-05-18 · reject · none · ref 56 · internal anchor

    A multi-country framework predicts individual parliamentary votes with up to 85% accuracy and bill outcomes with up to 84% accuracy, but the evaluation does not include trivial baselines.