Pith. sign in

REVIEW 4 major objections 5 minor 59 references

It Takes Nine to Smell a Rat: Neural Multi-Task Learning for Check-Worthiness Prediction

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper claims that training one model on the choices of nine fact-checking organizations improves per-source check-worthiness ranking and beats prior systems on the 2016 debate dataset.

desk verdict A sensible, well-scoped MTL application for check-worthiness with an honest but statistically fragile central claim. read the letter →

arxiv 1908.07912 v1 pith:QKJ4H442 submitted 2019-08-19 cs.CL cs.AI

classification cs.CLcs.AI
keywords check-worthinesspredictionmulti-tasklearninghardparametersharingfact-checkingpoliticaldebatesclaimrankingneuralnetworksCW-USPD-2016
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether a system that must imitate one fact-checking organization's choices can improve by simultaneously learning to imitate eight others. It claims yes: a multi-task neural network with one shared layer and nine task-specific outputs consistently outperforms single-task models on ranking which sentences in the 2016 US presidential and vice-presidential debates deserve fact-checking. The average gain is modest in absolute terms, with Mean Average Precision rising from 0.127 for the single-task baseline to 0.136 for multi-task, but it appears across most sources and all evaluation measures. The practical payoff is that scarce human fact-checking effort could be better prioritized by a model that internalizes several editorial judgments at once.

What carries the argument

The load-bearing object is a hard parameter sharing neural network: a 300-unit ReLU hidden layer shared by all tasks, followed by ten parallel task-specific layers and ten sigmoid outputs, one per fact-checking source plus the cumulative ANY task. Input features combine ClaimBuster-style cues such as TF.IDF bag of words, part-of-speech tags, named entities, sentiment, and sentence length with context features from prior work, including bias and assertiveness lexicons, position within the debate and intervention, LDA topics, discourse relations, and pretrained word embeddings. The shared layer is what carries multi-task transfer: backpropagation from every source updates the same weights, so the model is forced to represent check-worthiness in a way that serves all nine editors at once. Ablation experiments show that eliminating the embeddings produces the largest drop in performance.

What would settle it

Retrain the singleton and multi-task models on the same CW-USPD-2016 data with 30 random seeds and compute seed-level confidence intervals for each source; if the multi-task advantage on MAP is smaller than the within-model seed spread for most sources, the claimed transfer effect would not be distinguishable from noise. Alternatively, run both models on a second set of debates with the same nine sources; the claim predicts that multi-task should win on average across sources there too.

Watch

Extended reading notes

Core claim

The central claim is that hard parameter sharing across nine fact-checking sources transfers useful signal: the shared hidden layer learns general cues of check-worthiness while each task-specific layer adapts to a particular outlet's editorial taste. On the CW-USPD-2016 corpus, multi-task learning beats the per-source singleton model on averaged Mean Average Precision, R-Precision, and precision at 5, 10, 20, and 50. The New York Times is the exception, as its single-task scores remain highest, and the authors attribute this to distinctive features that blur under joint optimization. Including an extra ANY output that predicts whether any source would select a sentence does not help, because that information is already implicit in the nine tasks. The paper's conclusion is that it pays to learn from multiple fact-checkers simultaneously even when the final goal is to mimic only one.

Load-bearing premise

The central claim assumes that averaging over three random-seed runs on a 5,415-sentence, four-debate corpus makes small differences like 0.127 versus 0.136 in Mean Average Precision meaningful, even though no significance test or confidence interval is reported.

Editorial extensions

If this is right

  • A single deployed model can rank claims for several fact-checking organizations at once, rather than training and maintaining nine separate systems.
  • Sources with sparse labels can borrow signal from sources with different or more numerous selections; removing any one target from training, including NYT, worsens results for the others.
  • The ANY label is redundant once all nine individual labels are predicted, so resources need not be spent on a separate union-of-sources classifier.
  • The reported state-of-the-art result on CW-USPD-2016 gives automatic fact-checking pipelines a stronger first stage for prioritizing claims.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same architecture should transfer to other domains where multiple annotators disagree yet share a latent notion, such as news-worthiness, toxicity, or relevance judgment; the natural test is whether adding annotation sources that disagree with the target helps the target as it does here.
  • The NYT exception suggests that a variant with soft parameter sharing or per-source weighting could outperform hard sharing, and a direct comparison would quantify how much is lost by forcing one shared layer.
  • Because gains are averaged over only four debates and three seeds, a robustness study with more debates or more reruns would either confirm that the transfer effect is stable or show that it falls within seed variance.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a multi-task learning framework for check-worthiness prediction in political debates. Using the CW-USPD-2016 dataset, which contains 5,415 sentences from four 2016 US debates annotated by nine fact-checking organizations, the authors train a hard-parameter-sharing neural network that simultaneously predicts whether each source would select a sentence, plus an additional ANY task. The model uses a rich feature set adopted from prior work (TF.IDF, POS, named entities, sentiment, discourse, embeddings, etc.). Evaluation is performed with 4-fold leave-one-debate-out cross-validation, repeated over three random seeds, and the paper reports MAP, R-Precision, and P@k. The main result is that the multi-task model outperforms single-task baselines on most sources and for most measures, with average MAP increasing from .127 to .136, and the authors claim state-of-the-art results.

Significance. If the results are reliable, the paper provides evidence that multi-task learning across multiple fact-checking sources yields better per-source check-worthiness ranking than single-task models, and the feature and source ablations offer insight into positive and negative transfer between sources. The strengths of the paper include the use of a public dataset, comparison to an external baseline (ClaimBuster), and detailed ablation analysis. However, the empirical support is limited by the small dataset (four debates), the absence of significance tests or confidence intervals, and a per-source exception (NYT) that contradicts the unqualified headline claim. The methodological contribution is largely an application of existing hard-parameter-sharing MTL, so the paper's value rests on the soundness of the evaluation.

major comments (4)
  1. [Section 5, Tables 3 and 4] The central claim that multi-task learning 'pays' is based on small average improvements (e.g., MAP .127 to .136) with no reported variance, confidence intervals, or significance tests; with only four test debates and three seeds, these differences may well be noise. Please report per-fold and per-seed results and include a paired significance test (e.g., bootstrap over sentences or a signed-rank test over folds) to support the assertion of consistent improvement.
  2. [Section 5, Table 3 (NYT row) and Section 7] The abstract states that it pays to learn from multiple sources 'even when a particular source is chosen as a target to imitate,' but NYT's MAP drops from .187 (singleton) to .150 (multi). Although Section 5 acknowledges this exception, the abstract and conclusion repeat the unqualified claim, which is contradicted by the paper's own results; the claim should be tempered or the exception disclosed prominently.
  3. [Section 5, Tables 3-4] The 'state-of-the-art' claim is supported only by comparisons to ClaimBuster (2015) and the authors' own previous singleton system (Gencheva et al., 2017); no comparison is made to systems from the CLEF CheckThat! tasks (Atanasova et al., 2018, 2019) that address the same task on closely related data. Please justify the absence of such comparisons or soften the claim accordingly.
  4. [Section 6, Figure 2] The source ablation reports MAP differences (e.g., ABC drops .008 when CT is removed; FC improves after removing CT) without any indication of variance. Given the small evaluation corpus, these differences may not be reliable, and the conclusions about task conflicts and shared information should be backed with per-fold results or error bars.
minor comments (5)
  1. [Section 6] The feature 'Sim. to prev.' is not defined in this paper; please provide a definition or an explicit reference to the description in Gencheva et al. (2017), as it is part of the input to all compared models.
  2. [Section 5, footnote 7] The statement that Patwari et al. (2017) 'would perform similarly to ClaimBuster' is speculative and should be removed or substantiated with experiments.
  3. [Section 2] The word 'truthiness' appears to be a typo for 'truthfulness' (or is used intentionally as a play on words; if so, please clarify).
  4. [Table 5] Please add a caption note defining the column abbreviations 'N', 'Tgt', and '#' so the table is self-contained.
  5. [Figure 2] The figure is difficult to read in the provided version; please ensure high resolution and add a color scale and axis labels for the MAP differences.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation; the multi-task improvement claim rests on held-out evaluation against an external baseline, with only minor non-load-bearing self-citations.

full rationale

The paper's central claim is empirical: multi-task training on nine fact-checking sources improves per-source check-worthiness ranking over a single-task model. This is established by leave-one-debate-out evaluation on CW-USPD-2016, with labels defined by each source's actual fact-checking decisions, and by comparison to the external ClaimBuster system as well as to the authors' own singleton baseline. No equation defines a prediction in terms of a fitted parameter; no parameter is fitted to the held-out debates; and no uniqueness theorem or prior-work assertion is used to rule out alternatives. The self-citations to Gencheva et al. (2017) supply the public dataset and feature definitions, and the singletonG baseline, but the improvement of multi over singleton is measured by the authors' own controlled reimplementation, not taken on citation. The ANY task is a transparently derived union of the nine labels, and the paper explicitly reports that adding it does not help, so it is not used as independent evidence. The lack of significance testing and per-fold variance is a statistical robustness concern, not circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the validity of the dataset annotations, the reliability of the small evaluation setup, the construction of the 'Sim. to prev.' feature, and the modeling choice of hard parameter sharing. No new entities or fitted constants are introduced beyond standard neural network weights.

assumptions (4)
  • domain assumption The CW-USPD-2016 dataset provides valid gold labels for check-worthiness.
    The evaluation treats the nine fact-checking organizations' selections as ground truth. The dataset is from the authors' prior work (Gencheva et al., 2017).
  • domain assumption Leave-one-debate-out with four debates and three random seeds yields reliable performance estimates.
    No significance tests or confidence intervals are provided; the dataset has only 5,415 sentences and four debates.
  • domain assumption The 'Sim. to prev.' feature is computed without using gold labels from the test debate.
    The paper does not specify how this feature is constructed; if it uses target labels, the evaluation is compromised.
  • domain assumption Hard parameter sharing is an effective inductive bias for modeling multiple fact-checking sources.
    This is the paper's core modeling choice, and it is not proved but empirically tested.

how reviews work

0 comments
Cite this review

Pith. "Pith review of It Takes Nine to Smell a Rat: Neural Multi-Task Learning for Check-Worthiness Prediction." pith.science (2026). https://pith.science/paper/QKJ4H442

@misc{pith2026190807912,
  author       = {Pith},
  title        = {Pith review of: It Takes Nine to Smell a Rat: Neural Multi-Task Learning for Check-Worthiness Prediction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QKJ4H442}},
  note         = {Machine review of arXiv:1908.07912}
}
read the original abstract

We propose a multi-task deep-learning approach for estimating the check-worthiness of claims in political debates. Given a political debate, such as the 2016 US Presidential and Vice-Presidential ones, the task is to predict which statements in the debate should be prioritized for fact-checking. While different fact-checking organizations would naturally make different choices when analyzing the same debate, we show that it pays to learn from multiple sources simultaneously (PolitiFact, FactCheck, ABC, CNN, NPR, NYT, Chicago Tribune, The Guardian, and Washington Post) in a multi-task learning setup, even when a particular source is chosen as a target to imitate. Our evaluation shows state-of-the-art results on a standard dataset for the task of check-worthiness prediction.

Figures

Figures reproduced from arXiv: 1908.07912 by the authors.

Figure 1
Figure 1. The architecture of our neural multi-task [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Ablation experiment with the multi model. Each row is an experiment re￾moving one target. Each column is the MAP difference with respect to the multi model for the corresponding target. It would be interesting to investigate the reasons why the NYT source does not benefit from the multi-task architecture. In order to adapt to this situation with a single model, we plan to experi￾ment with a network with soft paramet… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

59 extracted references · 36 canonical work pages

  1. [1]

    Pepa Atanasova, Llu\' i s M\` a rquez, Alberto Barr\' o n-Cede\ n o, Tamer Elsayed, Reem Suwaileh, Wajdi Zaghouani, Spas Kyuchukov, Giovanni Da San Martino, and Preslav Nakov. 2018. Overview of the CLEF-2018 CheckThat! Lab on automatic identification and verification of political claims, T ask 1: Check-worthiness. In CLEF 2018 Working Notes. Working Notes...

  2. [2]

    Pepa Atanasova, Preslav Nakov, Georgi Karadzhov, Mitra Mohtarami, and Giovanni Da San Martino. 2019. Overview of the CLEF-2019 CheckThat! Lab on Automatic Identification and Verification of Claims. Task 1: Check-Worthiness . In CLEF 2019 Working Notes. Working Notes of CLEF 2019 - Conference and Labs of the Evaluation Forum\/ . CEUR-WS.org, Lugano, Switzerland

  3. [3]

    Mouhamadou Lamine Ba, Laure Berti-Equille, Kushal Shah, and Hossam M. Hammady. 2016. VERA : A platform for veracity estimation over web data. In Proceedings of the 25th International Conference Companion on World Wide Web\/ . Montr \'e al, Qu \'e bec, Canada, WWW '16, pages 159--162

  4. [4]

    Ramy Baly, Georgi Karadzhov, Dimitar Alexandrov, James Glass, and Preslav Nakov. 2018. Predicting factuality of reporting and bias of news media sources. In Proceedings of the Conference on Empirical Methods in Natural Language Processing\/ . Brussels, Belgium, EMNLP '18, pages 3528--3539

  5. [5]

    Ramy Baly, Georgi Karadzhov, Abdelrhman Saleh, James Glass, and Preslav Nakov. 2019. Multi-task ordinal regression for jointly predicting the trustworthiness and the leading political ideology of news media. In Proceedings of the 17th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolog...

  6. [6]

    Alberto Barr\' o n-Cede\ no, Giovanni Da San Martino, Israa Jaradat, and Preslav Nakov. 2019. Proppy: Organizing the news based on their propagandistic content . Information Processing & Management\/ 56(5):1849 -- 1864

  7. [7]

    Alberto Barr\' o n-Cede\ n o, Tamer Elsayed, Reem Suwaileh, Llu\' i s M\` a rquez, Pepa Atanasova, Wajdi Zaghouani, Spas Kyuchukov, Giovanni Da San Martino, and Preslav Nakov. 2018. Overview of the CLEF-2018 CheckThat! Lab on automatic identification and verification of political claims, T ask 2: Factuality. In CLEF 2018 Working Notes. Working Notes of CL...

  8. [8]

    David M Blei, Andrew Y Ng, and Michael I Jordan. 2003. Latent D irichlet allocation. Journal of Machine Learning Research\/ 3(Jan):993--1022

Show all 59 references
  1. [9]

    Ann M Brill. 2001. Online journalists embrace new marketing function. Newspaper Research Journal\/ 22(2):28

  2. [10]

    Voorhees

    Chris Buckley and Ellen M. Voorhees. 2000. Evaluating evaluation measure stability. In Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval\/ . Athens, Greece, SIGIR '00, pages 33--40

  3. [11]

    Canini, Bongwon Suh, and Peter L

    Kevin R. Canini, Bongwon Suh, and Peter L. Pirolli. 2011. Finding credible information sources in social networks based on content and social structure. In Proceedings of the IEEE International Conference on Privacy, Security, Risk, and Trust, and the IEEE International Confer...

  4. [12]

    Richard Caruana. 1993. Multitask learning: A knowledge-based source of inductive bias. In Proceedings of the International Conference on Machine Learning\/ . Amherst, MA, USA, ICML '13, pages 41--48

  5. [13]

    Carlos Castillo, Marcelo Mendoza, and Barbara Poblete. 2011. Information credibility on T witter. In Proceedings of the International Conference on World Wide Web\/ . Hyderabad, India, WWW '11, pages 675--684

  6. [14]

    Cheng Chen, Kui Wu, Venkatesh Srinivasan, and Xudong Zhang. 2013. Battling the I nternet W ater A rmy: detection of hidden paid posters. In Proceedings of the 2013 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining\/ . Niagara, Ontario, Canada...

  7. [15]

    Giovanni Da San Martino, Seunghak Yu, Alberto Barron-Cedeno, Rostislav Petrov, and Preslav Nakov. 2019. Fine-grained analysis of propaganda in news articles. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing\/ . Hong Kong, China, EMNLP '19

  8. [16]

    Ido Dagan, Bill Dolan, Bernardo Magnini, and Dan Roth. 2009. Recognizing textual entailment: Rational, evaluation and approaches. Natural Language Engineering\/ 15(4):i--xvii

  9. [17]

    Rafferty, and Christopher D

    Marie-Catherine de Marneffe, Anna N. Rafferty, and Christopher D. Manning. 2008. Finding contradictions in text. In Proceedings of the Annual Meeting of the Association for Computational Linguistics\/ . Columbus, OH, USA, ACL '08, pages 1039--1047

  10. [18]

    Leon Derczynski, Kalina Bontcheva, Maria Liakata, Rob Procter, Geraldine Wong Sak Hoi, and Arkaitz Zubiaga. 2017. SemEval-2017 Task 8: RumourEval : Determining rumour veracity and support for rumours. In Proceedings of the 11th International Workshop on Semantic Evaluation\/ ....

  11. [19]

    Long Duong, Trevor Cohn, Steven Bird, and Paul Cook. 2015. Low resource dependency parsing: Cross-lingual parameter sharing in a neural network parser. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Co...

  12. [20]

    Tamer Elsayed, Preslav Nakov, Alberto Barr\' o n-Cede\ n o, Maram Hasanain, Reem Suwaileh, Pepa Atanasova, and Giovanni Da San Martino. 2019 a . CheckThat ! at CLEF 2019: Automatic identification and verification of claims. In Proceedings of the 41st European Conference on Inf...

  13. [21]

    Tamer Elsayed, Preslav Nakov, Alberto Barr\' o n-Cede \ n o, Maram Hasanain, Reem Suwaileh, Giovanni Da San Martino , and Pepa Atanasova. 2019 b . Overview of the CLEF-2019 CheckThat! : Automatic identification and verification of claims. In Experimental IR Meets Multilinguali...

  14. [22]

    Rob Ennals, Dan Byler, John Mark Agosta, and Barbara Rosario. 2010 a . What is disputed on the web? In Proceedings of the 4th Workshop on Information Credibility\/ . New York, NY, USA, WICOW '10, pages 67--74

  15. [23]

    Rob Ennals, Beth Trushkowsky, and John Mark Agosta. 2010 b . Highlighting disputed claims on the web. In Proceedings of the International Conference on World Wide Web\/ . New York, NY, USA, WWW '10, pages 341--350

  16. [24]

    George Foster and Roland Kuhn. 2009. Stabilizing minimum error rate training. In Proceedings of the Fourth Workshop on Statistical Machine Translation\/ . Athens, Greece, StatMT '09, pages 242--249

  17. [25]

    Pepa Gencheva, Preslav Nakov, Llu\' i s M\` a rquez, Alberto Barr\' o n-Cede\ n o, and Ivan Koychev. 2017. A context-aware approach for detecting worth-checking claims in political debates. In Proceedings of the International Conference on Recent Advances in Natural Language P...

  18. [26]

    Genevieve Gorrell, Elena Kochkina, Maria Liakata, Ahmet Aker, Arkaitz Zubiaga, Kalina Bontcheva, and Leon Derczynski. 2019. S em E val-2019 task 7: R umour E val, determining rumour veracity and support for rumours. In Proceedings of the 13th International Workshop on Semantic...

  19. [27]

    Momchil Hardalov, Ivan Koychev, and Preslav Nakov. 2016. In search of credible news. In Proceedings of the 17th International Conference on Artificial Intelligence: Methodology, Systems, and Applications\/ . Varna, Bulgaria, AIMSA '16, pages 172--180

  20. [28]

    Maram Hasanain, Reem Suwaileh, Tamer Elsayed, Alberto Barr\' o n-Cede \ n o, and Preslav Nakov. 2019. Overview of the CLEF-2019 CheckThat! Lab on Automatic Identification and Verification of Claims. Task 2: Evidence and Factuality . In CLEF 2019 Working Notes. Working Notes of...

  21. [29]

    Hamilton, Chengkai Li, Mark Tremayne, Jun Yang, and Cong Yu

    Naeemul Hassan, Bill Adair, James T. Hamilton, Chengkai Li, Mark Tremayne, Jun Yang, and Cong Yu. 2015 a . The quest to automate fact-checking. In Proceedings of the Computation+Journalism Symposium\/

  22. [30]

    Naeemul Hassan, Chengkai Li, and Mark Tremayne. 2015 b . Detecting check-worthy factual claims in presidential debates. In Proceedings of the 24th ACM International Conference on Information and Knowledge Management\/ . Melbourne, Australia, CIKM '15, pages 1835--1838

  23. [31]

    Naeemul Hassan, Afroza Sultana, You Wu, Gensheng Zhang, Chengkai Li, Jun Yang, and Cong Yu. 2014. Data in, fact out: Automated monitoring of facts by FactWatcher . PVLDB\/ 7:1557--1560

  24. [32]

    Naeemul Hassan, Gensheng Zhang, Fatma Arslan, Josue Caraballo, Damian Jimenez, Siddhant Gawsane, Shohedul Hasan, Minumol Joseph, Aaditya Kulkarni, Anil Kumar Nayak, Vikas Sable, Chengkai Li, and Mark Tremayne. 2017. ClaimBuster : The first-ever end-to-end fact-checking system....

  25. [33]

    Joan B. Hooper. 1974. On Assertive Predicates\/ . Indiana University Linguistics Club

  26. [34]

    Shafiq Joty, Giuseppe Carenini, and Raymond T. Ng. 2015. CODRA : A novel discriminative framework for rhetorical analysis. Comput. Linguist.\/ 41(3):385--435

  27. [35]

    Vivek Kulkarni, Junting Ye, Steve Skiena, and William Yang Wang. 2018. Multi-view models for political ideology detection of news articles. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing\/ . Brussels, Belgium, EMNLP '18, pages 3518--3527

  28. [36]

    Lazer, Matthew A

    David M.J. Lazer, Matthew A. Baum, Yochai Benkler, Adam J. Berinsky, Kelly M. Greenhill, Filippo Menczer, Miriam J. Metzger, Brendan Nyhan, Gordon Pennycook, David Rothschild, Michael Schudson, Steven A. Sloman, Cass R. Sunstein, Emily A. Thorson, Duncan J. Watts, and Jonathan...

  29. [37]

    Dieu-Thu Le, Ngoc Thang Vu, and Andre Blessing. 2016. Towards a text analysis system for political debates. LaTeCH 2016\/ page 134

  30. [38]

    Yaliang Li, Jing Gao, Chuishi Meng, Qi Li, Lu Su, Bo Zhao, Wei Fan, and Jiawei Han. 2016. A survey on truth discovery. SIGKDD Explor. Newsl.\/ 17(2):1--16

  31. [39]

    Bing Liu, Minqing Hu, and Junsheng Cheng. 2005. Opinion O bserver: Analyzing and comparing opinions on the web. In Proceedings of the 14th International Conference on World Wide Web\/ . New York, NY, USA, WWW '05, pages 342--351

  32. [40]

    Jansen, Kam-Fai Wong, and Meeyoung Cha

    Jing Ma, Wei Gao, Prasenjit Mitra, Sejeong Kwon, Bernard J. Jansen, Kam-Fai Wong, and Meeyoung Cha. 2016. Detecting rumors from microblogs with recurrent neural networks. In Proceedings of the 25th International Joint Conference on Artificial Intelligence\/ . New York, New Yor...

  33. [41]

    Todor Mihaylov, Georgi Georgiev, and Preslav Nakov. 2015. Finding opinion manipulation trolls in news community forums. In Proceedings of the Conference on Computational Natural Language Learning\/ . Beijing, China, CoNLL '15, pages 310--314

  34. [42]

    Todor Mihaylov, Tsvetomila Mihaylova, Preslav Nakov, Llu\' i s M\` a rquez, Georgi Georgiev, and Ivan Koychev. 2018. The dark side of news community forums: Opinion manipulation trolls. Internet Research\/ 28(5):1292--1312

  35. [43]

    Tsvetomila Mihaylova, Georgi Karadzhov, Pepa Atanasova, Ramy Baly, Mitra Mohtarami, and Preslav Nakov. 2019. S em E val-2019 task 8: Fact checking in community question answering forums. In Proceedings of the 13th International Workshop on Semantic Evaluation\/ . Minneapolis, ...

  36. [44]

    Tomas Mikolov, Wen-Tau Yih, and Geoffrey Zweig. 2013. Linguistic regularities in continuous space word representations. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies\/ . Atlanta, GA,...

  37. [45]

    Preslav Nakov, Alberto Barr\' o n-Cede\ n o, Tamer Elsayed, Reem Suwaileh, Llu\' i s M\` a rquez, Wajdi Zaghouani, Pepa Atanasova, Spas Kyuchukov, and Giovanni Da San Martino. 2018. Overview of the CLEF-2018 CheckThat! lab on automatic identification and verification of politi...

  38. [46]

    Ayush Patwari, Dan Goldwasser, and Saurabh Bagchi. 2017. TATHYA: a multi-classifier system for detecting check-worthy statements in political debates. In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management\/ . Singapore, CIKM '17, pages 2259--2262

  39. [47]

    Kashyap Popat, Subhabrata Mukherjee, Jannik Str\" o tgen, and Gerhard Weikum. 2017. Where the truth lies: Explaining the credibility of emerging claims on the web and social media. In Proceedings of the 26th International Conference on World Wide Web Companion\/ . Perth, Austr...

  40. [48]

    Martin Potthast, Johannes Kiesel, Kevin Reinartz, Janek Bevendorff, and Benno Stein. 2018. A stylometric inquiry into hyperpartisan and fake news. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics\/ . Melbourne, Australia, ACL '18, page...

  41. [49]

    Marta Recasens, Cristian Danescu-Niculescu-Mizil, and Dan Jurafsky. 2013. Linguistic models for analyzing and detecting biased language. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics\/ . Sofia, Bulgaria, ACL '13, pages 1650--1659

  42. [50]

    Sebastian Ruder. 2017. An overview of multi-task learning in deep neural networks. CoRR\/ abs/1706.05098

  43. [51]

    Kai Shu, Amy Sliva, Suhang Wang, Jiliang Tang, and Huan Liu. 2017. Fake news detection on social media: A data mining perspective. SIGKDD Explor. Newsl.\/ 19(1):22--36

  44. [52]

    James Thorne and Andreas Vlachos. 2018. Automated fact checking: Task formulations, methods and future directions. In Proceedings of the 27th International Conference on Computational Linguistics\/ . Santa Fe, NM, USA, COLING '18, pages 3346--3359

  45. [53]

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER : a large-scale dataset for fact extraction and VER ification. In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human...

  46. [54]

    Andreas Vlachos and Sebastian Riedel. 2014. Fact checking: Task definition and dataset construction. In Proceedings of the ACL 2014 Workshop on Language Technologies and Computational Social Science\/ . Baltimore, MD, USA, pages 18--22

  47. [55]

    Soroush Vosoughi, Deb Roy, and Sinan Aral. 2018. The spread of true and false news online. Science\/ 359(6380):1146--1151

  48. [56]

    Yifan Zhang, Giovanni Da San Martino, Alberto Barrón-Cedeño, Salvatore Romeo, Jisun An, Haewoon Kwak, Todor Staykovski, Israa Jaradat, Georgi Karadzhov, Ramy Baly, Kareem Darwish, and Preslav Nakov James Glass. 2019. Tanbih: Get to know what you are reading. In Proceedings of ...

  49. [57]

    Arkaitz Zubiaga, Maria Liakata, Rob Procter, Geraldine Wong Sak Hoi, and Peter Tolmie. 2016. Analysing how people orient to and spread rumours in social media by looking at conversational threads. PLoS ONE\/ 11(3):1--29

  50. [58]

    , " * write output.state after.block = add.period write newline

    ENTRY address author booktitle chapter doi edition editor howpublished institution journal key month note number organization pages publisher school series title type url volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.se...

  51. [59]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.