REVIEW 5 major objections 5 minor 60 references
CrediBench: Building Web-Scale Network Datasets for Information Integrity
T0 review · 5 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read Hyperlink structure and webpage text both predict expert-rated credibility at web scale.
desk verdict A real pipeline and a genuinely new large graph, but the abstract promises results that aren't in the body and the artifact isn't out yet. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the temporal text-attributed graph (TAG): nodes are web domains, edges are hyperlinks, and each node carries an aggregated text embedding (Qwen3-0.6B) from its scraped pages. The graph is constructed from Common Crawl's WAT metadata and WET text, filtered to nodes with degree above 3, and annotated with DQR credibility labels (PC1 and MBFC). Random Node Initialization (RNI) is used as the GNN node feature, which the paper finds essential—zero initialization collapses to the mean baseline.
What would settle it
If a GNN trained on the December 2024 snapshot is evaluated on the binary classification set (662K domains) or on a sample of non-news domains, and its performance drops to the mean baseline, that would show the structural signal does not generalize beyond the news-biased DQR labels. Alternatively, randomly shuffling hyperlink edges (preserving node degrees) and observing no change in MAE would indicate the signal comes from degree or network statistics rather than genuine link structure.
Extended reading notes
Core claim
The central claim is that a domain's credibility can be predicted from two web-scale signals: its page text and its position in the hyperlink graph. On the December 2024 snapshot, a text-only MLP using Qwen3 embeddings achieves mean absolute error of 0.137 for the PC1 credibility score and 0.122 for MBFC, while GNNs with random node initialization do even better on the graph alone, with GAT reaching 0.129 and 0.114 on the same metrics—all outperforming the mean predictor (0.167 and 0.153). The paper further introduces a binary classification label set of 662,575 domains spanning misinformation, crowd-sourced, malware, and phishing, and reports that a multi-modal model improves accuracy from
Load-bearing premise
The DQR labels (PC1 and MBFC), covering only about 11.5K domains and biased toward news websites, are treated as ground truth credibility for a 45M-node web graph; if these labels are unrepresentative or noisy, the measured predictive power of text and structure will not generalize to web-wide credibility.
Editorial extensions
If this is right
- Hyperlink structure alone is a viable signal for ranking source credibility, without needing text.
- Text content alone also predicts credibility, justifying the inclusion of node features in graph models.
- Combining structure and text in one model should outperform either alone, motivating the multi-modal experiments claimed in the abstract.
- The pipeline's monthly snapshots enable study of temporal credibility dynamics, such as how domain connectivity evolves around elections.
- The binary label set of 662K domains provides a much larger training signal than the 11.5K DQR set for future classifiers.
Reading between the lines
- Because DQR labels skew heavily toward news domains, the measured text and structure signals may partly reflect genre or editorial style rather than credibility per se; applying the method to non-news domains is a key untested extension.
- The graph signal could be confounded by degree or popularity—domains with many links tend to be established institutions—so a degree-controlled ablation would sharpen the claim that structure itself carries credibility information.
- The multi-modal gains reported (MAE 0.107, accuracy 85%) appear only in the abstract; reproducing them with the released snapshot would confirm that text and structure are complementary rather than redundant.
- Since the pipeline is automated and public, it can be re-run on later crawls to track how credibility of domains shifts over time, enabling early-warning systems for misinformation ecosystems.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces CrediBench, a pipeline for constructing temporal web graphs from Common Crawl, and presents a December 2024 snapshot with about 45 million nodes and 1 billion edges after degree filtering. Nodes are web domains, edges are hyperlinks, and node text is scraped and embedded. For evaluation, the authors use the DQR dataset of roughly 11.5K expert-rated domains and perform two regression tasks (PC1 and MBFC scores) with a text-only MLP and with GNNs using random node initialization (RNI). They report that both text and graph structure improve MAE over a mean predictor. The abstract additionally claims a 662,575-domain binary label set and a multi-modal classifier achieving 85% accuracy and MAE 0.107, but the body contains no classification experiments, no multi-modal model, and no description of the 662K-label set. The dataset and code are promised for after the review period.
Significance. If fully realized, CrediBench could be a valuable community resource: it is among the largest web-domain graphs for misinformation research, with a plausible pipeline for temporal graphs and text features. The study also includes useful engineering details, such as distributed processing, neighbor sampling ablations, and compute costs. However, as submitted, the paper's headline contributions are neither present nor verifiable. The abstract promises results and data that the manuscript does not contain, and the body's experiments use a small, news-biased labeled set with only a mean predictor as baseline. The central claim that credential signals scale to the web is therefore unsupported.
major comments (5)
- [Abstract vs. §5] The abstract reports a 662,575-domain binary label set, a multi-modal classifier with 85% accuracy, and MAE 0.107 on regression. Section 5 only presents regression on the 11.5K DQR labels, with best MAE 0.129 (GAT, PC1) and 0.114 (GAT, MBFC). No classification experiments, no multi-modal model, and no 662K-label set appear anywhere in the paper. These are load-bearing claims of the abstract; their absence invalidates the paper's stated contributions.
- [§1 Reproducibility / Appendix A] The paper states that the dataset is available on HuggingFace and code is available 'here', but Appendix A says 'we will release our curated December 2024 snapshot after the review period.' No actual URLs or access information are provided. For a dataset paper, the artifact itself is the central contribution; deferring its release makes the results unreproducible and the benchmark unusable by the community.
- [§5 RQ2, Table 2] The GNN experiments compare only against a mean predictor. With random node initialization as the sole node feature, the models may exploit trivial structural statistics (e.g., degree, neighbor counts) rather than meaningful credibility propagation. Without a degree-based or label-propagation baseline, the claim that 'the hyperlink graph structure provides useful signal' is not established.
- [§3 / Appendix B / Figure 4] The DQR labels cover about 11.5K domains, roughly 0.03% of the processed graph's 45M nodes, and the authors acknowledge in Appendix B that the labels are biased toward news websites. The paper's motivation is web-scale credibility prediction, but the experiments and labels are confined to a small, non-representative subset. This undermines the generalization claim in the strongest claim and in the conclusion.
- [§4 Table 1] The text says nodes with degree strictly higher than 3 are kept, and Figure 2 and Section 4 describe a degree threshold of 3. Table 1, however, lists the processed graph as having min degree 0 and 28,857 leaves (degree 1). This contradiction suggests either the filtering is not as described or the table reports different statistics; in either case, the data description is not reliable.
minor comments (5)
- [Abstract] 'eights months' should be 'eight months'; 'Mean Average Error' should be 'Mean Absolute Error' (MAE) consistent with the body.
- [Figure 3] The y-axis label reads 'Frequancy' — typo for 'Frequency'.
- [§1 and §2] Several references are incomplete or lack retrieval dates; e.g., [30] and [27] have no year/venue. Please standardize.
- [Appendix D] The domain-text examples are lengthy and could be moved to an online appendix; consider trimming to two representative cases.
- [§5 RQ1] The MLP experiment uses scikit-learn MLP with two hidden layers, but no description of feature normalization or embedding pooling is given; this would help reproducibility.
Circularity Check
No circular derivation; central regression experiments are supervised on external DQR labels. Minor non-load-bearing self-citations and an abstract/body reporting gap are the only concerns.
full rationale
The claimed derivation chain is not circular. Credibility targets (PC1 and MBFC) come from the external Domain Quality Ratings dataset [35]; the paper attaches these labels to graph nodes by domain-name matching (Sec. 3, 'Domain credibility assignment task') and then trains supervised regressors (MLP on Qwen3 text embeddings; GCN/GraphSAGE/GAT/GATv2 with Random Node Initialization) on a 60/20/20 split. The test-set MAE comparisons against a mean predictor are genuine held-out evaluations, not quantities equal to their inputs by construction. RNI is a random feature, so any structural signal the GNN extracts must come from the adjacency structure and the external labels rather than from the model definition. Text embeddings are from an off-the-shelf model, not fitted to the DQR labels. The only self-citations in the paper ([42], [47]) appear in the related-work survey as examples of LLM/RAG approaches; they are not used to justify the dataset construction or the experimental conclusions, so they are not load-bearing. The paper's stated limitations (Appendix B: labels biased toward news websites; Appendix A: small labelled set relative to graph) are external-validity caveats, not circularity. The abstract's promises of a 662,575-domain binary label set and a multi-modal classifier with 85% accuracy / MAE 0.107 are not reproduced in the body (Sec. 5 reports only single-modality regression), which is a missing-support/omitted-proof issue, not a circular reduction. Score 2 reflects only the presence of minor self-citations that do not carry the argument.
Assumptions & free parameters
free parameters (3)
- degree threshold =
3
- neighbor sampling config =
[50,50,50]
- RNI distribution =
Normal(0,1)
assumptions (5)
- domain assumption Common Crawl coverage is representative enough of the web for credibility signals to generalize.
- domain assumption DQR scores (PC1 and MBFC) are valid ground-truth credibility labels.
- domain assumption Qwen3-0.6B text embeddings capture semantic content relevant to credibility.
- domain assumption Random Node Initialization lets GNNs learn structural signal rather than just random artifacts.
- domain assumption Concatenating the three longest and three shortest documents represents a domain's content.
Cite this review
Pith. "Pith review of CrediBench: Building Web-Scale Network Datasets for Information Integrity." pith.science (2026). https://pith.science/paper/4RXM3WGS
@misc{pith2026250923340,
author = {Pith},
title = {Pith review of: CrediBench: Building Web-Scale Network Datasets for Information Integrity},
year = {2026},
howpublished = {\url{https://pith.science/paper/4RXM3WGS}},
note = {Machine review of arXiv:2509.23340}
}
read the original abstract
Automatically assessing the credibility of online sources presents an invaluable tool for navigating today's information ecosystem. However, existing approaches either depend on scarce and costly human annotations, or focus exclusively on assessments at the level of individual claims. Misinformation often spreads via interlinked web domains, whose connections evolve over time. Focusing on claims alone ignores these structural and temporal credibility signals evident in the changing web topology. Existing datasets fail to capture these central modalities in web domain credibility prediction: namely, internet topology, temporality and text (webpage) content. To address this gap, we present CrediBench, a dataset containing eights months of web graph data; of which we analyze the three months surrounding the 2024 U.S. federal elections, a time of heightened misinformation propagation online. Each monthly snapshot contains over 40 million nodes, their scraped webpage content, and over 1 billion hyperlink edges. CrediBench supports credibility prediction as both a regression (continuous credibility score) and a binary classification task (credible or not). For classification, we curate a novel binary label set containing 662,575 web domains labelled for boolean credibility, spanning four areas (misinformation, crowd-sourced, malware and phishing). Our empirical experiments support that all task modalities-graph, text and time-contribute significantly to achieving the best performance. In particular, our multi-modal regression model trained on CrediBench outperforms other configurations and existing baselines, decreasing Mean Average Error from 0.162 to 0.107 on the regression task, while the multi-modal classifier improves accuracy from 56% to 85% on the classification one. CrediBench, our proposed web-scale multi-modal dataset, is available on Huggingface for future research.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Digital format description, Library of Congress,
WARC (Web ARChive) File Format. Digital format description, Library of Congress,
-
[2]
Media Bias/Fact Check. 2023. Media Bias/Fact Check methodology (Internet Archive), 2025. URL https://web.archive.org/web/20230502031920/https: //mediabiasfactcheck.com/methodology/. Last accessed: Sep 4, 2025
arXiv 2023
-
[3]
The Surprising Power of Graph Neural Networks with Random Node Initialization
Ralph Abboud, ˙Ismail ˙Ilkan Ceylan, Martin Grohe, and Thomas Lukasiewicz. The Surprising Power of Graph Neural Networks with Random Node Initialization. InProceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI 2021, Virtual Event / Montreal, Canada, 19-27 August 2021, pages 2112–2118. ijcai.org, 2021. doi: 10.24963/...
doi:10.24963/ijcai 2021
-
[4]
GraphLand: Evaluating Graph Machine Learning Models on Diverse Industrial Data, 2025
Gleb Bazhenov, Oleg Platonov, and Liudmila Prokhorenkova. GraphLand: Evaluating Graph Machine Learning Models on Diverse Industrial Data, 2025. URL https://arxiv.org/abs/ 2409.14500
arXiv 2025
-
[5]
Bronstein, Mathias Niepert, Bryan Perozzi, Mikhail Galkin, and Christopher Morris
Maya Bechler-Speicher, Ben Finkelshtein, Fabrizio Frasca, Luis Müller, Jan Tönshoff, Antoine Siraudin, Viktor Zaverkin, Michael M. Bronstein, Mathias Niepert, Bryan Perozzi, Mikhail Galkin, and Christopher Morris. Position: Graph Learning Will Lose Relevance Due To Poor Benchmarks.CoRR, abs/2502.14546, 2025. doi: 10.48550/ARXIV .2502.14546. URL https://do...
-
[6]
The CRAAP Test.LOEX Quarterly, 2004
Sarah Blakeslee. The CRAAP Test.LOEX Quarterly, 2004
2004
-
[7]
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts.CoRR, abs/2412.10510, 2024
Tobias Braun, Mark Rothermel, Marcus Rohrbach, and Anna Rohrbach. DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts.CoRR, abs/2412.10510, 2024. doi: 10.48550/ARXIV .2412.10510. URLhttps://doi.org/10.48550/arXiv.2412.10510
-
[8]
The Anatomy of a Large-Scale Hypertextual Web Search Engine.Comput
Sergey Brin and Lawrence Page. The Anatomy of a Large-Scale Hypertextual Web Search Engine.Comput. Networks, 30(1-7):107–117, 1998. doi: 10.1016/S0169-7552(98)00110-X. URLhttps://doi.org/10.1016/S0169-7552(98)00110-X. 7
Show all 60 references
-
[9]
How Attentive are Graph Attention Networks? In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022
Shaked Brody, Uri Alon, and Eran Yahav. How Attentive are Graph Attention Networks? In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum?id= F72ximsx7C1
2022
-
[10]
Information credibility on twitter
Carlos Castillo, Marcelo Mendoza, and Barbara Poblete. Information credibility on twitter. In Proceedings of the 20th International Conference on World Wide Web, WWW 2011, Hyderabad, India, March 28 - April 1, 2011, pages 675–684. ACM, 2011. doi: 10.1145/1963405.1963500. URLht...
2011
-
[11]
SIFT (The Four Moves).Hapgood, 2019
Mike Caulfield. SIFT (The Four Moves).Hapgood, 2019
2019
-
[12]
Can LLM-Generated Misinformation Be Detected? InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11,
Canyu Chen and Kai Shu. Can LLM-Generated Misinformation Be Detected? InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11,
2024
-
[13]
Combating misinformation in the age of LLMs: Opportunities and challenges.AI Mag., 45(3):354–368, 2024
Canyu Chen and Kai Shu. Combating misinformation in the age of LLMs: Opportunities and challenges.AI Mag., 45(3):354–368, 2024. doi: 10.1002/AAAI.12188. URL https: //doi.org/10.1002/aaai.12188
2024 doi
-
[14]
URLhttps://openreview.net/forum?id=ccxD4mtkTU
OpenReview.net, 2024. URLhttps://openreview.net/forum?id=ccxD4mtkTU
2024
-
[15]
Common Crawl PySpark Repository
Common Crawl. Common Crawl PySpark Repository. URL https://github.com/ commoncrawl/cc-pyspark
-
[16]
Linguistic feature based learning model for fake news detection and classification.Expert Syst
Anshika Choudhary and Anuja Arora. Linguistic feature based learning model for fake news detection and classification.Expert Syst. Appl., 169:114171, 2021. doi: 10.1016/J.ESW A.2020. 114171. URLhttps://doi.org/10.1016/j.eswa.2020.114171
2021
-
[17]
December 2024 Crawl Archive Now Available, 2025
Common Crawl. December 2024 Crawl Archive Now Available, 2025. URL https:// commoncrawl.org/blog/december-2024-crawl-archive-now-available
2024
-
[18]
Common Crawl Web Graphs, 2025
Common Crawl. Common Crawl Web Graphs, 2025. URL https://commoncrawl.org/ web-graphs
2025
-
[19]
Kenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzeminski, Genta Indra Winata, Saba Sturua, Saiteja Utpala, Mathieu Ciancone, Marion Schaeffer, Gabriel Sequeira, Diganta Misra, Shreeya Dhakal, Jona...
2025 doi
-
[20]
Global Risks Report 2025, Jan
Mark Elsner, Grace Atkinson, and Saadia Zahidi. Global Risks Report 2025, Jan. 2025. URL https://www.weforum.org/publications/global-risks-report-2025/ . 20th edition of the Global Risks Report
2025
-
[21]
Fast Graph Representation Learning with PyTorch Geomet- ric.CoRR, abs/1903.02428, 2019
Matthias Fey and Jan Eric Lenssen. Fast Graph Representation Learning with PyTorch Geomet- ric.CoRR, abs/1903.02428, 2019. URLhttp://arxiv.org/abs/1903.02428. 8
1903 arXiv
-
[22]
Infodemiology: the epidemiology of (mis)information
Gunther Eysenbauch. Infodemiology: the epidemiology of (mis)information. 113(9). URL https://www.amjmed.com/article/S0002-9343(02)01473-0/fulltext
-
[23]
TweetCred: Real- Time Credibility Assessment of Content on Twitter
Aditi Gupta, Ponnurangam Kumaraguru, Carlos Castillo, and Patrick Meier. TweetCred: Real- Time Credibility Assessment of Content on Twitter. InSocial Informatics - 6th International Conference, SocInfo 2014, Barcelona, Spain, November 11-13, 2014. Proceedings, volume 8851 ofLe...
2014 doi
-
[24]
PyG 2.0: Scalable Learning on Real World Graphs.CoRR, abs/2507.16991,
Matthias Fey, Jinu Sunil, Akihiro Nitta, Rishi Puri, Manan Shah, Blaz Stojanovic, Ramona Bendias, Alexandria Barghi, Vid Kocijan, Zecheng Zhang, Xinwei He, Jan Eric Lenssen, and Jure Leskovec. PyG 2.0: Scalable Learning on Real World Graphs.CoRR, abs/2507.16991,
-
[25]
Hamilton, Zhitao Ying, and Jure Leskovec
William L. Hamilton, Zhitao Ying, and Jure Leskovec. Inductive Representation Learning on Large Graphs. InAdvances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 1024–...
2017
-
[26]
Compare to the Knowledge: Graph Neural Fake News Detection with External Knowledge
Linmei Hu, Tianchi Yang, Luhao Zhang, Wanjun Zhong, Duyu Tang, Chuan Shi, Nan Duan, and Ming Zhou. Compare to the Knowledge: Graph Neural Fake News Detection with External Knowledge. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and ...
2021 doi
-
[27]
Suhaib Kh Hamed, Mohd Juzaiddin Ab Aziz, and Mohd Ridzwan Yaakub. A review of fake news detection approaches: A critical analysis of relevant studies and highlighting key challenges associated with the dataset, feature representation, and data fusion.Heliyon, 2023
2023
-
[28]
FakeBERT: Fake news detection in social media with a BERT-based deep learning approach.Multim
Rohit Kumar Kaliyar, Anurag Goswami, and Pratik Narang. FakeBERT: Fake news detection in social media with a BERT-based deep learning approach.Multim. Tools Appl., 80(8): 11765–11788, 2021. doi: 10.1007/S11042-020-10183-2. URL https://doi.org/10.1007/ s11042-020-10183-2
2021 doi
-
[29]
Representation learning for dynamic graphs: A survey.J
Seyed Mehran Kazemi, Rishab Goel, Kshitij Jain, Ivan Kobyzev, Akshay Sethi, Peter Forsyth, and Pascal Poupart. Representation learning for dynamic graphs: A survey.J. Mach. Learn. Res., 21:70:1–70:73, 2020. URLhttps://jmlr.org/papers/v21/19-447.html
2020
-
[30]
Disinformation Detection: An Evolving Challenge in the Age of LLMs
Bohan Jiang, Zhen Tan, Ayushi Nirmal, and Huan Liu. Disinformation Detection: An Evolving Challenge in the Age of LLMs. URL https://www.semanticscholar.org/ paper/Disinformation-Detection%3A-An-Evolving-Challenge-in-Jiang-Tan/ 6b1c431db1f7d10f0a55d51786d55ad6b6921730
-
[31]
Kipf and Max Welling
Thomas N. Kipf and Max Welling. Semi-Supervised Classification with Graph Convolutional Networks. In5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. URL https://openrevie...
2017
-
[32]
AMFB: Attention based multimodal Factorized Bilinear Pooling for multimodal Fake News Detection.Expert Syst
Rina Kumari and Asif Ekbal. AMFB: Attention based multimodal Factorized Bilinear Pooling for multimodal Fake News Detection.Expert Syst. Appl., 184:115412, 2021. doi: 10.1016/J. ESW A.2021.115412. URLhttps://doi.org/10.1016/j.eswa.2021.115412
2021
-
[33]
Evaluating credibility of social media information: current challenges, research directions and practical criteria
Hamid Keshavarz. Evaluating credibility of social media information: current challenges, research directions and practical criteria. URL https://www.emerald.com/insight/ content/doi/10.1108/idd-03-2020-0033/full/html
2020 doi
-
[34]
A Decision-Based Heteroge- nous Graph Attention Network for Multi-Class Fake News Detection.CoRR, abs/2501.03290,
Batool Lakzaei, Mostafa Haghir Chehreghani, and Alireza Bagheri. A Decision-Based Heteroge- nous Graph Attention Network for Multi-Class Fake News Detection.CoRR, abs/2501.03290,
-
[35]
High level of correspondence across different news domain quality rating sets.PNAS Nexus, 2(9):pgad286, 09 2023
Hause Lin, Jana Lasser, Stephan Lewandowsky, Rocky Cole, Andrew Gully, David G Rand, and Gordon Pennycook. High level of correspondence across different news domain quality rating sets.PNAS Nexus, 2(9):pgad286, 09 2023. ISSN 2752-6542. doi: 10.1093/pnasnexus/pgad286. URLhttps:...
2023 doi
-
[36]
Kakade, Prateek Jain, and Ali Farhadi
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham M. Kakade, Prateek Jain, and Ali Farhadi. Matryoshka Representation Learning. InAdvances in Neu- ral Information Processing Systems 35: A...
2022
-
[37]
Hejamadi Rama Moorthy, N. J. Avinash, Krishnaraj N. S. Rao, K. R. Raghunandan, Radhakr- ishna Dodmane, Jeremy Joseph Blum, and Lubna A Gabralla. Dual stream graph augmented transformer model integrating BERT and GNNs for context aware fake news detection. 15. URLhttps://www.na...
- [38]
-
[39]
Credibility in Context: An Analysis of Feature Distributions in Twitter
John O’Donovan, Byungkyu Kang, Greg Meyer, Tobias Höllerer, and Sibel Adali. Credibility in Context: An Analysis of Feature Distributions in Twitter. In2012 International Conference on Privacy, Security, Risk and Trust, PASSAT 2012, and 2012 International Confernece on Social ...
2012 doi
-
[40]
Fake news, rumor, information pollution in social media and web: A contemporary survey of state-of-the-arts, challenges and opportunities
Priyanka Meel and Dinesh Kumar Vishwakarma. Fake news, rumor, information pollution in social media and web: A contemporary survey of state-of-the-arts, challenges and opportunities. Expert Syst. Appl., 153:112986, 2020. doi: 10.1016/J.ESWA.2019.112986. URL https: //doi.org/10...
2020
-
[41]
Cross-SEAN: A cross-stitch semi-supervised neural attention model for COVID- 19 fake news detection.Appl
William Scott Paka, Rachit Bansal, Abhay Kaushik, Shubhashis Sengupta, and Tanmoy Chakraborty. Cross-SEAN: A cross-stitch semi-supervised neural attention model for COVID- 19 fake news detection.Appl. Soft Comput., 107:107393, 2021. doi: 10.1016/J.ASOC.2021. 107393. URLhttps:/...
2021
-
[42]
Fake News Detection with Retrieval Augmented Generative Artificial Intelligence
Mohammad Vatani Nezafat and Saeed Samet. Fake News Detection with Retrieval Augmented Generative Artificial Intelligence. In2nd International Conference on Foundation and Large Language Models, FLLM 2024, Dubai, United Arab Emirates, November 26-29, 2024, pages 160–167. IEEE, ...
2024
-
[43]
Adversarial Active Learning based Heterogeneous Graph Neural Network for Fake News Detection
Yuxiang Ren, Bo Wang, Jiawei Zhang, and Yi Chang. Adversarial Active Learning based Heterogeneous Graph Neural Network for Fake News Detection. In20th IEEE International Conference on Data Mining, ICDM 2020, Sorrento, Italy, November 17-20, 2020, pages 452–
2020
-
[44]
Web Credibility: Features Exploration and Credibility Prediction
Alexandra Olteanu, Stanislav Peshterliev, Xin Liu, and Karl Aberer. Web Credibility: Features Exploration and Credibility Prediction. InAdvances in Information Retrieval - 35th European Conference on IR Research, ECIR 2013, Moscow, Russia, March 24-27, 2013. Proceedings, volum...
2013 doi
-
[45]
Exploring Generalizability of Fine-Tuned Models for Fake News Detection
Abhijit Suprem, Sanjyot Vaidya, and Calton Pu. Exploring Generalizability of Fine-Tuned Models for Fake News Detection. In8th IEEE International Conference on Collaboration and Internet Computing, CIC 2022, Atlanta, GA, USA, December 14-16, 2022, pages 82–88. IEEE,
2022
-
[46]
Towards Reliable Misinformation Mitigation: Generalization, Uncertainty, and GPT-4
Kellin Pelrine, Anne Imouza, Camille Thibault, Meilina Reksoprodjo, Caleb Gupta, Joel Christoph, Jean-François Godbout, and Reihaneh Rabbany. Towards Reliable Misinformation Mitigation: Generalization, Uncertainty, and GPT-4. InProceedings of the 2023 Conference on Empirical M...
2023 doi
-
[47]
Web Retrieval Agents for Evidence-Based Misinformation Detection.CoRR, abs/2409.00009, 2024
Jacob-Junqi Tian, Hao Yu, Yury Orlovskiy, Tyler Vergho, Mauricio Rivera, Mayank Goel, Zachary Yang, Jean-François Godbout, Reihaneh Rabbany, and Kellin Pelrine. Web Retrieval Agents for Evidence-Based Misinformation Detection.CoRR, abs/2409.00009, 2024. doi: 10.48550/ARXIV .24...
-
[48]
Graph Attention Networks
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph Attention Networks. In6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. Op...
2018
-
[49]
CSI: A Hybrid Deep Model for Fake News Detection
Natali Ruchansky, Sungyong Seo, and Yan Liu. CSI: A Hybrid Deep Model for Fake News Detection. InProceedings of the 2017 ACM on Conference on Information and Knowledge Management, CIKM 2017, Singapore, November 06 - 10, 2017, pages 797–806. ACM, 2017. doi: 10.1145/3132847.3132...
2017
-
[50]
Information Credibility on Twitter in Emergency Situation
Xin Xia, Xiaohu Yang, Chao Wu, Shanping Li, and Linfeng Bao. Information Credibility on Twitter in Emergency Situation. InIntelligence and Security Informatics - Pacific Asia Workshop, PAISI 2012, Kuala Lumpur, Malaysia, May 29, 2012. Proceedings, volume 7299 ofLecture Notes i...
2012 doi
-
[51]
Qwen3 technical report.CoRR, abs/2505.09388, 2025
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jian Yang, Jiaxi ...
-
[52]
Della Vedova, Stefano Moret, and Luca de Al- faro
Eugenio Tacchini, Gabriele Ballarin, Marco L. Della Vedova, Stefano Moret, and Luca de Al- faro. Some Like it Hoax: Automated Fake News Detection in Social Networks.CoRR, abs/1704.07506, 2017. URLhttp://arxiv.org/abs/1704.07506
2017 arXiv
-
[53]
Fake news detection based on dual-channel graph convolutional attention network.J
Mengfan Zhao, Yutao Zhang, and Guozheng Rao. Fake news detection based on dual-channel graph convolutional attention network.J. Supercomput., 80(9):13250–13271, 2024. doi: 10.1007/S11227-024-05953-W. URL https://doi.org/10.1007/s11227-024-05953-w . A Limitations Relying on Com...
2024 doi
-
[55]
Liar, Liar Pants on Fire
William Yang Wang. "Liar, Liar Pants on Fire": A New Benchmark Dataset for Fake News Detection. InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 2: Short Papers, pages 422–426. As...
2017 doi
-
[58]
Jiawei Zhang, Bowen Dong, and Philip S. Yu. FakeDetector: Effective Fake News Detec- tion with Deep Diffusive Neural Network. In36th IEEE International Conference on Data Engineering, ICDE 2020, Dallas, TX, USA, April 20-24, 2020, pages 1826–1829. IEEE,
2020
-
[461]
doi: 10.1109/ICDM50108.2020.00054
IEEE, 2020. doi: 10.1109/ICDM50108.2020.00054. URL https://doi.org/10.1109/ ICDM50108.2020.00054
2020
-
[2020]
URL https://doi.org/10.1109/ICDE48307
doi: 10.1109/ICDE48307.2020.00180. URL https://doi.org/10.1109/ICDE48307. 2020.00180
2020
-
[2022]
URL https://doi.org/10.1109/CIC56439
doi: 10.1109/CIC56439.2022.00022. URL https://doi.org/10.1109/CIC56439. 2022.00022
2022
-
[2024]
URL https://www.loc.gov/preservation/digital/formats/fdd/fdd000236. shtml. Last significant FDD update: April 29, 2024
2024
- [2025]
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.