{"id":"a525e5ab-a176-4ff3-9118-7a7131a360be","arxiv_id":"2501.05292","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that unifies the literature on rumor detection, rumor source detection, and joint detection into one taxonomy with formal definitions.","lead":"This paper reviews published research on detecting false or unverified information in social networks and on tracing that information back to its original spreader. Because it covers both tasks in one place, it may serve as a reference map for researchers and engineers working on misinformation.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"First-joint-survey claim is unverified, and the joint category is internally inconsistent: [128] and [135] are the same paper, and Algorithm 3 does not implement Definition 5.","rationale":"I read the survey as a taxonomy and synthesis exercise, not as a new algorithm paper. Its formal definitions (Definitions 2, 4, 5) are clear and useful, and the taxonomy of snapshots and estimators is a legitimate service to the community. The paper deserves credit for organizing two separate literatures under one framework and for including the hide-and-seek material, which is rarely covered. The concern is not that the authors are wrong in their judgments; it is that the strongest advertised claim, the first-joint-survey claim, is precisely the kind of empirical literature claim that needs evidence, and the paper supplies none. The internal inconsistency around references [128]/[135] and the mismatch between Definition 5 and Algorithm 3 are objective, checkable problems rather than matters of taste. Because the reader's conditional verdict already asks for substantiation of the novelty claim and correction of Algorithm 3, my stress-test does not move the verdict; it sharpens the conditions. I mark agreement as partial: the reader focused on whether the coupling is practically meaningful, while I focus on whether the claimed survey-level novelty and category coherence are established.","tokens_in":36120,"tokens_out":7150,"duration_ms":73610,"concrete_test":"Run a single audit with two linked parts: (1) perform a defined pre-2025 literature search (Scopus/DBLP/Google Scholar) for survey papers whose title, abstract, or keywords contain both rumor/fake-news detection and source detection/localization, and compare every hit's scope with Section I contribution (i); (2) in the same corpus, deduplicate the reference list and reclassify all entries in Section III-D and Tables V–VI against Definition 5, recording whether each solves the joint function or only one component. If the search finds any prior survey with a joint treatment, or if [135] is not genuinely joint, the 'first survey' and 'joint consideration' formulations should be weakened to a more precise claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution (i) — that this is the first survey covering rumor detection, rumor-source detection, and their joint consideration — is load-bearing because the abstract and Section I organize the entire survey around it. The manuscript never supplies the comparison needed to establish it: there is no systematic search, no list of prior surveys checked, and the previously cited surveys that might overlap are not analyzed against the three-problem taxonomy. The internal evidence makes the novelty claim fragile. Section III-D identifies only two joint algorithms, [134] and [135]. But reference [128] and reference [135] are the same Seo, Mohapatra, and Abdelzaher paper, and [135] is also used in Section III-C and Tables V–VI as a single-source sensor-monitoring estimator. If [135] is genuinely joint, the taxonomy double-counts it; if it is not joint, only one paper ([134]) remains in the joint category. Additionally, Algorithm 3 does not implement Definition 5's f: (T, G_I) -> ({0,1}, S-hat): it sequentially runs source detection then rumor detection, and line 11 repeats the rumor condition, so the advertised joint inference problem is not represented by any described algorithm. These are concrete, checkable defects in the support for the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of rumor detection and rumor-source detection on social networks, organized around three problems: rumor detection, rumor-source detection, and their joint consideration. It offers formal definitions (Definitions 2, 4, and 5), taxonomies of algorithms, comparative summary tables, a discussion of hiding and obfuscating rumors and sources, and a list of research challenges and future directions. The paper claims to be the first survey covering all three problems together, with the joint problem formalized as a simultaneous classification-and-estimation task.","tokens_in":36382,"tokens_out":5882,"duration_ms":53377,"significance":"If the coverage and taxonomy are made reliable, the survey would be a useful entry point for researchers: the three-problem framing, the formal problem definitions, the snapshot classification (complete, partial, sensor-monitoring), and the inclusion of source-obfuscation literature are valuable organizational devices. The paper also credits prior work by explicitly presenting the joint problem as a combination of the two individual problems, rather than as a newly derived formalism, and it provides several structured summary tables that help the reader navigate the area. However, the central novelty claim and the integrity of the joint-detection taxonomy are not yet established, so the contribution as a 'first comprehensive survey' is currently conditional.","major_comments":[{"comment":"The claim that this is the first survey to cover rumor detection, rumor-source detection, and their joint consideration is load-bearing, but the manuscript does not provide the comparison needed to establish it. It reports no systematic search, no inclusion/exclusion criteria, and no analysis of the earlier surveys cited in the paper ([3], [7], [8], [9], [12], [25], [26], [99], [144]) against the three-problem taxonomy. Please add such a comparison or explicitly qualify the novelty claim.","section":"Section I, Contribution (i)"},{"comment":"Reference [128] and reference [135] are the same paper (Seo, Mohapatra, and Abdelzaher, SPIE Defense, Security, and Sensing, 2012). The survey uses it as a joint-detection algorithm in Section III-D and also as a single-source sensor-monitoring estimator in Section III-C and Tables V and VI. This double-counting undermines the joint-detection taxonomy: if the paper is genuinely joint, it should not be listed only as a single-source estimator; if it is not joint, the joint-algorithm section contains only [134]. Please deduplicate and correct the taxonomy accordingly.","section":"Section III-D and Tables V-VI"},{"comment":"Algorithm 3 does not implement Definition 5. Definition 5 defines a single function f: (T, G_I) -> ({0,1}, hat{S}) that returns both outputs simultaneously, whereas the algorithm first estimates sources from the infection graph and then runs a separate rumor-detection function on T. In addition, line 11 repeats the condition 'T is regarded as a rumor' from line 9, making the 'else if' branch unreachable and the non-rumor output (f(T)=1) impossible. Please correct the condition and either present the algorithm as a sequential pipeline or describe a genuinely joint inference loop.","section":"Algorithm 3"},{"comment":"The 'hybrid and other machine learning-based approaches' category includes web-spam papers [69], [70], and [72]. The surrounding text itself describes these as web-spam detection, not rumor detection, and no explicit bridge to rumors is provided. Including them as rumor-detection methods misrepresents the surveyed literature; either remove them or explain the intended connection (e.g., as feature/algorithmic precursors).","section":"Section III-A4 and Tables II-III"}],"minor_comments":[{"comment":"The row label 'Ruomor center' should read 'Rumor center'.","section":"Table VII"},{"comment":"The sentence 'It is known that [135] rumors are usually initiated by a small number of people' has an awkward citation placement; rephrase to 'It is known [135] that rumors are usually initiated...'.","section":"Section II-C"},{"comment":"References [137] and [138] describe steganographic file systems in social networks; the relevance to 'hiding rumors' specifically should be stated explicitly, since the current text describes hiding arbitrary information.","section":"Section IV-A"},{"comment":"The entry 'Deep-Faked' should be 'DEAP-FAKED' to match the cited method and the text in Section III-A1.","section":"Table IV"},{"comment":"Algorithm 1 applies f before describing what f does; consider a wording such as 'Run a rumor-detection algorithm f on T' to avoid giving the impression that f is assigned only afterward.","section":"Algorithm 1"}],"recommendation":"major_revision","confidential_remarks":"The duplicate reference [128]/[135] and the absence of a comparison with prior surveys make the novelty claim fragile, but both issues are fixable within the scope of a survey revision. I do not see grounds for rejection. The paper would also benefit from a careful pass over the taxonomy tables for out-of-scope entries such as the web-spam papers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know: this is a usable survey, but its central novelty claim is shaky and the joint-detection section has concrete errors. The paper maps rumor detection and rumor-source detection side by side, introduces formal definitions for both plus a joint problem, and organizes source-detection work by snapshot type and number of sources. That taxonomy is genuinely helpful. The section on hiding rumor sources is a nice addition, since most surveys stop at detection. The tables summarizing algorithms, evaluation metrics, and challenges are also a decent starting point for someone entering the area.\n\nThe soft spots are real but mostly fixable. First, the claim to be the first survey covering rumor detection, source detection, and their joint consideration is load-bearing but never demonstrated. There is no systematic comparison with earlier surveys; the authors just assert it. That needs to be substantiated or softened. Second, the stress-test note holds up: reference [128] and reference [135] are the same paper by Seo, Mohapatra, and Abdelzaher. It appears as a joint algorithm in Section III-D and also as a single-source sensor-monitoring estimator in Tables V and VI. That double-counting guts the joint-detection category down to one paper, SourceCR [134], unless the authors explain why the duplicate is genuinely joint. Third, Algorithm 3 does not implement Definition 5. It runs source detection and then rumor detection sequentially, not a joint function of (T, G_I), and line 11 repeats the rumor condition instead of the non-rumor branch, making the else-if unreachable. That is a concrete bug in a paper that advertises a formal joint problem. Fourth, the web spam papers [69], [70], [72] sit under rumor detection with no evident connection to rumors; that looks like careless categorization.\n\nOn the positive side, the formal definitions are clean restatements of existing concepts, not new theory, and the paper is honest that the joint problem is a combination. No result is fabricated, and the citation pattern is not self-serving. The defects are concentrated in the joint section and the novelty argument, not in the core survey material.\n\nWho is this for? A graduate student or researcher wanting a structured overview of rumor source detection and the joint problem could get value from the taxonomy and tables, but they should verify the entries. It deserves a serious referee because the errors are concrete and correctable, and the survey fills a real bibliographic gap if Duplicate reference and Algorithm 3 are fixed. I would send it to peer review with major revision, and I would not cite it in its current form.","headline":"Useful survey of two subfields and the joint idea, but the 'first joint survey' claim is unproven and the joint section has concrete bugs ([128]=[135], Algorithm 3).","tokens_in":36848,"tokens_out":1866,"would_cite":false,"duration_ms":20283,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims to be the first to treat rumor detection and rumor-source detection as one formal problem family, defining classification, inference, and joint detection.","keywords":["rumor detection","rumor source detection","joint inference","social networks","misinformation","source localization","graph neural networks","survey"],"falsifier":"A bibliographic search turning up a peer-reviewed survey published before 2025 that already treats rumor detection, rumor-source detection, and their joint consideration with formal definitions would falsify the paper's claim to be the first. Empirically, a study showing that knowing the set of estimated sources adds no predictive power for veracity, or that source multiplicity is unrelated to truthfulness, would undercut the joint problem's motivation.","tokens_in":35927,"feed_emoji":"🔎","tokens_out":4681,"duration_ms":46825,"temperature":0.7,"pith_summary":"The paper sets out to show that the question “is this information true?” and the question “who first spread it?” belong to a single problem family. Most prior surveys cover one or the other; this one organizes both plus their joint treatment, and gives each a formal definition. It matters because a rumor's credibility and its origin are connected in practice: reliable sources tend to spread truth, and true news tends to be reported by many independent people. If the joint framing holds, future systems could localize a rumor and assess its veracity from the same cascade data.","feed_headline":"First survey unites rumor spotting and source hunting","feed_subtitle":"Defines the joint problem formally and shows how detection, source inference, and the combined task relate.","key_machinery":"The load-bearing object is the joint detection function from Definition 5, written as $f: (T, G_I) \\to (\\{0,1\\}, \\hat{S})$, which wraps the rumor-detection classifier and the rumor-source estimator into a single simultaneous inference target. That function is supported by three smaller machinery pieces: the network structure and propagation model (SI, SIR, SIRS, or independent cascade) that generate the infection graph, the snapshot type that determines what data the estimator sees, and the estimator itself, either graph-based centrality such as the Jordan center or a probability-based rule such as the maximum-likelihood estimator. The survey's entire organization hangs on whether this joint function is a meaningful target rather than two tasks that happen to share data.","core_discovery":"On its own terms, the paper's discovery is taxonomic: the literature on misinformation in social networks separates into rumor detection (a classification problem $f: T \\to \\{0,1\\}$), rumor-source detection (an inference problem $f: G_I \\to \\hat{S}$), and joint detection $f: (T, G_I) \\to (\\{0,1\\}, \\hat{S})$, with the last almost unstudied. The paper assembles the existing algorithms under this scheme, classifies source-detection work by the number of sources and by snapshot type (complete, partial, or sensor-monitored), and identifies exactly two existing joint approaches, one driven by source reliability and one by the number of independent reporters. It also covers the opposite problem of hiding rumors and their sources, including protocols that keep the source at the leaf of the infection graph. The intended contribution is a map that lets a researcher see what has been solved and where the joint problem remains open.","pith_inferences":["A direct empirical test of the joint framing would benchmark the two existing joint approaches against their sequential counterparts on the same cascade data, measuring whether coupling actually raises detection probability or F1; the survey itself stops at cataloging rather than comparing.","Definition 5 suggests a modular architecture that the current joint papers only partially exploit: a source estimator can output a reliability-weighted candidate set, and a rumor classifier can consume that set as an additional feature, so the two tasks reinforce each other iteratively.","The source-hiding section implies that any rumor detector relying on source credibility is vulnerable to adversarial propagation timing, so robust detectors may need to treat estimated sources as adversarial inputs rather than trusted ground truth.","The number-of-sources argument would be testable on real event datasets: if a rumor is identified by a small number of independent reporters and true news by many, then a system could classify veracity from source multiplicity alone, a hypothesis the survey does not evaluate empirically."],"forward_implications":["Researchers can place any new rumor or source method into the survey's taxonomy: content-, propagation-, source-, or hybrid-based for rumor detection, and single- versus multiple-source with complete, partial, or sensor snapshots for source detection.","The two existing joint algorithms, one based on source reliability and one based on the number of sources, demonstrate that joint inference is technically possible and suggest it can outperform running the two tasks separately.","The adaptive-diffusion line of work implies that source detection methods must be evaluated against adversaries that actively obfuscate the infection graph, not just against benign cascades.","The challenges table points to concrete next steps: multimodal data integration, temporal alignment between rumor evolution and network dynamics, user-behavior modeling under adversarial manipulation, and scalable and interpretable joint algorithms."],"supporting_citations":[{"why":"Introduces rumor centrality and proves that the rumor center is the maximum-likelihood source estimator on regular trees under the SI model; this is the backbone of the single-source complete-snapshot line.","marker":"[92]–[94]"},{"why":"Establishes the Jordan center as a sample-path-based estimator under the SIR model and extends the approach to limited observations, grounding the probability-based estimator category.","marker":"[95]"},{"why":"Presents the SourceCR framework, one of only two joint-inference algorithms, which iterates credibility-reliability training for truth/rumor classification with a division-querying procedure for source inference.","marker":"[134]"},{"why":"First work to consider rumor detection and rumor-source detection simultaneously, using the number of independently reporting monitors to distinguish true information from rumors.","marker":"[135]"},{"why":"Proposes optimal Jordan cover for locating multiple diffusion sources from partial observations in the heterogeneous SIR model, a central example for the multi-source partial-snapshot category.","marker":"[121]"},{"why":"Studies infection-source estimation for the SI model when not all infected nodes can be observed, defining the partial-snapshot problem setting.","marker":"[110]"},{"why":"Introduces adaptive diffusion, the protocol that keeps the true source at a leaf of the infection graph and achieves near-optimal source obfuscation, grounding the hiding-rumor-sources section.","marker":"[139]"}],"fun_headline_variants":["Rumor detection and source tracing finally surveyed as one","Only two methods combine rumor detection and source finding","First joint survey of rumor spotting and source inference","Survey maps the rare joint study of rumors and their origins"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey assumes that rumor detection and source detection are genuinely coupled, through source reliability and the expected number of independently reporting sources, so that solving them jointly is a meaningful goal rather than two independent tasks.","fun_headline_variants_meta":{"raw":{"variants":["Rumor detection and source tracing finally surveyed as one","Only two methods combine rumor detection and source finding","First joint survey of rumor spotting and source inference","Survey maps the rare joint study of rumors and their origins"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000798,"raw_usage":{"total_tokens":3546,"prompt_tokens":1017,"completion_tokens":2529,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":633,"completion_tokens_details":{"reasoning_tokens":2467}},"tokens_in":633,"tokens_out":2529,"duration_ms":18087,"temperature":1.0,"reasoning_tokens":2467,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:13:15.421492+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A bibliographic search turning up a peer-reviewed survey published before 2025 that already treats rumor detection, rumor-source detection, and their joint consideration with formal definitions would falsify the paper's claim to be the first. Empirically, a study showing that knowing the set of estimated sources adds no predictive power for veracity, or that source multiplicity is unrelated to truthfulness, would undercut the joint problem's motivation.","supporting_citations":[{"cited_title":"Rumors in a Network: Who’s the Culpri t? IEEE Trans","cited_arxiv_id":null,"evidence_quote":"Establishes the Jordan center as a sample-path-based estimator under the SIR model and extends the approach to limited observations, grounding the probability-based estimator category."},{"cited_title":"Expert Systems With Applications February 16, 2024","cited_arxiv_id":null,"evidence_quote":"Presents the SourceCR framework, one of only two joint-inference algorithms, which iterates credibility-reliability training for truth/rumor classification with a division-querying procedure for source inference."},{"cited_title":"IEEE Confer- ence on Computer Communications 2020","cited_arxiv_id":null,"evidence_quote":"First work to consider rumor detection and rumor-source detection simultaneously, using the number of independently reporting monitors to distinguish true information from rumors."},{"cited_title":"Efﬁciently Spotting the Starting Points of an Epidemic in a Large Graph","cited_arxiv_id":null,"evidence_quote":"Proposes optimal Jordan cover for locating multiple diffusion sources from partial observations in the heterogeneous SIR model, a central example for the multi-source partial-snapshot category."},{"cited_title":"Conﬁdence Sets for Source of a Diff usion in Regular Trees","cited_arxiv_id":null,"evidence_quote":"Studies infection-source estimation for the SI model when not all infected nodes can be observed, defining the partial-snapshot problem setting."},{"cited_title":"SocialStegDisc: Appl ication of steganog- raphy in social networks to create a ﬁle system In In Proc","cited_arxiv_id":null,"evidence_quote":"Introduces adaptive diffusion, the protocol that keeps the true source at a leaf of the infection graph and achieves near-optimal source obfuscation, grounding the hiding-rumor-sources section."}],"review_version":1}