{"id":"e4fe4b04-d264-4761-a853-627ba62176f2","arxiv_id":"2607.12108","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A single-person, methodologically untested media rating site (MBFC) is commonly mischaracterized as an independent fact-checking organization in ~3.5–7.4% of misinformation papers, and this audit argues it should not be used as ground truth.","lead":"Media Bias/Fact Check (MBFC) claims to rate the bias and credibility of 10,000 media sources, but is actually one person's untested opinion. This paper audits MBFC's rubrics and a literature corpus to show it is widely misused and misdescribed in misinformation science.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'no papers adequately describe MBFC' claim rests on an incomplete snowball sample; a targeted full-text search of the unexamined 61% could falsify it.","rationale":"The reader's weakest_assumption is exactly the spot I would push on. The paper's own Section III B concedes incomplete text coverage and greedy sampling, yet the abstract states a universal negative. A single counterexample would not refute the claim that MBFC is widely misused, but it would force the conclusion from 'no papers' to 'few papers in our incomplete sample.' The central methodological critique stands on direct evidence from MBFC's archived pages, and it is right to credit that. I see no additional load-bearing flaw: the rubric analysis is concrete, the disclaimers are quoted, and the heuristic claim of undocumented weights/arbitrary bins is supported by Tables I–III. The main risk is overstating the literature audit as exhaustive. Hence keep the verdict CONDITIONAL: the paper should soften the unqualified 'no papers' claim or add the targeted check.","tokens_in":20958,"tokens_out":3071,"duration_ms":31682,"concrete_test":"Perform a targeted full-text search across all papers in the snowball database plus the 1,610 Google Scholar hits, retrieving PDFs via Unpaywall/Internet Archive and using full-text corpora (Semantic Scholar, PubMed Central, arXiv) for phrases like 'Dave Van Zandt', 'solely owned', 'one person', 'not a tested scientific method', 'hobby', and 'critically' near 'MBFC'. If any paper retrieved by this expanded search adequately describes MBFC as a single person's opinion or explicitly engages with its methodological limitations before using it, the abstract's 'no papers' claim is falsified; if none is found after this broader search, the claim is substantially strengthened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's most sweeping empirical assertion is in the abstract: 'We identified no papers that adequately describe MBFC as the opinions of a single person or critically engage with its methodology in order to justify proceeding with its use.' This negative existential is load-bearing for the claim that the literature 'rarely examines MBFC carefully' and for the conclusion that researchers should 'stop using' the dataset. But the audit that produces it is a snowball sample seeded only from Scopus metadata for 'Media Bias/Fact Check' and expanded through OpenAlex/Crossref/OpenCitations, with PDF fetch and OCR as a gate. Section III B reports that manuscript text was unavailable for 61% of the 10,642 papers in the database; any paper that uses MBFC without a citation in the expected format, sits behind a paywall, or has OCR failures is invisible to the search. The paper acknowledges these limitations ('we expect that 3.50% is an underestimate') but continues to use the unqualified 'no papers' phrase. The methodological critique of MBFC itself — unexplained rubric weights, inaccessible raw scores, arbitrary bins, single-owner disclaimer — is well evidenced from MBFC's own pages and archived versions, so this concern does not undermine the core 'do not treat as uncritical ground truth' conclusion; it only weakens the empirical claim that no paper anywhere in the literature has handled MBFC adequately.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper audits Media Bias/Fact Check (MBFC), a widely used source of media bias and credibility ratings, and argues that it is unsuitable as a basis for academic research. The authors document from MBFC's own pages and archives that its methodology uses unexplained weighting, arbitrary binning, inaccessible raw scores, nested credibility logic, and economically/politically contentious category definitions; they note MBFC's explicit disclaimer that it is not a tested scientific method. They then report a snowball literature search in which 372 of 10,642 papers (3.50%) in their database reference MBFC, with higher percentages in a misinformation-focused subset, and argue that the literature overwhelmingly misdescribes MBFC as an independent fact-checking organization rather than as one individual's opinion. They conclude by urging researchers to stop using MBFC and by interpreting the dataset's acceptance through a Gramscian/hegemony framework.","tokens_in":21286,"tokens_out":3542,"duration_ms":37422,"significance":"If the paper's central methodological critique stands, it is a valuable and overdue caution about a widely used dataset. The authors deserve credit for grounding the critique in primary sources: MBFC's own methodology pages, archived versions, and public disclaimers are quoted and dated, and the internal inconsistencies they identify (e.g., the three incompatible bias scales, the unexplained bin boundaries, the composite credibility logic) are concrete and verifiable. The paper also usefully documents how often MBFC is misdescribed in the literature. However, the strongest empirical claim—that no paper anywhere has adequately described or critically justified MBFC—rests on a sampling procedure with a large unexamined fraction, and the accuracy claim is supported by illustrative examples rather than a systematic validation. These issues limit the strength of the paper's abstract and conclusions, though they do not undermine the core point that MBFC's methodology is not transparent enough to serve as uncritical ground truth.","major_comments":[{"comment":"The abstract states: \"We identified no papers that adequately describe MBFC as the opinions of a single person or critically engage with its methodology.\" This negative existential is load-bearing for the paper's empirical claim that the literature \"rarely examines MBFC carefully,\" but the sampling method cannot support it. Section III B reports that manuscript text was unavailable for 61% of the 10,642 papers, and the snowball algorithm terminates branches whose PDFs cannot be fetched. Papers behind paywalls, with OCR failures, or citing MBFC in ways not captured by the seed/metadata pipeline are invisible to the search. The paper itself acknowledges \"we expect that 3.50% is an underestimate,\" yet the unqualified \"no papers\" phrase is retained. The conclusion should be restricted to \"no papers in our searchable sample,\" or the authors should supplement the snowball procedure with a targ","section":"Abstract; Section III B"},{"comment":"The abstract claims MBFC's data is \"not neutral or accurate.\" The non-neutrality and lack of reproducibility are well documented, but \"not accurate\" is a stronger empirical claim. The evidence in Section V C consists of selected internal inconsistencies and counterexamples (Saudi Arabia, Huffington Post, state-owned economies) rather than a systematic comparison of MBFC ratings against an independent benchmark or a formal test of rubric application. It is possible that the paper's intended meaning is that accuracy cannot be verified because raw scores are inaccessible and the methodology is inconsistent. That claim is supported and would be sufficient for the policy recommendation. The authors should either provide a systematic accuracy evaluation or reframe the conclusion as \"MBFC's accuracy cannot be established and its ratings exhibit internal inconsistencies,\" rather than asserting i","section":"Abstract; Section V C"},{"comment":"The Gramscian/hegemony interpretation is presented as an explanation for MBFC's design and scholarly acceptance, and the abstract includes the claim that MBFC is \"a computationally legible account of hegemony.\" This is a theoretically grounded interpretation, not an empirical result of the audit in Sections II–IV. The paper does not define observable criteria that would distinguish the hegemony explanation from alternative explanations (e.g., convenience, lack of better data, or simple negligence). Since the historical narrative and ideological framing are not required for the central methodological conclusion, they should be clearly labeled as an interpretive hypothesis rather than as a finding derived from the documented evidence. This separation would also make the paper's argument easier for readers who accept the methodology critique but do not subscribe to the specific theoretical","section":"Sections V B–V D"}],"minor_comments":[{"comment":"Typographical and formatting issues: \"outlet's\" should be \"outlets\"; \"wisely used\" in Section III B should presumably be \"widely used\"; \"plaintly\" in Section V D should be \"plainly\"; \"Hufftingon post\" in reference [48] should be \"Huffington Post\"; \"MNSIT\" in Section V D should be \"MNIST.\"","section":"Section II"},{"comment":"The phrase \"terminating the snowball sample for that branch\" is slightly misleading: if a PDF cannot be fetched, the algorithm stops following that paper's citations, so the sample is biased toward papers whose full text is openly available. This should be stated explicitly in the main text, not only in the limitations discussion.","section":"Section III B; Table IV"},{"comment":"The label \"MBFC would be here\" is unclear. It appears to indicate where MBFC-referencing papers would rank if they were highly cited, but the figure does not show a concrete position. Consider adding a marker or a more descriptive caption.","section":"Figure 4"},{"comment":"The list of \"countries with which the United States has strained or hostile relations\" is presented without selection criteria. Since the subsequent observation that almost all of these countries receive \"total oppression\" is used as evidence, the criteria for inclusion in this list should be specified to avoid the appearance of cherry-picking.","section":"Section II C; reference [26]"},{"comment":"The NLP context-window method excludes contexts containing a digit directly before an alias, which is reasonable for footnotes, but this also means that some relevant methodological discussions occurring in footnotes are not analyzed. The paper does rely on qualitative reading as well, but this limitation should be acknowledged near the NLP methods description.","section":"Section IV A"}],"recommendation":"major_revision","confidential_remarks":"The paper is best read as a critical dataset audit/commentary rather than a conventional empirical study. The methodological critique of MBFC is well evidenced and should be published in some form, but the abstract's strongest empirical and theoretical claims need to be brought in line with the evidence. I would advise the editor to consider whether the journal welcomes this type of field-critique paper, and to encourage the authors to add a reproducibility note or link to their snowball-search code, since it is currently not provided."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: the core critique of MBFC is solid and well documented from primary sources — unexplained rubric weights, uneven bins, inaccessible raw scores, nested logic, and MBFC's own disclaimer that the method is 'not a tested scientific method.' Against that, many papers describe MBFC as an 'independent fact-checking organization.' That mismatch is real and worth publishing.\n\nSecond: the paper's strongest empirical claim — that no paper adequately describes MBFC as one person's opinion or critically engages its methodology — is not established by the audit. The snowball sample covers 10,642 papers, but manuscript text was unavailable for 61% of them (Section III B). Many paywalled, non-OCR-able, or uncited uses would be invisible. The abstract and Section V A use 'no papers' without that caveat. Absence of evidence here is not evidence of absence. The core methodological critique does not depend on this claim, so the paper survives, but the headline 'no papers' should be softened to 'we found no papers among the 39% whose text we could examine.'\n\nWhat's good: the systematic audit is a useful contribution in itself; the tables of ratings and rubrics are carefully extracted; the discussion of MBFC's three scales and the 2025 rubric change is sharp; and the authors are honest about their sampling limitations in Section III B, even if the abstract doesn't carry the caveat. The historical section on the 'liberal media' campaign is clearly written and relevant, though it's interpretative rather than decisive.\n\nSoft spots beyond the census: no data or code is released, which weakens reproducibility for a paper whose topic is reproducibility. The 'stop using' conclusion is a normative leap; the evidence supports 'do not treat MBFC as uncritical ground truth' and 'describe it accurately if you use it.' A total ban is defensible but not forced by the data. Also, the paper does not directly benchmark MBFC against an independent gold standard; it argues the ratings' recognizable shape is itself the problem — fair as a critique of epistemic hygiene, less strong as a claim about accuracy.\n\nWho this is for: anyone in misinformation research who uses MBFC, and editors of venues that publish such work. It deserves a serious referee — it's an important, well-sourced intervention — even though the negative existential should be revised. I'd bring it to reading group.\n\nRecommendation: engage with it; require the sampling caveat to be carried through the abstract and the no-papers claim to be reworded, and ask for a data/code release.","headline":"A well-evidenced methodological warning about MBFC wrapped in a literature census whose strongest negative claim outruns the data.","tokens_in":21726,"tokens_out":3332,"would_cite":true,"duration_ms":33377,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that Media Bias/Fact Check, a widely used source of media-bias ratings, is methodologically unsound and that misinformation research should stop relying on it.","keywords":["media bias","fact-checking","misinformation","ground truth","data quality","methodology critique","media credibility","hegemony"],"falsifier":"A single peer-reviewed paper that (a) cites MBFC, (b) explicitly states that MBFC is the opinion of one individual, and (c) nonetheless justifies using it with a critical engagement with its limitations would falsify the paper's strongest negative claim. Alternatively, an archived pre-2025 version of MBFC's dataset with the raw scores would undercut the claim that the data is unauditable.","tokens_in":20862,"feed_emoji":"📰","tokens_out":4519,"duration_ms":41419,"temperature":0.7,"pith_summary":"This paper argues that Media Bias/Fact Check (MBFC) — a widely used source for rating the bias and credibility of news outlets — does not meet basic standards of academic rigor. A snowball survey of 10,642 papers found MBFC in 3.5% to 7.4% of the relevant literature, often as ground truth for machine-learning models, yet no paper in the sample adequately described MBFC as the opinion of a single individual or critically justified its use. The authors show that MBFC's methodology relies on unexplained weights, arbitrary bin sizes, hidden raw scores, and rubrics that contradict the site's own disclaimer that it is 'not a tested scientific method.' If the paper is right, results built on MBFC ratings rest on an unauditable and politically loaded account of media bias, and the field's crisis-of-trust conclusions may be reproducing the very discourse they study.","feed_headline":"One person's rating site feeds 372 research papers","feed_subtitle":"A methodology audit finds Media Bias/Fact Check's rubrics are arbitrary and hidden, yet papers treat it as an independent fact-checker.","key_machinery":"The argument rests on two mechanisms. First, a methodological audit of MBFC's public documentation and API output, exposing unexplained percentage weights, irregular binning of raw scores that are never exposed to users, nested override logic in the credibility rubric, and conceptual slippage between {bias} and {factual reporting}. Second, a snowball citation sample seeded from Scopus records, extended through metadata databases, that scans manuscript texts for MBFC mentions and analyzes the surrounding context with NLP to characterize how the literature describes the tool. The audit shows the data cannot be independently verified; the literature survey shows it is nevertheless used uncritic","core_discovery":"The paper's central claim is that MBFC's data is specious rather than neutral: it quantifies the results of a decades-long political campaign to discredit the press, then presents those results as simple facts about the world. The evidence comes from two sides: a close reading of MBFC's currently published rubrics (e.g., a {bias} score 35% determined by an {economic system} scale that would classify Nazi Germany or Saudi Arabia as leftist), and a literature analysis showing that 372 papers invoke MBFC, frequently describing it as an 'independent fact-checking organization' despite MBFC's own description as a one-person hobby with a disclaimer that its method is not tested science. The author","pith_inferences":["An editorial extension: the same structural critique may apply to other media-quality ratings (e.g., those built on small panels or proprietary rubrics) if they too hide their scoring procedures; a comparative audit of such systems would test whether MBFC is uniquely bad or representative of a wider problem.","If the paper's claim about hegemony is taken seriously, even an 'improved' MBFC that added transparency and citations could not fully fix the problem, because the underlying categories (liberal, biased, factual) are themselves politically contested; a purely procedural fix would not address the conceptual issue.","One testable extension: replicate the literature survey using full-text search through Google Scholar or publisher APIs rather than relying on PDFs and OCR; this should find papers that the snowball missed and directly test the paper's claim that no paper adequately describes MBFC as one person's opinion.","Another extension: examine whether downstream models trained on MBFC labels produce systematically different accuracy or harm outcomes when re-trained on alternative ratings, which would quantify the real-world cost of using MBFC."],"forward_implications":["Any study that uses MBFC ratings as ground truth for training or validating fake-news classifiers inherits a single individual's untested, politically situated judgments and risks encoding them into automated moderation systems.","Highly cited papers that rely on MBFC—including several that define 'reliable' and 'questionable' sources—should be re-examined for how their conclusions change if MBFC's labels are not accepted.","Composite datasets and commercial products that bundle MBFC propagate the same specious ground truth, sometimes beyond the point where provenance is traceable.","Existing misinformation research that uses MBFC cannot simply be corrected; the underlying measurements are unarchived for pre-2025 data, so comparisons across time are impossible.","The field needs alternative, transparent, and independently auditable source-rating schemes before MBFC is retired."],"fun_headline_variants":["372 studies lean on one person's media ratings","Media Bias/Fact Check: a one-person hobby, not science","Research leans on a site that calls itself 'not tested science'","One biased rubric drives 372 misinformation papers"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The strong empirical claim that no paper in the literature adequately describes MBFC as one person's opinion rests on the snowball sample being complete, but manuscript text was unavailable for 61% of the papers in the database, so the claim could be an artifact of missing data rather than the true state of the literature.","fun_headline_variants_meta":{"raw":{"variants":["372 studies lean on one person's media ratings","Media Bias/Fact Check: a one-person hobby, not science","Research leans on a site that calls itself 'not tested science'","One biased rubric drives 372 misinformation papers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000795,"raw_usage":{"total_tokens":3346,"prompt_tokens":761,"completion_tokens":2585,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":2528}},"tokens_in":505,"tokens_out":2585,"duration_ms":17325,"temperature":1.0,"reasoning_tokens":2528,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T06:40:29.183174+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A single peer-reviewed paper that (a) cites MBFC, (b) explicitly states that MBFC is the opinion of one individual, and (c) nonetheless justifies using it with a critical engagement with its limitations would falsify the paper's strongest negative claim. Alternatively, an archived pre-2025 version of MBFC's dataset with the raw scores would undercut the claim that the data is unauditable.","supporting_citations":[],"review_version":2}