Pith. sign in

REVIEW 1 cited by

Accuracy and Political Bias of News Source Credibility Ratings by Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.00228 v3 pith:IDWRPIFU submitted 2023-04-01 cs.CL cs.CYcs.IR

classification cs.CLcs.CYcs.IR
keywords llmsratingsinformationmodelsnewssourcesbiaspolitical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Search engines increasingly leverage large language models (LLMs) to generate direct answers, and AI chatbots now access the Internet for fresh data. As information curators for billions of users, LLMs must assess the accuracy and reliability of different sources. This paper audits nine widely used LLMs from three leading providers -- OpenAI, Google, and Meta -- to evaluate their ability to discern credible and high-quality information sources from low-credibility ones. We find that while LLMs can rate most tested news outlets, larger models more frequently refuse to provide ratings due to insufficient information, whereas smaller models are more prone to making errors in their ratings. For sources where ratings are provided, LLMs exhibit a high level of agreement among themselves (average Spearman's $\rho = 0.79$), but their ratings align only moderately with human expert evaluations (average $\rho = 0.50$). Analyzing news sources with different political leanings in the US, we observe a liberal bias in credibility ratings yielded by all LLMs in default configurations. Additionally, assigning partisan roles to LLMs consistently induces strong politically congruent bias in their ratings. These findings have important implications for the use of LLMs in curating news and political information.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 41 citations worldwide. Full citation record

  1. Fisher Random Walk: Automatic Debiasing Contextual Preference Inference for Large Language Model Evaluation

    stat.ML 2025-09 conditional novelty 7.0 of 10

    A Fisher random walk weighted residual estimator achieves semiparametric efficient confidence intervals for contextual Bradley-Terry-Luce preference comparisons with flexible score estimators.

Pith tools