Pith. sign in

REVIEW 4 major objections 5 minor 22 references

Language models mediate politics well until the evidence goes murky — then they fail systematically.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 03:22 UTC pith:BOHHIG64

load-bearing objection A genuinely useful and unusually transparent political-mediation benchmark, but every headline number flows through unvalidated LLM judges — human re-annotation should be the condition for treating it as a certified audit. the 4 major comments →

arxiv 2607.25953 v2 pith:BOHHIG64 submitted 2026-07-28 cs.CL cs.CY

Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections

classification cs.CL cs.CY
keywords LLM evaluationpolitical information mediationEpistemic Modestyelection integrityinformation environmentsLLM-as-judgeparty priorsvoting advice
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish a standard for judging whether LLMs responsibly mediate political information for voters, and shows that high aggregate scores hide systematic failures. It proposes Epistemic Modesty — faithful, impartial, and well-calibrated handling of evidence — and measures it with a benchmark that varies the clarity and consistency of the evidence. Across three frontier LLMs and the 2025 German and Dutch elections, models stay reliable when evidence is clear (about 97% adherence) but drop to 85–86% under absent or vague evidence and to 80% under contradictory evidence. They also scrub the intensity of political rhetoric, and party-specific patterns (like sanitizing a far-right party or falling back on priors for a left-wing party) persist across conditions. The point: good average scores are not evidence of safe mediation, because the failures cluster exactly where voters most need caution.

Core claim

POLISTEMICS, the paper's benchmark, treats LLM mediation as a transformation from a party-position query plus a controlled evidence context into a free-form answer, and scores each answer on Faithfulness (does it represent the evidence?), Impartiality (does it avoid steering?), and Epistemic Calibration (does its certainty match the evidence?). The central finding is that the models' aggregate scores mask localized breakdowns: under the clear-evidence baseline all three evaluated models reach 96–98% adherence, but absent, vague, or contradictory evidence pushes adherence down to 74–86% for most model-environment combinations, with the hardest cases being contradictory evidence (80% on averag

What carries the argument

The load-bearing device is a set of six controlled Information Environments built from standardized, real party-position evidence: Baseline, Absent (no evidence), Vague (rewritten to obscure the stance), Contradictory (two opposing passages), Noisy (distractors), and Counterfactual (evidence conflicting with priors). Each environment is scored by a three-model judge panel answering yes/no sub-questions, aggregated into an Adherence Index per environment and an Epistemic Modesty Index (geometric mean across environments). The environments isolate which informational property causes breakdowns, and the party-level breakdowns turn a single average score into a diagnostic.

Load-bearing premise

The entire benchmark rests on three LLM judges' yes/no verdicts being accurate measures of Faithfulness, Impartiality, and Epistemic Calibration — but those judges were never checked against human ratings (the paper concedes this in its Limitations).

What would settle it

Take a random sample of the benchmark's outputs and have a panel of human experts score them under the same rubrics; if the LLM judge panel and human raters disagree materially (or if judges' verdicts shift when the 'expected stance' priming line is removed), then the reported pass rates are properties of the judges, not the mediated answers.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • The 97% baseline ceiling is the relevant benchmark: any model that scores at or near that ceiling on clear evidence should be assumed fragile until tested on inconclusive evidence.
  • Deploying LLMs as voting assistants without guarding absent, vague, or contradictory evidence is unsafe; developers and regulators should require reporting under inconclusive evidence.
  • The party-prior effect implies models will systematically misrepresent smaller or newer parties (e.g., BSW, BBB) when evidence is weak, and will drift as parties change positions after training cutoffs.
  • Because anonymizing party labels shifts behavior, much of what these models express as knowledge is a learned association between a party name and a stance, not reasoning from the evidence at hand.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The judge panel is the whole measurement chain; until human-validated, the absolute pass rates (80–98%) are best read as relative rankings, not calibrated facts.
  • The first-chunk preference under contradictory evidence in Germany (but not the Netherlands) suggests the models anchor on recency or salience; this is testable by reversing chunk order and re-measuring.
  • The English-output ablation shrinking sanitization gaps hints the effect is partly a translation artifact toward a neutral register, implying single-language evaluations may overstate language compression.
  • A practical extension: present the same evidence with the stance removed entirely (like Vague) and test whether refusal or hedging improves with simple prompt instructions, which would suggest the failure is trainable.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces Polistemics, a theory-grounded benchmark for evaluating LLMs as mediators of political information in elections. It defines three rubrics — Faithfulness, Impartiality, and Epistemic Calibration — and tests them across six controlled information environments built from standardized Wahl-O-Mat and StemWijzer evidence for the 2025 German and Dutch elections. Three frontier LLMs are queried under each condition and scored by a three-judge LLM panel with majority voting. The main empirical claims are that aggregate scores are high but mask a breakdown under absent, vague, or contradictory evidence, and that political language is flattened throughout, with party-specific disparities suggesting model priors. A Dutch replication is included.

Significance. If the findings hold, the benchmark is a valuable step toward standardized, theory-grounded evaluation of political information mediation. The paper has real strengths: controlled construction of information environments, fully disclosed scoring rules, prevalence-robust inter-judge agreement (AC1), quantified nondeterminism at temperature 0, and a cross-country replication that openly reports non-replications. The repository and evaluation pipeline are also concrete assets. However, the quantitative validity of every headline number currently depends on three LLM judges that were never checked against human raters, and at least one headline pattern is substantially encoded in the scoring rules. The paper is a promising diagnostic instrument, but its central empirical claims need human validation and a clearer separation between normative scoring and behavioral measurement before they can be accepted as stated.

major comments (4)
  1. [§5.1, §10, Table 11] All headline numbers are produced by three LLM judges whose verdicts are never checked against human ratings; the paper explicitly concedes this in §10. High agreement (App. G.1: AC1 = 0.76–0.95) shows consistency, not validity. Two design choices compound the risk: for Baseline/Noisy/Counterfactual the judge prompt tells the judge the 'expected stance' (Table 11), potentially anchoring F1 and Impartiality judgments; and the Gemini 3 Flash judge is also the model that generated the standardized evidence and all Vague passages (§4.2, App. E.1), so a family-affinity effect is not excluded by the Panickssery et al. discussion. I recommend a human reannotation of a stratified sample (e.g., 200–400 items across rubrics/IEs) with reported agreement against the panel, and re-estimation of the main per-IE scores on items where humans and judges agree.
  2. [§5.1, Fig. 19] The Adherence Index is defined as the average of the three Rubric Scores, but under Absent only Epistemic Calibration is scored (Fig. 19 explicitly says 'Scored on EC only'). The Absent row in Fig. 3 is therefore an EC-only number, not an average of three rubrics like the other IEs. This makes the cross-IE comparison 'Baseline 97% → Absent 86%' partly a comparison of different score types. The paper should either define and report an Absent-specific composite, or present the Absent result solely as an EC score and adjust the abstract's aggregate claim accordingly.
  3. [Table 10, §7.2] The headline 'break down when evidence is absent, vague, or contradictory' is largely encoded in the EC scoring rules: under inconclusive IEs, EC1 passes only if the model does not take a definitive stance, EC2 passes only if it hedges, and EC3 requires it to state the context's limits. A model that answers from knowledge is therefore scored as failing by construction. The paper's behavior-level SQ analyses (e.g., Figs. 24–27) partly address this, but the abstract and §7.2 phrase the pattern as an empirical discovery. I recommend reporting unconditional behavior frequencies (e.g., abstention, hedging, fallback rates) alongside pass rates and softening the causal/descriptive wording so the normative scoring and the empirical behavior are kept distinct.
  4. [App. D.4, Table 6, §10] The 'flattening the intensity of political language throughout' claim rests on I4 Sanitization, but I4 is judged against standardized evidence that is 1.6–2.2× longer and roughly twice as intensifier-dense as the raw rationales (Table 6). The standardization step may therefore pre-intensify the reference, making any model output look sanitized. The paper's own correlation checks (DE ρ = −0.05, NL ρ = 0.06) and matched-tercile persistence are reassuring for party differences, but they do not validate the absolute 'flattening throughout' claim. Also, I4 has the lowest inter-judge agreement (AC1 = 0.79; §10, App. G.1). Please qualify the claim as relative to the standardized evidence and, if the abstract's language is retained, report comparisons against raw rationales.
minor comments (5)
  1. [§5.1] The aggregation description should explicitly state the Absent exception (EC-only scoring); the current wording implies three rubric scores are always averaged.
  2. [Figures 2 and 3] The color scale is said to be 0.60–1.00, but several appendix heatmaps contain values below 0.60 (e.g., Fig. 14, GPT Contradictory μ = 0.26). Clarify whether the scale applies only to the main figures or also to appendices.
  3. [Abstract and body] The benchmark name is inconsistent (Polistemics vs. POLISTEMICS); pick one convention.
  4. [App. E.1, Table 8] The caption says n = 485 observations total, but the sample is pooled across DE and NL; please report per-country counts so the reader can see how the mode-level adherence estimates are distributed.
  5. [§10] The Limitations section acknowledges the lack of human validation, but the abstract and conclusion do not hedge the headline claims accordingly. A one-sentence caveat in the abstract would better reflect the paper's own stated limitation.

Circularity Check

1 steps flagged

Partial circularity: the headline 'breakdown under inconclusive evidence' is substantially written into the EC scoring rules (Table 10), though model-level and party-level differences are empirical.

specific steps
  1. self definitional [§5.1, Table 10 (Epistemic Calibration Scoring), interpreted in §7.2 Inconclusive IEs and the Abstract]
    "Epistemic Calibration Scoring Unlike other rubrics, Epistemic Calibration defines normatively 'good' behavior differently depending on the environment. An output is calibrated if it follows the logic in Table 10. Table 10: Inconclusive IEs (Absent, Vague, Contradictory): EC1 Certainty Pass if No; EC2 Hedging Pass if Yes; EC3 Transparency Pass if Yes; EC4 Fallback Pass if No."

    The abstract's central finding — 'Models mediate reliably under clear evidence but break down when it is absent, vague, or contradictory' — is, in its Epistemic Calibration component, a restatement of Table 10 rather than an independent discovery. Under conclusive IEs, committing to a definitive stance and not hedging passes; under Absent/Vague/Contradictory, the same behaviors fail EC1/EC2 by definition, and any output using outside knowledge fails EC4. Since the paper's own results show only GPT abstains perfectly under Absent, the remaining models are scored as 'breaking down' essentially because their answers are definitive or knowledge-based — exactly the behavior the rubric defines as failure. The magnitudes and model/party differences (e.g., Qwen's Die Linke fallback) are empirical,

full rationale

Most of the paper is a self-contained benchmark rather than a fit-then-predict derivation. I found no fitted parameter renamed as prediction, no imported uniqueness theorem, and no load-bearing self-citation. The primary circularity concern is the EC scoring protocol: the IE difficulty ranking (Inconclusive hardest) is directionally encoded in Table 10, so the headline claim should be read as an evaluation against a normative rule, not an emergent empirical law; the paper is transparent about this by presenting Epistemic Modesty as a derived standard. §10 also concedes 'The scoring relies on LLM judges without additional human validation' and 'The party-prior mechanism ... is inferred from converging behavioral patterns rather than measured directly'; these are validity limitations rather than circularity, but they weaken the force of any independent empirical claim. The empirical content that remains (EC3 transparency bottleneck, GPT's perfect abstention, I4 Sanitization gaps, Dutch replication) does not reduce to the rubric. Overall, one central claim is partially definitional, so score 4 rather than 0 or 6.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The benchmark is an empirical instrument rather than a derivation, so the ledger tracks hand-chosen thresholds, normative scoring rules, and unvalidated measurement assumptions instead of fitted constants. The paper discloses almost all of these; the three that matter most are the EC rubric's abstain-always rule (which by construction converts 'answers from prior knowledge' into 'breakdown'), the LLM-judge-as-truth premise, and the standardization rules that amplify the rhetoric baseline against which sanitization is judged.

free parameters (4)
  • Judge majority threshold (≥2/3) = ≥2/3 of three judges
    Hand-chosen in §5.1 for panel verdicts; a different threshold would change pass/fail assignments and all downstream scores.
  • Party-spread reporting heuristic = 0.20 (max–min spread)
    Explicitly "a reporting heuristic rather than a significance criterion" (§7); drives which party disparities are reported as findings.
  • Evidence standardization padding rules = 4–6 sentences; rhetorical emphasis expansion
    Standardization lengthens evidence 1.6–2.2× and more than doubles intensifier density (App. D.4, Table 6); I4 Sanitization is judged against this synthetic, amplified baseline rather than the party's original wording.
  • EC rubric pass conditions (inconclusive IEs) = abstain/hedge = pass; parametric fallback = always fail
    Table 10: under Absent/Vague/Contradictory, EC1/EC2 flip to require uncertainty and EC4 requires no outside knowledge. This normative encoding is what makes "breakdown under inconclusive evidence" a measurable drop.
axioms (5)
  • domain assumption Epistemic Modesty, operationalized as Faithfulness + Impartiality + Epistemic Calibration, is the correct normative standard for political information mediation.
    §3 derives the framework from epistemic agency (Coeckelbergh), but the specific choice to score abstention as always epistemically modest and any flagged parametric fallback as a failure is a normative decision imposed by the rubric, not a theorem.
  • domain assumption LLM judges produce valid measurements of the rubrics without human validation.
    §5.1, §10: agreement coefficients (AC1 0.76–0.95, App. G.1) measure consistency, not accuracy; no human ratings anchor the scale.
  • standard math City-block distance over VAA stance codes is a valid ideological proximity metric for pair selection in Contradictory/Counterfactual.
    App. E.3, Louwerse & Rosema (2014): D(P,Q) = mean |p_i − q_i| with Disagree=0, Neutral=0.5, Agree=1. Reasonable, but categorical stance differences are a coarse proxy for ideological proximity.
  • domain assumption Standardization preserves stance content and rhetoric intensity.
    §4.2/App. D: validation is programmatic only (sentence count 4–6, [PARTY] placeholder, pronoun exclusion); semantic fidelity to the original rationale is asserted, not verified, and D.4 shows intensity is deliberately amplified.
  • domain assumption Single-turn, T=0, persona-free 'helpful assistant' queries capture ecologically relevant mediation behavior.
    §4, §10: the paper explicitly narrows to this setting; deployed VAA and multi-turn contexts are excluded, so all conclusions are conditional on this setup.

pith-pipeline@v1.3.0-alltime-deepseek · 25992 in / 22497 out tokens · 301502 ms · 2026-08-04T03:22:51.716648+00:00 · methodology

0 comments
read the original abstract

As LLMs increasingly shape the political information citizens rely on, no standard exists to assess whether they do so responsibly. We introduce Polistemics, a theory-grounded diagnostic benchmark for evaluating LLMs as mediators of political information in elections. Prior work has treated this task as reproduction rather than mediation, leaving its epistemic dimensions and interaction with imperfect information unaddressed. We ground the evaluation in Epistemic Modesty, a normative standard derived from citizens' epistemic agency, and test it across controlled settings that vary the clarity, noise, and consistency of the available evidence. Applying the benchmark to three state-of-the-art LLMs across the 2025 German and Dutch elections, we find that high aggregate scores mask systematic failures. Models mediate reliably under clear evidence but break down when it is absent, vague, or contradictory, while flattening the intensity of political language throughout. These failures point to party priors, shifting with party labels and output language. Reliable mediation appears achievable, but no model delivers it consistently.

Figures

Figures reproduced from arXiv: 2607.25953 by Baran Peters, Gabor Hollbeck, Kevin O'Sullivan, Robert Jakob.

Figure 1
Figure 1. Figure 1: The POLISTEMICS pipeline. Standardized evidence, unaltered as the Baseline or manipulated into five controlled Information Environments ( [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overall model performance (Germany). Adherence Index per model × rubric, plus EMI. Scores in %; colour scale 0.60–1.00; rows ordered by EMI (NL: App. I). 7.2 Information Robustness Which IEs are hardest? The Baseline sets the effective ceiling at 97%. The interfering IEs remain within 1 pp of the Baseline. The inconclusive IEs are substantially more demanding, with Absent and Vague both falling to 87% and … view at source ↗
Figure 3
Figure 3. Figure 3: Robustness across IEs (Germany). Adher￾ence Index per model × IE; columns ordered baseline￾first then by difficulty. Scale as [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Even-Handedness across IEs (DE + NL). Party spread (max − min Adherence Index across parties) per model × IE; columns ordered by difficulty within each country. Shared scale, 0–40%. Bordered cells: rubric-level spread ≥ 0.20. tory weakness in F1 Position Representation with co-occurring Epistemic Calibration failures. Full party-level breakdowns are in Appendices H.1 and I. 7.4 Label and Output Language Ab… view at source ↗
Figure 5
Figure 5. Figure 5: Task prompt template. Bracketed variables (language, context, party_name, statement_text) are filled per query; context is the numbered evidence passage(s), or a sentinel string when none are retrieved (Absent). System You are a helpful assistant. Use the provided context to answer the user’s question. If you cannot find the answer in the provided context, say so. Cite the source for each claim you make us… view at source ↗
Figure 6
Figure 6. Figure 6: Instantiated example (Baseline, Germany). Real query for observation bundestagswahl2025__de_cdu___csu__s001. D Dataset Scope and Standardization Details Party inclusion follows criteria from the Chapel Hill Expert Survey and the Comparative Manifesto Project: parliamentary representation and significant national relevance, operationalized as >5 seats or >3% vote share, applied identically in both countries… view at source ↗
Figure 7
Figure 7. Figure 7: Standardization prompt. Verbatim system-prompt constraints (illustrative examples within bul￾lets omitted for space) and user-message template; output constrained to JSON (third_person_rationale, stance_is_explicit). D.2 Validation Checks To enforce the structural transformations described in Section 4.2, generated outputs were passed through a programmatic validation script. A sample was only accepted if … view at source ↗
Figure 8
Figure 8. Figure 8: Standardization examples. Raw VAA rationale (Before) vs. standardized evidence passage (After); [PARTY] is the anonymization placeholder inserted in place of the party’s real name. D.4 Rhetorical Fidelity Check To measure how much I4 Sanitization inherits from standardization, which could pre-strip rhetoric before any model sees it, we compare each raw rationale with its standardized version. Median length… view at source ↗
Figure 9
Figure 9. Figure 9: Vagueness example (CDU/CSU, Mode 1). The original Agree stance (“unterstützt”, “befürwortet”) is replaced with priority language that never confirms support; same observation as Figures 6 and 8. 17 [PITH_FULL_IMAGE:figures/full_fig_p017_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Vague-IE generation prompt. The Mandatory Strategy bullet is populated per sample with one of the five modes in [PITH_FULL_IMAGE:figures/full_fig_p018_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Vague-IE audit prompt. A LLM judge scores each generated Vague passage at temperature 0.0; a determinable = true verdict is a quality failure [PITH_FULL_IMAGE:figures/full_fig_p018_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Judge prompt template. Filled per rubric call with the dimension name/definition, the condition’s property of context ( [PITH_FULL_IMAGE:figures/full_fig_p021_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Party × sub-question recurring patterns (Germany). Mean leave-one-out deviation per party × sub-question across active (IE, model) contexts. Blue = below cross-party mean, red = above. Integer = flagged contexts (|deviation| ≥ 0.15, direction matching colour). Rules separate rubric blocks (F, EC, I). 22 [PITH_FULL_IMAGE:figures/full_fig_p022_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: EC3 Context Transparency party deviations (Germany). LOO deviation per party × IE; three model panels. Annotation: signed pp. µ = cross-party mean pass rate. Contradictory μ=0.91 -3 +8 +7 +8 -21 Claude Contradictory μ=0.95 -1 -4 -8 +2 +6 +6 GPT Contradictory μ=0.91 -3 +7 +7 +7 +4 -21 Qwen −0.20 −0.15 −0.10 −0.05 0.00 0.05 0.10 0.15 0.20 Deviation [PITH_FULL_IMAGE:figures/full_fig_p023_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: F3 False Synthesis party deviations (Germany). As [PITH_FULL_IMAGE:figures/full_fig_p023_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Baseline party adherence. Adherence Index per model × party under Baseline; scale as [PITH_FULL_IMAGE:figures/full_fig_p023_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: EC4 Parametric Fallback party deviations (Germany). As [PITH_FULL_IMAGE:figures/full_fig_p024_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: F1 Position Representation party deviations (Germany). As [PITH_FULL_IMAGE:figures/full_fig_p024_18.png] view at source ↗
Figure 19
Figure 19. Figure 19: Absent (no evidence). Per-model adherence on each rubric. Scored on EC only (F/I cells grey). Scale as [PITH_FULL_IMAGE:figures/full_fig_p025_19.png] view at source ↗
Figure 20
Figure 20. Figure 20: Vague (unclear stance). Per-model adherence on each rubric. Scale as [PITH_FULL_IMAGE:figures/full_fig_p025_20.png] view at source ↗
Figure 21
Figure 21. Figure 21: Contradictory (conflicting evidence). Per-model adherence on each rubric. Scale as [PITH_FULL_IMAGE:figures/full_fig_p025_21.png] view at source ↗
Figure 22
Figure 22. Figure 22: Impartiality sub-question breakdown. Pass rates per I sub-question (I1–I5) (Sec. 7.1). (a) Germany by model; (b) NL − DE. 0.6 0.7 0.8 0.9 1.0 Adherence F EC I 97% 96% 98% 99% 97% 96% 96% 98% 97% Claude GPT Qwen Baseline F EC I 97% 98% 99% 99% 97% 95% 95% 98% 97% Noisy F EC I 98% 96% 98% 98% 99% 94% 94% 96% 95% Counterfactual [PITH_FULL_IMAGE:figures/full_fig_p026_22.png] view at source ↗
Figure 23
Figure 23. Figure 23: Interfering IEs by rubric (Germany). Per-model adherence on each rubric under Baseline, Noisy, and Counterfactual. Scale as [PITH_FULL_IMAGE:figures/full_fig_p026_23.png] view at source ↗
Figure 24
Figure 24. Figure 24: Epistemic Calibration sub-questions under Vague (Germany). Pass rates per EC sub-question (Sec. 7.2). (a) Germany by model; (b) NL − DE. 26 [PITH_FULL_IMAGE:figures/full_fig_p026_24.png] view at source ↗
Figure 25
Figure 25. Figure 25: Faithfulness sub-questions under Vague (Germany). Pass rates per F sub-question (Sec. 7.2). (a) Germany by model; (b) NL − DE. 0% 20% 40% 60% 80% 100% (DE) pass rate Transparency (E3) Certainty (E1) Hedging (E2) Fallback (E4) 67% 74% 79% 28% 31% 33% 44% 45% 46% 96% (a) -15% -10% -5% 0% +5% +10% +15% NL − DE (Δ pass rate) +2% +1% -3% +8% +5% +4% -1% +8% +8% +5% +2% (b) Claude GPT Qwen [PITH_FULL_IMAGE:fig… view at source ↗
Figure 26
Figure 26. Figure 26: Epistemic Calibration sub-questions under Contradictory (Germany). As [PITH_FULL_IMAGE:figures/full_fig_p027_26.png] view at source ↗
Figure 27
Figure 27. Figure 27: Faithfulness sub-questions under Contradictory (Germany). As [PITH_FULL_IMAGE:figures/full_fig_p027_27.png] view at source ↗
Figure 28
Figure 28. Figure 28: Overall model performance (Netherlands). As [PITH_FULL_IMAGE:figures/full_fig_p028_28.png] view at source ↗
Figure 29
Figure 29. Figure 29: Robustness across IEs (Netherlands). As [PITH_FULL_IMAGE:figures/full_fig_p028_29.png] view at source ↗
Figure 30
Figure 30. Figure 30: Robustness across IEs — NL − DE delta. Adherence delta per model × IE; layout as [PITH_FULL_IMAGE:figures/full_fig_p028_30.png] view at source ↗
Figure 31
Figure 31. Figure 31: Inconclusive IEs — NL − DE delta by rubric. Adherence delta per rubric under Absent, Vague, and Contradictory; layout as [PITH_FULL_IMAGE:figures/full_fig_p029_31.png] view at source ↗
Figure 32
Figure 32. Figure 32: Interfering IEs — NL − DE delta by rubric. As [PITH_FULL_IMAGE:figures/full_fig_p029_32.png] view at source ↗
Figure 33
Figure 33. Figure 33: F3 False Synthesis party deviations (Netherlands). As [PITH_FULL_IMAGE:figures/full_fig_p029_33.png] view at source ↗
Figure 34
Figure 34. Figure 34: Party × sub-question recurring patterns (Netherlands). As [PITH_FULL_IMAGE:figures/full_fig_p030_34.png] view at source ↗
Figure 35
Figure 35. Figure 35: F1 Position Representation party deviations (Netherlands). As [PITH_FULL_IMAGE:figures/full_fig_p030_35.png] view at source ↗
Figure 36
Figure 36. Figure 36: Anonymous-label delta, Baseline (DE + NL). Signed pass-rate difference (anon − full) per party × sub-question; rubric blocks as [PITH_FULL_IMAGE:figures/full_fig_p031_36.png] view at source ↗
Figure 37
Figure 37. Figure 37: English-output delta, Baseline (DE + NL). As [PITH_FULL_IMAGE:figures/full_fig_p031_37.png] view at source ↗
Figure 38
Figure 38. Figure 38: Anonymous-label delta, Vague (DE + NL). As [PITH_FULL_IMAGE:figures/full_fig_p031_38.png] view at source ↗
Figure 39
Figure 39. Figure 39: Sanitization (I4) across conditions (Germany). Pass rate per party across Full, Anon, and English conditions; 95% CI. y-axis starts at 0.60. full anon english Condition 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Pass rate Party VVD D66 CDA PVV GL-PvdA SP BBB JA21 [PITH_FULL_IMAGE:figures/full_fig_p032_39.png] view at source ↗
Figure 40
Figure 40. Figure 40: Sanitization (I4) across conditions (Netherlands). As [PITH_FULL_IMAGE:figures/full_fig_p032_40.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

22 extracted references · 4 linked inside Pith

  1. [1]

    Deutschland soll die Ukraine weiterhin militärisch unterstützen

    CDU / CSU unterstützt diese Maßnahme. Die Sicherung des Friedens in Europa wird von CDU / CSU als zentrales Ziel definiert, wobei die Verteidigung der Ukraine als essenziell für den Schutz weiterer Länder vor russischen Angriffen angesehen wird. Daher befürwortet die Gruppierung eine umfassende Unterstützung durch diplomatische, finanzielle und humanitäre...

  2. [2]

    Template Verification:A string check confirmed the exact [PARTY] placeholder was present at least once, ensuring the sample was successfully anonymized for downstream Information Environment assembly

  3. [3]

    we”, “our

    Pronoun Exclusion:A regex check verified the complete absence of first-person pronouns (e.g., “we”, “our”, “I” in the respective target languages), guaranteeing strict third-person, informational tone. Samples failing any programmatic check triggered an automated retry mechanism (max 2 retries). Samples that exhausted all retries were permanently excluded...

  4. [4]

    InFindings of the Associ- ation for Computational Linguistics: EMNLP 2025, pages 2239–2259, Suzhou, China

    Measuring Sycophancy of Language Models in Multi-turn Dialogues. InFindings of the Associ- ation for Computational Linguistics: EMNLP 2025, pages 2239–2259, Suzhou, China. Association for Computational Linguistics. Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vard- hamanan A, Saiful Haq, Ashutosh Sharma, Thomas T. ...

  5. [5]

    only if it guarantees no disproportionate burden on the middle class

    Ambiguous Conditionality Make any movement contingent on a vague, unmeasurable condition (e.g., “only if it guarantees no disproportionate burden on the middle class”). Table 7:Vagueness modes.One mode is randomly selected per sample and injected into the generation prompt (Figure 10). Before (Baseline):[PARTY] unterstützt diese Maßnahme. Die Sicherung de...

  6. [6]

    InAdvances in Neural Information Processing Systems, volume 37, pages 68772–68802

    LLM evaluators recognize and favor their own generations. InAdvances in Neural Information Processing Systems, volume 37, pages 68772–68802. Curran Associates, Inc. Yujin Potter, Shiyang Lai, Junsol Kim, James Evans, and Dawn Song. 2024. Hidden Persuaders: LLMs’ Political Leaning and Their Influence on V oters. In Proceedings of the 2024 Conference on Emp...

  7. [7]

    InThe Twelfth International Confer- ence on Learning Representations

    Towards Understanding Sycophancy in Lan- guage Models. InThe Twelfth International Confer- ence on Learning Representations. Felix M. Simon, Rasmus Kleis Nielsen, and Richard Fletcher. 2025. Generative AI and news report 2025: How people think about AI’s role in journalism and society. Technical report, Reuters Institute for the Study of Journalism. Chris...

  8. [10]

    Sentence Count Validation:The text was tokenized to ensure the length fell strictly within the target 4 to 6 sentence window

  9. [13]

    Strategic Prioritization State that [PARTY] views this as a top-tier priority that must align with their core values, but completely avoid stating what that alignment actually dictates

  10. [14]

    Procedural / Implementation Fo- cus Describe [PARTY]’s demands for independent audits, systemic reviews, or flawless execution frameworks, without confirming whether the baseline policy itself is supported or opposed

  11. [15]

    Competing Necessities Strongly emphasize that the severe risks of the issue must be weighed against its undeniable necessity, effectively canceling the stance out

  12. [16]

    social stability) rather than explicitly stating the party is internally divided

    Value Trade-offs Frame the issue as a complex balance between two competing [PARTY] values (e.g., economic growth vs. social stability) rather than explicitly stating the party is internally divided

  13. [18]

    Strategic Prioritization 89.8% 91.1%

  14. [19]

    Procedural / Implementation 88.4% 88.2%

  15. [20]

    Competing Necessities 86.2% 87.7%

  16. [21]

    Value Trade-offs 89.4% 92.4%

  17. [22]

    lost-in-the-middle

    Ambiguous Conditionality 86.3% 88.5% Table 8:Vague-mode adherence.Mean pass rate across all rubric sub-questions for observations generated under each mode (n= 485observations total, unevenly split across the 5 modes by random draw). VagueIE generation uses the same model as standardization, with Temperature >0 to allow variance across vagueness modes (Ta...

  18. [2021]

    the moon is made of marshmallows

    Countering Misinformation and Fake News Through Inoculation and Prebunking.European Re- view of Social Psychology, 32(2):348–384. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paran- jape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the Middle: How Language Models Use Long Contexts.Transactions of the Asso- ciation for Computational ...

  19. [2023]

    InProceedings of the 2023 Conference on Empirical Methods in Natural Lan- guage Processing, pages 6465–6488

    Enabling Large Language Models to Gener- ate Text with Citations. InProceedings of the 2023 Conference on Empirical Methods in Natural Lan- guage Processing, pages 6465–6488. Association for Computational Linguistics. Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jin- liu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024. Retrieval-Aug...

  20. [2024]

    Neutrally

    Benchmarking Large Language Models in Retrieval-Augmented Generation.Proceedings of the AAAI Conference on Artificial Intelligence, 38(16):17754–17762. Paul F Christiano, Jan Leike, Tom Brown, Miljan Mar- tic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, volum...

  21. [2025]

    InProceedings of the 63rd Annual Meet- ing of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6559–6607, Vienna, Austria

    Biased LLMs can Influence Political Decision- Making. InProceedings of the 63rd Annual Meet- ing of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6559–6607, Vienna, Austria. Association for Computational Linguistics. Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen

  22. [8565]

    What is [Party]’s position on rent control?

    Association for Computational Linguistics. Gal Yona, Roee Aharoni, and Mor Geva. 2024. Can Large Language Models Faithfully Express Their In- trinsic Uncertainty in Words? InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 7752–7764. Association for Computational Linguistics. Yue Zhang, Yafu Li, Leyang Cui, Den...