Pith. sign in

REVIEW 3 major objections 7 minor 69 references

Large Language Models Can Be a Viable Substitute for Expert Political Surveys When a Shock Disrupts Traditional Measurement Approaches

T0 review · 3 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read LLM scores replicate pre-shock expert ideology ratings and predict which federal agencies DOGE targeted.

desk verdict AIPS is a genuinely useful replication of expert ideology with LLMs, but the central 'viable substitute' claim leans on KIPS, a new construct with no external benchmark. read the letter →

arxiv 2506.06540 v1 pith:F2V7XU4F submitted 2025-06-06 cs.CY cs.AIcs.CL

classification cs.CYcs.AIcs.CL
keywords largelanguagemodelsexpertsurveysubstitutelatentpoliticalattributesagencyideologyDOGElayoffspairwisecomparisonsBradley-Terrymodelknowledgeinstitutions
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that when a shock like the 2025 DOGE federal layoffs contaminates expert judgment with hindsight, large language models can be used in place of expert political surveys to recover the pre-shock perceptions researchers need. The case study asks three open-weight LLMs to make pairwise comparisons between federal agencies—first on liberal-conservative ideology, then on whether agencies are perceived as knowledge institutions—and scales the answers into continuous scores with the Bradley-Terry model. The resulting ideology scores track the 2014 expert survey of agency ideology and replicate the finding that perceived ideology predicted which agencies were targeted by DOGE. A new knowledge-institution score predicts DOGE layoffs even after ideology, staffing, and budget are held constant, a construct no expert survey could now measure without post-event bias. The paper concludes with a two-part criterion for when LLM substitution is justified: a substantive theory that the construct matters, and a measurement gap that makes pre-shock expert data unobtainable.

What carries the argument

The load-bearing mechanism is pairwise comparison prompting combined with the Bradley-Terry model. The LLM is asked a forced-choice question—'WHICH AGENCY IS PERCEIVED TO BE MORE LIBERAL: [A] OR [B]?'—for all 7,503 agency pairs, repeated three times per model, and the same structure is reused with a knowledge-institution prompt. The Bradley-Terry model converts the collection of pairwise wins, losses, and ties (ties credited as half a win to each side) into a continuous score per agency, with quasi-standard errors giving confidence intervals. This two-step design supplies the temporal flexibility the paper relies on: open-weight models with pre-shock release dates are treated as frozen repositories of pre-event aggregate perceptions that a survey administered today cannot provide.

What would settle it

Concrete test: take a model with a verified pre-February-2025 checkpoint and re-run the pairwise ideology and knowledge-institution prompts with and without the words 'DOGE layoffs' or '2025 federal firings' added to the prompt; if the agency rankings or the predicted layoff probabilities change materially, the scores are not a stable pre-shock measure. A second decisive check is an audit or probing of the model's training data for any post-shock documents—if the model can reproduce details of the layoffs, such as naming agencies fired before February 2025, the cutoff assumption fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that LLMs, trained on large corpora of digital media, can serve as a viable substitute for expert political surveys whenever a shock makes post-event expert judgments unreliable. In the DOGE case, the paper derives Agency Ideology Pairwise Scores (AIPS) from pairwise LLM comparisons scaled with a Bradley-Terry model; these scores correlate with Richardson et al.'s pre-layoff perceived ideology measures (0.80–0.90 across three models) and reproduce Bonica's finding that ideology predicts which agencies experienced layoffs after controlling for budget and staff. It then introduces Knowledge Institution Pairwise Scores (KIPS), a new construct measuring whether an agency is perceived to produce, distribute, or legitimize knowledge, and shows that KIPS predicts DOGE layoffs even when ideology, staffing, and budget are held constant, in models using either LLM-based ideology or the earlier expert measure. The author interprets this as evidence that the approach can rapidly and cheaply test hypothesized factors behind a shock when traditional measurement is impossible, and proposes a two-part theory-gap criterion for when such substitution is appropriate.

Load-bearing premise

The argument stands on the assumption that an open-weight model's release date before February 2025 guarantees its training information ends before the DOGE layoffs, so the answers it gives are a clean pre-shock snapshot rather than hindsight.

Editorial extensions

If this is right

  • If the claim holds, researchers can reconstruct pre-shock latent attributes for any shock for which LLM training data exists, without waiting for or depending on expert panels.
  • The DOGE-specific finding—that perceived ideology predicts layoffs—is confirmed using three independent open-weight models, suggesting the earlier survey-based result is not an artifact of one dataset.
  • The new knowledge-institution measure offers a testable mechanism for the 'Cathedral' thesis: agencies perceived as knowledge producers were targeted net of ideology, staff, and budget.
  • The two-part theory-gap criterion gives practitioners a rule for when LLM substitution is legitimate and when it is not, separating retrospective measurement from forward-looking 'silicon sample' simulations.
  • Because the method is fast and cheap, it supports piloting and iterative hypothesis testing during unfolding shocks, before traditional surveys become feasible.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The strongest test of the paper's retrospective logic would be to run the same pairwise protocol on a shock with a fully observable pre-shock benchmark—e.g., public-health agency trust before a pandemic—and check whether LLM scores match archived pre-shock survey data; the paper does not itself do this.
  • The KIPS result suggests that anti-knowledge-institution rhetoric is not reducible to standard left-right ideology; a natural extension is to test whether KIPS also predicts cuts to universities, media grants, and scientific funding beyond federal agencies.
  • Because the approach depends on training-data coverage, its reliability will likely be highest for well-documented agencies and actors and lowest for obscure ones; researchers should report per-item uncertainty or coverage diagnostics as a routine companion to LLM-derived scores.
  • An untested boundary condition is prompt sensitivity—whether slightly different wording of the knowledge-institution prompt changes the agency ranking enough to alter the layoff-prediction result.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper argues that large language models (LLMs) can serve as a viable substitute for expert political surveys when a shock makes post-event expert judgments uninformative. As a case study, it estimates Agency Ideology Pairwise Scores (AIPS) for 123 federal agencies using pairwise comparison prompts with three open-weight LLMs (Llama 3.3 70B Instruct, Mistral Small 3, Phi-4), scales the comparisons with a Bradley-Terry model, shows that AIPS correlates highly with the 2014 Richardson et al. expert ideology measure, and replicates Bonica's finding that perceived agency ideology predicts DOGE layoffs controlling for budget and staff. It then introduces a new Knowledge Institution Pairwise Score (KIPS), estimated with Llama 3.3 70B Instruct, and reports that KIPS predicts DOGE layoffs even when controlling for ideology. The paper concludes by proposing a two-part 'theory-gap criterion' for when LLM-based measurement can substitute for expert surveys.

Significance. The AIPS portion of the paper is a credible and largely reproducible measurement exercise: it uses openly available models, a transparent pairwise-comparison design, and reports correlations of 0.80-0.90 with an external 2014 expert survey (Table 1), while replicating the substantive DOGE finding from Bonica (Table 2). This alone is a useful contribution to the emerging literature on LLM-based latent measurement. The paper is also commendably explicit about limitations, notably in Section 5.2 where it concedes there is no benchmark for KIPS. However, the broader 'viable substitute' claim depends substantially on KIPS, a new construct with no external validation and whose only evidence is an outcome regression that may be contaminated by the prompt's reliance on the very 'Cathedral' discourse used to justify the layoffs. If the KIPS validation gap is addressed, the paper would make a stronger methodological contribution; as it stands, the central claim exceeds what the evidence establishes.

major comments (3)
  1. [Section 5.1] The statement that 'using open weight models with verified checkpoint dates preceding the shock guarantees that the model's knowledge cutoff is pre-event' is not supported by the evidence in the paper. A checkpoint or release date is not the same as a training-data cutoff date, and for instruction-tuned models the knowledge cutoff must be verified directly rather than inferred from the release date. Moreover, even with a verified cutoff, the unanchored prompt in Section 3.1 does not ask the model to adopt a pre-February 2025 perspective, so the output is an amalgam of training corpora rather than a temporally indexed pre-shock measurement. The paper should provide a direct knowledge-cutoff check (for example, control prompts asking about post-shock facts, or comparison of responses from models with known cutoffs on either side of the shock) before claiming that LLM estimates are ex ante measures. This is load-bearing because the paper's core 'temporal flexibility' advantage rests on this guarantee.
  2. [Section 4 and Section 5.2] KIPS is a new construct with no external benchmark, and the paper concedes in Section 5.2 that 'we do not have a benchmark against which to compare KIPS.' The only validity evidence is the outcome regression in Table 3. This is insufficient because the KIPS prompt itself contains the words 'ACADEMIC AND EDUCATIONAL INSTITUTIONS, THE MEDIA, AND CIVIL SOCIETY ORGANIZATIONS' and defines knowledge institutions in terms of producing, distributing, and legitimizing knowledge—language that directly mirrors the 'Cathedral' discourse described in Section 4 as motivating DOGE. The observed association between KIPS and layoffs may therefore reflect the model's encoding of the anti-knowledge-institution narrative that drives the outcome, rather than an independent pre-shock expert perception. The 'viable substitute' claim for new constructs requires additional validation, such as using multiple model families, prompt variations that do not embed the outcome-related theory, construct validation against agency behaviors (e.g., collaborations with universities or civil society), or an expert survey administered post-shock as an imperfect benchmark, which the paper itself acknowledges as possible.
  3. [Section 4, Table 3] KIPS is estimated using a single LLM (Llama 3.3 70B Instruct) and a single prompt, so there is no sensitivity analysis for the central new measure. In contrast, the AIPS analysis uses three models and reports cross-model correlations (Table 1); the KIPS analysis should be held to the same standard, or the paper should explicitly state that KIPS is exploratory. Without model- or prompt-robustness evidence, the regression results in Table 3 may be specific to one model's priors, and the general claim that 'this same approach' can measure knowledge-institution perceptions is not established.
minor comments (7)
  1. [Section 3.1, Section 4, Appendix A.3] The prompts contain typographical errors and missing spaces (e.g., '[AGENCY1]OR[AGENCY2]?'), which makes them difficult to reproduce exactly; please provide clean, copy-pasteable prompt text.
  2. [Section 1] The phrase 'perception as aknowledge institution' contains a typo; 'aknowledge' should be 'a knowledge'.
  3. [Table 3] The regression table would benefit from reporting the number of agencies and a goodness-of-fit statistic (e.g., McFadden R^2), and from discussing the scale of KIPS so the coefficients can be interpreted substantively.
  4. [Figure 2] The legend entry 'a aNo Layoffs' contains a stray 'a', and the labels in Figure 1 overlap substantially; these should be cleaned up for readability.
  5. [General] The paper does not include a data or code availability statement. Given the paper's emphasis on reproducibility and open-weight models, releasing the prompts, raw outputs, and scaling code would strengthen the replication claim.
  6. [Section 5.2] The sentence about uncertainty contains a missing space ('uncertaintyof'); minor grammar and spacing issues like this appear in several places and should be corrected throughout.
  7. [Section 5.3] The two-part theory-gap criterion is sensible but would benefit from explicit operationalization: what counts as 'strong theory,' and how should a researcher assess whether pre-shock data collection was 'theoretically possible' in practice?

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: AIPS is externally benchmarked, and KIPS, while lacking an independent benchmark, is not fitted to the layoff outcome.

full rationale

The main AIPS chain is self-contained and externally anchored: pairwise LLM comparisons are converted via the Bradley-Terry model into ideology scores that are then compared with Richardson et al. (2018) and used to replicate Bonica (2025). Neither step fits the DOGE layoff outcome into the measurement itself, and the 2014 expert survey provides an independent external benchmark. The KIPS chain is the closest thing to a circularity concern because the paper concedes in Section 5.2 that 'we do not have a benchmark against which to compare KIPS' and its principal support is the Table 3 regression on DOGE layoffs. However, KIPS is not statistically constructed from that outcome: the scores are assembled from pairwise LLM responses before the layoff regression, so the regression is a criterion-validity test, not a fitted input renamed as a prediction. The validity gap is real and explicitly acknowledged, but it is a construct-validity limitation rather than a reduction of the prediction to the input by construction. The Section 5.1 claim that pre-shock checkpoint dates guarantee a pre-event knowledge cutoff is an empirical premise about training data, not a circular step. The paper cites Wu et al. (2023) for the pairwise-comparison protocol, which is a self-citation, but that protocol is also supported by standard references such as Bradley and Terry (1952) and Turner and Firth (2012), and the non-reproducible Napolio preprint is explicitly not relied upon for the new results. No uniqueness theorem, ansatz, or renaming of a known result is used to force the conclusion. Therefore, no load-bearing circular step is present; the low score reflects only the acknowledged absence of an external KIPS benchmark and the minor self-citation.

Assumptions & free parameters 2 free parameters · 6 assumptions · 1 invented entities

The central measurement approach rests on LLM-derived scores (AIPS and KIPS) fitted from pairwise comparisons, and on the assumptions that LLM training data is a faithful, pre-shock sample of aggregate expert perceptions and that open-weight release dates guarantee pre-shock knowledge cutoffs. The knowledge-institution construct is new and has no external validation.

free parameters (2)
  • AIPS agency ideology scores (123 agencies) = Estimated by Bradley-Terry from LLM pairwise comparisons; values in Appendix figures
    Each agency gets a scalar ideology score fitted to 7,503 pairwise comparisons, repeated three times per LLM. These scores are the basis for the ideology replication and layoff prediction.
  • KIPS knowledge-institution scores (123 agencies) = Estimated by Bradley-Terry from Llama 3.3 70B comparisons; values in Appendix Figure 6
    Same construction as AIPS but for the knowledge-institution construct. Only one LLM (Llama 3.3 70B) is used, and there is no external benchmark to validate the scale.
assumptions (6)
  • domain assumption LLM training corpora contain pre-shock information about federal agencies that reflects aggregate expert perceptions, and pairwise responses recover those perceptions.
    The entire method assumes the LLM is a valid substitute for experts because it was trained on digital media before the shock. Invoked in Sections 2.3 and 5.1.
  • domain assumption The release date of an open-weight model is a reliable proxy for a pre-shock knowledge cutoff.
    Section 5.1 claims that 'verified checkpoint dates preceding the shock guarantees that the model's knowledge cutoff is pre-event.' Release dates do not always equal training-data cutoff dates.
  • domain assumption The 2014 Richardson et al. survey is a valid benchmark for the true pre-layoff perceived ideology of federal agencies.
    Used as ground truth for correlation in Section 3.3.1 and as a comparison in Tables 2 and 3.
  • domain assumption The substantive theory that knowledge institutions are targets of the administration, drawn from Yarvin and Vance, is credible enough to justify the KIPS construct.
    Section 4 relies on this theory to give KIPS meaning and to justify its inclusion in the layoff regressions.
  • standard math The Bradley-Terry model with 0.5 wins for ties produces interval-scale latent scores from pairwise comparisons.
    Standard model, but the tie-handling rule is a modeling choice that affects the resulting scores.
  • domain assumption Expert surveys cannot be collected post-shock without outcome bias, so ex post validation of KIPS is impossible now.
    This justifies the measurement-gap part of the proposed criterion and explains the absence of a benchmark for KIPS.
invented entities (1)
  • Knowledge Institution Pairwise Scores (KIPS) construct
    purpose: Measures the perception of each agency as a producer, distributor, or legitimizer of knowledge, to test the 'Cathedral' theory of DOGE targeting.
    No external benchmark exists for KIPS. The paper explicitly states in Section 5.2 that there is limited external validation for new constructs and that no benchmark is available. Without an independent measure, the construct's validity rests on its correlation with the outcome it is meant to predict.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large Language Models Can Be a Viable Substitute for Expert Political Surveys When a Shock Disrupts Traditional Measurement Approaches." pith.science (2026). https://pith.science/paper/F2V7XU4F

@misc{pith2026250606540,
  author       = {Pith},
  title        = {Pith review of: Large Language Models Can Be a Viable Substitute for Expert Political Surveys When a Shock Disrupts Traditional Measurement Approaches},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F2V7XU4F}},
  note         = {Machine review of arXiv:2506.06540}
}
read the original abstract

After a disruptive event or shock, such as the Department of Government Efficiency (DOGE) federal layoffs of 2025, expert judgments are colored by knowledge of the outcome. This can make it difficult or impossible to reconstruct the pre-event perceptions needed to study the factors associated with the event. This position paper argues that large language models (LLMs), trained on vast amounts of digital media data, can be a viable substitute for expert political surveys when a shock disrupts traditional measurement. We analyze the DOGE layoffs as a specific case study for this position. We use pairwise comparison prompts with LLMs and derive ideology scores for federal executive agencies. These scores replicate pre-layoff expert measures and predict which agencies were targeted by DOGE. We also use this same approach and find that the perceptions of certain federal agencies as knowledge institutions predict which agencies were targeted by DOGE, even when controlling for ideology. This case study demonstrates that using LLMs allows us to rapidly and easily test the associated factors hypothesized behind the shock. More broadly, our case study of this recent event exemplifies how LLMs offer insights into the correlational factors of the shock when traditional measurement techniques fail. We conclude by proposing a two-part criterion for when researchers can turn to LLMs as a substitute for expert political surveys.

Figures

Figures reproduced from arXiv: 2506.06540 by the authors.

Figure 1
Figure 1. From a perspective of face validity, the results made sense: the most liberal federal agencies [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 1
Figure 1. The perceived agency ideology scores from Richardson et al. [2018] versus our Agency [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Total staff by agency (on a log scale) versus our Knowledge Institution Pairwise Scores [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figures from the paper (4 more)
Figure 3
Figure 3. Figure 3: AIPS for Llama 3.3 70B Instruct. 95% confidence intervals are based on quasi-standard [PITH_FULL_IMAGE:figures/full_fig_p016_3.png]
Figure 4
Figure 4. Figure 4: AIPS for Mistral Small 3. 95% confidence intervals are based on quasi-standard errors. [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: AIPS for Phi-4. 95% confidence intervals are based on quasi-standard errors. [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]
Figure 6
Figure 6. Figure 6: KIPS for Llama 3.3 70B Instruct. 95% confidence intervals are based on quasi-standard [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

69 extracted references · 37 canonical work pages

  1. [1]

    Hewett, Mojan Javaheripi, Piero Kauffmann, James R

    Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J. Hewett, Mojan Javaheripi, Piero Kauffmann, James R. Lee, Yin Tat Lee, Yuanzhi Li, Weishung Liu, Caio C. T. Mendes, Anh Nguyen, Eric Price, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Xin Wang, Rachel Ward, Yue Wu, Dingli Yu,...

  2. [2]

    Measurement validity: A shared standard for qualitative and quantitative research

    Robert Adcock and David Collier. Measurement validity: A shared standard for qualitative and quantitative research. American Political Science Review, 95 0 (3): 0 529–546, 2001. doi:10.1017/S0003055401003100

  3. [3]

    Snyder, and Charles Stewart

    Stephen Ansolabehere, James M. Snyder, and Charles Stewart. Candidate positioning in U.S. house elections. American Journal of Political Science, 45 0 (1): 0 136--159, 2001. ISSN 00925853, 15405907. URL http://www.jstor.org/stable/2669364

  4. [4]

    Argyle, Ethan C

    Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. Out of one, many: Using language models to simulate human samples. Political Analysis, 31 0 (3): 0 337–351, 2023. doi:10.1017/pan.2023.2

  5. [5]

    Uncertainty in natural language generation: From theory to applications, 2023

    Joris Baan, Nico Daheim, Evgenia Ilia, Dennis Ulmer, Haau-Sing Li, Raquel Fernández, Barbara Plank, Rico Sennrich, Chrysoula Zerva, and Wilker Aziz. Uncertainty in natural language generation: From theory to applications, 2023. URL https://arxiv.org/abs/2307.15703

  6. [6]

    Birds of the same feather tweet together: Bayesian ideal point estimation using twitter data

    Pablo Barberá. Birds of the same feather tweet together: Bayesian ideal point estimation using twitter data. Political Analysis, 23 0 (1): 0 76–91, 2015. doi:10.1093/pan/mpu011

  7. [7]

    Larry M. Bartels. Economic Inequality and Political Representation . In The Unsustainable American State . Oxford University Press, 10 2009. ISBN 9780195392135. doi:10.1093/acprof:oso/9780195392135.003.0007. URL https://doi.org/10.1093/acprof:oso/9780195392135.003.0007

  8. [8]

    Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M

    James Bisbee, Joshua D. Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M. Larson. Synthetic replacements for human survey data? the perils of large language models. Political Analysis, 32 0 (4): 0 401–416, 2024. doi:10.1017/pan.2024.5

Show all 69 references
  1. [9]

    Mapping the ideological marketplace

    Adam Bonica. Mapping the ideological marketplace. American Journal of Political Science, 58 0 (2): 0 367--386, 2014. ISSN 00925853, 15405907. URL http://www.jstor.org/stable/24363491

  2. [10]

    The DOGE purge: E mpirical E vidence of P olitically M otivated F irings, February 2025

    Adam Bonica. The DOGE purge: E mpirical E vidence of P olitically M otivated F irings, February 2025. URL https://data4democracy.substack.com/p/the-doge-purge-empirical-evidence. On Data and Democracy, Substack blog

  3. [11]

    Elmendorf, and Scott A

    Cheryl Boudreau, Christopher S. Elmendorf, and Scott A. MacKenzie. Racial or spatial voting? the effects of candidate ethnicity and ethnic group endorsements in local elections. American Journal of Political Science, 63 0 (1): 0 5--20, 2019. doi:https://doi.org/10.1111/ajps.12...

  4. [12]

    Ralph Allan Bradley and Milton E. Terry. Rank analysis of incomplete block designs: The method of paired comparisons. Biometrika, 39 0 (3-4): 0 324--345, 12 1952. ISSN 0006-3444. doi:10.1093/biomet/39.3-4.324. URL https://doi.org/10.1093/biomet/39.3-4.324

  5. [13]

    Montgomery

    David Carlson and Jacob M. Montgomery. A pairwise comparison framework for fast, flexible, and reliable human coding of political texts. American Political Science Review, 111 0 (4): 0 835–843, 2017. doi:10.1017/S0003055417000302

  6. [14]

    Carpenter

    Daniel P. Carpenter. The Forging of Bureaucratic Autonomy: Reputations, Networks, and Policy Innovation in Executive Agencies, 1862-1928, volume 173. Princeton University Press, 2001. ISBN 9780691070100

  7. [15]

    Lewis, James Lo, Keith T

    Royce Carroll, Jeffrey B. Lewis, James Lo, Keith T. Poole, and Howard Rosenthal. Measuring bias and uncertainty in DW-NOMINATE ideal point estimates via the parametric bootstrap. Political Analysis, 17 0 (3): 0 261--275, 2009. ISSN 10471987, 14764989. URL http://www.jstor.org/...

  8. [16]

    Federal employee unionization and presidential control of the bureaucracy: Estimating and explaining ideological change in executive agencies

    Jowei Chen and Tim Johnson. Federal employee unionization and presidential control of the bureaucracy: Estimating and explaining ideological change in executive agencies. Journal of Theoretical Politics, 27 0 (1): 0 151--174, 2015. doi:10.1177/0951629813518126. URL https://doi...

  9. [17]

    Clinton and David E

    Joshua D. Clinton and David E. Lewis. Expert opinion, agency characteristics, and agency preferences. Political Analysis, 16 0 (1): 0 3–20, 2008. doi:10.1093/pan/mpm009

  10. [18]

    Clinton, Anthony Bertelli, Christian R

    Joshua D. Clinton, Anthony Bertelli, Christian R. Grose, David E. Lewis, and David C. Nixon. Separated powers in the united states: The ideology of agencies, presidents, and congress. American Journal of Political Science, 56 0 (2): 0 341--354, 2012. doi:https://doi.org/10.111...

  11. [19]

    Can ai language models replace human participants? Trends in Cognitive Sciences, 27 0 (7): 0 597--600, 2023

    Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. Can ai language models replace human participants? Trends in Cognitive Sciences, 27 0 (7): 0 597--600, 2023. ISSN 1364-6613. doi:https://doi.org/10.1016/j.tics.2023.04.008

  12. [20]

    Tucker, and Jonathan Nagler

    Gregory Eady, Richard Bonneau, Joshua A. Tucker, and Jonathan Nagler. News sharing on social media: Mapping the ideology of news media, politicians, and the mass public. Political Analysis, 33 0 (2): 0 73–90, 2025. doi:10.1017/pan.2024.19

  13. [21]

    Detecting hallucinations in large language models using semantic entropy

    Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. Detecting hallucinations in large language models using semantic entropy. Nature, 630 0 (8017): 0 625--630, 2024. ISSN 1476-4687. doi:10.1038/s41586-024-07421-0. URL https://doi.org/10.1038/s41586-024-07421-0

  14. [22]

    De Menezes

    David Firth and Renée X. De Menezes. Quasi‐variances . Biometrika, 91 0 (1): 0 65--80, 03 2004. ISSN 0006-3444. doi:10.1093/biomet/91.1.65. URL https://doi.org/10.1093/biomet/91.1.65

  15. [23]

    Shapiro, and Matt Taddy

    Matthew Gentzkow, Jesse M. Shapiro, and Matt Taddy. Measuring group differences in high-dimensional choices: Method and application to congressional speech. Econometrica, 87 0 (4): 0 1307--1340, 2019. doi:https://doi.org/10.3982/ECTA16566

  16. [24]

    Gilmour and David E

    John B. Gilmour and David E. Lewis. Political appointees and the competence of federal program management. American Politics Research, 34 0 (1): 0 22--50, 2006. doi:10.1177/1532673X04271905

  17. [25]

    `What Elon Musk Said Is a Bold-Faced Lie'

    Michael Hirsh. `What Elon Musk Said Is a Bold-Faced Lie' . Politico Magazine, February 2025. URL https://www.politico.com/news/magazine/2025/02/04/elon-musk-usaid-00202409

  18. [26]

    Exhaustive or exhausting? evidence on respondent fatigue in long surveys

    Dahyeon Jeong, Shilpa Aggarwal, Jonathan Robinson, Naresh Kumar, Alan Spearot, and David Sungho Park. Exhaustive or exhausting? evidence on respondent fatigue in long surveys. Journal of Development Economics, 161: 0 102992, 2023. ISSN 0304-3878. doi:https://doi.org/10.1016/j....

  19. [27]

    From NeoReactionary Theory to the Alt-Right, pages 101--120

    Andrew Jones. From NeoReactionary Theory to the Alt-Right, pages 101--120. Springer International Publishing, Cham, 2019. ISBN 978-3-030-18753-8. doi:10.1007/978-3-030-18753-8_6

  20. [28]

    Positioning political texts with large language models by asking and averaging

    Gaël Le Mens and Aina Gallego. Positioning political texts with large language models by asking and averaging. Political Analysis, page 1–9, 2025. doi:10.1017/pan.2024.29

  21. [29]

    Teaching models to express their uncertainty in words

    Stephanie Lin, Jacob Hilton, and Owain Evans. Teaching models to express their uncertainty in words. Transactions on Machine Learning Research, 2022. ISSN 2835-8856. URL https://openreview.net/forum?id=8s8K2UZGTZ

  22. [30]

    Generating with confidence: Uncertainty quantification for black-box large language models

    Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. Generating with confidence: Uncertainty quantification for black-box large language models. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. URL https://openreview.net/forum?id=DWkJCSxKU5

  23. [31]

    Uncertainty estimation and quantification for llms: A simple supervised approach, 2024

    Linyu Liu, Yu Pan, Xiaocheng Li, and Guanting Chen. Uncertainty estimation and quantification for llms: A simple supervised approach, 2024. URL https://arxiv.org/abs/2404.15993

  24. [32]

    The L lama 3 herd of models, 2024

    Llama Team, AI @ Meta . The L lama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783

  25. [33]

    Testing the power of arguments in referendums: A Bradley–Terry approach

    Peter John Loewen, Daniel Rubenson, and Arthur Spirling. Testing the power of arguments in referendums: A Bradley–Terry approach. Electoral Studies, 31 0 (1): 0 212--221, 2012. ISSN 0261-3794. doi:https://doi.org/10.1016/j.electstud.2011.07.003. URL https://www.sciencedirect.c...

  26. [34]

    Robert C. Luskin. Measuring political sophistication. American Journal of Political Science, 31 0 (4): 0 856--899, 1987. ISSN 00925853, 15405907. URL http://www.jstor.org/stable/2111227

  27. [35]

    Brown, Joshua A

    Maggie Macdonald, Megan A. Brown, Joshua A. Tucker, and Jonathan Nagler. To moderate, or not to moderate: Strategic domain sharing by congressional campaigns. Electoral Studies, 95: 0 102907, 2025. ISSN 0261-3794. doi:https://doi.org/10.1016/j.electstud.2025.102907. URL https:...

  28. [36]

    Curtis yarvin says democracy is done

    David Marchese. Curtis yarvin says democracy is done. powerful conservatives are listening., 2025. URL https://www.nytimes.com/2025/01/18/magazine/curtis-yarvin-interview.html

  29. [37]

    Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau

    Sabrina J. Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. Reducing conversational agents' overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics, 10: 0 857--872, 2022. doi:10.1162/tacl_a_00494. URL https://aclantholo...

  30. [38]

    Mistral small 3, January 2025

    Mistral AI Team . Mistral small 3, January 2025. URL https://mistral.ai/news/mistral-small-3

  31. [39]

    Terry M. Moe. Control and feedback in economic regulation: The case of the nlrb. American Political Science Review, 79 0 (4): 0 1094–1116, 1985. doi:10.2307/1956250

  32. [40]

    Virtual personas for language models via an anthology of backstories

    Suhong Moon, Marwa Abdulhai, Minwoo Kang, Joseph Suh, Widyadewi Soedarmadji, Eran Kohen Behar, and David Chan. Virtual personas for language models via an anthology of backstories. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conferenc...

  33. [41]

    Measuring executive agency ideology using large language models, n.d

    Nicholas G Napolio. Measuring executive agency ideology using large language models, n.d. URL https://static1.squarespace.com/static/59b80bc4d2b857d5b7da854d/t/665d6980e716480726e732b4/1717397889454/Napolio+Agency+Ideology.pdf

  34. [42]

    Crowdsourcing subjective annotations using pairwise comparisons reduces bias and error compared to the majority-vote method

    Hasti Narimanzadeh, Arash Badie-Modiri, Iuliia G Smirnova, and Ted Hsuan Yun Chen. Crowdsourcing subjective annotations using pairwise comparisons reduces bias and error compared to the majority-vote method. Proceedings of the ACM on Human-Computer Interaction, 7 0 (CSCW2): 0 ...

  35. [43]

    David C. Nixon. Separation of powers and appointee ideology. The Journal of Law, Economics, and Organization, 20 0 (2): 0 438--457, 10 2004. ISSN 8756-6222. doi:10.1093/jleo/ewh041

  36. [44]

    Measurement in the age of llms: An application to ideological scaling, 2024

    Sean O'Hagan and Aaron Schein. Measurement in the age of llms: An application to ideological scaling, 2024. URL https://arxiv.org/abs/2312.09203

  37. [45]

    When do politicians grandstand? measuring message politics in committee hearings

    Ju Yeon Park. When do politicians grandstand? measuring message politics in committee hearings. The Journal of Politics, 83 0 (1): 0 214--228, 2021. doi:10.1086/709147

  38. [46]

    Federal workers express shock, anger over mass firings: ``you are not fit for continued employment''

    Aimee Picchi. Federal workers express shock, anger over mass firings: ``you are not fit for continued employment''. CBS News, 2025. URL https://www.cbsnews.com/news/trump-federal-employees-probationary-firings-layoffs-workers-impact/

  39. [47]

    Poole and Howard Rosenthal

    Keith T. Poole and Howard Rosenthal. A spatial model for legislative roll call analysis. American Journal of Political Science, 29 0 (2): 0 357--384, 1985. ISSN 00925853, 15405907. URL http://www.jstor.org/stable/2111172

  40. [48]

    Poole and Howard Rosenthal

    Keith T. Poole and Howard Rosenthal. Ideology and Congress: A Political Economic History of Roll Call Voting. Yale University Press, New Haven, CT, 1997

  41. [49]

    Porter, Michael E

    Stephen R. Porter, Michael E. Whitcomb, and William H. Weitzer. Multiple surveys of students and survey fatigue. New Directions for Institutional Research, 2004 0 (121): 0 63--73, 2004. doi:https://doi.org/10.1002/ir.101

  42. [50]

    Richardson, Joshua D

    Mark D. Richardson, Joshua D. Clinton, and David E. Lewis. Elite perceptions of agency ideology and workforce skill. The Journal of Politics, 80 0 (1): 0 303--308, 2018. doi:10.1086/694846. URL https://doi.org/10.1086/694846

  43. [51]

    Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models

    Paul R \"o ttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schuetze, and Dirk Hovy. Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models. In Lun-Wei Ku, Andre Martins, and Vive...

  44. [52]

    Minority opposition and asymmetric parties? S enators’ partisan rhetoric on twitter

    Annelise Russell. Minority opposition and asymmetric parties? S enators’ partisan rhetoric on twitter. Political Research Quarterly, 74 0 (3): 0 615--627, 2021. doi:10.1177/1065912920921239. URL https://doi.org/10.1177/1065912920921239

  45. [53]

    Wu, Kristina Miler, Alexander Miserlis Hoyle, and Philip Resnik

    Rupak Sarkar, Patrick Y. Wu, Kristina Miler, Alexander Miserlis Hoyle, and Philip Resnik. P air S cale: Analyzing attitude change with pairwise comparisons. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Findings of the Association for Computational Linguistics: NAACL 20...

  46. [54]

    Adler, Lea Rau, and Bernd Schmitt

    Marko Sarstedt, Susanne J. Adler, Lea Rau, and Bernd Schmitt. Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing, 41 0 (6): 0 1254--1270, 2024. doi:https://doi.org/10.100...

  47. [55]

    Schmidt and Michael C

    Michael S. Schmidt and Michael C. Bender. Trump administration halts harvard's ability to enroll international students, May 2025. URL https://www.nytimes.com/2025/05/22/us/politics/trump-harvard-international-students.html

  48. [56]

    Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting

    Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Representations, 2024. URL...

  49. [57]

    The A merican system of shared powers: the P resident, C ongress, and the NLRB

    SK Snyder and BR Weingast. The A merican system of shared powers: the P resident, C ongress, and the NLRB . The Journal of Law, Economics, and Organization, 16 0 (2): 0 269--305, 10 2000. ISSN 8756-6222. doi:10.1093/jleo/16.2.269

  50. [58]

    Language model fine-tuning on scaled survey data for predicting distributions of public opinions, 2025

    Joseph Suh, Erfan Jahanparast, Suhong Moon, Minwoo Kang, and Serina Chang. Language model fine-tuning on scaled survey data for predicting distributions of public opinions, 2025. URL https://arxiv.org/abs/2502.16761

  51. [59]

    Do llms exhibit human-like response biases? a case study in survey design

    Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkwar, and Graham Neubig. Do llms exhibit human-like response biases? a case study in survey design. Transactions of the Association for Computational Linguistics, 12: 0 1011--1026, 09 2024. ISSN 2307-387X. doi:10.1162/ta...

  52. [60]

    Bradley-terry models in r: The bradleyterry2 package

    Heather Turner and David Firth. Bradley-terry models in r: The bradleyterry2 package. Journal of Statistical Software, 48 0 (9): 0 1–21, 2012. doi:10.18637/jss.v048.i09. URL https://www.jstatsoft.org/index.php/jss/article/view/v048i09

  53. [61]

    He's anti-democracy and pro-trump: the obscure 'dark enlightenment' blogger influencing the next us administration, 2024

    Jason Wilson. He's anti-democracy and pro-trump: the obscure 'dark enlightenment' blogger influencing the next us administration, 2024. URL https://www.theguardian.com/us-news/2024/dec/21/curtis-yarvin-trump

  54. [62]

    Cultural M arxism and the C athedral: Two alt-right perspectives on critical theory

    Alan Woods. Cultural M arxism and the C athedral: Two alt-right perspectives on critical theory. In Caroline Battista and Marcelle Sande, editors, Critical Theory and the Humanities in the Age of the Alt-Right, pages 45--64. Palgrave Macmillan, Cham, 2019. doi:10.1007/978-3-03...

  55. [63]

    Patrick Y. Wu. Using semantically unrelated and opposite terms for in-context learning: A case study in identifying political aversion in tweets. In Proceedings of the 17th ACM Web Science Conference 2025, Websci '25, page 528–533, New York, NY, USA, 2025. Association for Comp...

  56. [64]

    Wu, Jonathan Nagler, Joshua A

    Patrick Y. Wu, Jonathan Nagler, Joshua A. Tucker, and Solomon Messing. Large language models can be used to estimate the latent positions of politicians, 2023. URL https://arxiv.org/abs/2303.12057

  57. [65]

    Wu, Jonathan Nagler, Joshua A

    Patrick Y. Wu, Jonathan Nagler, Joshua A. Tucker, and Solomon Messing. Concept-guided chain-of-thought prompting for pairwise comparison scoring of texts with large language models. In 2024 IEEE International Conference on Big Data (BigData), pages 7232--7241, 2024. doi:10.110...

  58. [66]

    Can LLM s express their uncertainty? an empirical evaluation of confidence elicitation in LLM s

    Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. Can LLM s express their uncertainty? an empirical evaluation of confidence elicitation in LLM s. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview....

  59. [67]

    Navigating the grey area: How expressions of uncertainty and overconfidence affect language models

    Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto. Navigating the grey area: How expressions of uncertainty and overconfidence affect language models. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural La...

  60. [68]

    P ro SA : Assessing and understanding the prompt sensitivity of LLM s

    Jingming Zhuo, Songyang Zhang, Xinyu Fang, Haodong Duan, Dahua Lin, and Kai Chen. P ro SA : Assessing and understanding the prompt sensitivity of LLM s. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EM...

  61. [69]

    Can large language models transform computational social science? Computational Linguistics, 50 0 (1): 0 237--291, 03 2024

    Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. Can large language models transform computational social science? Computational Linguistics, 50 0 (1): 0 237--291, 03 2024. ISSN 0891-2017. doi:10.1162/coli_a_00502. URL https://doi.org/10.1162/co...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.