REVIEW 3 major objections 7 minor 69 references
Large Language Models Can Be a Viable Substitute for Expert Political Surveys When a Shock Disrupts Traditional Measurement Approaches
T0 review · 3 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read LLM scores replicate pre-shock expert ideology ratings and predict which federal agencies DOGE targeted.
desk verdict AIPS is a genuinely useful replication of expert ideology with LLMs, but the central 'viable substitute' claim leans on KIPS, a new construct with no external benchmark. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is pairwise comparison prompting combined with the Bradley-Terry model. The LLM is asked a forced-choice question—'WHICH AGENCY IS PERCEIVED TO BE MORE LIBERAL: [A] OR [B]?'—for all 7,503 agency pairs, repeated three times per model, and the same structure is reused with a knowledge-institution prompt. The Bradley-Terry model converts the collection of pairwise wins, losses, and ties (ties credited as half a win to each side) into a continuous score per agency, with quasi-standard errors giving confidence intervals. This two-step design supplies the temporal flexibility the paper relies on: open-weight models with pre-shock release dates are treated as frozen repositories of pre-event aggregate perceptions that a survey administered today cannot provide.
What would settle it
Concrete test: take a model with a verified pre-February-2025 checkpoint and re-run the pairwise ideology and knowledge-institution prompts with and without the words 'DOGE layoffs' or '2025 federal firings' added to the prompt; if the agency rankings or the predicted layoff probabilities change materially, the scores are not a stable pre-shock measure. A second decisive check is an audit or probing of the model's training data for any post-shock documents—if the model can reproduce details of the layoffs, such as naming agencies fired before February 2025, the cutoff assumption fails.
Extended reading notes
Core claim
The paper's central claim is that LLMs, trained on large corpora of digital media, can serve as a viable substitute for expert political surveys whenever a shock makes post-event expert judgments unreliable. In the DOGE case, the paper derives Agency Ideology Pairwise Scores (AIPS) from pairwise LLM comparisons scaled with a Bradley-Terry model; these scores correlate with Richardson et al.'s pre-layoff perceived ideology measures (0.80–0.90 across three models) and reproduce Bonica's finding that ideology predicts which agencies experienced layoffs after controlling for budget and staff. It then introduces Knowledge Institution Pairwise Scores (KIPS), a new construct measuring whether an agency is perceived to produce, distribute, or legitimize knowledge, and shows that KIPS predicts DOGE layoffs even when ideology, staffing, and budget are held constant, in models using either LLM-based ideology or the earlier expert measure. The author interprets this as evidence that the approach can rapidly and cheaply test hypothesized factors behind a shock when traditional measurement is impossible, and proposes a two-part theory-gap criterion for when such substitution is appropriate.
Load-bearing premise
The argument stands on the assumption that an open-weight model's release date before February 2025 guarantees its training information ends before the DOGE layoffs, so the answers it gives are a clean pre-shock snapshot rather than hindsight.
Editorial extensions
If this is right
- If the claim holds, researchers can reconstruct pre-shock latent attributes for any shock for which LLM training data exists, without waiting for or depending on expert panels.
- The DOGE-specific finding—that perceived ideology predicts layoffs—is confirmed using three independent open-weight models, suggesting the earlier survey-based result is not an artifact of one dataset.
- The new knowledge-institution measure offers a testable mechanism for the 'Cathedral' thesis: agencies perceived as knowledge producers were targeted net of ideology, staff, and budget.
- The two-part theory-gap criterion gives practitioners a rule for when LLM substitution is legitimate and when it is not, separating retrospective measurement from forward-looking 'silicon sample' simulations.
- Because the method is fast and cheap, it supports piloting and iterative hypothesis testing during unfolding shocks, before traditional surveys become feasible.
Reading between the lines
- The strongest test of the paper's retrospective logic would be to run the same pairwise protocol on a shock with a fully observable pre-shock benchmark—e.g., public-health agency trust before a pandemic—and check whether LLM scores match archived pre-shock survey data; the paper does not itself do this.
- The KIPS result suggests that anti-knowledge-institution rhetoric is not reducible to standard left-right ideology; a natural extension is to test whether KIPS also predicts cuts to universities, media grants, and scientific funding beyond federal agencies.
- Because the approach depends on training-data coverage, its reliability will likely be highest for well-documented agencies and actors and lowest for obscure ones; researchers should report per-item uncertainty or coverage diagnostics as a routine companion to LLM-derived scores.
- An untested boundary condition is prompt sensitivity—whether slightly different wording of the knowledge-institution prompt changes the agency ranking enough to alter the layoff-prediction result.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that large language models (LLMs) can serve as a viable substitute for expert political surveys when a shock makes post-event expert judgments uninformative. As a case study, it estimates Agency Ideology Pairwise Scores (AIPS) for 123 federal agencies using pairwise comparison prompts with three open-weight LLMs (Llama 3.3 70B Instruct, Mistral Small 3, Phi-4), scales the comparisons with a Bradley-Terry model, shows that AIPS correlates highly with the 2014 Richardson et al. expert ideology measure, and replicates Bonica's finding that perceived agency ideology predicts DOGE layoffs controlling for budget and staff. It then introduces a new Knowledge Institution Pairwise Score (KIPS), estimated with Llama 3.3 70B Instruct, and reports that KIPS predicts DOGE layoffs even when controlling for ideology. The paper concludes by proposing a two-part 'theory-gap criterion' for when LLM-based measurement can substitute for expert surveys.
Significance. The AIPS portion of the paper is a credible and largely reproducible measurement exercise: it uses openly available models, a transparent pairwise-comparison design, and reports correlations of 0.80-0.90 with an external 2014 expert survey (Table 1), while replicating the substantive DOGE finding from Bonica (Table 2). This alone is a useful contribution to the emerging literature on LLM-based latent measurement. The paper is also commendably explicit about limitations, notably in Section 5.2 where it concedes there is no benchmark for KIPS. However, the broader 'viable substitute' claim depends substantially on KIPS, a new construct with no external validation and whose only evidence is an outcome regression that may be contaminated by the prompt's reliance on the very 'Cathedral' discourse used to justify the layoffs. If the KIPS validation gap is addressed, the paper would make a stronger methodological contribution; as it stands, the central claim exceeds what the evidence establishes.
major comments (3)
- [Section 5.1] The statement that 'using open weight models with verified checkpoint dates preceding the shock guarantees that the model's knowledge cutoff is pre-event' is not supported by the evidence in the paper. A checkpoint or release date is not the same as a training-data cutoff date, and for instruction-tuned models the knowledge cutoff must be verified directly rather than inferred from the release date. Moreover, even with a verified cutoff, the unanchored prompt in Section 3.1 does not ask the model to adopt a pre-February 2025 perspective, so the output is an amalgam of training corpora rather than a temporally indexed pre-shock measurement. The paper should provide a direct knowledge-cutoff check (for example, control prompts asking about post-shock facts, or comparison of responses from models with known cutoffs on either side of the shock) before claiming that LLM estimates are ex ante measures. This is load-bearing because the paper's core 'temporal flexibility' advantage rests on this guarantee.
- [Section 4 and Section 5.2] KIPS is a new construct with no external benchmark, and the paper concedes in Section 5.2 that 'we do not have a benchmark against which to compare KIPS.' The only validity evidence is the outcome regression in Table 3. This is insufficient because the KIPS prompt itself contains the words 'ACADEMIC AND EDUCATIONAL INSTITUTIONS, THE MEDIA, AND CIVIL SOCIETY ORGANIZATIONS' and defines knowledge institutions in terms of producing, distributing, and legitimizing knowledge—language that directly mirrors the 'Cathedral' discourse described in Section 4 as motivating DOGE. The observed association between KIPS and layoffs may therefore reflect the model's encoding of the anti-knowledge-institution narrative that drives the outcome, rather than an independent pre-shock expert perception. The 'viable substitute' claim for new constructs requires additional validation, such as using multiple model families, prompt variations that do not embed the outcome-related theory, construct validation against agency behaviors (e.g., collaborations with universities or civil society), or an expert survey administered post-shock as an imperfect benchmark, which the paper itself acknowledges as possible.
- [Section 4, Table 3] KIPS is estimated using a single LLM (Llama 3.3 70B Instruct) and a single prompt, so there is no sensitivity analysis for the central new measure. In contrast, the AIPS analysis uses three models and reports cross-model correlations (Table 1); the KIPS analysis should be held to the same standard, or the paper should explicitly state that KIPS is exploratory. Without model- or prompt-robustness evidence, the regression results in Table 3 may be specific to one model's priors, and the general claim that 'this same approach' can measure knowledge-institution perceptions is not established.
minor comments (7)
- [Section 3.1, Section 4, Appendix A.3] The prompts contain typographical errors and missing spaces (e.g., '[AGENCY1]OR[AGENCY2]?'), which makes them difficult to reproduce exactly; please provide clean, copy-pasteable prompt text.
- [Section 1] The phrase 'perception as aknowledge institution' contains a typo; 'aknowledge' should be 'a knowledge'.
- [Table 3] The regression table would benefit from reporting the number of agencies and a goodness-of-fit statistic (e.g., McFadden R^2), and from discussing the scale of KIPS so the coefficients can be interpreted substantively.
- [Figure 2] The legend entry 'a aNo Layoffs' contains a stray 'a', and the labels in Figure 1 overlap substantially; these should be cleaned up for readability.
- [General] The paper does not include a data or code availability statement. Given the paper's emphasis on reproducibility and open-weight models, releasing the prompts, raw outputs, and scaling code would strengthen the replication claim.
- [Section 5.2] The sentence about uncertainty contains a missing space ('uncertaintyof'); minor grammar and spacing issues like this appear in several places and should be corrected throughout.
- [Section 5.3] The two-part theory-gap criterion is sensible but would benefit from explicit operationalization: what counts as 'strong theory,' and how should a researcher assess whether pre-shock data collection was 'theoretically possible' in practice?
Circularity Check
No circular derivation: AIPS is externally benchmarked, and KIPS, while lacking an independent benchmark, is not fitted to the layoff outcome.
full rationale
The main AIPS chain is self-contained and externally anchored: pairwise LLM comparisons are converted via the Bradley-Terry model into ideology scores that are then compared with Richardson et al. (2018) and used to replicate Bonica (2025). Neither step fits the DOGE layoff outcome into the measurement itself, and the 2014 expert survey provides an independent external benchmark. The KIPS chain is the closest thing to a circularity concern because the paper concedes in Section 5.2 that 'we do not have a benchmark against which to compare KIPS' and its principal support is the Table 3 regression on DOGE layoffs. However, KIPS is not statistically constructed from that outcome: the scores are assembled from pairwise LLM responses before the layoff regression, so the regression is a criterion-validity test, not a fitted input renamed as a prediction. The validity gap is real and explicitly acknowledged, but it is a construct-validity limitation rather than a reduction of the prediction to the input by construction. The Section 5.1 claim that pre-shock checkpoint dates guarantee a pre-event knowledge cutoff is an empirical premise about training data, not a circular step. The paper cites Wu et al. (2023) for the pairwise-comparison protocol, which is a self-citation, but that protocol is also supported by standard references such as Bradley and Terry (1952) and Turner and Firth (2012), and the non-reproducible Napolio preprint is explicitly not relied upon for the new results. No uniqueness theorem, ansatz, or renaming of a known result is used to force the conclusion. Therefore, no load-bearing circular step is present; the low score reflects only the acknowledged absence of an external KIPS benchmark and the minor self-citation.
Assumptions & free parameters
free parameters (2)
- AIPS agency ideology scores (123 agencies) =
Estimated by Bradley-Terry from LLM pairwise comparisons; values in Appendix figures
- KIPS knowledge-institution scores (123 agencies) =
Estimated by Bradley-Terry from Llama 3.3 70B comparisons; values in Appendix Figure 6
assumptions (6)
- domain assumption LLM training corpora contain pre-shock information about federal agencies that reflects aggregate expert perceptions, and pairwise responses recover those perceptions.
- domain assumption The release date of an open-weight model is a reliable proxy for a pre-shock knowledge cutoff.
- domain assumption The 2014 Richardson et al. survey is a valid benchmark for the true pre-layoff perceived ideology of federal agencies.
- domain assumption The substantive theory that knowledge institutions are targets of the administration, drawn from Yarvin and Vance, is credible enough to justify the KIPS construct.
- standard math The Bradley-Terry model with 0.5 wins for ties produces interval-scale latent scores from pairwise comparisons.
- domain assumption Expert surveys cannot be collected post-shock without outcome bias, so ex post validation of KIPS is impossible now.
invented entities (1)
-
Knowledge Institution Pairwise Scores (KIPS) construct
Cite this review
Pith. "Pith review of Large Language Models Can Be a Viable Substitute for Expert Political Surveys When a Shock Disrupts Traditional Measurement Approaches." pith.science (2026). https://pith.science/paper/F2V7XU4F
@misc{pith2026250606540,
author = {Pith},
title = {Pith review of: Large Language Models Can Be a Viable Substitute for Expert Political Surveys When a Shock Disrupts Traditional Measurement Approaches},
year = {2026},
howpublished = {\url{https://pith.science/paper/F2V7XU4F}},
note = {Machine review of arXiv:2506.06540}
}
read the original abstract
After a disruptive event or shock, such as the Department of Government Efficiency (DOGE) federal layoffs of 2025, expert judgments are colored by knowledge of the outcome. This can make it difficult or impossible to reconstruct the pre-event perceptions needed to study the factors associated with the event. This position paper argues that large language models (LLMs), trained on vast amounts of digital media data, can be a viable substitute for expert political surveys when a shock disrupts traditional measurement. We analyze the DOGE layoffs as a specific case study for this position. We use pairwise comparison prompts with LLMs and derive ideology scores for federal executive agencies. These scores replicate pre-layoff expert measures and predict which agencies were targeted by DOGE. We also use this same approach and find that the perceptions of certain federal agencies as knowledge institutions predict which agencies were targeted by DOGE, even when controlling for ideology. This case study demonstrates that using LLMs allows us to rapidly and easily test the associated factors hypothesized behind the shock. More broadly, our case study of this recent event exemplifies how LLMs offer insights into the correlational factors of the shock when traditional measurement techniques fail. We conclude by proposing a two-part criterion for when researchers can turn to LLMs as a substitute for expert political surveys.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Hewett, Mojan Javaheripi, Piero Kauffmann, James R
Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J. Hewett, Mojan Javaheripi, Piero Kauffmann, James R. Lee, Yin Tat Lee, Yuanzhi Li, Weishung Liu, Caio C. T. Mendes, Anh Nguyen, Eric Price, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Xin Wang, Rachel Ward, Yue Wu, Dingli Yu,...
arXiv 2024
-
[2]
Measurement validity: A shared standard for qualitative and quantitative research
Robert Adcock and David Collier. Measurement validity: A shared standard for qualitative and quantitative research. American Political Science Review, 95 0 (3): 0 529–546, 2001. doi:10.1017/S0003055401003100
-
[3]
Stephen Ansolabehere, James M. Snyder, and Charles Stewart. Candidate positioning in U.S. house elections. American Journal of Political Science, 45 0 (1): 0 136--159, 2001. ISSN 00925853, 15405907. URL http://www.jstor.org/stable/2669364
-
[4]
Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. Out of one, many: Using language models to simulate human samples. Political Analysis, 31 0 (3): 0 337–351, 2023. doi:10.1017/pan.2023.2
-
[5]
Uncertainty in natural language generation: From theory to applications, 2023
Joris Baan, Nico Daheim, Evgenia Ilia, Dennis Ulmer, Haau-Sing Li, Raquel Fernández, Barbara Plank, Rico Sennrich, Chrysoula Zerva, and Wilker Aziz. Uncertainty in natural language generation: From theory to applications, 2023. URL https://arxiv.org/abs/2307.15703
arXiv 2023
-
[6]
Birds of the same feather tweet together: Bayesian ideal point estimation using twitter data
Pablo Barberá. Birds of the same feather tweet together: Bayesian ideal point estimation using twitter data. Political Analysis, 23 0 (1): 0 76–91, 2015. doi:10.1093/pan/mpu011
-
[7]
Larry M. Bartels. Economic Inequality and Political Representation . In The Unsustainable American State . Oxford University Press, 10 2009. ISBN 9780195392135. doi:10.1093/acprof:oso/9780195392135.003.0007. URL https://doi.org/10.1093/acprof:oso/9780195392135.003.0007
arXiv 2009
-
[8]
Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M
James Bisbee, Joshua D. Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M. Larson. Synthetic replacements for human survey data? the perils of large language models. Political Analysis, 32 0 (4): 0 401–416, 2024. doi:10.1017/pan.2024.5
Show all 69 references
-
[9]
Mapping the ideological marketplace
Adam Bonica. Mapping the ideological marketplace. American Journal of Political Science, 58 0 (2): 0 367--386, 2014. ISSN 00925853, 15405907. URL http://www.jstor.org/stable/24363491
2014
-
[10]
The DOGE purge: E mpirical E vidence of P olitically M otivated F irings, February 2025
Adam Bonica. The DOGE purge: E mpirical E vidence of P olitically M otivated F irings, February 2025. URL https://data4democracy.substack.com/p/the-doge-purge-empirical-evidence. On Data and Democracy, Substack blog
2025
-
[11]
Elmendorf, and Scott A
Cheryl Boudreau, Christopher S. Elmendorf, and Scott A. MacKenzie. Racial or spatial voting? the effects of candidate ethnicity and ethnic group endorsements in local elections. American Journal of Political Science, 63 0 (1): 0 5--20, 2019. doi:https://doi.org/10.1111/ajps.12...
2019 doi
-
[12]
Ralph Allan Bradley and Milton E. Terry. Rank analysis of incomplete block designs: The method of paired comparisons. Biometrika, 39 0 (3-4): 0 324--345, 12 1952. ISSN 0006-3444. doi:10.1093/biomet/39.3-4.324. URL https://doi.org/10.1093/biomet/39.3-4.324
1952 doi
-
[13]
Montgomery
David Carlson and Jacob M. Montgomery. A pairwise comparison framework for fast, flexible, and reliable human coding of political texts. American Political Science Review, 111 0 (4): 0 835–843, 2017. doi:10.1017/S0003055417000302
2017 doi
-
[14]
Carpenter
Daniel P. Carpenter. The Forging of Bureaucratic Autonomy: Reputations, Networks, and Policy Innovation in Executive Agencies, 1862-1928, volume 173. Princeton University Press, 2001. ISBN 9780691070100
1928
-
[15]
Lewis, James Lo, Keith T
Royce Carroll, Jeffrey B. Lewis, James Lo, Keith T. Poole, and Howard Rosenthal. Measuring bias and uncertainty in DW-NOMINATE ideal point estimates via the parametric bootstrap. Political Analysis, 17 0 (3): 0 261--275, 2009. ISSN 10471987, 14764989. URL http://www.jstor.org/...
2009
-
[16]
Federal employee unionization and presidential control of the bureaucracy: Estimating and explaining ideological change in executive agencies
Jowei Chen and Tim Johnson. Federal employee unionization and presidential control of the bureaucracy: Estimating and explaining ideological change in executive agencies. Journal of Theoretical Politics, 27 0 (1): 0 151--174, 2015. doi:10.1177/0951629813518126. URL https://doi...
2015 doi
-
[17]
Clinton and David E
Joshua D. Clinton and David E. Lewis. Expert opinion, agency characteristics, and agency preferences. Political Analysis, 16 0 (1): 0 3–20, 2008. doi:10.1093/pan/mpm009
2008 doi
-
[18]
Clinton, Anthony Bertelli, Christian R
Joshua D. Clinton, Anthony Bertelli, Christian R. Grose, David E. Lewis, and David C. Nixon. Separated powers in the united states: The ideology of agencies, presidents, and congress. American Journal of Political Science, 56 0 (2): 0 341--354, 2012. doi:https://doi.org/10.111...
2012
-
[19]
Can ai language models replace human participants? Trends in Cognitive Sciences, 27 0 (7): 0 597--600, 2023
Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. Can ai language models replace human participants? Trends in Cognitive Sciences, 27 0 (7): 0 597--600, 2023. ISSN 1364-6613. doi:https://doi.org/10.1016/j.tics.2023.04.008
2023 doi
-
[20]
Tucker, and Jonathan Nagler
Gregory Eady, Richard Bonneau, Joshua A. Tucker, and Jonathan Nagler. News sharing on social media: Mapping the ideology of news media, politicians, and the mass public. Political Analysis, 33 0 (2): 0 73–90, 2025. doi:10.1017/pan.2024.19
2025 doi
-
[21]
Detecting hallucinations in large language models using semantic entropy
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. Detecting hallucinations in large language models using semantic entropy. Nature, 630 0 (8017): 0 625--630, 2024. ISSN 1476-4687. doi:10.1038/s41586-024-07421-0. URL https://doi.org/10.1038/s41586-024-07421-0
2024 doi
-
[22]
De Menezes
David Firth and Renée X. De Menezes. Quasi‐variances . Biometrika, 91 0 (1): 0 65--80, 03 2004. ISSN 0006-3444. doi:10.1093/biomet/91.1.65. URL https://doi.org/10.1093/biomet/91.1.65
2004 doi
-
[23]
Shapiro, and Matt Taddy
Matthew Gentzkow, Jesse M. Shapiro, and Matt Taddy. Measuring group differences in high-dimensional choices: Method and application to congressional speech. Econometrica, 87 0 (4): 0 1307--1340, 2019. doi:https://doi.org/10.3982/ECTA16566
2019 doi
-
[24]
Gilmour and David E
John B. Gilmour and David E. Lewis. Political appointees and the competence of federal program management. American Politics Research, 34 0 (1): 0 22--50, 2006. doi:10.1177/1532673X04271905
2006 doi
-
[25]
`What Elon Musk Said Is a Bold-Faced Lie'
Michael Hirsh. `What Elon Musk Said Is a Bold-Faced Lie' . Politico Magazine, February 2025. URL https://www.politico.com/news/magazine/2025/02/04/elon-musk-usaid-00202409
2025
-
[26]
Exhaustive or exhausting? evidence on respondent fatigue in long surveys
Dahyeon Jeong, Shilpa Aggarwal, Jonathan Robinson, Naresh Kumar, Alan Spearot, and David Sungho Park. Exhaustive or exhausting? evidence on respondent fatigue in long surveys. Journal of Development Economics, 161: 0 102992, 2023. ISSN 0304-3878. doi:https://doi.org/10.1016/j....
2023
-
[27]
From NeoReactionary Theory to the Alt-Right, pages 101--120
Andrew Jones. From NeoReactionary Theory to the Alt-Right, pages 101--120. Springer International Publishing, Cham, 2019. ISBN 978-3-030-18753-8. doi:10.1007/978-3-030-18753-8_6
2019 doi
-
[28]
Positioning political texts with large language models by asking and averaging
Gaël Le Mens and Aina Gallego. Positioning political texts with large language models by asking and averaging. Political Analysis, page 1–9, 2025. doi:10.1017/pan.2024.29
2025 doi
-
[29]
Teaching models to express their uncertainty in words
Stephanie Lin, Jacob Hilton, and Owain Evans. Teaching models to express their uncertainty in words. Transactions on Machine Learning Research, 2022. ISSN 2835-8856. URL https://openreview.net/forum?id=8s8K2UZGTZ
2022
-
[30]
Generating with confidence: Uncertainty quantification for black-box large language models
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. Generating with confidence: Uncertainty quantification for black-box large language models. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. URL https://openreview.net/forum?id=DWkJCSxKU5
2024
-
[31]
Uncertainty estimation and quantification for llms: A simple supervised approach, 2024
Linyu Liu, Yu Pan, Xiaocheng Li, and Guanting Chen. Uncertainty estimation and quantification for llms: A simple supervised approach, 2024. URL https://arxiv.org/abs/2404.15993
2024 arXiv
-
[32]
The L lama 3 herd of models, 2024
Llama Team, AI @ Meta . The L lama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783
2024 arXiv
-
[33]
Testing the power of arguments in referendums: A Bradley–Terry approach
Peter John Loewen, Daniel Rubenson, and Arthur Spirling. Testing the power of arguments in referendums: A Bradley–Terry approach. Electoral Studies, 31 0 (1): 0 212--221, 2012. ISSN 0261-3794. doi:https://doi.org/10.1016/j.electstud.2011.07.003. URL https://www.sciencedirect.c...
2012 doi
-
[34]
Robert C. Luskin. Measuring political sophistication. American Journal of Political Science, 31 0 (4): 0 856--899, 1987. ISSN 00925853, 15405907. URL http://www.jstor.org/stable/2111227
1987
-
[35]
Brown, Joshua A
Maggie Macdonald, Megan A. Brown, Joshua A. Tucker, and Jonathan Nagler. To moderate, or not to moderate: Strategic domain sharing by congressional campaigns. Electoral Studies, 95: 0 102907, 2025. ISSN 0261-3794. doi:https://doi.org/10.1016/j.electstud.2025.102907. URL https:...
2025
-
[36]
Curtis yarvin says democracy is done
David Marchese. Curtis yarvin says democracy is done. powerful conservatives are listening., 2025. URL https://www.nytimes.com/2025/01/18/magazine/curtis-yarvin-interview.html
2025
-
[37]
Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau
Sabrina J. Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. Reducing conversational agents' overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics, 10: 0 857--872, 2022. doi:10.1162/tacl_a_00494. URL https://aclantholo...
2022 doi
-
[38]
Mistral small 3, January 2025
Mistral AI Team . Mistral small 3, January 2025. URL https://mistral.ai/news/mistral-small-3
2025
-
[39]
Terry M. Moe. Control and feedback in economic regulation: The case of the nlrb. American Political Science Review, 79 0 (4): 0 1094–1116, 1985. doi:10.2307/1956250
1985 doi
-
[40]
Virtual personas for language models via an anthology of backstories
Suhong Moon, Marwa Abdulhai, Minwoo Kang, Joseph Suh, Widyadewi Soedarmadji, Eran Kohen Behar, and David Chan. Virtual personas for language models via an anthology of backstories. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conferenc...
2024 doi
-
[41]
Measuring executive agency ideology using large language models, n.d
Nicholas G Napolio. Measuring executive agency ideology using large language models, n.d. URL https://static1.squarespace.com/static/59b80bc4d2b857d5b7da854d/t/665d6980e716480726e732b4/1717397889454/Napolio+Agency+Ideology.pdf
-
[42]
Crowdsourcing subjective annotations using pairwise comparisons reduces bias and error compared to the majority-vote method
Hasti Narimanzadeh, Arash Badie-Modiri, Iuliia G Smirnova, and Ted Hsuan Yun Chen. Crowdsourcing subjective annotations using pairwise comparisons reduces bias and error compared to the majority-vote method. Proceedings of the ACM on Human-Computer Interaction, 7 0 (CSCW2): 0 ...
2023
-
[43]
David C. Nixon. Separation of powers and appointee ideology. The Journal of Law, Economics, and Organization, 20 0 (2): 0 438--457, 10 2004. ISSN 8756-6222. doi:10.1093/jleo/ewh041
2004 doi
-
[44]
Measurement in the age of llms: An application to ideological scaling, 2024
Sean O'Hagan and Aaron Schein. Measurement in the age of llms: An application to ideological scaling, 2024. URL https://arxiv.org/abs/2312.09203
2024 arXiv
-
[45]
When do politicians grandstand? measuring message politics in committee hearings
Ju Yeon Park. When do politicians grandstand? measuring message politics in committee hearings. The Journal of Politics, 83 0 (1): 0 214--228, 2021. doi:10.1086/709147
2021 doi
-
[46]
Federal workers express shock, anger over mass firings: ``you are not fit for continued employment''
Aimee Picchi. Federal workers express shock, anger over mass firings: ``you are not fit for continued employment''. CBS News, 2025. URL https://www.cbsnews.com/news/trump-federal-employees-probationary-firings-layoffs-workers-impact/
2025
-
[47]
Poole and Howard Rosenthal
Keith T. Poole and Howard Rosenthal. A spatial model for legislative roll call analysis. American Journal of Political Science, 29 0 (2): 0 357--384, 1985. ISSN 00925853, 15405907. URL http://www.jstor.org/stable/2111172
1985
-
[48]
Poole and Howard Rosenthal
Keith T. Poole and Howard Rosenthal. Ideology and Congress: A Political Economic History of Roll Call Voting. Yale University Press, New Haven, CT, 1997
1997
-
[49]
Porter, Michael E
Stephen R. Porter, Michael E. Whitcomb, and William H. Weitzer. Multiple surveys of students and survey fatigue. New Directions for Institutional Research, 2004 0 (121): 0 63--73, 2004. doi:https://doi.org/10.1002/ir.101
2004 doi
-
[50]
Richardson, Joshua D
Mark D. Richardson, Joshua D. Clinton, and David E. Lewis. Elite perceptions of agency ideology and workforce skill. The Journal of Politics, 80 0 (1): 0 303--308, 2018. doi:10.1086/694846. URL https://doi.org/10.1086/694846
2018 doi
-
[51]
Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models
Paul R \"o ttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schuetze, and Dirk Hovy. Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models. In Lun-Wei Ku, Andre Martins, and Vive...
2024
-
[52]
Minority opposition and asymmetric parties? S enators’ partisan rhetoric on twitter
Annelise Russell. Minority opposition and asymmetric parties? S enators’ partisan rhetoric on twitter. Political Research Quarterly, 74 0 (3): 0 615--627, 2021. doi:10.1177/1065912920921239. URL https://doi.org/10.1177/1065912920921239
2021 doi
-
[53]
Wu, Kristina Miler, Alexander Miserlis Hoyle, and Philip Resnik
Rupak Sarkar, Patrick Y. Wu, Kristina Miler, Alexander Miserlis Hoyle, and Philip Resnik. P air S cale: Analyzing attitude change with pairwise comparisons. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Findings of the Association for Computational Linguistics: NAACL 20...
2025
-
[54]
Adler, Lea Rau, and Bernd Schmitt
Marko Sarstedt, Susanne J. Adler, Lea Rau, and Bernd Schmitt. Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing, 41 0 (6): 0 1254--1270, 2024. doi:https://doi.org/10.100...
2024 doi
-
[55]
Schmidt and Michael C
Michael S. Schmidt and Michael C. Bender. Trump administration halts harvard's ability to enroll international students, May 2025. URL https://www.nytimes.com/2025/05/22/us/politics/trump-harvard-international-students.html
2025
-
[56]
Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Representations, 2024. URL...
2024
-
[57]
The A merican system of shared powers: the P resident, C ongress, and the NLRB
SK Snyder and BR Weingast. The A merican system of shared powers: the P resident, C ongress, and the NLRB . The Journal of Law, Economics, and Organization, 16 0 (2): 0 269--305, 10 2000. ISSN 8756-6222. doi:10.1093/jleo/16.2.269
-
[58]
Language model fine-tuning on scaled survey data for predicting distributions of public opinions, 2025
Joseph Suh, Erfan Jahanparast, Suhong Moon, Minwoo Kang, and Serina Chang. Language model fine-tuning on scaled survey data for predicting distributions of public opinions, 2025. URL https://arxiv.org/abs/2502.16761
2025 arXiv
-
[59]
Do llms exhibit human-like response biases? a case study in survey design
Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkwar, and Graham Neubig. Do llms exhibit human-like response biases? a case study in survey design. Transactions of the Association for Computational Linguistics, 12: 0 1011--1026, 09 2024. ISSN 2307-387X. doi:10.1162/ta...
2024 doi
-
[60]
Bradley-terry models in r: The bradleyterry2 package
Heather Turner and David Firth. Bradley-terry models in r: The bradleyterry2 package. Journal of Statistical Software, 48 0 (9): 0 1–21, 2012. doi:10.18637/jss.v048.i09. URL https://www.jstatsoft.org/index.php/jss/article/view/v048i09
2012 doi
-
[61]
He's anti-democracy and pro-trump: the obscure 'dark enlightenment' blogger influencing the next us administration, 2024
Jason Wilson. He's anti-democracy and pro-trump: the obscure 'dark enlightenment' blogger influencing the next us administration, 2024. URL https://www.theguardian.com/us-news/2024/dec/21/curtis-yarvin-trump
2024
-
[62]
Cultural M arxism and the C athedral: Two alt-right perspectives on critical theory
Alan Woods. Cultural M arxism and the C athedral: Two alt-right perspectives on critical theory. In Caroline Battista and Marcelle Sande, editors, Critical Theory and the Humanities in the Age of the Alt-Right, pages 45--64. Palgrave Macmillan, Cham, 2019. doi:10.1007/978-3-03...
2019 doi
-
[63]
Patrick Y. Wu. Using semantically unrelated and opposite terms for in-context learning: A case study in identifying political aversion in tweets. In Proceedings of the 17th ACM Web Science Conference 2025, Websci '25, page 528–533, New York, NY, USA, 2025. Association for Comp...
2025
-
[64]
Wu, Jonathan Nagler, Joshua A
Patrick Y. Wu, Jonathan Nagler, Joshua A. Tucker, and Solomon Messing. Large language models can be used to estimate the latent positions of politicians, 2023. URL https://arxiv.org/abs/2303.12057
2023 arXiv
-
[65]
Wu, Jonathan Nagler, Joshua A
Patrick Y. Wu, Jonathan Nagler, Joshua A. Tucker, and Solomon Messing. Concept-guided chain-of-thought prompting for pairwise comparison scoring of texts with large language models. In 2024 IEEE International Conference on Big Data (BigData), pages 7232--7241, 2024. doi:10.110...
2024
-
[66]
Can LLM s express their uncertainty? an empirical evaluation of confidence elicitation in LLM s
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. Can LLM s express their uncertainty? an empirical evaluation of confidence elicitation in LLM s. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview....
2024
-
[67]
Navigating the grey area: How expressions of uncertainty and overconfidence affect language models
Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto. Navigating the grey area: How expressions of uncertainty and overconfidence affect language models. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural La...
2023 doi
-
[68]
P ro SA : Assessing and understanding the prompt sensitivity of LLM s
Jingming Zhuo, Songyang Zhang, Xinyu Fang, Haodong Duan, Dahua Lin, and Kai Chen. P ro SA : Assessing and understanding the prompt sensitivity of LLM s. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EM...
2024 doi
-
[69]
Can large language models transform computational social science? Computational Linguistics, 50 0 (1): 0 237--291, 03 2024
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. Can large language models transform computational social science? Computational Linguistics, 50 0 (1): 0 237--291, 03 2024. ISSN 0891-2017. doi:10.1162/coli_a_00502. URL https://doi.org/10.1162/co...
2024 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.