{"id":"36acef25-bd49-4e40-885b-264973fca53c","arxiv_id":"2607.07915","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"In 50 LLM measurement tasks from 27 top-journal papers, LLM outputs are often central to claims yet validation is limited, mostly convergent, and frequently incomplete.","lead":"Social scientists are using LLMs to measure concepts like ideology and sentiment in top journals, but validation of those measures is inconsistent and narrow. The paper maps current practices and argues for multi-lens construct validity plus transparent reporting to protect research credibility.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The paper’s strongest claim is a carefully scoped empirical map of early practice, not a universal causal claim about all social science. Methods §5.1–5.2 and the inclusion criteria make the corpus construction auditable; the public OSF release allows direct verification of the counts that underwrite Figure 1 and the “8 tasks / 6 papers with no validation” statement. The reader correctly flags journal selection and analytic-code subjectivity as residual limitations, yet these are standard for this genre and do not invert the observed dominance of convergent validity or the centrality of LLM measures. Because the descriptive pattern is robust under ordinary qualitative-content-analysis standards and the recommendations follow directly from measurement theory without hidden assumptions, no adjustment to the ACCEPT verdict is warranted.","tokens_in":20114,"tokens_out":412,"duration_ms":4538,"concrete_test":"Independently re-apply the released codebook (OSF) to a random 10-task subsample and recompute the UpSet counts for ‘no validation’ and ‘convergent-only’; if the 8/50 and 28/42 figures shift by more than 2–3 tasks, the headline descriptive claim would need qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central descriptive claim—that LLM-generated measurements are often central yet validation is inconsistent and dominated by convergent validity (8/50 tasks with none)—is supported by a transparent, comprehensive corpus of 27 papers/50 tasks from the selected flagship journals, a released coding dataset, and an explicit codebook. The reader’s weakest assumption (representativeness of this first-wave corpus and reliability of analytic codes) is real but ordinary for qualitative content analysis of an emerging practice; the paper itself frames the set as the complete early wave rather than a statistical sample of all social science, and residual coder subjectivity does not reverse the observed pattern of narrow validation. No internal inconsistency or load-bearing error undermines the strongest claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper systematically documents how large language models are used as measurement instruments in a comprehensive early corpus of 27 articles (50 tasks) drawn from eight flagship social-science journals (2023–late 2025). After dual qualitative coding of prompts, models, answer-extraction procedures, and validation practices, the authors report that LLM-generated measurements frequently occupy a central role in empirical claims, yet validation is inconsistent: eight tasks report none, most remaining tasks assess only convergent validity, and other facets of construct validity (face, content, predictive, discriminant, hypothesis) are rare. Conceptual definitions in prompts are often underspecified, and many instrument components (decoding, non-compliance handling, exact model versions) are incompletely reported. Drawing on measurement theory, the paper maps complementary validation strategies and recommends stronger reporting and multi-lens validation norms for authors, reviewers, and journals. A 33-dimension coding dataset is released on OSF.","tokens_in":20283,"tokens_out":1000,"duration_ms":30231,"significance":"If the descriptive pattern holds, the paper supplies timely, field-level evidence that an increasingly central measurement practice is outrunning the validation norms needed to underwrite credible social-science claims. Strengths include an explicit, publicly released codebook and coding dataset, dual coding of descriptive fields with documented disagreement resolution, transparent UpSet and supplementary tables, and constructive, actionable recommendations grounded in established construct-validity frameworks rather than ad-hoc checklists. The work is well positioned to shape emerging journal guidelines and reviewer expectations around LLM-based measurement.","major_comments":[{"comment":"§5.2 and Figure 1: Descriptive codes were dual-coded with primary-coder review, but analytic codes for aspects of construct validity (the basis of the UpSet plot and the claim that convergent validity dominates, 38/50 tasks) were developed by the primary coder after team discussion, without a reported second-coder reliability check or quantitative agreement metric on those codes. Because the central claim about narrow validation rests on these classifications, the manuscript should either (a) report a second-coder audit on a substantial subset of tasks for the validity-aspect codes or (b) more explicitly document the consensus procedure used for each task’s validity classification so readers can assess residual subjectivity.","section":"§5.2 / Figure 1"},{"comment":"§2.1, §5.1, and Conclusion: The corpus is framed as a “comprehensive” first wave whose practices illuminate emerging field norms. That framing is defensible for the selected journals and keyword filter, yet three sociology journals yielded zero papers and the bulk of tasks come from PNAS, Nature Human Behaviour, and Political Analysis. The Discussion/Conclusion should more explicitly treat this venue concentration as a scope limitation when generalizing from “these 27 articles” to norms across social science, so that the normative recommendations are not over-read as already field-wide.","section":"§2.1 / §5.1 / Conclusion"}],"minor_comments":[{"comment":"Table S2 and the accompanying text state that LLM measurements are often central; a brief cross-reference in the main Results §2.1 to the exact counts (e.g., 12 papers / 17 tasks as part of main analysis) would help readers without opening the supplement.","section":"§2.1 / Table S2"},{"comment":"Figure 2 is a useful running example of construct-validity checks, but the dense multi-column layout is hard to parse in print; consider a slightly larger type or a two-panel layout so the guiding questions remain legible.","section":"Figure 2"},{"comment":"Table S3 reports that only 4 papers / 6 tasks document answer-extraction procedures; the main text §2.3 already notes this, but a single sentence quantifying non-reporting of decoding and non-compliance handling would make the transparency gap more immediately visible.","section":"§2.3 / Table S3"},{"comment":"In §2.2 the taxonomy of concept specification (single word / dictionary / stipulative) is adapted from Halterman & Keith; a brief parenthetical reminder of the source at first use in the main text (beyond the footnote) would aid readers who skip the supplement.","section":"§2.2"},{"comment":"A few references appear with future or near-future years (e.g., 2026 conference proceedings); confirm final bibliographic details at proof stage so DOIs and page ranges are stable.","section":"References"}],"recommendation":"minor_revision","confidential_remarks":"Strong fit for a methods- or computational-social-science venue. The descriptive core is solid and the OSF release is a genuine contribution; the two major points are clarification/reliability documentation rather than redesign. I would not require new data collection. No concerns about novelty disclosure or citation pattern."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is the first systematic coding of how top social-science journals actually use LLMs as measurement instruments. They pulled every research article from eight flagship outlets (2022–late 2025), screened to 27 papers / 50 tasks, dual-coded instrument design and validation, and released the 33-dimension dataset on OSF. That inventory is the real contribution.\n\nWhat they show is clear and useful: LLM outputs are often central to the main analysis, yet validation is thin and uneven—convergent validity dominates, eight tasks report none, concepts are frequently underspecified, and key design choices (exact prompt, decoding, answer extraction, noncompliance handling) are under-reported. They map this onto standard construct-validity lenses (Jacobs & Wallach / Adcock & Collier) and give concrete examples of what face, content, predictive, etc. checks could look like for LLM tasks. The UpSet plot and tables make the pattern easy to see.\n\nSoft spots are ordinary for this kind of work, not load-bearing. The corpus is the complete early wave in those journals, not a probability sample of all social science; analytic codes for “aspects of validity” still involve judgment; and the recommendations are sensible restatements of measurement best practice rather than a new theory. None of that reverses the descriptive claim. Citation pattern is appropriate; methods and inclusion criteria are transparent.\n\nThis is for methodologists, journal editors, and anyone now treating LLMs as annotation or survey instruments. It deserves a serious referee and should be engaged. I would bring it to reading group and expect to cite the empirical map when discussing LLM measurement norms.","headline":"Solid first empirical map of LLM-as-measurement practices in flagship social-science journals; the descriptive gaps are real and the recommendations are usable.","tokens_in":20827,"tokens_out":427,"would_cite":true,"duration_ms":18400,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"LLM measurements already drive top social-science claims, yet validation is thin and mostly limited to one check.","keywords":["large language models","measurement validity","construct validity","social science methods","prompting","annotation","reporting standards","epistemic threats"],"falsifier":"A re-coding of the same corpus by independent raters, or expansion to the next two years of the same journals, that finds most tasks already employing multiple validity lenses and precise concept definitions would undermine the claim of limited, inconsistent practice.","tokens_in":21034,"feed_emoji":"📏","tokens_out":584,"duration_ms":6477,"temperature":0.7,"pith_summary":"Social scientists are increasingly prompting large language models to produce quantitative measures of concepts such as ideology, sentiment, or offensiveness. This paper systematically reviews every such measurement task published in eight flagship journals from 2023 to 2025. It finds that these LLM-generated numbers frequently sit at the center of a paper’s main empirical analysis, yet the practices used to check whether those numbers are valid remain uneven and incomplete. Most validation efforts compare LLM outputs only to a single gold-standard source (convergent validity); other aspects of construct validity are rarely examined, and eight of fifty tasks report no validation at all. Concepts are often left underspecified, and key design choices—prompt wording, model version, how answers are extracted—are inconsistently documented. The authors argue that without precise concept definitions, transparent reporting, and multi-lens validation, researchers risk handing conceptual control to opaque models and importing systematic bias into published claims.","feed_headline":"LLM measures drive top journal claims, yet validation stays thin","feed_subtitle":"Fifty tasks in flagship journals show central use but mostly one validity check","key_machinery":"A systematic qualitative coding of conceptualization, operationalization, and the full suite of construct-validity aspects (face, convergent, content, predictive, hypothesis, discriminant) applied to every qualifying LLM measurement task in the eight journals.","core_discovery":"Across a complete corpus of 50 measurement tasks in 27 papers from eight leading social-science journals, LLM-generated measurements commonly serve as inputs to primary analyses, yet validation is dominated by a single form of evidence (convergent validity) and is missing entirely for eight tasks; concept definitions and instrument details are frequently underspecified or unreported.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["LLM measures fuel top journal claims while validation remains limited","Flagship journals use LLM measures for key claims with thin validation","Fifty LLM measurement tasks in top journals show mostly weak validation","LLM tools drive social science findings yet validation practices lag","Central LLM measures in elite papers often lack full validity evidence"],"cache_read_input_tokens":5504,"weakest_assumption_plain":"That the 27 papers and 50 tasks drawn from these eight journals between 2023 and 2025 form a reliable first-wave sample whose coding patterns can stand for emerging field-wide norms.","fun_headline_variants_meta":{"raw":{"variants":["LLM measures fuel top journal claims while validation remains limited","Flagship journals use LLM measures for key claims with thin validation","Fifty LLM measurement tasks in top journals show mostly weak validation","LLM tools drive social science findings yet validation practices lag","Central LLM measures in elite papers often lack full validity evidence"]},"model":"grok-4.5","effort":"low","cost_usd":0.007184,"raw_usage":{"total_tokens":1638,"prompt_tokens":661,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":71840000,"prompt_tokens_details":{"text_tokens":661,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":914,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":661,"tokens_out":63,"duration_ms":8615,"temperature":1.0,"reasoning_tokens":914,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T15:28:42.785339+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A re-coding of the same corpus by independent raters, or expansion to the next two years of the same journals, that finds most tasks already employing multiple validity lenses and precise concept definitions would undermine the claim of limited, inconsistent practice.","supporting_citations":[],"review_version":1}