{"id":"46c84680-fd68-4aed-a390-b740ca327c78","arxiv_id":"2606.13670","paper_version":2,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"LLMs match original qualitative conclusions in 80% of 180 studies and effect sizes in 24%, performing similarly to humans in a tested subset, positioning them as a screening tool rather than a full replacement.","lead":"This paper tests large language models on automating reproducibility checks for 180 published social and behavioral science studies by comparing LLM reanalyses to original claims. If scalable, the approach could reduce the cost of verifying empirical findings across many papers.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Representativeness of the 180 studies with predefined claims is the load-bearing assumption for generalizing the 80% qualitative agreement to a screening tool","rationale":"The reader's weakest_assumption directly identifies the condition required for the headline performance numbers to underwrite the 'scalable screening tool' claim. No more technical internal inconsistency (e.g., in effect-size tolerance or LLM pipeline) is visible from the supplied abstract, and the full-text placeholder does not alter the selection-bias risk. The low reader confidence is therefore appropriate; the concern is not manufactured.","tokens_in":1798,"tokens_out":390,"duration_ms":12316,"concrete_test":"In the methods section, locate the paragraph describing how the N=180 studies were assembled (search for terms like 'selection', 'sample', 'predefined claims', or 'inclusion criteria'). Extract the exact procedure; if it is non-random curation or relies on papers already known to have extractable claims, draw a fresh random sample of 30 papers from the same journals/years without that filter and rerun the LLM pipeline to measure qualitative agreement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that LLM performance (80% same qualitative conclusion, 24% effect-size recovery within ±0.05 Cohen's d) on these studies demonstrates LLMs can serve as a scalable screening tool rather than substitute for experts. This requires the 180 studies to be sufficiently representative of typical social/behavioral science papers. The abstract states only that the studies have 'predefined claims' and gives no selection criteria, sampling frame, or exclusion rules. If the set was curated toward papers with clear, easily extractable claims (or excluded ambiguous analyses), the reported rates would not support the broader usefulness conclusion. The 11 failed cases and the human-reanalysis subset do not address this selection issue.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that large language models can automate reproducibility assessments for empirical studies in the social and behavioral sciences. On a set of N=180 published studies with predefined claims, an LLM pipeline produced viable effect-size estimates for 169 studies and matched the original study's qualitative conclusion in 80% of those cases while recovering the original effect size (within ±0.05 Cohen's d) in 24%. In a human-reanalysis subset the LLM matched originals at 95% (vs. 83% for humans) and recovered effect sizes at 40% (vs. 28% for humans). The authors conclude that LLMs can serve as a scalable screening tool to support systematic audits rather than substitute for expert judgment.","tokens_in":1970,"tokens_out":553,"duration_ms":11499,"significance":"If the performance numbers generalize beyond the tested sample and the pipeline details are made reproducible, the work would demonstrate a practical, low-cost method for large-scale reproducibility screening. The direct empirical comparison to both original claims and human reanalyses is a strength, as is the explicit framing that LLMs are not a full substitute. The absence of selection criteria and implementation details, however, prevents a firm assessment of how far the result can be extrapolated to typical papers in the field.","major_comments":[{"comment":"Methods (or equivalent section describing the LLM pipeline): the manuscript supplies no information on the specific LLM used, the prompting strategy, the procedure for extracting effect sizes from text or code, or how missing data and non-viable outputs were handled. These omissions make the reported 80% qualitative agreement and 24% effect-size recovery rates impossible to evaluate or replicate, directly undermining the central performance claims.","section":"Methods"},{"comment":"Data and sample description: the 180 studies are described only as 'published studies with predefined claims'; no sampling frame, inclusion/exclusion criteria, or justification for representativeness is provided. Because the screening-tool conclusion rests on the assumption that performance on this set predicts usefulness on typical social/behavioral-science papers, the lack of selection details is load-bearing for the generalizability claim.","section":"Data and sample description"}],"minor_comments":[{"comment":"The abstract states that the LLM 'reached the same qualitative conclusion as the original study in 95% of studies' in the human-reanalysis subset; it would be clearer to specify the exact size of that subset and whether the 95% figure is computed on the same denominator as the 80% figure.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive comments. We agree that additional details on the LLM pipeline and study sample are needed to strengthen the manuscript and will revise accordingly.","responses":[{"response":"We agree that the manuscript currently lacks these implementation details. In the revised version we will add a dedicated Methods section specifying the LLM model, prompting strategy, effect-size extraction procedure, and handling of missing or non-viable outputs. This will directly address the replicability concern.","revision_made":"yes","referee_comment":"[Methods] Methods (or equivalent section describing the LLM pipeline): the manuscript supplies no information on the specific LLM used, the prompting strategy, the procedure for extracting effect sizes from text or code, or how missing data and non-viable outputs were handled. These omissions make the reported 80% qualitative agreement and 24% effect-size recovery rates impossible to evaluate or replicate, directly undermining the central performance claims."},{"response":"We acknowledge that the current description of the 180 studies is insufficient. We will expand the Data section in revision to include the sampling frame, explicit inclusion/exclusion criteria, and any justification for representativeness that can be provided from the study selection process.","revision_made":"yes","referee_comment":"[Data and sample description] Data and sample description: the 180 studies are described only as 'published studies with predefined claims'; no sampling frame, inclusion/exclusion criteria, or justification for representativeness is provided. Because the screening-tool conclusion rests on the assumption that performance on this set predicts usefulness on typical social/behavioral-science papers, the lack of selection details is load-bearing for the generalizability claim."}],"tokens_in":1498,"tokens_out":371,"duration_ms":12736,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper gives new numbers on LLM performance for full reproducibility checks: 80% qualitative agreement with originals on 169 studies, 24% effect-size recovery within ±0.05 Cohen's d, and similar rates to human reanalysts on a smaller subset. That head-to-head benchmark against both originals and humans is the concrete addition.\n\nIt is useful to see the pipeline run end-to-end on real published claims and to get the failure count (11 studies) plus the human comparison. The authors are clear that this is meant as a screen, not a replacement.\n\nThe soft spot is the study set itself. The abstract says only that the 180 papers have \"predefined claims\"; there is no sampling frame, exclusion criteria, or description of how ambiguous or complex analyses were handled. If the collection was filtered toward papers with clean, extractable results, the 80% and 24% figures do not tell us what would happen on a typical social-science paper. That selection step is load-bearing for the claim that LLMs can serve as a scalable audit tool.\n\nDetails on the specific LLM, prompting, effect-size extraction rules, and missing-data handling are also missing from the abstract, which makes it hard to judge how much of the result is method versus model.\n\nThis is worth referee time for anyone working on reproducibility infrastructure. Readers who need a documented benchmark on a defined set of studies will get something from it; readers hoping for evidence that LLMs are ready for routine use across the field will not. I would send it to review with a request for the selection protocol and full pipeline description.","headline":"LLMs matched originals on qualitative conclusions for 80% of these 180 studies and did roughly as well as humans on a subset, but the sample selection is too opaque to treat this as evidence for a general screening tool.","tokens_in":2521,"tokens_out":420,"would_cite":false,"duration_ms":9290,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Large language models can automate reproducibility assessments by reanalyzing published studies and matching original conclusions in 80% of cases.","keywords":["reproducibility","large language models","social sciences","behavioral sciences","automation","effect sizes","reanalysis","empirical research"],"falsifier":"Apply the same LLM pipeline to a fresh, randomly sampled set of 100 studies without predefined claims and measure whether qualitative agreement stays near 80%.","tokens_in":2716,"feed_emoji":"🤖","tokens_out":671,"duration_ms":17539,"temperature":0.7,"pith_summary":"The paper demonstrates that LLMs can generate reanalyses of data from published studies to check if findings hold up. Across 180 studies from the social and behavioral sciences, the LLM pipeline produced viable effect sizes for most cases and agreed with the original qualitative conclusions 80% of the time. Effect size recovery within a narrow tolerance occurred in 24% of studies. In a human-comparison subset, LLM performance on qualitative agreement reached 95%, close to the 83% rate for human reanalysts. The work positions LLMs as a scalable screening aid rather than a full replacement for expert review.","feed_headline":"LLMs match original conclusions in 80% of reproducibility checks","feed_subtitle":"Pipeline tested on 180 social and behavioral science studies reaches human-level agreement on qualitative findings and recovers effect sizes","key_machinery":"An LLM pipeline that takes study data and generates statistical reanalyses to compute effect sizes and compare conclusions to the originals.","core_discovery":"Using an LLM pipeline on N=180 published studies with predefined claims, the model reached the same qualitative conclusion as the original study in 80% of the 169 cases with viable effect size estimates and recovered the original effect sizes within +/-0.05 Cohen's d in 24% of studies. In the human-reanalysis subset, the LLM matched the original qualitative conclusion in 95% of studies (compared to 83% for humans) and recovered effect sizes in 40% of cases (compared to 28% for humans).","pith_inferences":["The method could be tested on studies from other disciplines to see if agreement rates transfer.","Integration into journal submission systems might allow automated reproducibility flags before publication.","Improvements in LLM reasoning could raise the rate of exact effect-size recovery above the current 24%.","A hybrid workflow where LLMs handle initial screening and humans resolve ambiguous cases might increase overall throughput."],"forward_implications":["LLMs can support systematic audits of empirical results at larger scales than manual reanalysis allows.","LLMs can act as a first-pass screening tool to flag studies for deeper expert review.","Performance on qualitative conclusions is comparable to that of human reanalysts.","LLMs should augment rather than replace expert judgment given current limitations on precise effect-size recovery."],"fun_headline_variants":["LLMs reach same conclusion as original in 80% of studies","LLM pipeline matches 95% of original findings in reanalysis subset","24% of studies see LLM recover original effect sizes","In subset LLM matches 40% effect sizes versus 28% for humans"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 180 studies with predefined claims represent the broader range of social and behavioral science research well enough for the results to generalize.","fun_headline_variants_meta":{"raw":{"variants":["LLMs reach same conclusion as original in 80% of studies","LLM pipeline matches 95% of original findings in reanalysis subset","24% of studies see LLM recover original effect sizes","In subset LLM matches 40% effect sizes versus 28% for humans"]},"model":"grok-4.3","cost_usd":0.005847,"raw_usage":{"total_tokens":2811,"prompt_tokens":729,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":58474500,"prompt_tokens_details":{"text_tokens":729,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2009,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":729,"tokens_out":73,"duration_ms":12084,"temperature":1.0,"reasoning_tokens":2009,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T06:28:45.408335+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Apply the same LLM pipeline to a fresh, randomly sampled set of 100 studies without predefined claims and measure whether qualitative agreement stays near 80%.","supporting_citations":[],"review_version":1}