{"id":"2b49f0a2-a1b9-4205-8feb-3806af44d22b","arxiv_id":"2506.15041","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"GPT-4o extracts inflation narratives with 44% accuracy against an expert gold standard, well below the 67-74% of human annotators, but in valid structured form.","lead":"This paper tests whether the GPT-4o language model can pull economic narratives out of Wall Street Journal and New York Times articles about inflation, using a detailed rulebook and an expert-made answer key. It finds the model can output structured narratives, but it misses and invents narratives more often than expert human coders do.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gold standard built by the same experts used as the human baseline may inflate the human–model gap; this self-referential comparison is the central load-bearing concern.","rationale":"The reader identified the same weakest assumption: the gold standard is built by the same experts who are later evaluated against it. My stress-test agrees and takes this as the primary load-bearing concern. The paper's central empirical claim is not merely that GPT-4o can extract narratives, but that it 'falls short of expert-level performance' by a specific margin. That margin is credible only if the human baseline is a fair measure of expert-level performance. Because the gold standard was produced through consensus among the very experts who are then scored against it, and because those experts were involved in codebook development, the baseline is likely to overstate what an independent expert would achieve. This does not invalidate the paper's qualitative conclusion that LLM extraction requires human validation, but it undermines the quantitative precision of the claimed gap. The proposed test—using independent annotators without consensus exposure—would directly resolve whether the gap is real or inflated. The reader's CONDITIONAL verdict remains appropriate: the paper is methodologically honest and addresses many limitations, but the central comparison needs this validation. I therefore recommend no change to the reader's verdict.","tokens_in":29096,"tokens_out":3193,"duration_ms":35728,"concrete_test":"Recruit two or three independent economics PhD annotators with no affiliation to this project, give them the final codebook and the same 80 test documents, and collect their individual annotations before any consensus discussion. Then compute their accuracy and major-deviation rates against the existing gold standard. If independent experts score within the original 67–74% accuracy and 0.35–0.49 major-deviation ranges, the human baseline is confirmed and the paper's central claim is supported. If they fall toward the model's 44% accuracy or 1.25 major deviations, the original human–model gap is largely an artifact of a self-referential gold standard.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline comparison (GPT-4o at 44% accuracy vs. experts at 67–74%, and 1.25 vs. 0.35–0.49 major deviations per document, Table 2) rests on a gold standard that the same three expert annotators constructed in consensus sessions (Section 4.2). Those same experts then serve as the human baseline, and their individual annotations are scored against that gold standard. Because the gold standard is the set of narratives these three experts unanimously agreed on after discussing every individually coded narrative, each expert's own annotations are likely to be closer to it than an independent, non-author annotator's would be. The 'expected deviation' baseline is therefore a within-group consistency measure, not an estimate of expert-level performance in the general sense. The model, by contrast, is a single outsider with no access to the consensus discussions. If the consensus process embedded the authors' interpretive preferences into the gold standard, the observed human–model gap is inflated by construction. This concern is reinforced by the fact that the annotators helped develop the codebook in workshops (Section 4.1), so they are not independent of the measurement instrument. The paper itself acknowledges that 'no narrative extraction can be truly objective' (Section 4.1), which makes the choice of baseline especially consequential: the quantitative magnitude of the human–model gap cannot be separated from the gold standard's provenance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops and evaluates a prompt-based LLM pipeline for extracting economic narratives from newspaper text. It defines an economic narrative as a causal link between two events, operationalizes this definition in a detailed codebook, builds a gold standard from three expert annotators on 100 WSJ/NYT excerpts, and compares GPT-4o, prompted with a few-shot Chain-of-Thought protocol, against that gold standard. The headline findings are that GPT-4o reaches 44% accuracy and 1.25 major deviations per document, while the three expert annotators reach 67–74% accuracy and 0.35–0.49 major deviations. The paper concludes that LLM-based extraction is feasible but not yet expert-level, and it recommends hybrid workflows in which human experts retain interpretive oversight.","tokens_in":29330,"tokens_out":4238,"duration_ms":43744,"significance":"If the central comparison were fully credible, this paper would be a useful methodological contribution: it provides a transferable codebook, a transparent prompt-engineering procedure, a publicly released annotation dataset, and a structured way to quantify narrative-extraction errors. The paper is also refreshingly honest about the subjectivity of narrative coding and about the exploratory nature of the aggregation step. Its main limitation is measurement: the human baseline is built by the same experts who are later scored against it, the accuracy metric is never defined, and no uncertainty quantification accompanies the headline numbers. These gaps affect the central claim that GPT-4o 'falls short of expert-level performance' and need to be addressed before the quantitative magnitude of the human–model gap can be taken at face value.","major_comments":[{"comment":"The gold standard is constructed by the same three experts who are later evaluated against it: after individual coding, every narrative was discussed by all three annotators and included only upon unanimous agreement. The expert accuracy figures in Table 2 are therefore a measure of within-group consistency after deliberation, not an estimate of independent expert-level performance. Because GPT-4o had no access to the consensus discussions or the codebook workshops (Section 4.1), the human–model gap could be inflated by shared interpretive preferences that are embedded in the gold standard. The paper acknowledges that 'no narrative extraction can be truly objective' (Section 4.1), but it does not address this self-referential baseline. Please provide an external validation, for example by computing each expert's accuracy against a gold standard built from the other two experts' unanimous agreements, or by comparing the model against an independent annotator who was not involved in codebook development or consensus sessions.","section":"Section 4.2, Table 2"},{"comment":"The central quantitative claim rests on an 'accuracy' metric that is never defined. The reader cannot tell whether 44% and 67–74% refer to exact narrative-level matches, partial credit for near-miss narratives, token-level overlap, or some other rule. In addition, the paper reports only point estimates from a single run at temperature 0.2 on 80 documents, with no confidence intervals, significance tests, or multiple runs. The difference between 44% and 67% could plausibly be within sampling variation given such a small test set. Please define the accuracy metric precisely, report bootstrap confidence intervals or a significance test for the expert–model gap, and either run the model multiple times or use a temperature-0 protocol with multiple seeds to assess stability.","section":"Section 6.1, Table 2"},{"comment":"The final prompt and the number of few-shots were selected by evaluating 1–9 few-shots on a 20-document validation split, and the final few-shot set was then chosen by rotating combinations of seven examples. This is a legitimate model-selection procedure, but the reported test performance is conditional on a selection process with very few validation documents, so the expert–model gap could change under a different validation split. Please report the variability of test performance across few-shot sets or validate the chosen prompt on multiple random splits, and clarify explicitly that the 80 test documents were completely untouched during prompt selection.","section":"Section 5.2, Section 5.3"}],"minor_comments":[{"comment":"There is a typo in the sentence beginning 'Morover, since an LLM's conceptual understanding...'; it should read 'Moreover'.","section":"Section 5.1"},{"comment":"The sentence 'This translate to an unexpected major-deviation rate...' should read 'This translates to...'.","section":"Section 6.1"},{"comment":"The validation description is internally inconsistent: the paper first says 20 documents are used for cross-validation, then says the model is evaluated on the 'remaining 11 examples' when using 1–9 few-shots, and later states that performance peaks with 7 few-shots. With 20 validation documents and 7 few-shots, 13 evaluation examples would remain. Please clarify the exact split and evaluation protocol.","section":"Section 5.3"},{"comment":"The Jaccard similarity metric is described only as 'lexical overlap between predicted and reference token sets'; the tokenization scheme, whether events are compared separately or as full narratives, and how coreference-resolved variants are handled should be specified.","section":"Table 2"},{"comment":"The paper states that 100 documents were 'randomly sampled' from the corpus, but it does not report the sampling seed, any stratification, or the total corpus size. This information would help readers assess generalizability and reproducibility.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The self-referential gold-standard issue is the key risk for this paper. The expert baseline is a within-group consensus measure, and the headline human–model gap may be inflated by construction. If the authors add an independent-expert check or reframe the comparison as a within-group consistency benchmark, the paper could become publishable after revision. The undefined accuracy metric and missing uncertainty quantification should also be fixed before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a worthwhile, honestly reported benchmark of GPT-4o on a hard annotation task. The codebook and gold-standard protocol are the real contributions; anyone building LLM annotation pipelines for narrative work will learn something. My main caveat is the one flagged in the stress test: the human baseline is scored against a gold standard that the same three experts built in consensus sessions. That doesn't make the result wrong, but it means the 44% vs. 67–74% gap is probably partly a within-group consistency advantage rather than a clean measure of expert-level performance. I would want an independent external coder, or at least a leave-one-out calibration, before trusting the magnitudes.\n\nWhat's new and good: the paper defines narratives as causal A-causes-B links, operationalizes that definition into a detailed codebook, and tests GPT-4o with few-shot chain-of-thought prompting against an expert-constructed gold standard. The \"expected deviation\" concept—measuring how far the experts themselves deviate from the gold standard—is a sensible way to calibrate the model's errors. They also document concrete failure modes (causal forks, chains, over-identification in sparse texts, hallucinations) and release their data and code. That is real, reproducible evidence.\n\nSoft spots: Table 2 reports \"Accuracy\" but never defines how it is computed—that is a basic omission. Jaccard similarity is defined, but there are no confidence intervals or significance tests, and the model is run once at temperature 0.2. The aggregation/clustering section is admittedly exploratory and a bit ad hoc; it is fine as a sketch but should not be taken as a validated method. The paper cites several of the authors' own prior works, which is natural in this niche and not by itself a problem, though the LLM-vs-pipeline comparison would benefit from a more critical reading of those earlier results.\n\nWho this is for: computational social scientists and economists who want to use LLMs for structured annotation and need a realistic picture of what a frontier model can and cannot do. The paper deserves a serious referee, and I would accept it for review, but with requested revisions: define the accuracy metric, address the gold-standard/independence issue, add uncertainty quantification, and ideally benchmark against at least one open-weights model so the results are less tied to a single proprietary snapshot.","headline":"Solid, honestly reported benchmark of GPT-4o for economic narrative extraction, with a genuinely reusable codebook; the headline human–model gap is plausible but its magnitude is not fully trustworthy because the human baseline is scored against a gold standard built by the same three experts.","tokens_in":29894,"tokens_out":2016,"would_cite":true,"duration_ms":23446,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that GPT-4o can extract economic narratives from news text in a structured format, but that the model falls short of expert-level performance on complex documents, reaching only 44% accuracy against a gold standard where…","keywords":["economic narratives","large language models","GPT-4o","narrative extraction","inflation","few-shot learning","chain-of-thought prompting","content analysis"],"falsifier":"Re-annotate the 80 test documents with a fresh panel of experts who use only the written codebook and were not part of the consensus sessions, then compute their accuracy against the published gold standard. If their accuracy is similar to GPT-4o's (around 44%) rather than to the original annotators' 67–74%, the human-model gap is substantially an artifact of consensus construction; if fresh experts land at 67–74%, the gap reflects a genuine shortfall in the model.","tokens_in":28884,"feed_emoji":"📈","tokens_out":5286,"duration_ms":45248,"temperature":0.7,"pith_summary":"The paper asks whether a general-purpose large language model can replace multi-stage NLP pipelines previously used to extract economic narratives from text, and whether the model's output can match expert human coders. Using a corpus of Wall Street Journal and New York Times articles about inflation, the authors define an economic narrative as a causal link between two events, build a detailed codebook for human annotation, and create a gold standard through consensus among three expert annotators. They then prompt GPT-4o with a few-shot chain-of-thought prompt and compare its structured extractions to that gold standard. The central result is that GPT-4o produces valid, machine-readable narratives and follows the codebook well, but it underperforms experts on complex documents: 44% accuracy against 67–74% for individual experts, with more than twice the major-deviation rate. A sympathetic reading is that LLM-based extraction is feasible for scalable narrative measurement, provided human validation remains part of the workflow.","feed_headline":"GPT-4o trails expert coders at spotting inflation narratives: 44% vs 74%","feed_subtitle":"LLM-based narrative coding is feasible but needs human oversight for complex causal stories","key_machinery":"The central machinery is a two-part extraction procedure: an operational definition of an economic narrative as a causal connection between two temporally and semantically distinct events, expressed in the form 'A causes B' or 'A is caused by B'; and a few-shot chain-of-thought prompt for GPT-4o that guides the model through five discrete steps—focused excerpt, sequence of interest, causal restatement, coreference resolution, and event rephrasing—before checking for further narratives. The gold standard is built from consensus discussions among three expert annotators using a detailed codebook, and deviations are classified as major or minor. This machinery lets the authors compare model output and human deviations against a common benchmark and enables post-processing such as event decomposition, topic clustering, and valence assignment.","core_discovery":"The paper's central claim is that a state-of-the-art instruction-tuned LLM, prompted with a concise codebook plus seven hand-coded input-output examples and an explicit chain-of-thought procedure, can extract economic narratives from newspaper excerpts in a structured, aggregable form, but that this capability falls short of expert-level performance. On the 80-document test set, GPT-4o achieved 44% accuracy against the gold standard, while the three expert annotators scored 72%, 74%, and 67%; the model produced 1.25 major deviations per document versus 0.35–0.49 for the experts. The model also shows a systematic bias in narrative density, averaging 2.32 narratives per document with a standard deviation of 1.15 versus 2.22–2.36 with standard deviations near 2 for humans, reflecting a tendency to over-identify narratives in sparse texts and under-identify in dense ones, particularly forked and chained causal structures. The authors interpret these results as evidence that LLMs are promising tools for scaling narrative research, but that expert judgment remains necessary for reliable annotation.","pith_inferences":["If the consensus gold standard embeds the original annotators' interpretive preferences, the model's relative gap might narrow against a majority-vote or independently derived gold standard; this could be tested by comparing GPT-4o to a fresh panel's labels.","The same prompting and evaluation approach could be applied to other economically salient topics—labor markets, housing, climate—where narrative density and structural complexity might differ.","The observed inductive bias toward an average narrative count suggests that prompting strategies that explicitly vary the expected number of narratives (e.g., providing a per-document rarity prior or using a 'reject if none' option) could reduce both false positives in sparse texts and false negatives in dense ones.","A testable extension is to measure whether adding decomposed fork/chain training signals in few-shots—rather than just instructions—improves the model's handling of complex causal structures."],"forward_implications":["LLM-based narrative extraction can be deployed as a scalable screening tool, but every output should be reviewed by a human coder for complex documents.","Few-shot chain-of-thought prompting is an effective format for translating a human codebook into an LLM prompt for structured annotation tasks in economics.","The model's tendency to compress forked or chained narratives into single statements means downstream counts of narrative density will be biased unless post-processing explicitly forks compound events.","The aggregation pipeline of topic-valence clustering can turn raw extractions into recurring macro narrative arcs, such as 'loose monetary policy causes rising inflation'.","Narrative extraction remains inherently subjective even among experts, so accuracy metrics should be interpreted relative to expert disagreement, not as absolute truth."],"supporting_citations":[{"why":"Supplies the initial concept of economic narratives as stories that shape expectations and spread through the public.","marker":"Shiller (2017)"},{"why":"Frames narratives as simplified causal models (DAGs), grounding the paper's requirement that a narrative assert a causal link.","marker":"Eliaz and Spiegler (2020)"},{"why":"Provides the empirical model of backward-looking causal inflation narratives that the paper's definition and corpus are built to capture.","marker":"Andre et al. (2024)"},{"why":"Defines the prior SRL-based pipeline (RELATIO) that the paper argues falls short because it omits causality, motivating the LLM approach.","marker":"Ash et al. (2024)"},{"why":"Shows the error-cascade problem of multi-component pipelines, the key motivation for an integrated LLM.","marker":"Lange et al. (2022a)"},{"why":"Supplies the chain-of-thought prompting mechanism that the paper adapts into its five-step extraction procedure.","marker":"Wei et al. (2022)"},{"why":"Documents the GPT-4o model used as the test subject.","marker":"OpenAI et al. (2024b)"},{"why":"Justifies the consensus-based gold-standard procedure as more valid than majority voting.","marker":"Burla et al. (2008)"}],"fun_headline_variants":["GPT-4o scores 44% on inflation narratives, experts 74%","LLM narrative mining lags expert coders on inflation stories","GPT-4o identifies inflation narratives but trails expert accuracy","Inflation narrative extraction: GPT-4o vs experts, 44% to 74%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The gold standard is produced through consensus sessions in which the same three experts who are later evaluated against it discuss and decide which narratives count, and the paper assumes this consensus yields an unbiased, valid measure of the true narratives.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4o scores 44% on inflation narratives, experts 74%","LLM narrative mining lags expert coders on inflation stories","GPT-4o identifies inflation narratives but trails expert accuracy","Inflation narrative extraction: GPT-4o vs experts, 44% to 74%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1471,"prompt_tokens":963,"completion_tokens":508,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":427}},"tokens_in":579,"tokens_out":508,"duration_ms":4837,"temperature":1.0,"reasoning_tokens":427,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:44:45.494749+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate the 80 test documents with a fresh panel of experts who use only the written codebook and were not part of the consensus sessions, then compute their accuracy against the published gold standard. If their accuracy is similar to GPT-4o's (around 44%) rather than to the original annotators' 67–74%, the human-model gap is substantially an artifact of consensus construction; if fresh experts land at 67–74%, the gap reflects a genuine shortfall in the model.","supporting_citations":[],"review_version":2}