{"id":"0cf1514e-8373-4b94-bb3d-db72b571de0a","arxiv_id":"2412.15375","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"GPT-4 can extract the four concepts of a proportional metaphoric analogy from short literary texts with 77% frame-wise head-noun accuracy, but full quadruple accuracy is 61%.","lead":"This paper creates a small expert-labeled dataset of 204 literary metaphors and asks large language models to recover the four concepts that make up each metaphor's analogy. The best model, GPT-4, finds the right head noun for about 77 percent of individual slots, but only 61 percent of complete four-part analogies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-annotator gold labels with weak quadruple-level agreement undermine the reliability of the reported full-quadruple accuracy, especially the GPT-4 Q=0.61 score.","rationale":"The reader's weakest assumption—that the gold labels are correct enough to serve as ground truth—is the most load-bearing concern for the paper's central claim. The reported accuracies (0.77 frame-wise, 0.61 full-quadruple) are only as trustworthy as the labels they are measured against. The paper's own IAA data show low agreement on complete quadruples, which is exactly the metric used for the headline Q score. This is not a fatal flaw because per-frame agreement is moderate to substantial and A5 is an expert with high agreement with the majority on the 20-instance sample, but it introduces enough uncertainty that the absolute numbers cannot be taken at face value. A concrete re-annotation study would directly quantify how much the model's accuracy depends on the choice of annotator. If the variance is small, the concern is resolved; if large, the central quantitative claim requires qualification. The paper's conclusion that model performance is 'in line with human annotators' is also unsupported because no human accuracy on the test set is reported; the IAA is not a performance measure. This reinforces the need for the conditional verdict. I considered memorization/leakage of famous literary quotes as an alternative concern, but the dataset includes less famous 18th-century texts and the model still performs well, and the paper acknowledges the possibility without fully addressing it. The gold-label reliability issue is more direct and more easily testable.","tokens_in":14847,"tokens_out":7254,"duration_ms":66712,"concrete_test":"Sample 50 instances from the 204-example dataset and have two new expert annotators independently label T1:T2::S1:S2 following the same instructions as the original study. Recompute GPT-4's frame-wise and full-quadruple accuracy against each annotator's labels and against the original gold. If full-quadruple accuracy varies by more than ±0.10 across label sets, or if the new annotators' agreement with the original gold is below kappa=0.5, the reported 0.61 Q score is not a stable estimator of model ability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim rests on a gold standard labeled almost entirely by one expert (A5), with only 20 instances cross-annotated for IAA. Table 2 shows pairwise Cohen's kappa on complete quadruples (column Q) ranging from 0.28 to 0.54, averaging near 0.40—weak agreement by conventional thresholds. Full-quadruple accuracies (e.g., GPT-4 Q=0.61 in Table 6) are computed against A5's labels, so it is unclear whether these scores reflect robust analogy-structuring ability or agreement with one particular interpretation. The Limitations section explicitly concedes that most instances receive a single annotation and that no manual evaluation of frame extraction was performed. Per-frame agreement is higher (0.57–0.87), which partially supports the frame-wise claim, but the headline 'full-quadruple' accuracy is directly threatened. Without multi-annotator validation of the full dataset, the reported numbers are not stable under reasonable variation in expert judgment.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces a new task and dataset for extracting four-term metaphoric analogies of the form T1:T2::S1:S2 from short literary texts. The dataset contains 204 sourced instances labeled by an expert annotator (A5) and double-checked by a second author, with a 20-instance inter-annotator agreement study. The authors evaluate several out-of-the-box instructed LLMs (GPT-3.5, GPT-4, Llama-3-70B, Mixtral 8x7B/8x22B) in a few-shot prompting setup, measuring both extraction of the explicit frames and generation of implicit terms. The central results are that GPT-4 achieves the highest frame-wise accuracy (0.77 lemmatized head-noun match) and full-quadruple accuracy (0.61), and that GPT-4's generated implicit terms receive an average human relevance score of 1.21 out of 2, outperforming Mixtral 8x22B. The paper also analyzes performance by input/output frame, number of implicit terms, and sentence length.","tokens_in":14969,"tokens_out":3314,"duration_ms":30280,"significance":"If the main results hold, this is a useful contribution: it introduces a novel, publicly released benchmark for a task that has not been systematically evaluated before, and it provides evidence that large language models can recover the structure of proportional metaphoric analogies from natural text without task-specific training, including inferring implicit terms. The evaluation protocol is a clean, independent prompt-and-measure setup with no fitted parameters, and the authors release the dataset and scripts, which supports reproducibility. However, the load-bearing assumptions about the quality of the gold standard and the validity of the automatic metric need to be addressed before the quantitative claims can be taken at face value.","major_comments":[{"comment":"The headline full-quadruple accuracies (e.g., GPT-4 Q=0.61 in Table 6) are computed against gold labels produced almost entirely by one expert annotator (A5) on 204 instances, with only 20 instances independently annotated for IAA. The pairwise Cohen's kappa on complete quadruples (column Q in Table 2) is weak, ranging from 0.28 to 0.54 (average ~0.40). The authors acknowledge this in the Limitations section, but the central claim of the paper rests on these numbers. To make the claim robust, the manuscript should either provide multi-annotator validation for a larger portion of the dataset or report model accuracies against alternative gold labels (e.g., the majority vote or adjudicated labels on the IAA set) to show that the relative model rankings and the GPT-4 Q score are stable under reasonable annotator variation. As it stands, the Q=0.61 figure may reflect agreement with one particular interpretation rather than a robust structural ability.","section":null},{"comment":"All accuracy numbers use the lemmatized head-noun matching metric, but the paper states explicitly that no manual evaluation of analogy frame extraction was performed. The authors describe the metric as a 'lower bound' on true accuracy, but no evidence is provided to support this direction or to quantify the gap. This is a load-bearing issue because the model comparisons in Tables 4 and 6 rely entirely on this automatic metric. The manuscript should include a human evaluation of a sample of model extractions (for example, the same 50-sentence protocol used for implicit-term evaluation) to estimate the metric's precision and recall relative to human judgment, and to confirm that the ranking of models is not an artifact of the matching function.","section":null},{"comment":"The comparison between GPT-4 and Mixtral 8x22B is presented as a close call (0.77 vs 0.75 frame-wise; 0.61 vs 0.58 Q), but no statistical significance testing or confidence intervals are reported. Only GPT-3.5 was run on 10 batches (std=0.03); the other models were averaged over only 3 batches. The claim that GPT-4 is the best model, and the more general ordering of models, would be more credible if the authors reported variance across batches or instances, or ran a paired significance test. This is specific to the comparison in Table 4 and Table 6.","section":null},{"comment":"The conclusion states that the performance of the models is 'remarkably high, in line with human annotators.' This comparison is not supported by the reported metrics: human agreement is measured as Cohen's kappa (chance-corrected) in Table 2, while model performance is reported as raw accuracy in Tables 4 and 6. These numbers are not directly comparable. To support the claim, the authors should compute a model-human agreement measure on a common set of instances (e.g., the 20-instance IAA set) or rephrase the claim to avoid implying direct comparability.","section":null}],"minor_comments":[{"comment":"The phrase 'even including the frames that are have implicit values' should read 'even including the frames that have implicit values.'","section":null},{"comment":"The caption reads 'with on example sentence'; this should be 'with one example sentence.'","section":null},{"comment":"The column header 'In Acc. for fields' is unclear; consider using 'Input' and 'Frame accuracy' explicitly.","section":null},{"comment":"The publisher name 'John Hopking University Press' should be 'Johns Hopkins University Press.'","section":null}],"recommendation":"major_revision","confidential_remarks":"The central result is plausible and the dataset is a useful resource, but the weak IAA on complete quadruples is a substantive concern. I would encourage the editor to weigh whether the authors can reasonably strengthen the gold-standard validation within a revision; if not, the quantitative model-comparison claims may need to be presented more cautiously as preliminary. The paper is within the scope of the journal, but a conditional acceptance with the required revisions seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers what it says: a new 204-instance dataset of literary proportional metaphors with four-slot (T1:T2::S1:S2) annotations, including implicit terms, and a prompt-based evaluation of five LLMs. The task formulation is clean, the dataset is sourced from established literary collections, and the authors publish data and code. They also report a two-annotator human rating of generated implicit terms with substantial agreement (Spearman 0.7), which is a genuine plus. The head-noun matching metric is honestly described as a lower bound, and the appendices give enough detail to reproduce the prompts and the IAA procedure.\n\nThe main soft spot is exactly where the stress-test note points: the gold labels for the full dataset come almost entirely from one expert (A5), with only 20 instances cross-annotated. Pairwise kappa on complete quadruples is weak (Q=0.28–0.54, average ~0.40), so the reported full-quadruple accuracies, like GPT-4's Q=0.61, are not stable under plausible variation in expert judgment. That said, the per-frame agreement is considerably better (T2, S1, S2 kappas mostly in the 0.57–0.87 range), and A5's labels align with the majority vote 95% of the time. This means the frame-wise results are reasonably solid, while the full-quadruple numbers should be read with caution. The paper's Limitations section concedes most of this, which is to its credit.\n\nA separate, clearer flaw is in the conclusion: \"performance ... in line with human annotators.\" That comparison is not actually established, because human agreement is reported as kappa while model performance is accuracy on a different (single-gold) reference. The two metrics are not directly comparable, and the statement oversells what the data show. The stress-test concern about the single annotator is real, but it is not fatal; it primarily weakens the strong form of the claim (full-quadruple accuracy) rather than the core finding that LLMs can structure these analogies reasonably well.\n\nThe in-context examples are sampled from the same dataset, which is a mild contamination worry, but the paper's own variance check over batches shows small standard deviation. The dataset is small (204 instances), so subset analyses (e.g., two-implicit-term cases, n=10) are underpowered; the authors note this.\n\nWho is this for? Anyone building metaphor or analogy extraction benchmarks, and NLP researchers interested in probing LLM reasoning about figurative language. It deserves a serious referee. I would recommend sending it to review, with the request that the authors either expand the IAA set or soften the human-comparison claim, and that they report model accuracy against multiple annotator golds where possible.","headline":"A useful, honest small benchmark for metaphoric analogy extraction, with a real but manageable annotation-reliability caveat and one overreach in the conclusion.","tokens_in":15584,"tokens_out":2176,"would_cite":true,"duration_ms":22141,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-4 recovers 77% of literary metaphor analogy frames","keywords":["metaphor extraction","analogical reasoning","proportional analogy","large language models","literary texts","implicit concept inference","dataset annotation","GPT-4"],"falsifier":"Have a panel of fresh experts independently label a sample of the same texts, then compare the model-extracted quadruples to the panel's majority vote; if model agreement with the new panel is no better than the original annotators' agreement with each other, the reported accuracies would reflect task ambiguity rather than reliable extraction of structure.","tokens_in":14606,"feed_emoji":"🔍","tokens_out":11667,"duration_ms":85445,"temperature":0.7,"pith_summary":"The paper introduces the task of extracting four-term proportional analogies ($T_1:T_2::S_1:S_2$) from short literary texts, and builds a 204-instance dataset of sourced metaphors annotated by experts. It claims that out-of-the-box large language models can recover this structure, with GPT-4 reaching 77% frame-wise accuracy and 61% full-quadruple accuracy, and can generate relevant concepts for implicit terms with a mean human relevance score of 1.21 out of 2. If correct, this would let unstructured metaphors be converted into structured analogical mappings automatically, reducing the need for expert annotation and enabling large-scale metaphor knowledge bases.","feed_headline":"GPT-4 recovers 77% of literary metaphor analogy frames","feed_subtitle":"Out-of-the-box LLMs can extract the four parts of a metaphor and infer unstated concepts.","key_machinery":"The load-bearing object is the four-slot analogy frame $T_1:T_2::S_1:S_2$, which the task forces the model to fill by outputting one noun phrase per slot, with $T_1$ and $T_2$ in the target domain (the topic) and $S_1$ and $S_2$ in the source domain (the metaphoric image). The dataset supplies gold frames for 204 sourced literary metaphors, about half with at least one implicit slot. The evaluation counts an answer correct when the lemmatized head noun of the model's phrase matches the gold concept, ignoring modifier variation and plural/singular differences. The single-step prompt provides seven in-context examples and asks the model to complete the remaining slots, testing raw analogical extraction without fine-tuning.","core_discovery":"Given a short literary text containing a metaphor and one of its four concept slots, the paper claims that recent instructed LLMs can identify the three remaining concepts forming the proportional analogy $T_1:T_2::S_1:S_2$, distinguishing target-domain from source-domain terms. On the new 204-instance dataset, GPT-4 achieves 0.77 frame-wise accuracy and 0.61 full-quadruple accuracy under lemmatized head-noun matching, with Mixtral 8*22 close behind at 0.75 and 0.58. When a slot is implicit in the text, GPT-4 generates a concept that human raters score on average 1.21 out of 2, meaning relevant but imperfect inference. Accuracy stays stable as the noun count in the text rises, which the paper reads as a sign the method can extend to longer passages.","pith_inferences":["A likely next step is full open extraction, where the model must also locate the anchor concept itself rather than being given one slot; the paper's experimental design does not test this, but the high frame-wise scores suggest it could work on curated short texts.","The weak inter-annotator agreement on complete quadruples (kappa 0.28-0.54) implies that the 0.61 full-quadruple accuracy for GPT-4 may be near the level of agreement among competent humans, so absolute comparisons to human agreement would be a more informative benchmark than raw accuracy.","The same prompting recipe could be tried on non-literary metaphor-rich text such as news headlines or social media posts, where surface forms are less curated; the paper's finding that accuracy is robust to noun count gives some reason to expect transfer.","The dataset deliberately excludes nested and multi-relation metaphors, so the current result covers only single-relation proportional analogies; a broader claim about metaphor extraction in general would require extending the frame to triples or multi-relation mappings."],"forward_implications":["The single-step prompting approach can convert any short text containing a metaphor into a structured four-term mapping in one model call, without domain-specific training.","Because accuracy does not decline as the number of nouns in the sentence increases, the same method is a candidate for scaling to longer texts such as paragraphs or pages.","The released 204-instance dataset and its evaluation protocol give future work a benchmark for metaphor extraction that goes beyond single-pair tagging.","The gap between frame-blind and frame-wise scores suggests a two-step pipeline (extract terms, then assign frames) could close part of the remaining error, a direction the paper explicitly suggests."],"supporting_citations":[{"why":"This theory frames metaphors as a subset of analogies, which justifies casting the task as proportional analogy extraction.","marker":"(Bowdle and Gentner, 2005)"},{"why":"This collection supplies 60 dataset instances from the Metaphor of Mind repository.","marker":"(Pasanek, 2015)"},{"why":"This compilation supplies 124 dataset instances, the largest share of the evaluation data.","marker":"(Grothe, 2008)"},{"why":"This book of similes provides additional sourced texts for the dataset.","marker":"(Baldwin and Paris, 1982)"},{"why":"This work shows that instructed LLMs can solve a broad range of analogies when prompted, motivating the single-step prompting design.","marker":"(Webb et al., 2023)"},{"why":"This work extracts metaphor mappings with GPT-3 from annotated source/target data, serving as the direct predecessor for LLM-based mapping extraction.","marker":"(Wachowiak and Gromann, 2023)"},{"why":"This corpus of metaphoric analogies in quadruple form is the closest existing resource the new dataset extends.","marker":"(Czinczoll et al., 2022)"}],"fun_headline_variants":["LLMs decode metaphoric analogies from literary texts","GPT-4 maps metaphor components in prose automatically","New dataset shows LLMs infer missing metaphor parts","AI extracts four-part analogies from literature with 77% accuracy","LLMs match human skill at spotting metaphor structure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The assumption that the gold labels, produced mostly by one expert and double-checked by a second author, are correct enough to score model outputs, even though the best pairwise annotator agreement on a complete quadruple is only a Cohen's kappa of 0.54.","fun_headline_variants_meta":{"raw":{"variants":["LLMs decode metaphoric analogies from literary texts","GPT-4 maps metaphor components in prose automatically","New dataset shows LLMs infer missing metaphor parts","AI extracts four-part analogies from literature with 77% accuracy","LLMs match human skill at spotting metaphor structure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000137,"raw_usage":{"total_tokens":1107,"prompt_tokens":858,"completion_tokens":249,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":474,"completion_tokens_details":{"reasoning_tokens":173}},"tokens_in":474,"tokens_out":249,"duration_ms":3151,"temperature":1.0,"reasoning_tokens":173,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:28:25.275707+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a panel of fresh experts independently label a sample of the same texts, then compare the model-extracted quadruples to the panel's majority vote; if model agreement with the new panel is no better than the original annotators' agreement with each other, the reported accuracies would reflect task ambiguity rather than reliable extraction of structure.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This theory frames metaphors as a subset of analogies, which justifies casting the task as proportional analogy extraction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This collection supplies 60 dataset instances from the Metaphor of Mind repository."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This compilation supplies 124 dataset instances, the largest share of the evaluation data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This book of similes provides additional sourced texts for the dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This corpus of metaphoric analogies in quadruple form is the closest existing resource the new dataset extends."}],"review_version":1}