Pith. sign in

REVIEW 4 major objections 4 minor 87 references

Historical analogy for foresight is a causal inference problem: analogies must be matched on hidden structural positions, not surface descriptions.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 04:42 UTC pith:IFRCHD4R

load-bearing objection Genuinely new benchmark and a structural-analogy agent that beats baselines, but the central foresight claim rests on an unmeasured transfer assumption and a rubric that partly bakes in the agent's design. the 4 major comments →

arxiv 2607.13602 v1 pith:IFRCHD4R submitted 2026-07-15 cs.CL cs.LGstat.ML

Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis

classification cs.CL cs.LGstat.ML
keywords analogical deep researchhistorical analogyforesight analysiscausal structureLLM agentsstructural alignmenthidden factor inferenceADR-bench
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that using historical analogies for foresight is fundamentally a causal problem: an analogy is useful only when the two events share underlying structural positions in their causal graphs, not when their descriptions look alike. It proves that surface-level matching cannot identify hidden structural positions, no matter how many analogies are retrieved, and that independent analogies confirming the same position multiply the evidence that the position is real. On that basis it proposes two principles—retrieve by mechanism, and require at least two independent analogies per position—and an agent, CANA, that decomposes events into preconditions, mechanisms, and outcomes and iteratively confirms analogies. On its benchmark, CANA raises cross-analogy structural claims from at most about one per event to 3.2–16.6 per event and hidden-factor identifications from 0 of 42 to as many as 16 of 42.

Core claim

The paper's central claim is that historical-analogy retrieval for foresight should be treated as causal inference over structural positions. Its surface non-identifiability theorem shows that two worlds with identical surface observations but different hidden positions are indistinguishable to any surface-level method, and even an infinite number of surface-matched analogies cannot reduce foresight risk below half the separation gap between the worlds. Its cross-analogy confirmation theorem shows that when analogies independently confirm a position, each confirmation multiplies the posterior odds by q/p; under the paper's calibrated values (prior 0.5, q=1, p=0.2), two confirmations suffice

What carries the argument

The load-bearing device is the distinction between an event's descriptive representation D(E) and its mechanistic representation M(E), a directed causal graph whose factors can be assigned structural positions (for example trigger, enabler, amplifier, mediator, outcome). Retrieval operates on M(E) by aligning positions across events. The second device is a Bayes-factor confirmation rule: independent analogies that confirm the same position update posterior odds multiplicatively, which yields a required number of confirmations per position (two in the calibrated regime). CANA operationalizes both by decomposing events into preconditions, temporal chains, mechanisms, and outcomes and by reflec

Load-bearing premise

The load-bearing premise is that if two events occupy the same structural position and one has progressed further, the source event's observed trajectory is a good prediction of the target's future trajectory, with bounded error—this transfer is assumed, not measured.

What would settle it

Run CANA on the five forward events using only pre-cutoff information, without giving it the oracle analogies, then check its L4 hidden-factor predictions against what actually happened after each cutoff. If the hidden-factor hit rate does not exceed a surface-matching baseline, or if the aligned analogies' trajectory forecasts are no closer to realized outcomes than randomly chosen historical events, the mechanism-transfer assumption fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Because surface matching is information-theoretically blind to hidden positions, foresight reports that do not attempt structural alignment cannot be expected to uncover hidden factors, regardless of model scale or retrieval budget.
  • Two independent structural confirmations per position are enough, under the paper's calibrated regime, to treat a position as necessary rather than coincidental.
  • An agent that adds a structural analogy brief to a general deep-research pipeline can match or exceed commercial deep-research agents even with a weaker backbone.
  • Structural decomposition of events into roles, rather than topical summaries, changes which analogies are retrieved and reduces self-analogy and surface-match errors.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Inference: The paper implies that analogy retrieval quality should be measured by posterior coverage of hidden positions, not by similarity rankings; a practical extension is to have agents output position-level coverage and stop at the two-confirmation threshold.
  • Inference: The transferability assumption is the empirical crux: if the benchmark's oracle analogies already presuppose that aligned positions transfer, then the method's gains on hidden-factor hits may partly reflect benchmark construction; a stronger test would let the agent discover analogies without oracle hints and then score predicted hidden factors against actual post-cutoff outcomes.
  • Inference: The same two-principle recipe could generalize to other partial-observation domains, such as medical case comparison or geopolitical risk, where multiple historical cases with different surface features share structural roles.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper defines a new task, Analogical Deep Research (ADR), in which LLM agents must retrieve historical analogies and integrate them for foresight analysis. It argues that ADR is inherently causal: analogy matching should align mechanistic structure M(E) rather than surface descriptions D(E). A short theory section states a surface non-identifiability theorem (Theorem 4/13), a cross-analogy confirmation theorem (Theorem 5/16), and derives two principles: retrieve on mechanisms, and require at least two independent confirming analogies. The authors introduce CANA, a framework that decomposes events into preconditions, temporal chains, mechanisms, and outcomes, and iteratively refines analogy candidates using structural feedback. They also construct ADR-bench with 15 events (10 historical, 5 forward) and compare commercial deep-research agents, vanilla MiroFlow, and MiroFlow+CANA. Results on the Li et al. (2025) analogy generation benchmark and ADR-bench show consistent gains for CANA across several LLM backbones.

Significance. If the central claims hold, this is a useful step: the paper gives a concrete task formulation, a transparent theoretical argument for why surface matching is insufficient, and an agent design that consistently outperforms strong deep-research baselines across multiple backbones. The empirical gains in Tables 1–3 are large and coherent, and the connection between analogy retrieval and causal identifiability is a good framing. However, the load-bearing claims are not yet independently validated. The Mechanism Transfer assumption (Assumption 3/12) is never tested, and the ADR-bench evaluation relies on an author-built LLM-judged rubric, single runs without confidence intervals, and forward events whose outcome-based FQS is scored while outcomes are still unfolding. The theoretical 'two analogies suffice' result also depends on hand-set probabilities that are not calibrated. These are fixable with additional experiments and analysis; with the current evidence, the contribution is promising but not fully established.

major comments (4)
  1. [§3.1 / Appendix C.2, Assumption 12 and Corollary 15] The scientific value of historical analogies for foresight rests entirely on the Mechanism Transfer assumption: if source and target factors occupy the same structural position and the source has progressed further, then TV(bP_s^E_S, P_s^T) ≤ α_s^tr. Corollary 15 states that mechanism matching strictly dominates surface matching only when Σ μ_s α_s^tr < Σ μ_s Δ_s, yet neither α_s^tr nor Δ_s is ever estimated or bounded empirically. No experiment tests whether dynamics at an aligned structural position actually transfer to the target. This is load-bearing because ADR-bench's oracle analogies are selected on the basis of documented shared mechanisms, and the L4 rubric credits only cross-analogy inference from ≥2 events — both presuppose the same transferability. Consequently Table 3 cannot independently validate the transfer assumption. Please add a direct transfer test (e.g., using histor
  2. [§3.3 / Theorem 16 and Appendix D] The claim that 'two independent analogies suffice' is derived from the hand-set values π=0.5, q=1, p=0.2, δ=0.05, and the assumption of conditional independence across analogies. No calibration of p, q, or the independence assumption is provided; LLM-generated analogies are not independent, and q/p is never measured. In addition, the ADR-bench L4 rubric explicitly requires citing ≥2 analogies for full credit, and L3-S requires ≥2 events. CANA is specifically designed to produce ≥2 crossing analogies, so a substantial part of its L3-S/L4 advantage over baselines is by construction. The paper should report an ablation in which the same evidence is scored without the two-analogy requirement, or calibrate p and q on a held-out set, to show that the advantage is not merely rubric-induced.
  3. [§5.2, Table 3] The headline ADR-bench results are based on a single run per configuration (no confidence intervals or variance), only 15 events, and an LLM judge (Claude Sonnet 4.5) using a rubric designed by the authors. For the five forward events, FQS is scored against outcomes that are still unfolding: the FQS prompt instructs the judge to compare against 'what actually happened', but those events have no settled outcome. This makes the FQS numbers unverifiable for the forward split. Please report multiple runs with standard deviations, provide human–LLM agreement on claim decomposition and scoring, and either restrict FQS to historical events or defer forward-event FQS until outcomes resolve.
  4. [§5.2, 'Connecting to theory'] The paper states that 'HF@L4=0/42 for all commercial agents validates Theorem 4.' This overstates what the theorem shows. Theorem 4 is a worst-case information-theoretic result conditioned on identical surface observations; commercial DR agents' failure to find hidden factors may be due to retrieval, planning, prompt design, or evaluation granularity rather than surface-level identifiability. A more direct test would compare a surface-level retriever against a mechanism-aligned retriever on the same surface observations, or vary the amount of surface evidence systematically. Please temper this claim or add such a controlled experiment.
minor comments (4)
  1. [References] References [9] and [10] are duplicates: both are Clement and Gentner, 'Systematicity as a selection constraint in analogical mapping.' Please merge.
  2. [§4.1] The phrase 'As mentioned in Def. 3.1' refers to a definition from the main text but the actual formal definition is Definition 6 in Appendix C.1. Also, 'MiroFlow' appears with inconsistent markup across the paper.
  3. [§1 / §6] The contributions state 'more than 10% performance' in analogy retrieval, while the abstract and conclusion say 'up to 10% improvements.' Please clarify the exact setting and metric supporting the 10% figure; Tables 1 and 2 show different relative gains depending on backbone and rubric.
  4. [Appendix B / Evaluation] The paper candidly acknowledges in 'Limitation and Future Works' that the benchmark scale is limited and evaluation relies on LLMs. This is valuable, but the abstract and conclusion should carry a corresponding caveat so that readers are not misled about the strength of the ADR-bench evidence.

Circularity Check

2 steps flagged

CANA's 'two analogies suffice' result follows from hand-set q=1, p=0.2, and ADR-bench's L4 rubric defines success as citing ≥2 analogies—the exact behavior CANA is prompted to emit; partial circularity.

specific steps
  1. self definitional [Sec. 3.2 (Evaluation design), Appendix D (Call 3 rubric), Sec. 4.1 (Structural reflective generation)]
    "L4 (Hidden Factor Inference): An inference about a SPECIFIC, CURRENTLY UNRECOGNIZED factor in the current situation, justified by cross-analogy evidence from ≥2 historical events. ... A claim can only be L3-S or L4 if it references MULTIPLE (≥2) historical events. ... iteratively retrieve cross-confirming analogies until each hidden position is supported by ≥ 2 independent analogies, realizing Principle 2."

    ADR-bench defines the headline metric L4 as a claim justified by ≥2 analogies, and L3-S likewise requires ≥2 events with a structural role. CANA's core loop is prompted to collect ≥2 confirming analogies per position before producing the Structural Analogy Brief. A CANA report that follows its own prompt automatically satisfies the citation-count condition for L3-S/L4, while agents not given that instruction cannot. The Table 3 gap therefore partly measures the rubric matching the method's design, not an independent confirmation of the cross-analogy confirmation principle.

  2. fitted input called prediction [Sec. 3.3 (Theoretical Discussions) and Theorem 16 (Appendix C.2); Sec. 4.1]
    "For example, given π_s=0.5, q_s=1, p_s=0.2, δ=0.05, one requires 2 independent analogies per position to distinguish the structural necessity. ... Under the calibrated regime (π_s=0.5, q_s=1, p_s=0.2), this gives n*_s=2 for δ=0.05."

    n*_s=2 is the closed-form solution of the Bayes-factor inequality evaluated at hand-chosen q_s=1, p_s=0.2, π_s=0.5, δ=0.05; no data calibrate these values. The paper labels them a 'calibrated regime', converts the arithmetic consequence into Principle 2 (K*_s ≥ 2), and hardwires it into CANA's stopping criterion ('until each hidden position is supported by ≥ 2 independent analogies'). The claim that Table 3 'validates Theorem 5' is therefore partly circular: the system was built to satisfy the theorem's hand-set threshold, and the rubric rewards reaching it.

full rationale

The formal theorems are not themselves circular: Theorem 13/4 is a standard indistinguishability argument from O∞_D(w0)=O∞_D(w1), and Theorem 16/5 is a correct Bayes-factor calculation conditional on assumed q_k,s, p_k,s and conditional independence. Corollary 15 is also valid conditionally on Assumption 12. The circularity enters the empirical validation pipeline. The 'two analogies suffice' threshold is an arithmetic consequence of q=1, p=0.2, π=0.5, δ=0.05; CANA is then instructed to retrieve until each position has ≥2 analogies; and the ADR-bench rubric defines L4 as a claim justified by ≥2 historical events and denies L3-S/L4 to claims with fewer than 2 events. Hence a substantial part of CANA's L3-S/L4 advantage is by construction. This is partial, not total: CANA also improves analogy generation on the pre-existing Li et al. benchmark (Tables 1-2), and hidden-factor hits require matching the annotated reference set, which provides some independent signal. Assumption 12's transfer bound is load-bearing but acknowledged as an assumption; its lack of empirical estimation is a correctness risk rather than a circular reduction. No load-bearing self-citation chain was found. Score 5.

Axiom & Free-Parameter Ledger

3 free parameters · 6 axioms · 1 invented entities

The central contribution adds one invented abstraction (structural positions in a shared causal pattern), assumes transferability of dynamics across aligned events, assumes conditional independence of confirmations, and treats LLM decomposition and LLM judging as faithful proxies. The free parameters p, q, π, δ are hand-set to produce the 'two analogies' rule that is then encoded in the evaluation rubric.

free parameters (3)
  • coincidence probability p_s in Theorem 5/16 = 0.2
    Chosen by hand in the 'calibrated regime' to make n*=2; no empirical estimation. This value enters the two-analogy principle that CANA enforces.
  • confirmation probability q_s = 1
    Assumed perfect confirmation when a position is active; yields per-analogy Bayes factor 5. No empirical basis provided.
  • prior π_s and target error δ = 0.5 and 0.05
    Used together with p and q to derive n*=2; standard choices but still hand-set.
axioms (6)
  • domain assumption Mechanism Transfer (Assumption 3/12): aligned factors at the same structural position have bounded TV transfer error α_s.
    Load-bearing premise for the entire ADR enterprise; no estimation or empirical test of α_s is given.
  • domain assumption Events instantiate a common abstract causal pattern via edge-preserving graph homomorphisms (Def. 7, Appendix C.1).
    Assumes a canonical alignment exists across domains; this is structure-mapping theory, not established for the 15 benchmark events.
  • domain assumption X_{k,s} are conditionally independent across analogies given Z_s (Theorem 5/16).
    Needed for the multiplicative Bayes factor and n*=2; in practice retrieved analogies are correlated, so independence is suspect.
  • ad hoc to paper LLM-extracted structural decomposition (preconditions, temporal chains, mechanisms, outcomes) faithfully recovers M(E).
    CANA's Principle 1 rests on this; no guarantee and no control verifies that the LLM decomposition is not another surface representation.
  • ad hoc to paper LLM judge (Claude Sonnet 4.5 / GPT-5.4) scores align with expert human judgment.
    All benchmark conclusions are mediated by LLM grading; no human calibration or inter-annotator agreement is reported.
  • domain assumption For forward events, hidden factors and eventual outcomes are sufficiently known/foreseeable to score FQS.
    Five events are 'currently unfolding', yet FQS is computed against known outcomes, which is temporally questionable.
invented entities (1)
  • Abstract causal pattern 𝔓 with structural positions S_𝔓 no independent evidence
    purpose: Formal backbone for aligning target and source events by shared causal roles; used in Def. 7 and in CANA's structural-decomposition prompts.
    The paper assumes events instantiate a common abstract pattern via graph homomorphism; no independent falsifiable handle is given beyond the authors' curated oracle analogies.

pith-pipeline@v1.3.0-alltime-deepseek · 35695 in / 15884 out tokens · 149194 ms · 2026-08-02T04:42:04.294225+00:00 · methodology

0 comments
read the original abstract

Systematic comparisons between current situations and structurally similar past events in the historical, i.e., historical analogies, is among the most powerful tools for foresight analysis. In this work, we present a new task called Analogical Deep Research (ADR) to Large Language Model (LLM) agents and construct the first ADR benchmark ADR-bench to study whether LLM agents are able to find and leverage historical analogies when doing foresight analysis. Our investigation reveals a key obstacle: LLM agents are poor at finding analogies because they match on surface features rather than underlying mechanisms. We argue that ADR is inherently a causal question as it requires understanding why the event occurred. Based on our theoretical analysis, we propose two principles required for ADR, including the mechanism alignment and cross-analogy confirmation. Built upon our theoretical results, we propose a new agentic framework called Causal Analogical Researcher (CANA) that guides LLMs to find and integrate historical analogies. CANA incorporates a simple yet effective structural decomposition representation, and integrates structural feedback for reflective improvements of historical analogy identification and integration. We show that CANA brings up to 10% improvements in historical analogy generation, and surpasses the state-of-the-art deep research agents in the ADR-bench. Case studies with the ongoing events confirm the effectiveness of CANA in leveraging historical analogies.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

87 extracted references · 1 canonical work pages

  1. [1]

    The making of an applied historian: Stage two.The Public Historian, 5 (2):21–46, 1983

    W Andrew Achenbaum. The making of an applied historian: Stage two.The Public Historian, 5 (2):21–46, 1983

  2. [2]

    Identification of partially observed linear causal models: Graphical conditions for the non-Gaussian and heterogeneous cases

    Jeffrey Adams, Niels Hansen, and Kun Zhang. Identification of partially observed linear causal models: Graphical conditions for the non-Gaussian and heterogeneous cases. InAdvances in Neural Information Processing Systems (NeurIPS), 2021

  3. [3]

    Claude Sonnet 4.5

    Anthropic. Claude Sonnet 4.5. https://www.anthropic.com/news/ claude-sonnet-4-5, September 2025. Large language model

  4. [4]

    Claude Sonnet 4.6 system card

    Anthropic. Claude Sonnet 4.6 system card. Technical report, Anthropic, February 2026. URL https://anthropic.com/claude-sonnet-4-6-system-card

  5. [5]

    Paul F. A. Bartha.By Parallel Reasoning: The Construction and Evaluation of Analogical Arguments. Oxford University Press, 2010

  6. [6]

    Weakly supervised causal representation learning

    Johann Brehmer, Pim de Haan, Phillip Lippe, and Taco Cohen. Weakly supervised causal representation learning. InAdvances in Neural Information Processing Systems (NeurIPS), 2022

  7. [7]

    Darren C. Brunk. Curing the Somalia syndrome: Analogy, foreign policy decision making, and the Rwandan genocide.Foreign Policy Analysis, 4(3):301–320, 2008

  8. [8]

    Spears, Derya Unutmaz, Kevin Weil, Steven Yin, and Nikita Zhivotovskiy

    Sébastien Bubeck, Christian Coester, Ronen Eldan, Timothy Gowers, Yin Tat Lee, Alexandru Lupsasca, Mehtaab Sawhney, Robert Scherrer, Mark Sellke, Brian K. Spears, Derya Unutmaz, Kevin Weil, Steven Yin, and Nikita Zhivotovskiy. Early science acceleration experiments with gpt-5.ArXiv, abs/2511.16072, 2025

  9. [10]

    Clement and Dedre Gentner

    Catherine A. Clement and Dedre Gentner. Systematicity as a selection constraint in analogical mapping.Cognitive Science, 15(1):89–132, 1991

  10. [11]

    Simon & Schuster, New York, 2017

    Ray Dalio.Principles: Life and Work. Simon & Schuster, New York, 2017. ISBN 9781501124020

  11. [12]

    How scientists really reason: Scientific reasoning in real-world laboratories

    Kevin Dunbar. How scientists really reason: Scientific reasoning in real-world laboratories. In Robert J. Sternberg and Janet E. Davidson, editors,The Nature of Insight, pages 365–395. MIT Press, 1995

  12. [13]

    Forbus, and Dedre Gentner

    Brian Falkenhainer, Kenneth D. Forbus, and Dedre Gentner. The structure-mapping engine: Algorithm and examples.Artificial Intelligence, 41(1):1–63, 1989

  13. [14]

    Structure-mapping: A theoretical framework for analogy.Cognitive Science, 7 (2):155–170, 1983

    Dedre Gentner. Structure-mapping: A theoretical framework for analogy.Cognitive Science, 7 (2):155–170, 1983

  14. [15]

    Looking forward to the past: An interdisciplinary discussion on the use of historical analogies and their effects.Memory Studies, 10(3):274–285, 2017

    Djouaria Ghilani, Olivier Luminet, Hans-Peter Erb, Christine Flassbeck, Valérie Rosoux, Ismee Tames, and Olivier Klein. Looking forward to the past: An interdisciplinary discussion on the use of historical analogies and their effects.Memory Studies, 10(3):274–285, 2017. 12 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Fores...

  15. [16]

    Try Deep Research and our new experimental model in Gemini, your AI assistant

    Google. Try Deep Research and our new experimental model in Gemini, your AI assistant. https://blog.google/products/gemini/google-gemini-deep-research/, Decem- ber 2024. Accessed: 2026-05-07

  16. [17]

    Green and J

    Kesten C. Green and J. Scott Armstrong. Structured analogies for forecasting.International Journal of Forecasting, 23(3):365–376, 2007

  17. [18]

    Cambridge University Press, 2014

    Jo Guldi and David Armitage.The history manifesto. Cambridge University Press, 2014

  18. [19]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025

  19. [20]

    The role of analogical reasoning in novel foreign-policy situations

    David Patrick Houghton. The role of analogical reasoning in novel foreign-policy situations. British Journal of Political Science, 26(4):523–552, 1996

  20. [21]

    Causal discovery from multiple data sets with non-identical variable sets

    Biwei Huang, Kun Zhang, Mingming Gong, and Clark Glymour. Causal discovery from multiple data sets with non-identical variable sets. InProceedings of AAAI, 2020

  21. [22]

    Deep research agents: A systematic examination and roadmap.arXiv preprint arXiv:2506.18096, 2025

    Yuxuan Huang, Yihang Chen, Haozheng Zhang, Kang Li, Huichi Zhou, Meng Fang, Linyi Yang, Xiaoguang Li, Lifeng Shang, Songcen Xu, et al. Deep research agents: A systematic examination and roadmap.arXiv preprint arXiv:2506.18096, 2025

  22. [23]

    Openai o1 system card.arXiv preprint arXiv:2412.16720, 2024

    Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card.arXiv preprint arXiv:2412.16720, 2024

  23. [24]

    StoryAnalogy: Deriving story-level analogies from large language models to unlock analogical understanding

    Cheng Jiayang, Lin Qiu, Tsz Ho Chan, Tianqing Fang, Weiqi Wang, Chunkit Chan, Dongyu Ru, Qipeng Guo, Hongming Zhang, Yangqiu Song, Yue Zhang, and Zheng Zhang. StoryAnalogy: Deriving story-level analogies from large language models to unlock analogical understanding. InProceedings of EMNLP, pages 11518–11537, 2023

  24. [25]

    Ezra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs, Danny Halawi, Fred Zhang, and Philip E. Tetlock. ForecastBench: A dynamic benchmark of AI forecasting capabilities. In Proceedings of ICLR, 2025

  25. [26]

    Historical analogies: Functions, limitations and the correct use of histori- cal analogies in applied history.Journal of Applied History, 5(2):111 – 131, 2023

    Sjoerd Keulen. Historical analogies: Functions, limitations and the correct use of histori- cal analogies in applied history.Journal of Applied History, 5(2):111 – 131, 2023. doi: 10.1163/25895893-bja10036. URL https://brill.com/view/journals/joah/5/2/ article-p111_2.xml

  26. [27]

    Princeton University Press, 1992

    Yuen Foong Khong.Analogies at War: Korea, Munich, Dien Bien Phu, and the Vietnam Decisions of 1965. Princeton University Press, 1992

  27. [28]

    Past meets present: Creating historical analogy with large language models

    Nianqi Li, Siyu Yuan, Jiangjie Chen, Jiaqing Liang, Feng Wei, Zujie Liang, Deqing Yang, and Yanghua Xiao. Past meets present: Creating historical analogy with large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), pages 3942–3957, 2025

  28. [29]

    WebThinker: Empowering large reasoning models with deep research capability

    Xiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian, Yutao Zhu, Yongkang Wu, Ji-Rong Wen, and Zhicheng Dou. WebThinker: Empowering large reasoning models with deep research capability. arXiv preprint arXiv:2504.21776, 2025

  29. [30]

    GAIA: A benchmark for general AI assistants.arXiv preprint arXiv:2311.12983, 2023

    Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. GAIA: A benchmark for general AI assistants.arXiv preprint arXiv:2311.12983, 2023. 13 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis

  30. [31]

    Landsness, Dániel L

    Ludovico Mitchener, Angela Yiu, Benjamin Chang, Mathieu Bourdenx, Tyler Nadolski, Arvis Sulovari, Eric C. Landsness, Dániel L. Barabási, Siddharth Narayanan, Nicky Evans, Shriya Reddy, Martha S. Foiani, Aizad Kamal, Leah P. Shriver, Fang Cao, Asmamaw T. Wassie, Jon M. Laurent, Edwin Melville-Green, Mayk Caldas Ramos, Albert Bou, Kaleigh F. Roberts, Sladja...

  31. [32]

    Nersessian.Creating Scientific Concepts

    Nancy J. Nersessian.Creating Scientific Concepts. MIT Press, 2008

  32. [33]

    Neustadt and Ernest R

    Richard E. Neustadt and Ernest R. May.Thinking in Time: The Uses of History for Decision-Makers. Free Press, 1986

  33. [34]

    Chatgpt.https://chat.openai.com/chat/, 2022

    OpenAI. Chatgpt.https://chat.openai.com/chat/, 2022

  34. [35]

    Introducing deep research

    OpenAI. Introducing deep research. https://openai.com/index/ introducing-deep-research/, February 2025. Accessed: 2026-05-07

  35. [36]

    Introducing GPT-5.4

    OpenAI. Introducing GPT-5.4. https://openai.com/index/introducing-gpt-5-4/, March 2026. Large language model

  36. [37]

    Introducing GPT-5.4 mini and nano

    OpenAI. Introducing GPT-5.4 mini and nano. https://openai.com/index/ introducing-gpt-5-4-mini-and-nano/, March 2026

  37. [38]

    Stevenson

    Gustaw Opiełka, Hannes Rosenbusch, and Claire E. Stevenson. Analogical reasoning inside large language models: Concept vectors and the limits of abstraction.arXiv preprint arXiv:2503.03666, 2025

  38. [39]

    Historical analogies as tools in understanding transformation

    Meg Parsons and Johanna Nalau. Historical analogies as tools in understanding transformation. Global Environmental Change, 38:82–96, 2016

  39. [40]

    Humanity’s last exam.arXiv preprint arXiv:2501.14249, 2025

    Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al. Humanity’s last exam.arXiv preprint arXiv:2501.14249, 2025

  40. [41]

    Can LLMs aid analogical reasoning for strategic decisions? a comparative study.Strategy Science, 11:118–136, 2026

    Prothit Sen, Maciej Workiewicz, and Phanish Puranam. Can LLMs aid analogical reasoning for strategic decisions? a comparative study.Strategy Science, 11:118–136, 2026

  41. [42]

    ARN: Analogical reasoning on narratives.Transactions of the Association for Computational Linguistics, 12:1063–1086, 2024

    Zhivar Sourati, Filip Ilievski, Pia Sommerauer, and Yifan Jiang. ARN: Analogical reasoning on narratives.Transactions of the Association for Computational Linguistics, 12:1063–1086, 2024

  42. [43]

    Miroflow: Towards high-performance and robust open-source agent framework for general deep research tasks.arXiv preprint arXiv:2602.22808, 2026

    MiroMind Team et al. Miroflow: Towards high-performance and robust open-source agent framework for general deep research tasks.arXiv preprint arXiv:2602.22808, 2026

  43. [44]

    Tongyi DeepResearch technical report.arXiv preprint arXiv:2510.24701, 2025

    Tongyi DeepResearch Team. Tongyi DeepResearch technical report.arXiv preprint arXiv:2510.24701, 2025

  44. [45]

    Holyoak, and Hongjing Lu

    Taylor Webb, Keith J. Holyoak, and Hongjing Lu. Emergent analogical reasoning in large language models.Nature Human Behaviour, 7:1526–1541, 2023

  45. [46]

    BrowseComp: A simple yet challenging benchmark for browsing agents.arXiv preprint arXiv:2504.12516, 2025

    Jason Wei et al. BrowseComp: A simple yet challenging benchmark for browsing agents.arXiv preprint arXiv:2504.12516, 2025

  46. [47]

    Acomprehensivesurveyofdeepresearch: Systems,methodologies, and applications.arXiv preprint arXiv:2506.12594, 2025

    RenjunXuandJingwenPeng. Acomprehensivesurveyofdeepresearch: Systems,methodologies, and applications.arXiv preprint arXiv:2506.12594, 2025. 14 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis

  47. [48]

    Qwen3technicalreport.arXivpreprintarXiv:2505.09388, 2025

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, ChengenHuang, ChenxuLv, etal. Qwen3technicalreport.arXivpreprintarXiv:2505.09388, 2025

  48. [49]

    Multi-view causal representation learning with partial observability

    Dingling Yao, Danru Xu, Sébastien Lachapelle, Sara Magliacane, Perouz Taslakian, Georg Martius, Julius von Kügelgen, and Francesco Locatello. Multi-view causal representation learning with partial observability. InProceedings of ICLR, 2024

  49. [50]

    Chi, and Denny Zhou

    Michihiro Yasunaga, Xinyun Chen, Yujia Li, Panupong Pasupat, Jure Leskovec, Percy Liang, Ed H. Chi, and Denny Zhou. Large language models as analogical reasoners. InProceedings of ICLR, 2024

  50. [51]

    AnaloBench: Benchmarking the identification of abstract and long-context analogies

    Xiao Ye, Andrew Wang, Jacob Choi, Yining Lu, Shreya Sharma, Lingfeng Shen, Vijay Murari Tiyyala, Nicholas Andrews, and Daniel Khashabi. AnaloBench: Benchmarking the identification of abstract and long-context analogies. InProceedings of EMNLP, 2024

  51. [52]

    Thought propagation: An analogical approach to complex reasoning with large language models

    Junchi Yu, Ran He, and Rex Ying. Thought propagation: An analogical approach to complex reasoning with large language models. InProceedings of ICLR, 2024

  52. [53]

    Suchow, and Khaldoun Khashanah

    Yangyang Yu, Haohang Li, Zhi Chen, Yuechen Jiang, Yang Li, Denghui Zhang, Rong Liu, Jordan W. Suchow, and Khaldoun Khashanah. FinMem: A performance-enhanced LLM trading agent with layered memory and character design.arXiv preprint arXiv:2311.13743, 2023

  53. [54]

    ANALOGYKB:Unlockinganalogicalreasoningoflanguagemodelswithamillion-scaleknowledge base

    Siyu Yuan, Jiangjie Chen, Changzhi Sun, Jiaqing Liang, Yanghua Xiao, and Deqing Yang. ANALOGYKB:Unlockinganalogicalreasoningoflanguagemodelswithamillion-scaleknowledge base. InProceedings of ACL (Long Papers), pages 1249–1265, 2024

  54. [55]

    Futurex: An advanced live benchmark for LLM agents in future prediction

    Zhiyuan Zeng, Jiashuo Liu, Siyuan Chen, Tianci He, Yali Liao, Yixiao Tian, wangjinpeng.levi, Zaiyuan Wang, YangYang, Lingyue Yin, Mingren Yin, Zhu Zhenwei, Tianle Cai, Xinjie Chen, Zehui Chen, Jiecao Chen, Yantao Du, Xiang Gao, Jiacheng Guo, LIANG HU, Jianpeng Jiao, Xiangsheng Li, Jingkai Liu, nishuang, Zhoufutu Wen, Ge Zhang, Kaiyuan Zhang, xin zhou, Jos...

  55. [56]

    A multimodal foundation agent for financial trading: Tool-augmented, diversified, and generalist.arXiv preprint arXiv:2402.18485, 2024

    Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, et al. A multimodal foundation agent for financial trading: Tool-augmented, diversified, and generalist.arXiv preprint arXiv:2402.18485, 2024. 15 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight An...

  56. [57]

    PASS” if real historical event, “FAIL

    factual_existence:“PASS” if real historical event, “FAIL” if hallucinated

  57. [58]

    both involve banks

    structural_relevance (0–4)— how complete is the causal mechanism mapping? 0 = No structural relevance; pure surface match (same country/topic only). 1 = Shares one superficial feature (“both involve banks”). 2 = Identifies 1–2 genuine shared mechanisms but without causal chain depth. 3 = Identifies 3+ shared mechanisms with causal chain articulation (trig...

  58. [59]

    YES” if matches a reference event, “PARTIAL

    is_in_reference_set:“YES” if matches a reference event, “PARTIAL” if sibling/related, “NO” otherwise

  59. [60]

    NOT_NOVEL

    novelty:“NOT_NOVEL” if in reference set; “NOVEL_VALID” if real event with structural_relevance≥ 2 not in reference set; “NOVEL_INVALID” otherwise

  60. [61]

    the scale is different

    difference_awareness (0–2): 0 = No limitations discussed. 1 = Mentions differences superficially (“the scale is different”). 2 = Identifies WHERE the analogy breaks down and WHY it matters for the analysis. </part-A> <part-B: cross-analogy reasoning (CARS)> Evaluate how the report uses multiple analogies TOGETHER. This is the critical test — most reports ...

  61. [62]

    L4 identifies a HIDDEN FACTOR, not an OUTCOME

    Probabilistic forecasts and scenario predictions are NEVER L4, regardless of historical references. L4 identifies a HIDDEN FACTOR, not an OUTCOME

  62. [63]

    Name-dropping a historical event without mechanism mapping is L1, not L2

  63. [64]

    Noting a shared feature across events without specifying its structural role is L3-D, not L3-S

  64. [65]

    27 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis

    A claim can only be L3-S or L4 if it references MULTIPLE (≥2) historical events. 27 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis

  65. [66]

    claims": [{

    If unsure between L3-D and L3-S, check: does the claim name or clearly imply a structural role (trigger, amplifier, mediator, enabler, outcome)? If yes→L3-S. If no→L3-D. </critical-rules> <input> REPORT: {report} HIDDEN FACTORS FROM REFERENCE SET (use to check if any claims — at ANY level — match known hidden factors): {hidden_factors} </input> <output> {...

  66. [67]

    ALIAS— candidate is a different NAME for the SAME specific historical event. YES patterns (templated; angle brackets denote slots, not specific events): –⟨event under codename⟩≡⟨same event under descriptive popular name⟩ 34 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis –⟨event under regional/local name⟩≡⟨...

  67. [68]

    ⟨generic ceremony class⟩

    SUB-INSTANCE of GENERIC input— input names a generic CATEGORY, candidate is one specific instance. YES patterns: – input “⟨generic ceremony class⟩” + candidate “⟨a regional or cultural variant of that ceremony⟩” – input “⟨class of recurring institutional events⟩” + candidate “⟨one dated instance of that class⟩” – input “⟨class of accidents/disasters⟩” + c...

  68. [69]

    ⟨multi-year war⟩

    PROPER SUBSET— input is a CONTAINER event (a multi-year war, movement, era, crisis, or campaign that comprises many sub-events) and the candidate is one named sub-event WITHIN it. YES patterns (true proper subsets): – input “⟨multi-year war⟩” + candidate “⟨one named battle within that war⟩” – input “⟨decade-long geopolitical confrontation⟩” + candidate “⟨...

  69. [70]

    Are they the same real-world event under different names?

    Identify the specific historical event each name refers to. Are they the same real-world event under different names?

  70. [71]

    If different, check: is one a SUB-INSTANCE of a generic category named by the other?

  71. [72]

    If different, check: is one a PROPER SUBSET (named sub-event within a container event)?

  72. [73]

    ” (empty string). - If a dimension matches with no important gap, set GAP to “

    State your final decision. </reasoning> <output> Reasoning:⟨2–3 sentences⟩ Answer: YES or NO </output> Per-Candidate 5-Dimension Match/Gap Evaluation prompt <task> You will compare an input event against a candidate analogy along 5 causal dimensions. For each dimension, identify both what MATCHES and what does NOT match. </task> <input-event> Input Event:...

  73. [74]

    incumbent power

    ACTORS— principal agents by structural ROLE (e.g., “incumbent power”, “rising challenger”), NOT proper names. 1–2 sentences

  74. [75]

    1–2 sentences

    RELATIONSHIPS— how actors relate (alliance, rivalry, dependence, hierarchy). 1–2 sentences

  75. [76]

    Short phrases

    ACTIONS— JSON array of 4–8 key actions in CAUSAL ORDER. Short phrases

  76. [77]

    1–2 sentences

    GOALS— what each actor aims to achieve. 1–2 sentences

  77. [78]

    1–2 sentences

    LOCATION— setting (geographical + institutional + domain). 1–2 sentences

  78. [79]

    imperial overreach leads to quagmire

    THEME— highest-order causal pattern in one phrase (e.g., “imperial overreach leads to quagmire”)

  79. [80]

    step": "⟨short phrase⟩

    CAUSAL_CHAIN— JSON array of ordered steps with VALENCE. Each step:{"step": "⟨short phrase⟩", "valence": -1|0|1} where+1= building/positive trajectory,−1= declining/nega- tive,0= neutral. Use STRICTLY{−1,0,1}— no±2

  80. [81]

    imperial_overextension

    CAUSAL_ELEMENTS— JSON array of 5–8 specific causal factors, snake_case short phrases. ONLY factors that play a CAUSAL role (drove or sustained the dynamic). NOT descriptions, dates, or proper names. GOOD: ["imperial_overextension", "asymmetric_warfare", "domestic_antiwar_pressure"] BAD:["nineteen sixty four", "Saigon", "war"] </fields> <input> Event: {nam...

Showing first 80 references.