REVIEW 4 major objections 4 minor 87 references
Historical analogy for foresight is a causal inference problem: analogies must be matched on hidden structural positions, not surface descriptions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 04:42 UTC pith:IFRCHD4R
load-bearing objection Genuinely new benchmark and a structural-analogy agent that beats baselines, but the central foresight claim rests on an unmeasured transfer assumption and a rubric that partly bakes in the agent's design. the 4 major comments →
Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that historical-analogy retrieval for foresight should be treated as causal inference over structural positions. Its surface non-identifiability theorem shows that two worlds with identical surface observations but different hidden positions are indistinguishable to any surface-level method, and even an infinite number of surface-matched analogies cannot reduce foresight risk below half the separation gap between the worlds. Its cross-analogy confirmation theorem shows that when analogies independently confirm a position, each confirmation multiplies the posterior odds by q/p; under the paper's calibrated values (prior 0.5, q=1, p=0.2), two confirmations suffice
What carries the argument
The load-bearing device is the distinction between an event's descriptive representation D(E) and its mechanistic representation M(E), a directed causal graph whose factors can be assigned structural positions (for example trigger, enabler, amplifier, mediator, outcome). Retrieval operates on M(E) by aligning positions across events. The second device is a Bayes-factor confirmation rule: independent analogies that confirm the same position update posterior odds multiplicatively, which yields a required number of confirmations per position (two in the calibrated regime). CANA operationalizes both by decomposing events into preconditions, temporal chains, mechanisms, and outcomes and by reflec
Load-bearing premise
The load-bearing premise is that if two events occupy the same structural position and one has progressed further, the source event's observed trajectory is a good prediction of the target's future trajectory, with bounded error—this transfer is assumed, not measured.
What would settle it
Run CANA on the five forward events using only pre-cutoff information, without giving it the oracle analogies, then check its L4 hidden-factor predictions against what actually happened after each cutoff. If the hidden-factor hit rate does not exceed a surface-matching baseline, or if the aligned analogies' trajectory forecasts are no closer to realized outcomes than randomly chosen historical events, the mechanism-transfer assumption fails.
If this is right
- Because surface matching is information-theoretically blind to hidden positions, foresight reports that do not attempt structural alignment cannot be expected to uncover hidden factors, regardless of model scale or retrieval budget.
- Two independent structural confirmations per position are enough, under the paper's calibrated regime, to treat a position as necessary rather than coincidental.
- An agent that adds a structural analogy brief to a general deep-research pipeline can match or exceed commercial deep-research agents even with a weaker backbone.
- Structural decomposition of events into roles, rather than topical summaries, changes which analogies are retrieved and reduces self-analogy and surface-match errors.
Where Pith is reading between the lines
- Inference: The paper implies that analogy retrieval quality should be measured by posterior coverage of hidden positions, not by similarity rankings; a practical extension is to have agents output position-level coverage and stop at the two-confirmation threshold.
- Inference: The transferability assumption is the empirical crux: if the benchmark's oracle analogies already presuppose that aligned positions transfer, then the method's gains on hidden-factor hits may partly reflect benchmark construction; a stronger test would let the agent discover analogies without oracle hints and then score predicted hidden factors against actual post-cutoff outcomes.
- Inference: The same two-principle recipe could generalize to other partial-observation domains, such as medical case comparison or geopolitical risk, where multiple historical cases with different surface features share structural roles.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper defines a new task, Analogical Deep Research (ADR), in which LLM agents must retrieve historical analogies and integrate them for foresight analysis. It argues that ADR is inherently causal: analogy matching should align mechanistic structure M(E) rather than surface descriptions D(E). A short theory section states a surface non-identifiability theorem (Theorem 4/13), a cross-analogy confirmation theorem (Theorem 5/16), and derives two principles: retrieve on mechanisms, and require at least two independent confirming analogies. The authors introduce CANA, a framework that decomposes events into preconditions, temporal chains, mechanisms, and outcomes, and iteratively refines analogy candidates using structural feedback. They also construct ADR-bench with 15 events (10 historical, 5 forward) and compare commercial deep-research agents, vanilla MiroFlow, and MiroFlow+CANA. Results on the Li et al. (2025) analogy generation benchmark and ADR-bench show consistent gains for CANA across several LLM backbones.
Significance. If the central claims hold, this is a useful step: the paper gives a concrete task formulation, a transparent theoretical argument for why surface matching is insufficient, and an agent design that consistently outperforms strong deep-research baselines across multiple backbones. The empirical gains in Tables 1–3 are large and coherent, and the connection between analogy retrieval and causal identifiability is a good framing. However, the load-bearing claims are not yet independently validated. The Mechanism Transfer assumption (Assumption 3/12) is never tested, and the ADR-bench evaluation relies on an author-built LLM-judged rubric, single runs without confidence intervals, and forward events whose outcome-based FQS is scored while outcomes are still unfolding. The theoretical 'two analogies suffice' result also depends on hand-set probabilities that are not calibrated. These are fixable with additional experiments and analysis; with the current evidence, the contribution is promising but not fully established.
major comments (4)
- [§3.1 / Appendix C.2, Assumption 12 and Corollary 15] The scientific value of historical analogies for foresight rests entirely on the Mechanism Transfer assumption: if source and target factors occupy the same structural position and the source has progressed further, then TV(bP_s^E_S, P_s^T) ≤ α_s^tr. Corollary 15 states that mechanism matching strictly dominates surface matching only when Σ μ_s α_s^tr < Σ μ_s Δ_s, yet neither α_s^tr nor Δ_s is ever estimated or bounded empirically. No experiment tests whether dynamics at an aligned structural position actually transfer to the target. This is load-bearing because ADR-bench's oracle analogies are selected on the basis of documented shared mechanisms, and the L4 rubric credits only cross-analogy inference from ≥2 events — both presuppose the same transferability. Consequently Table 3 cannot independently validate the transfer assumption. Please add a direct transfer test (e.g., using histor
- [§3.3 / Theorem 16 and Appendix D] The claim that 'two independent analogies suffice' is derived from the hand-set values π=0.5, q=1, p=0.2, δ=0.05, and the assumption of conditional independence across analogies. No calibration of p, q, or the independence assumption is provided; LLM-generated analogies are not independent, and q/p is never measured. In addition, the ADR-bench L4 rubric explicitly requires citing ≥2 analogies for full credit, and L3-S requires ≥2 events. CANA is specifically designed to produce ≥2 crossing analogies, so a substantial part of its L3-S/L4 advantage over baselines is by construction. The paper should report an ablation in which the same evidence is scored without the two-analogy requirement, or calibrate p and q on a held-out set, to show that the advantage is not merely rubric-induced.
- [§5.2, Table 3] The headline ADR-bench results are based on a single run per configuration (no confidence intervals or variance), only 15 events, and an LLM judge (Claude Sonnet 4.5) using a rubric designed by the authors. For the five forward events, FQS is scored against outcomes that are still unfolding: the FQS prompt instructs the judge to compare against 'what actually happened', but those events have no settled outcome. This makes the FQS numbers unverifiable for the forward split. Please report multiple runs with standard deviations, provide human–LLM agreement on claim decomposition and scoring, and either restrict FQS to historical events or defer forward-event FQS until outcomes resolve.
- [§5.2, 'Connecting to theory'] The paper states that 'HF@L4=0/42 for all commercial agents validates Theorem 4.' This overstates what the theorem shows. Theorem 4 is a worst-case information-theoretic result conditioned on identical surface observations; commercial DR agents' failure to find hidden factors may be due to retrieval, planning, prompt design, or evaluation granularity rather than surface-level identifiability. A more direct test would compare a surface-level retriever against a mechanism-aligned retriever on the same surface observations, or vary the amount of surface evidence systematically. Please temper this claim or add such a controlled experiment.
minor comments (4)
- [References] References [9] and [10] are duplicates: both are Clement and Gentner, 'Systematicity as a selection constraint in analogical mapping.' Please merge.
- [§4.1] The phrase 'As mentioned in Def. 3.1' refers to a definition from the main text but the actual formal definition is Definition 6 in Appendix C.1. Also, 'MiroFlow' appears with inconsistent markup across the paper.
- [§1 / §6] The contributions state 'more than 10% performance' in analogy retrieval, while the abstract and conclusion say 'up to 10% improvements.' Please clarify the exact setting and metric supporting the 10% figure; Tables 1 and 2 show different relative gains depending on backbone and rubric.
- [Appendix B / Evaluation] The paper candidly acknowledges in 'Limitation and Future Works' that the benchmark scale is limited and evaluation relies on LLMs. This is valuable, but the abstract and conclusion should carry a corresponding caveat so that readers are not misled about the strength of the ADR-bench evidence.
Circularity Check
CANA's 'two analogies suffice' result follows from hand-set q=1, p=0.2, and ADR-bench's L4 rubric defines success as citing ≥2 analogies—the exact behavior CANA is prompted to emit; partial circularity.
specific steps
-
self definitional
[Sec. 3.2 (Evaluation design), Appendix D (Call 3 rubric), Sec. 4.1 (Structural reflective generation)]
"L4 (Hidden Factor Inference): An inference about a SPECIFIC, CURRENTLY UNRECOGNIZED factor in the current situation, justified by cross-analogy evidence from ≥2 historical events. ... A claim can only be L3-S or L4 if it references MULTIPLE (≥2) historical events. ... iteratively retrieve cross-confirming analogies until each hidden position is supported by ≥ 2 independent analogies, realizing Principle 2."
ADR-bench defines the headline metric L4 as a claim justified by ≥2 analogies, and L3-S likewise requires ≥2 events with a structural role. CANA's core loop is prompted to collect ≥2 confirming analogies per position before producing the Structural Analogy Brief. A CANA report that follows its own prompt automatically satisfies the citation-count condition for L3-S/L4, while agents not given that instruction cannot. The Table 3 gap therefore partly measures the rubric matching the method's design, not an independent confirmation of the cross-analogy confirmation principle.
-
fitted input called prediction
[Sec. 3.3 (Theoretical Discussions) and Theorem 16 (Appendix C.2); Sec. 4.1]
"For example, given π_s=0.5, q_s=1, p_s=0.2, δ=0.05, one requires 2 independent analogies per position to distinguish the structural necessity. ... Under the calibrated regime (π_s=0.5, q_s=1, p_s=0.2), this gives n*_s=2 for δ=0.05."
n*_s=2 is the closed-form solution of the Bayes-factor inequality evaluated at hand-chosen q_s=1, p_s=0.2, π_s=0.5, δ=0.05; no data calibrate these values. The paper labels them a 'calibrated regime', converts the arithmetic consequence into Principle 2 (K*_s ≥ 2), and hardwires it into CANA's stopping criterion ('until each hidden position is supported by ≥ 2 independent analogies'). The claim that Table 3 'validates Theorem 5' is therefore partly circular: the system was built to satisfy the theorem's hand-set threshold, and the rubric rewards reaching it.
full rationale
The formal theorems are not themselves circular: Theorem 13/4 is a standard indistinguishability argument from O∞_D(w0)=O∞_D(w1), and Theorem 16/5 is a correct Bayes-factor calculation conditional on assumed q_k,s, p_k,s and conditional independence. Corollary 15 is also valid conditionally on Assumption 12. The circularity enters the empirical validation pipeline. The 'two analogies suffice' threshold is an arithmetic consequence of q=1, p=0.2, π=0.5, δ=0.05; CANA is then instructed to retrieve until each position has ≥2 analogies; and the ADR-bench rubric defines L4 as a claim justified by ≥2 historical events and denies L3-S/L4 to claims with fewer than 2 events. Hence a substantial part of CANA's L3-S/L4 advantage is by construction. This is partial, not total: CANA also improves analogy generation on the pre-existing Li et al. benchmark (Tables 1-2), and hidden-factor hits require matching the annotated reference set, which provides some independent signal. Assumption 12's transfer bound is load-bearing but acknowledged as an assumption; its lack of empirical estimation is a correctness risk rather than a circular reduction. No load-bearing self-citation chain was found. Score 5.
Axiom & Free-Parameter Ledger
free parameters (3)
- coincidence probability p_s in Theorem 5/16 =
0.2
- confirmation probability q_s =
1
- prior π_s and target error δ =
0.5 and 0.05
axioms (6)
- domain assumption Mechanism Transfer (Assumption 3/12): aligned factors at the same structural position have bounded TV transfer error α_s.
- domain assumption Events instantiate a common abstract causal pattern via edge-preserving graph homomorphisms (Def. 7, Appendix C.1).
- domain assumption X_{k,s} are conditionally independent across analogies given Z_s (Theorem 5/16).
- ad hoc to paper LLM-extracted structural decomposition (preconditions, temporal chains, mechanisms, outcomes) faithfully recovers M(E).
- ad hoc to paper LLM judge (Claude Sonnet 4.5 / GPT-5.4) scores align with expert human judgment.
- domain assumption For forward events, hidden factors and eventual outcomes are sufficiently known/foreseeable to score FQS.
invented entities (1)
-
Abstract causal pattern 𝔓 with structural positions S_𝔓
no independent evidence
read the original abstract
Systematic comparisons between current situations and structurally similar past events in the historical, i.e., historical analogies, is among the most powerful tools for foresight analysis. In this work, we present a new task called Analogical Deep Research (ADR) to Large Language Model (LLM) agents and construct the first ADR benchmark ADR-bench to study whether LLM agents are able to find and leverage historical analogies when doing foresight analysis. Our investigation reveals a key obstacle: LLM agents are poor at finding analogies because they match on surface features rather than underlying mechanisms. We argue that ADR is inherently a causal question as it requires understanding why the event occurred. Based on our theoretical analysis, we propose two principles required for ADR, including the mechanism alignment and cross-analogy confirmation. Built upon our theoretical results, we propose a new agentic framework called Causal Analogical Researcher (CANA) that guides LLMs to find and integrate historical analogies. CANA incorporates a simple yet effective structural decomposition representation, and integrates structural feedback for reflective improvements of historical analogy identification and integration. We show that CANA brings up to 10% improvements in historical analogy generation, and surpasses the state-of-the-art deep research agents in the ADR-bench. Case studies with the ongoing events confirm the effectiveness of CANA in leveraging historical analogies.
Reference graph
Works this paper leans on
-
[1]
The making of an applied historian: Stage two.The Public Historian, 5 (2):21–46, 1983
W Andrew Achenbaum. The making of an applied historian: Stage two.The Public Historian, 5 (2):21–46, 1983
1983
-
[2]
Identification of partially observed linear causal models: Graphical conditions for the non-Gaussian and heterogeneous cases
Jeffrey Adams, Niels Hansen, and Kun Zhang. Identification of partially observed linear causal models: Graphical conditions for the non-Gaussian and heterogeneous cases. InAdvances in Neural Information Processing Systems (NeurIPS), 2021
2021
-
[3]
Claude Sonnet 4.5
Anthropic. Claude Sonnet 4.5. https://www.anthropic.com/news/ claude-sonnet-4-5, September 2025. Large language model
2025
-
[4]
Claude Sonnet 4.6 system card
Anthropic. Claude Sonnet 4.6 system card. Technical report, Anthropic, February 2026. URL https://anthropic.com/claude-sonnet-4-6-system-card
2026
-
[5]
Paul F. A. Bartha.By Parallel Reasoning: The Construction and Evaluation of Analogical Arguments. Oxford University Press, 2010
2010
-
[6]
Weakly supervised causal representation learning
Johann Brehmer, Pim de Haan, Phillip Lippe, and Taco Cohen. Weakly supervised causal representation learning. InAdvances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[7]
Darren C. Brunk. Curing the Somalia syndrome: Analogy, foreign policy decision making, and the Rwandan genocide.Foreign Policy Analysis, 4(3):301–320, 2008
2008
-
[8]
Spears, Derya Unutmaz, Kevin Weil, Steven Yin, and Nikita Zhivotovskiy
Sébastien Bubeck, Christian Coester, Ronen Eldan, Timothy Gowers, Yin Tat Lee, Alexandru Lupsasca, Mehtaab Sawhney, Robert Scherrer, Mark Sellke, Brian K. Spears, Derya Unutmaz, Kevin Weil, Steven Yin, and Nikita Zhivotovskiy. Early science acceleration experiments with gpt-5.ArXiv, abs/2511.16072, 2025
arXiv 2025
-
[10]
Clement and Dedre Gentner
Catherine A. Clement and Dedre Gentner. Systematicity as a selection constraint in analogical mapping.Cognitive Science, 15(1):89–132, 1991
1991
-
[11]
Simon & Schuster, New York, 2017
Ray Dalio.Principles: Life and Work. Simon & Schuster, New York, 2017. ISBN 9781501124020
2017
-
[12]
How scientists really reason: Scientific reasoning in real-world laboratories
Kevin Dunbar. How scientists really reason: Scientific reasoning in real-world laboratories. In Robert J. Sternberg and Janet E. Davidson, editors,The Nature of Insight, pages 365–395. MIT Press, 1995
1995
-
[13]
Forbus, and Dedre Gentner
Brian Falkenhainer, Kenneth D. Forbus, and Dedre Gentner. The structure-mapping engine: Algorithm and examples.Artificial Intelligence, 41(1):1–63, 1989
1989
-
[14]
Structure-mapping: A theoretical framework for analogy.Cognitive Science, 7 (2):155–170, 1983
Dedre Gentner. Structure-mapping: A theoretical framework for analogy.Cognitive Science, 7 (2):155–170, 1983
1983
-
[15]
Looking forward to the past: An interdisciplinary discussion on the use of historical analogies and their effects.Memory Studies, 10(3):274–285, 2017
Djouaria Ghilani, Olivier Luminet, Hans-Peter Erb, Christine Flassbeck, Valérie Rosoux, Ismee Tames, and Olivier Klein. Looking forward to the past: An interdisciplinary discussion on the use of historical analogies and their effects.Memory Studies, 10(3):274–285, 2017. 12 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Fores...
2017
-
[16]
Try Deep Research and our new experimental model in Gemini, your AI assistant
Google. Try Deep Research and our new experimental model in Gemini, your AI assistant. https://blog.google/products/gemini/google-gemini-deep-research/, Decem- ber 2024. Accessed: 2026-05-07
2024
-
[17]
Green and J
Kesten C. Green and J. Scott Armstrong. Structured analogies for forecasting.International Journal of Forecasting, 23(3):365–376, 2007
2007
-
[18]
Cambridge University Press, 2014
Jo Guldi and David Armitage.The history manifesto. Cambridge University Press, 2014
2014
-
[19]
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025
Pith/arXiv arXiv 2025
-
[20]
The role of analogical reasoning in novel foreign-policy situations
David Patrick Houghton. The role of analogical reasoning in novel foreign-policy situations. British Journal of Political Science, 26(4):523–552, 1996
1996
-
[21]
Causal discovery from multiple data sets with non-identical variable sets
Biwei Huang, Kun Zhang, Mingming Gong, and Clark Glymour. Causal discovery from multiple data sets with non-identical variable sets. InProceedings of AAAI, 2020
2020
-
[22]
Deep research agents: A systematic examination and roadmap.arXiv preprint arXiv:2506.18096, 2025
Yuxuan Huang, Yihang Chen, Haozheng Zhang, Kang Li, Huichi Zhou, Meng Fang, Linyi Yang, Xiaoguang Li, Lifeng Shang, Songcen Xu, et al. Deep research agents: A systematic examination and roadmap.arXiv preprint arXiv:2506.18096, 2025
Pith/arXiv arXiv 2025
-
[23]
Openai o1 system card.arXiv preprint arXiv:2412.16720, 2024
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card.arXiv preprint arXiv:2412.16720, 2024
Pith/arXiv arXiv 2024
-
[24]
StoryAnalogy: Deriving story-level analogies from large language models to unlock analogical understanding
Cheng Jiayang, Lin Qiu, Tsz Ho Chan, Tianqing Fang, Weiqi Wang, Chunkit Chan, Dongyu Ru, Qipeng Guo, Hongming Zhang, Yangqiu Song, Yue Zhang, and Zheng Zhang. StoryAnalogy: Deriving story-level analogies from large language models to unlock analogical understanding. InProceedings of EMNLP, pages 11518–11537, 2023
2023
-
[25]
Ezra Karger, Houtan Bastani, Yueh-Han Chen, Zachary Jacobs, Danny Halawi, Fred Zhang, and Philip E. Tetlock. ForecastBench: A dynamic benchmark of AI forecasting capabilities. In Proceedings of ICLR, 2025
2025
-
[26]
Sjoerd Keulen. Historical analogies: Functions, limitations and the correct use of histori- cal analogies in applied history.Journal of Applied History, 5(2):111 – 131, 2023. doi: 10.1163/25895893-bja10036. URL https://brill.com/view/journals/joah/5/2/ article-p111_2.xml
-
[27]
Princeton University Press, 1992
Yuen Foong Khong.Analogies at War: Korea, Munich, Dien Bien Phu, and the Vietnam Decisions of 1965. Princeton University Press, 1992
1965
-
[28]
Past meets present: Creating historical analogy with large language models
Nianqi Li, Siyu Yuan, Jiangjie Chen, Jiaqing Liang, Feng Wei, Zujie Liang, Deqing Yang, and Yanghua Xiao. Past meets present: Creating historical analogy with large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), pages 3942–3957, 2025
2025
-
[29]
WebThinker: Empowering large reasoning models with deep research capability
Xiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian, Yutao Zhu, Yongkang Wu, Ji-Rong Wen, and Zhicheng Dou. WebThinker: Empowering large reasoning models with deep research capability. arXiv preprint arXiv:2504.21776, 2025
Pith/arXiv arXiv 2025
-
[30]
GAIA: A benchmark for general AI assistants.arXiv preprint arXiv:2311.12983, 2023
Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. GAIA: A benchmark for general AI assistants.arXiv preprint arXiv:2311.12983, 2023. 13 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis
Pith/arXiv arXiv 2023
-
[31]
Ludovico Mitchener, Angela Yiu, Benjamin Chang, Mathieu Bourdenx, Tyler Nadolski, Arvis Sulovari, Eric C. Landsness, Dániel L. Barabási, Siddharth Narayanan, Nicky Evans, Shriya Reddy, Martha S. Foiani, Aizad Kamal, Leah P. Shriver, Fang Cao, Asmamaw T. Wassie, Jon M. Laurent, Edwin Melville-Green, Mayk Caldas Ramos, Albert Bou, Kaleigh F. Roberts, Sladja...
Pith/arXiv arXiv 2025
-
[32]
Nersessian.Creating Scientific Concepts
Nancy J. Nersessian.Creating Scientific Concepts. MIT Press, 2008
2008
-
[33]
Neustadt and Ernest R
Richard E. Neustadt and Ernest R. May.Thinking in Time: The Uses of History for Decision-Makers. Free Press, 1986
1986
-
[34]
Chatgpt.https://chat.openai.com/chat/, 2022
OpenAI. Chatgpt.https://chat.openai.com/chat/, 2022
2022
-
[35]
Introducing deep research
OpenAI. Introducing deep research. https://openai.com/index/ introducing-deep-research/, February 2025. Accessed: 2026-05-07
2025
-
[36]
Introducing GPT-5.4
OpenAI. Introducing GPT-5.4. https://openai.com/index/introducing-gpt-5-4/, March 2026. Large language model
2026
-
[37]
Introducing GPT-5.4 mini and nano
OpenAI. Introducing GPT-5.4 mini and nano. https://openai.com/index/ introducing-gpt-5-4-mini-and-nano/, March 2026
2026
-
[38]
Gustaw Opiełka, Hannes Rosenbusch, and Claire E. Stevenson. Analogical reasoning inside large language models: Concept vectors and the limits of abstraction.arXiv preprint arXiv:2503.03666, 2025
Pith/arXiv arXiv 2025
-
[39]
Historical analogies as tools in understanding transformation
Meg Parsons and Johanna Nalau. Historical analogies as tools in understanding transformation. Global Environmental Change, 38:82–96, 2016
2016
-
[40]
Humanity’s last exam.arXiv preprint arXiv:2501.14249, 2025
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al. Humanity’s last exam.arXiv preprint arXiv:2501.14249, 2025
Pith/arXiv arXiv 2025
-
[41]
Can LLMs aid analogical reasoning for strategic decisions? a comparative study.Strategy Science, 11:118–136, 2026
Prothit Sen, Maciej Workiewicz, and Phanish Puranam. Can LLMs aid analogical reasoning for strategic decisions? a comparative study.Strategy Science, 11:118–136, 2026
2026
-
[42]
ARN: Analogical reasoning on narratives.Transactions of the Association for Computational Linguistics, 12:1063–1086, 2024
Zhivar Sourati, Filip Ilievski, Pia Sommerauer, and Yifan Jiang. ARN: Analogical reasoning on narratives.Transactions of the Association for Computational Linguistics, 12:1063–1086, 2024
2024
-
[43]
MiroMind Team et al. Miroflow: Towards high-performance and robust open-source agent framework for general deep research tasks.arXiv preprint arXiv:2602.22808, 2026
arXiv 2026
-
[44]
Tongyi DeepResearch technical report.arXiv preprint arXiv:2510.24701, 2025
Tongyi DeepResearch Team. Tongyi DeepResearch technical report.arXiv preprint arXiv:2510.24701, 2025
Pith/arXiv arXiv 2025
-
[45]
Holyoak, and Hongjing Lu
Taylor Webb, Keith J. Holyoak, and Hongjing Lu. Emergent analogical reasoning in large language models.Nature Human Behaviour, 7:1526–1541, 2023
2023
-
[46]
Jason Wei et al. BrowseComp: A simple yet challenging benchmark for browsing agents.arXiv preprint arXiv:2504.12516, 2025
Pith/arXiv arXiv 2025
-
[47]
RenjunXuandJingwenPeng. Acomprehensivesurveyofdeepresearch: Systems,methodologies, and applications.arXiv preprint arXiv:2506.12594, 2025. 14 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis
Pith/arXiv arXiv 2025
-
[48]
Qwen3technicalreport.arXivpreprintarXiv:2505.09388, 2025
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, ChengenHuang, ChenxuLv, etal. Qwen3technicalreport.arXivpreprintarXiv:2505.09388, 2025
Pith/arXiv arXiv 2025
-
[49]
Multi-view causal representation learning with partial observability
Dingling Yao, Danru Xu, Sébastien Lachapelle, Sara Magliacane, Perouz Taslakian, Georg Martius, Julius von Kügelgen, and Francesco Locatello. Multi-view causal representation learning with partial observability. InProceedings of ICLR, 2024
2024
-
[50]
Chi, and Denny Zhou
Michihiro Yasunaga, Xinyun Chen, Yujia Li, Panupong Pasupat, Jure Leskovec, Percy Liang, Ed H. Chi, and Denny Zhou. Large language models as analogical reasoners. InProceedings of ICLR, 2024
2024
-
[51]
AnaloBench: Benchmarking the identification of abstract and long-context analogies
Xiao Ye, Andrew Wang, Jacob Choi, Yining Lu, Shreya Sharma, Lingfeng Shen, Vijay Murari Tiyyala, Nicholas Andrews, and Daniel Khashabi. AnaloBench: Benchmarking the identification of abstract and long-context analogies. InProceedings of EMNLP, 2024
2024
-
[52]
Thought propagation: An analogical approach to complex reasoning with large language models
Junchi Yu, Ran He, and Rex Ying. Thought propagation: An analogical approach to complex reasoning with large language models. InProceedings of ICLR, 2024
2024
-
[53]
Suchow, and Khaldoun Khashanah
Yangyang Yu, Haohang Li, Zhi Chen, Yuechen Jiang, Yang Li, Denghui Zhang, Rong Liu, Jordan W. Suchow, and Khaldoun Khashanah. FinMem: A performance-enhanced LLM trading agent with layered memory and character design.arXiv preprint arXiv:2311.13743, 2023
Pith/arXiv arXiv 2023
-
[54]
ANALOGYKB:Unlockinganalogicalreasoningoflanguagemodelswithamillion-scaleknowledge base
Siyu Yuan, Jiangjie Chen, Changzhi Sun, Jiaqing Liang, Yanghua Xiao, and Deqing Yang. ANALOGYKB:Unlockinganalogicalreasoningoflanguagemodelswithamillion-scaleknowledge base. InProceedings of ACL (Long Papers), pages 1249–1265, 2024
2024
-
[55]
Futurex: An advanced live benchmark for LLM agents in future prediction
Zhiyuan Zeng, Jiashuo Liu, Siyuan Chen, Tianci He, Yali Liao, Yixiao Tian, wangjinpeng.levi, Zaiyuan Wang, YangYang, Lingyue Yin, Mingren Yin, Zhu Zhenwei, Tianle Cai, Xinjie Chen, Zehui Chen, Jiecao Chen, Yantao Du, Xiang Gao, Jiacheng Guo, LIANG HU, Jianpeng Jiao, Xiangsheng Li, Jingkai Liu, nishuang, Zhoufutu Wen, Ge Zhang, Kaiyuan Zhang, xin zhou, Jos...
2026
-
[56]
Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, et al. A multimodal foundation agent for financial trading: Tool-augmented, diversified, and generalist.arXiv preprint arXiv:2402.18485, 2024. 15 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight An...
Pith/arXiv arXiv 2024
-
[57]
PASS” if real historical event, “FAIL
factual_existence:“PASS” if real historical event, “FAIL” if hallucinated
-
[58]
both involve banks
structural_relevance (0–4)— how complete is the causal mechanism mapping? 0 = No structural relevance; pure surface match (same country/topic only). 1 = Shares one superficial feature (“both involve banks”). 2 = Identifies 1–2 genuine shared mechanisms but without causal chain depth. 3 = Identifies 3+ shared mechanisms with causal chain articulation (trig...
1907
-
[59]
YES” if matches a reference event, “PARTIAL
is_in_reference_set:“YES” if matches a reference event, “PARTIAL” if sibling/related, “NO” otherwise
-
[60]
NOT_NOVEL
novelty:“NOT_NOVEL” if in reference set; “NOVEL_VALID” if real event with structural_relevance≥ 2 not in reference set; “NOVEL_INVALID” otherwise
-
[61]
the scale is different
difference_awareness (0–2): 0 = No limitations discussed. 1 = Mentions differences superficially (“the scale is different”). 2 = Identifies WHERE the analogy breaks down and WHY it matters for the analysis. </part-A> <part-B: cross-analogy reasoning (CARS)> Evaluate how the report uses multiple analogies TOGETHER. This is the critical test — most reports ...
1907
-
[62]
L4 identifies a HIDDEN FACTOR, not an OUTCOME
Probabilistic forecasts and scenario predictions are NEVER L4, regardless of historical references. L4 identifies a HIDDEN FACTOR, not an OUTCOME
-
[63]
Name-dropping a historical event without mechanism mapping is L1, not L2
-
[64]
Noting a shared feature across events without specifying its structural role is L3-D, not L3-S
-
[65]
27 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis
A claim can only be L3-S or L4 if it references MULTIPLE (≥2) historical events. 27 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis
-
[66]
claims": [{
If unsure between L3-D and L3-S, check: does the claim name or clearly imply a structural role (trigger, amplifier, mediator, enabler, outcome)? If yes→L3-S. If no→L3-D. </critical-rules> <input> REPORT: {report} HIDDEN FACTORS FROM REFERENCE SET (use to check if any claims — at ANY level — match known hidden factors): {hidden_factors} </input> <output> {...
1907
-
[67]
ALIAS— candidate is a different NAME for the SAME specific historical event. YES patterns (templated; angle brackets denote slots, not specific events): –⟨event under codename⟩≡⟨same event under descriptive popular name⟩ 34 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis –⟨event under regional/local name⟩≡⟨...
-
[68]
⟨generic ceremony class⟩
SUB-INSTANCE of GENERIC input— input names a generic CATEGORY, candidate is one specific instance. YES patterns: – input “⟨generic ceremony class⟩” + candidate “⟨a regional or cultural variant of that ceremony⟩” – input “⟨class of recurring institutional events⟩” + candidate “⟨one dated instance of that class⟩” – input “⟨class of accidents/disasters⟩” + c...
-
[69]
⟨multi-year war⟩
PROPER SUBSET— input is a CONTAINER event (a multi-year war, movement, era, crisis, or campaign that comprises many sub-events) and the candidate is one named sub-event WITHIN it. YES patterns (true proper subsets): – input “⟨multi-year war⟩” + candidate “⟨one named battle within that war⟩” – input “⟨decade-long geopolitical confrontation⟩” + candidate “⟨...
-
[70]
Are they the same real-world event under different names?
Identify the specific historical event each name refers to. Are they the same real-world event under different names?
-
[71]
If different, check: is one a SUB-INSTANCE of a generic category named by the other?
-
[72]
If different, check: is one a PROPER SUBSET (named sub-event within a container event)?
-
[73]
” (empty string). - If a dimension matches with no important gap, set GAP to “
State your final decision. </reasoning> <output> Reasoning:⟨2–3 sentences⟩ Answer: YES or NO </output> Per-Candidate 5-Dimension Match/Gap Evaluation prompt <task> You will compare an input event against a candidate analogy along 5 causal dimensions. For each dimension, identify both what MATCHES and what does NOT match. </task> <input-event> Input Event:...
2024
-
[74]
incumbent power
ACTORS— principal agents by structural ROLE (e.g., “incumbent power”, “rising challenger”), NOT proper names. 1–2 sentences
-
[75]
1–2 sentences
RELATIONSHIPS— how actors relate (alliance, rivalry, dependence, hierarchy). 1–2 sentences
-
[76]
Short phrases
ACTIONS— JSON array of 4–8 key actions in CAUSAL ORDER. Short phrases
-
[77]
1–2 sentences
GOALS— what each actor aims to achieve. 1–2 sentences
-
[78]
1–2 sentences
LOCATION— setting (geographical + institutional + domain). 1–2 sentences
-
[79]
imperial overreach leads to quagmire
THEME— highest-order causal pattern in one phrase (e.g., “imperial overreach leads to quagmire”)
-
[80]
step": "⟨short phrase⟩
CAUSAL_CHAIN— JSON array of ordered steps with VALENCE. Each step:{"step": "⟨short phrase⟩", "valence": -1|0|1} where+1= building/positive trajectory,−1= declining/nega- tive,0= neutral. Use STRICTLY{−1,0,1}— no±2
-
[81]
imperial_overextension
CAUSAL_ELEMENTS— JSON array of 5–8 specific causal factors, snake_case short phrases. ONLY factors that play a CAUSAL role (drove or sustained the dynamic). NOT descriptions, dates, or proper names. GOOD: ["imperial_overextension", "asymmetric_warfare", "domestic_antiwar_pressure"] BAD:["nineteen sixty four", "Saigon", "war"] </fields> <input> Event: {nam...
1989
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.