Pith. sign in

REVIEW 3 major objections 4 minor 59 references

The paper argues that inconsistent definitions of "hypothesis" across NLP tasks are now a critical obstacle to a machine-interpretable scholarly record.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A position paper documenting inconsistent and often implicit definitions of 'hypothesis' across NLP hypothesis mining tasks and calling for standardization.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection A useful descriptive survey of how 'hypothesis' is operationalized across NLP tasks, but its central claim that variability is now a critical problem is asserted, not demonstrated. the 3 major comments →

arxiv 2509.00185 v1 pith:7KSRCWJ6 submitted 2025-08-29 cs.CL cs.AI

What Are Research Hypotheses?

classification cs.CL cs.AI
keywords natural language processingnatural language understandinghypothesis extractionhypothesis generationscientific claim verificationhypothesis definitionmachine-readable scholarly record
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the term "hypothesis" no longer has a settled meaning in natural language processing: different tasks—entailment, extraction, verification, and generation—define and operationalize hypotheses in incompatible ways, and many papers never define the term at all, relying on benchmark datasets as implicit definitions. The authors survey definitions from Greek philosophy and the philosophy of science through to current large-language-model systems, and they sort modern uses into four types: ideas, claims, structured proposals, and formal expressions such as context–variable–relationship trees. Their central claim is that this ambiguity, once tolerable, has become a critical concern as the field moves toward a machine-interpretable scholarly record, because quantitative operationalization of hypotheses requires explicit, well-structured definitions. A sympathetic reader would care because the same word is being used as the input, output, and evaluation target of systems meant to extract, verify, and generate science, so incompatible definitions block comparison and knowledge assembly.

Core claim

The paper's central claim, stated in its introduction and repeated in its conclusion, is that the historic ambiguity around the definition of a hypothesis—tolerable, even productive, when human communities supplied shared context—becomes critical now that NLP tasks require quantitative operationalization of hypotheses. Across recent tasks, hypotheses function as ideas, as untested claims, as structured research proposals, or as formal expressions decomposable into contexts, variables, and relationships. The paper documents that datasets implicitly define the term through their annotation schemes: a hypothesis can be a single declarative sentence, a high-level research question in scientific

What carries the argument

The paper's central object is the term "hypothesis" as an operationalized unit in NLU tasks, examined through a four-way taxonomy: ideas, claims, hypothesis-proposals, and formal expressions such as context–variable–relationship semantic trees and entailment trees. The taxonomy carries the argument: it shows that the same word anchors four different objects across extraction, verification, and generation, and that benchmark datasets rather than explicit definitions provide the de facto semantics of the term.

Load-bearing premise

The load-bearing premise is that definitional ambiguity is harmful rather than benign—that incompatible uses of "hypothesis" actually block comparison, aggregation, and progress; if per-task implicit definitions work adequately, the paper's motivation collapses.

What would settle it

Take two benchmarks with different hypothesis formats (for example, declarative sentences versus context–variable–relationship trees), and have both human coders and trained systems map evidence judgments from one format to the other. If judgments transfer without loss, the claim that ambiguity blocks progress is weakened; if they degrade, the claim is supported.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Benchmark scores in hypothesis-related tasks measure dataset-specific operationalizations, not a shared construct, so comparisons across systems inherit the definitional mismatch.
  • A computable scholarly record depends on machine-readable hypotheses in predictable formats; at minimum, papers should state explicit definitions instead of relying on benchmark datasets as implicit definitions.
  • Standardized or explicit definitions would enable corpus-level tools such as knowledge graphs linking interdisciplinary hypotheses and hypothesis-generation models that transfer across domains.
  • Because "hypothesis generation" currently takes four input types and three output types, it is not one task; progress requires specifying which variant a system addresses.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Inference: A concrete remedy the paper leaves implicit is a shared hypothesis schema—statement type, test status, context, variables, relationships—that all hypothesis-related tasks could map onto.
  • Inference: The argument implies a measurable prediction: cross-dataset evaluation, training on one benchmark's hypotheses and testing on another's, should reveal performance loss if the definitional mismatch matters; the paper does not run that experiment.
  • Inference: Large-language-model hypothesis generators may amplify definitional drift because they are trained on corpora where "hypothesis" is used inconsistently, making explicit definitions more urgent as these systems scale.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This position/review paper surveys how the term 'hypothesis' is defined and operationalized across philosophy of science and, more centrally, across recent NLP/NLU tasks including natural language inference, hypothesis/claim extraction, scientific hypothesis evidencing/claim verification, and hypothesis generation. The authors document that the term has migrated from classical notions (e.g., Plato, Popper, Kuhn) into task-specific forms: declarative sentences in SciTail, questions in SHE, proposals in SciAgents, and ideas in SciMon. They argue that while terminological ambiguity was historically acceptable, it is now a 'critical concern' for building a machine-interpretable scholarly record, and they call for well-structured, machine-readable hypothesis definitions or explicit definitions accompanying datasets. The paper is primarily descriptive, with a prescriptive conclusion about standardization.

Significance. If the central claim is accepted, the paper provides a useful map of a fragmented area and motivates a research program in formalizing scientific hypotheses for NLP. The descriptive survey is a strength: it synthesizes definitions and operationalizations across many recent datasets, and the illustrative table (Table 1) and figure (Figure 1) are helpful for readers entering this area. The paper is transparent in its ambitions, and the recommendation to include explicit definitions in datasets is a sensible, low-cost suggestion. However, the significance of the paper's core claim—that ambiguity is harmful and standardization is critical—rests on an untested normative premise, so the contribution is currently more a research agenda than an established finding.

major comments (3)
  1. [Section 1 and Section 4] The central claim that ambiguity and variability in hypothesis definitions 'is now a critical concern' is asserted but not supported by empirical evidence. The paper itself concedes that shared community understanding often makes such ambiguity harmless, yet it does not provide a concrete failure case for hypotheses: no example of incompatible dataset labels blocking cross-dataset learning, no annotator-disagreement result, no evidence that implicit definitions have degraded model performance. Since the prescriptive conclusion in Section 4 ('machine-readable hypotheses... are critical to this mission') depends entirely on this premise, the paper needs either concrete evidence of harm or a reframed, weaker claim (e.g., 'may become a concern' or 'can impede aggregation').
  2. [Section 2.2, Section 3.3, Figure 2] The proposed taxonomy is not clean, and this weakens the paper's descriptive contribution. In particular, the distinction between 'Claims as hypotheses' and 'Hypothesis-proposals' is not operationally defined, and Figure 2 lists H2 and H3 as identical sentences, undercutting the claim that SHE hypotheses differ from lower-level SCV claims in a consistent way. The reader is left with no principled criterion for when a statement is a hypothesis, a claim, or an idea, which is problematic because the paper's stated goal is to 'delineate various definitions of hypothesis.'
  3. [Section 4] The proposed remedy—standardization or explicit formal definitions—is not developed enough to assess feasibility. The paper does not discuss what a 'well-structured machine-readable hypothesis' should look like, how it would be enforced or validated, or how it would coexist with task-specific variation. Without a concrete proposal (e.g., a minimal schema or a set of desiderata), the recommendation remains a slogan. This is not fatal for a position paper, but it should be acknowledged as an open design problem, and at least one illustrative schema (e.g., the context-variable-relationship tree from DiscoveryBench) could be discussed.
minor comments (4)
  1. [Section 4, last paragraph] The sentence 'standardizing some of to support corpus-level tools' is incomplete; the phrase 'some of' should be removed or completed.
  2. [Table 1, SciMon row] Typo: 'lanuage' should be 'language'.
  3. [Figure 2] H2 and H3 are identical; if intentional, it should be explained; if not, one instance should be corrected or labeled differently.
  4. [Section 2.2, 'Formal expressions'] The transition from philosophical definitions to NLU task-specific forms would benefit from a summary table or enumerated criteria, so that readers can see how the categories relate.

Circularity Check

0 steps flagged

No significant circularity: the paper is a descriptive survey whose conclusions are independent of its authors' prior work.

full rationale

This is a position/review paper, not a derivation: there are no equations, no fitted parameters, and no quantities predicted from inputs. The central claim—that definitional variability around 'hypothesis' is now critical for NLP/NLU operationalization—is an argument supported by a survey of external datasets (SciTail, EntailmentBank, SciFact, DiscoveryBench, etc.), not by the authors' own results. The self-citations that appear ([21], [30], [43]) are used only as examples of existing work (open-science precedents, the SHE dataset, and claim extraction respectively); they do not provide the taxonomy, the evidence of variability, or the prescriptive conclusion. Even the paper's own disclaimer that it is 'more descriptive than prescriptive' reinforces that no circular derivation is attempted. The normative premise that variability is harmful may be underevidenced, but that is an evidentiary/correctness concern, explicitly outside circularity per the review rules. No self-definitional, fitted-input, uniqueness-imported, or ansatz-via-citation pattern is present, so the appropriate finding is no circularity (score 0).

Axiom & Free-Parameter Ledger

0 free parameters · 2 axioms · 0 invented entities

This is a review; the central claim rests on the assumption that the surveyed papers are representative and that the observed variation is consequential. No free parameters or invented entities are introduced.

axioms (2)
  • domain assumption The surveyed NLU tasks and datasets are representative of hypothesis mining research.
    The paper selects examples without a systematic search, so coverage is assumed.
  • ad hoc to paper Variation in hypothesis definitions across tasks undermines knowledge assembly.
    This is the paper's motivating premise, asserted in Section 1 and Section 4 without empirical demonstration.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of What Are Research Hypotheses?." pith.science (2026). https://pith.science/paper/7KSRCWJ6

@misc{pith2026250900185,
  author       = {Pith},
  title        = {Pith review of: What Are Research Hypotheses?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7KSRCWJ6}},
  note         = {Machine review of arXiv:2509.00185}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Over the past decades, alongside advancements in natural language processing, significant attention has been paid to training models to automatically extract, understand, test, and generate hypotheses in open and scientific domains. However, interpretations of the term \emph{hypothesis} for various natural language understanding (NLU) tasks have migrated from traditional definitions in the natural, social, and formal sciences. Even within NLU, we observe differences defining hypotheses across literature. In this paper, we overview and delineate various definitions of hypothesis. Especially, we discern the nuances of definitions across recently published NLU tasks. We highlight the importance of well-structured and well-defined hypotheses, particularly as we move toward a machine-interpretable scholarly record.

Figures

Figures reproduced from arXiv: 2509.00185 by Jian Wu, Sarah Rajtmajer.

Figure 1
Figure 1. Figure 1: An example in the EntailmentBank dataset, demonstrating the entailment tree structure to support the hypothesis (in a shaded box) generated based on the question and answer. The figure is adopted from [35]. Claims as hypotheses. The classical definition of claim is the conclusion or assertion that you want your audience to accept [26]. Adopting this definition, in scientific literature, a claim can be defi… view at source ↗
Figure 2
Figure 2. Figure 2: The collaborative review document structure for an example question and the hypotheses derived from the (question, answer) pairs. [30]. 3.2. Hypothesis and claim extraction Here, the goal is to automatically identify hypotheses from a scientific document. Because hypotheses can be viewed as claims prior to testing, the input, output, and methods for extracting hypotheses and claims are similar. The input d… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

59 extracted references · 46 canonical work pages · 2 internal anchors

  1. [1]

    T. S. Kuhn, D. Hawkins, The structure of scientific revolutions, American Journal of Physics 31 (1963) 554–555. URL: https://doi.org/10.1119/1.1969660. doi: 10.1119/1.1969660. arXiv:https://pubs.aip.org/aapt/ajp/article-pdf/31/7/554/12111921/554_1_online.pdf

  2. [2]

    K. R. Popper, Science as falsification, Conjectures and refutations 1 (1963) 33–39

  3. [3]

    I. Lakatos, History of science and its rational reconstructions, in: PSA: Proceedings of the biennial meeting of the philosophy of science association, volume 1970, Cambridge University Press, 1970, pp. 91–136

  4. [4]

    R. A. Fisher, Statistical methods for research workers, 5, Oliver and Boyd, 1928

  5. [5]

    Fisher, Statistical methods and scientific induction, Journal of the Royal Statistical Society Series B: Statistical Methodology 17 (1955) 69–78

    R. Fisher, Statistical methods and scientific induction, Journal of the Royal Statistical Society Series B: Statistical Methodology 17 (1955) 69–78

  6. [6]

    D. J. Glass, N. Hall, A brief history of the hypothesis, Cell 134 (2008) 378–381

  7. [7]

    Fauconnier, Mappings in thought and language, Cambridge University Press, 1997

    G. Fauconnier, Mappings in thought and language, Cambridge University Press, 1997

  8. [8]

    B. C. Malt, A. Majid, How thought is mapped into words, Wiley Interdisciplinary Reviews: Cognitive Science 4 (2013) 583–597

  9. [9]

    Yarkoni, The generalizability crisis, Behavioral and Brain Sciences 45 (2022) e1

    T. Yarkoni, The generalizability crisis, Behavioral and Brain Sciences 45 (2022) e1

  10. [10]

    Stodden, On emergent limits to knowledge—or, how to trust the robot researchers: A pocket guide, Harvard Data Science Review 6 (2024)

    V. Stodden, On emergent limits to knowledge—or, how to trust the robot researchers: A pocket guide, Harvard Data Science Review 6 (2024)

  11. [11]

    Karasman¯es, V

    V. Karasman¯es, V. Karasmanis, The hypothetical method in Plato’s middle dialogues, Ph.D. thesis, University of Oxford, 1987

  12. [12]

    H. W. Ausland, Socrates’ dialectical use of hypothesis, in: New Perspectives on Platonic Dialectic, Routledge, 2022, pp. 25–50

  13. [13]

    Barnes, Posterior analytics (1994)

    J. Barnes, Posterior analytics (1994)

  14. [14]

    Galilei, Dialogue concerning the two chief world systems, Berkeley: University of California Press.·[1638] (1914)

    G. Galilei, Dialogue concerning the two chief world systems, Berkeley: University of California Press.·[1638] (1914)

  15. [15]

    Newton, Philosophiae naturalis principia mathematica, volume 1, G

    I. Newton, Philosophiae naturalis principia mathematica, volume 1, G. Brookman, 1833

  16. [16]

    Descartes, A discourse on method, JM Dent & Sons Limited, 1912

    R. Descartes, A discourse on method, JM Dent & Sons Limited, 1912

  17. [17]

    Sakellariadis, Descartes’s use of empirical data to test hypotheses, Isis 73 (1982) 68–76

    S. Sakellariadis, Descartes’s use of empirical data to test hypotheses, Isis 73 (1982) 68–76

  18. [18]

    J. F. Quinn, A. E. Dunham, On hypothesis testing in ecology and evolution, The American Naturalist 122 (1983) 602–617

  19. [19]

    Leplin, A novel defense of scientific realism, Oxford University Press, 1997

    J. Leplin, A novel defense of scientific realism, Oxford University Press, 1997

  20. [20]

    Shapere, The structure of scientific revolutions, The Philosophical Review 73 (1964) 383–394

    D. Shapere, The structure of scientific revolutions, The Philosophical Review 73 (1964) 383–394

  21. [21]

    S. M. Rajtmajer, T. M. Errington, F. G. Hillary, How failure to falsify in high-volume science contributes to the replication crisis, Elife 11 (2022) e78830

  22. [22]

    B. A. Nosek, C. R. Ebersole, A. C. DeHaven, D. T. Mellor, The preregistration revolution, Proceedings of the National Academy of Sciences 115 (2018) 2600–2606

  23. [23]

    Mueller, S

    R. Mueller, S. Abdullaev, Deepcause: Hypothesis extraction from information systems papers with deep learning for theory ontology learning (2019)

  24. [24]

    Kumar, T

    S. Kumar, T. Ghosal, V. Goyal, A. Ekbal, Can large language models unlock novel scientific research ideas?, CoRR abs/2409.06185 (2024). URL: https://doi.org/10.48550/arXiv.2409.06185. doi:10.48550/ARXIV.2409.06185. arXiv:2409.06185

  25. [25]

    Q. Wang, D. Downey, H. Ji, T. Hope, Scimon: Scientific inspiration machines optimized for novelty, in: L. Ku, A. Martins, V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, Association for Computational Linguistics, 2024, pp. ...

  26. [26]

    S. E. Toulmin, The uses of argument, Cambridge university press, 2003

  27. [27]

    Heger, A

    T. Heger, A. Algergawy, M. Brinner, J. M. Jeschke, B. König-Ries, D. Mietchen, S. Zarrieß, Nat- ural language hypotheses in scientific papers and how to tame them: Suggested steps for formalizing complex scientific claims, in: Robust Argumentation Machines: First Interna- tional Conference, RATIO 2024, Bielefeld, Germany, June 5–7, 2024, Proceedings, Spri...

  28. [28]

    Alipourfard, B

    N. Alipourfard, B. Arendt, D. M. Benjamin, N. Benkler, M. Bishop, M. Burstein, M. Bush, J. Caverlee, Y. Chen, C. Clark, et al., Systematizing confidence in open research and evidence (score), SocArXiv (2021)

  29. [29]

    R. Vasu, C. Sarasua, A. Bernstein, Scihyp: A fine-grained dataset describing hypotheses and their components from scientific articles, in: International Semantic Web Conference, Springer, 2024, pp. 134–152

  30. [30]

    S. D. Koneru, J. Wu, S. Rajtmajer, Can large language models discern evidence for scientific hypotheses? case studies in the social sciences, in: N. Calzolari, M. Kan, V. Hoste, A. Lenci, S. Sakti, N. Xue (Eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC/COLING 2024, 20-25 ...

  31. [31]

    Wadden, S

    D. Wadden, S. Lin, K. Lo, L. L. Wang, M. van Zuylen, A. Cohan, H. Hajishirzi, Fact or fiction: Verifying scientific claims, in: B. Webber, T. Cohn, Y. He, Y. Liu (Eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, Association for Computational Linguistics, 2020, pp. 7534...

  32. [32]

    Gottweis, W

    J. Gottweis, W. Weng, A. N. Daryin, T. Tu, A. Palepu, P. Sirkovic, A. Myaskovsky, F. Weissenberger, K. Rong, R. Tanno, K. Saab, D. Popovici, J. Blum, F. Zhang, K. Chou, A. Hassidim, B. Gokturk, A. Vahdat, P. Kohli, Y. Matias, A. Carroll, K. Kulkarni, N. Tomasev, Y. Guan, V. Dhillon, E. D. Vaishnav, B. Lee, T. R. D. Costa, J. R. Penadés, G. Peltz, Y. Xu, A...

  33. [33]

    Ghafarollahi, M

    A. Ghafarollahi, M. J. Buehler, Sciagents: Automating scientific discovery through bioinspired multi-agent intelligent graph reasoning, Advanced Ma- terials 37 (2025) 2413523. URL: https://advanced.onlinelibrary.wiley.com/doi/abs/ 10.1002/adma.202413523. doi: https://doi.org/10.1002/adma.202413523. arXiv:https://advanced.onlinelibrary.wiley.com/doi/pdf/10...

  34. [34]

    T. Khot, A. Sabharwal, P. Clark, SciTaiL: A Textual Entailment Dataset from Science Question Answering, Proceedings of the AAAI Conference on Artificial Intelligence 32 (2018). URL: https: //ojs.aaai.org/index.php/AAAI/article/view/12022. doi:10.1609/aaai.v32i1.12022

  35. [35]

    Dalvi, P

    B. Dalvi, P. Jansen, O. Tafjord, Z. Xie, H. Smith, L. Pipatanangkura, P. Clark, Explaining answers with entailment trees, in: M. Moens, X. Huang, L. Specia, S. W. Yih (Eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, Association f...

  36. [36]

    Dagan, D

    I. Dagan, D. Roth, F. Zanzotto, M. Sammons, Recognizing textual entailment: Models and applica- tions, Springer Nature, 2022

  37. [37]

    Storks, Q

    S. Storks, Q. Gao, J. Y. Chai, Recent advances in natural language inference: A survey of benchmarks, resources, and approaches, arXiv preprint arXiv:1904.01172 (2019)

  38. [38]

    e-SNLI-VE: Corrected Visual-Textual Entailment with Natural Language Explanations

    V. Do, O.-M. Camburu, Z. Akata, T. Lukasiewicz, e-snli-ve: Corrected visual-textual entailment with natural language explanations, arXiv preprint arXiv:2004.03744 (2020)

  39. [39]

    Bentivogli, P

    L. Bentivogli, P. Clark, I. Dagan, D. Giampiccolo, The fifth pascal recognizing textual entailment challenge., TAC 7 (2009) 1

  40. [40]

    Shivade, Mednli—a natural language inference dataset for the clinical domain, (No Title) (2017)

    C. Shivade, Mednli—a natural language inference dataset for the clinical domain, (No Title) (2017)

  41. [41]

    Bastan, M

    M. Bastan, M. Surdeanu, N. Balasubramanian, BioNLI: Generating a biomedical NLI dataset using lexico-semantic constraints for adversarial examples, in: Y. Goldberg, Z. Kozareva, Y. Zhang (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2022, Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 2022, pp. 5093–...

  42. [42]

    Claim Extraction in Biomedical Publications using Deep Discourse Model and Transfer Learning

    T. Achakulvisut, C. Bhagavatula, D. Acuna, K. Kording, Claim extraction in biomedical publications using deep discourse model and transfer learning, arXiv preprint arXiv:1907.00962 (2019)

  43. [43]

    X. Wei, M. R. U. Hoque, J. Wu, J. Li, Claimdistiller: Scientific claim extraction with supervised contrastive learning, in: Proceedings of Joint Workshop of the 4th Extraction and Evaluation of Knowledge Entities from Scientific Documents (EEKE2023) and the 3rd AI + Informetrics (All2023) co-located with the JCDL 2023, Santa Fe, New Mexico, United States,...

  44. [44]

    White, K

    E. White, K. B. Cohen, L. Hunter, The cisp annotation schema uncovers hypotheses and explana- tions in full-text scientific journal articles, in: Proceedings of BioNLP 2011 Workshop, BioNLP ’11, Association for Computational Linguistics, USA, 2011, p. 134–135

  45. [45]

    Lancet, Information for Authors, https://www.thelancet.com/pb-assets/Lancet/authors/ tl-info-for-authors-1740074875577.pdf, 2025

  46. [46]

    Saakyan, T

    A. Saakyan, T. Chakrabarty, S. Muresan, COVID-fact: Fact extraction and verification of real-world claims on COVID-19 pandemic, in: C. Zong, F. Xia, W. Li, R. Navigli (Eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Pap...

  47. [47]

    Sarrouti, A

    M. Sarrouti, A. Ben Abacha, Y. Mrabet, D. Demner-Fushman, Evidence-based fact-checking of health-related claims, in: M.-F. Moens, X. Huang, L. Specia, S. W.-t. Yih (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2021, Association for Computational Linguistics, Punta Cana, Dominican Republic, 2021, pp. 3499–3512. URL: https://aclan...

  48. [48]

    B. P. Majumder, H. Surana, D. Agarwal, B. D. Mishra, A. Meena, A. Prakhar, T. Vora, T. Khot, A. Sab- harwal, P. Clark, Discoverybench: Towards data-driven discovery with large language models, in: The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, OpenReview.net, 2025. URL: https://openreview.net/...

  49. [49]

    BahadarKhan, A

    K. BahadarKhan, A. A Khaliq, M. Shahid, A morphological hessian based approach for retinal blood vessels segmentation and denoising using region based otsu thresholding, PLOS ONE 11 (2016) 1–

  50. [50]

    doi:10.1371/journal.pone.0158996

    URL: https://doi.org/10.1371/journal.pone.0158996. doi:10.1371/journal.pone.0158996

  51. [51]

    Y. Zhou, H. Liu, T. Srivastava, H. Mei, C. Tan, Hypothesis generation with large language models, in: L. Peled-Cohen, N. Calderon, S. Lissak, R. Reichart (Eds.), Proceedings of the 1st Workshop on NLP for Science (NLP4Science), Association for Computational Linguistics, Miami, FL, USA, 2024, pp. 117–139. URL: https://aclanthology.org/2024.nlp4science-1.10...

  52. [52]

    C. Tan, L. Lee, B. Pang, The effect of wording on message propagation: Topic- and author- controlled natural experiments on Twitter, in: K. Toutanova, H. Wu (Eds.), Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, Baltimore, Maryland, 2014, pp. 175–1...

  53. [53]

    Touvron, L

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhar- gava, S. Bhosale, D. Bikel, L. Blecher, C. Canton-Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Ko...

  54. [54]

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, D. Amodei, ...

  55. [55]

    C. Si, D. Yang, T. Hashimoto, Can llms generate novel research ideas? A large-scale human study with 100+ NLP researchers, in: The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, OpenReview.net, 2025. URL: https: //openreview.net/forum?id=M23dTGWCZy

  56. [56]

    Y. Pu, T. Lin, H. Chen, Piflow: Principle-aware scientific discovery with multi-agent collabora- tion, CoRR abs/2505.15047 (2025). URL: https://doi.org/10.48550/arXiv.2505.15047. doi: 10.48550/ ARXIV.2505.15047. arXiv:2505.15047

  57. [57]

    M. Chai, E. Herron, E. Cervantes, T. Ghosal, Exploring scientific hypothesis generation with mamba, in: L. Peled-Cohen, N. Calderon, S. Lissak, R. Reichart (Eds.), Proceedings of the 1st Workshop on NLP for Science (NLP4Science), Association for Computational Linguistics, Miami, FL, USA, 2024, pp. 197–207. URL: https://aclanthology.org/2024.nlp4science-1....

  58. [58]

    Z. Yang, X. Du, J. Li, J. Zheng, S. Poria, E. Cambria, Large language models for automated open- domain scientific hypotheses discovery, in: L.-W. Ku, A. Martins, V. Srikumar (Eds.), Findings of the Association for Computational Linguistics: ACL 2024, Association for Computational Linguistics, Bangkok, Thailand, 2024, pp. 13545–13565. URL: https://aclanth...

  59. [59]

    B. Qi, K. Zhang, H. Li, K. Tian, S. Zeng, Z.-R. Chen, B. Zhou, Large language models are zero shot hypothesis proposers, arXiv preprint arXiv:2311.05965 (2023)

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.