Pith. sign in

Paper Citation Record · LEDGER

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models

As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2505.14599.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14599 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:35:44.323361Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:33:46.563630Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:28:44.597412Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact6
  • verified fuzzy28
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5197324c-4360-41e6-9a7a-06d6e70faee5 · outbound

This paper cites Scientific Hypothesis Generation by a Large Language Model: Laboratory Validation in Breast Cancer Treatment.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Scientific Hypothesis Generation by a Large Language Model: Laboratory Validation in Breast Cancer Treatment

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:35:44.781859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.097720Z digest=sha256:85769922b5c76c822f35ee6768be4c9e457e0bc38841ce49eadf2a3f1605a3e9

Observation 98fee1f0-be9b-47d0-b0d5-203e00306dc6 · outbound

This paper cites GPT-4 Technical Report.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.102339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.102339Z digest=sha256:e4c5a533b98a7a6c617b7439d9ff3fe3d5570ca7ef4fdc1d788267f66240f8f4

Observation d75a0f51-d44e-4073-9724-2b0c754c90ad · outbound

This paper cites ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.106103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.106103Z digest=sha256:67abe4877d241dd1811fadb9c873c1e769d1292259cb102acad3ae4626d80669

Observation b968f917-a696-46f1-ae1f-d11dcaa7c6e1 · outbound

This paper cites Harnessing the Power of Adversarial Prompting and Large Language Models for Robust Hypothesis Generation in Astronomy.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Harnessing the Power of Adversarial Prompting and Large Language Models for Robust Hypothesis Generation in Astronomy

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.109864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.109864Z digest=sha256:0ce80f8d0a51887838efb5dfb8cc2f180737bb2407a035240610e4bd398bd472

Observation 21bc12ea-1b9a-4e55-b724-de1c9e31d1d0 · outbound

This paper cites MARG: Multi-Agent Review Generation for Scientific Papers.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models MARG: Multi-Agent Review Generation for Scientific Papers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.114139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.114139Z digest=sha256:e1edacd83b0f7383801539fd1ad13bb78b9fb3bff985f58c3a2e7cf2fda2aae9

Observation d6b4a4ee-51b0-425e-bbdc-a078bd72cb90 · outbound

This paper cites Towards A Rigorous Science of Interpretable Machine Learning.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Towards A Rigorous Science of Interpretable Machine Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.117891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.117891Z digest=sha256:fe111d0a9673b2ac7aa9909746dbd5dc89f4800e4aca056109396239a52984dd

Observation 821683c7-0cae-4cae-aff7-03eae1b46040 · outbound

This paper cites The Llama 3 Herd of Models.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.122791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.122791Z digest=sha256:8c51acfc9b3f7e5cd3ea4d953e4808d51413fb4eb47646d4937587175f9cebb7

Observation 052e5752-5177-406d-9e1e-279669a86186 · outbound

This paper cites On the creativity of large language models.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models On the creativity of large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.167874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.126501Z digest=sha256:1bf9566ff20a245b59d19a201661d881213fae71614b867baac336e278990535

Observation 2c560203-9948-4f9e-be1c-1a8d1fd1837c · outbound

This paper cites Forecasting high-impact research topics via machine learning on evolving knowledge graphs.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Forecasting high-impact research topics via machine learning on evolving knowledge graphs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.130125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.130125Z digest=sha256:400f7c6b697bd6e55964e1df0b5b79dd676a2a92558d83ce4c6354cc8d655735

Observation 078e0045-d96a-4f5b-ba3d-4d2ab23a8e3b · outbound

This paper cites Embracing foundation models for advancing scientific discovery.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Embracing foundation models for advancing scientific discovery

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.155434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.134128Z digest=sha256:003626e022afe697f84ed58ce96ec55b5e742073ce368c80a92c840522e9e16d

Observation ce3f7d09-122c-4d91-8f0b-3a17a80fa916 · outbound

This paper cites Williams, Stefan Bekiranov, and Aidong Zhang.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Williams, Stefan Bekiranov, and Aidong Zhang

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.141899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.137708Z digest=sha256:d1fe005eb70bff499998ad16690aa670da4eb395b5024d12371835f2741f0e79

Observation e72d0f18-290b-4063-9601-5531f8ea2455 · outbound

This paper cites Nova: An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Nova: An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.141222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.141222Z digest=sha256:724ac9d58317e8af26e68d088e4b3c3df9608696c0ee492849c2019f73632ef3

Observation 92c8de35-1753-4ea2-b776-27d4d7a51810 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.145049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.145049Z digest=sha256:e7def702ab353f20c3477e216dd97fc9d72b10ae8b5ad6e5a487208cca6b8873

Observation 0f3e372c-3071-4cf0-904a-7b26a0249b90 · outbound

This paper cites Autonomous llm-driven research—from data to human-verifiable research papers.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Autonomous llm-driven research—from data to human-verifiable research papers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.129310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.148710Z digest=sha256:d8d12e9e137de3bcae9192c5ca0c8b8f614d7b9ec8d7bf2616ebc7c1c8cdfe48

Observation 9a6072e4-059b-4e8d-b8f1-60a953be7027 · outbound

This paper cites A survey on knowledge graphs: Representation, acquisition, and applications.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models A survey on knowledge graphs: Representation, acquisition, and applications

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.116440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.152298Z digest=sha256:bfc8e18d6d09c58494212d894ef80f05f6c79681c92ec7f697707b218162a2d9

Observation dcc46d8b-f53b-4709-a0b4-9691795dea8b · outbound

This paper cites Entry-level guide to the use of large language models for medical research.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Entry-level guide to the use of large language models for medical research

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:35:44.642806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.155536Z digest=sha256:77d48e3298db9d14fea78cdd8e7bb124bcbcc659ecac2ecccaf6575c403b44c1

Observation 503862d8-5a8a-48d2-b160-cc04042c0288 · outbound

This paper cites Large language models versus natural language understanding and generation.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Large language models versus natural language understanding and generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.103276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.159783Z digest=sha256:5fc12cffb089256c4cfd5306c417c469b1d437a81b21d9d468247e23c0721a00

Observation 145eb18e-99cd-41ff-9318-6d895e73f6bc · outbound

This paper cites Forecasting the future of artificial intelligence with machine learning-based link prediction in an exponentially growing knowledge network.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Forecasting the future of artificial intelligence with machine learning-based link prediction in an exponentially growing knowledge network

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.089819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.164595Z digest=sha256:d3f0ef7a2c589babf8fc187a0ce8d3cadbea4833beb1dadd7cd2e7a81ec8c191

Observation 3b121a38-41eb-4888-8e39-c167c0f2ccec · outbound

This paper cites MyCrunchGPT: A chatGPT assisted framework for scientific machine learning.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models MyCrunchGPT: A chatGPT assisted framework for scientific machine learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:35:44.624537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.168145Z digest=sha256:c9c12ff004dc6f5286174e98cc7aaa383dd7a3e1f3f49f0d5c12925b7673374a

Observation a58a1df5-fa8b-4564-81cc-b1a410462f83 · outbound

This paper cites PaperQA: Retrieval-Augmented Generative Agent for Scientific Research.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models PaperQA: Retrieval-Augmented Generative Agent for Scientific Research

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.171915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.171915Z digest=sha256:e263c8311b0f565cf593b84177b69f79118bb51e0f317311644a0ef147f161b4

Observation c74ed565-c5e3-4cc6-b800-1924174ea04b · outbound

This paper cites u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.176053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.176053Z digest=sha256:67a89e904e97a0193868070aeba077ff887a8a2fa881cc12b138086c7c9c5da5

Observation a5522196-0f18-4a30-afeb-0cd1dd39e4bf · outbound

This paper cites Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.179758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.179758Z digest=sha256:729fcff9d22afb43f449e621968ddd403946e8afeff534763e5a8e1723e180ff

Observation 65eece7d-9619-4b6d-b092-ef1cbcf8df27 · outbound

This paper cites Learning entity and relation embeddings for knowledge graph completion.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Learning entity and relation embeddings for knowledge graph completion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.067795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.184160Z digest=sha256:d0f44ec19b42ebd738a077ed44dcb38574ecfafc73dfb98ef8d8bd42bacb513e

Observation a34e5c1d-b18a-41b8-bbb1-7aff8bfd6039 · outbound

This paper cites A Survey on Graph Classification and Link Prediction based on GNN.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models A Survey on Graph Classification and Link Prediction based on GNN

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.187587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.187587Z digest=sha256:02d40f01f07619e38f1fc845a64c124a18ea616f98e996eaa4b703642d63f680

Observation b550870d-7087-4183-90c4-a14ffc558b16 · outbound

This paper cites Conversational drug editing using retrieval and domain feedback.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Conversational drug editing using retrieval and domain feedback

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.054613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.191267Z digest=sha256:d8deaffd89a9b49ed68e6a091884bf7945a9bc8aa36503e5132f91b883de9c42

Observation b949a6cd-c77f-416c-96e9-f17a3fcd3dbe · outbound

This paper cites Application of explainable artificial intelligence for healthcare: A systematic review of the last decade (2011--2022).

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Application of explainable artificial intelligence for healthcare: A systematic review of the last decade (2011--2022)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.041748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.194938Z digest=sha256:ef22437c546e43edf9e120aa5342d72fb84530ab30ec1b47364e60fd0dd156fb

Observation 22fa719a-227b-4dfd-9a5a-2c7acaaf6b93 · outbound

This paper cites Improving biomedical information retrieval with neural retrievers.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Improving biomedical information retrieval with neural retrievers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.029617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.198460Z digest=sha256:7d29b47b647e3a926d6cdcd7b6ddad096a0627218719177895305f3f2f7a6945

Observation 818de6a1-ae14-483f-bbc2-8d6a5fbedddc · outbound

This paper cites Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.017323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.202006Z digest=sha256:5349eeadfe3b9216da8db50e47ba5030d595347b6b3854ddc421465f34b3cd03

Observation 380e70f7-1d1a-4986-974f-11593da7ecb7 · outbound

This paper cites Think-on-graph 2.0: Deep and interpretable large language model reasoning with knowledge graph-guided retrieval.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Think-on-graph 2.0: Deep and interpretable large language model reasoning with knowledge graph-guided retrieval

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:45.005412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.205450Z digest=sha256:d3f82dfdc0d9320f9ef32c62c9c2642b5d45991795361ea39d9ca15d427e6a6b

Observation d03f8af5-cd5b-43c4-b2cd-eed17ec8cbd8 · outbound

This paper cites Explainable ai is dead, long live explainable ai! hypothesis-driven decision support using evaluative ai.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Explainable ai is dead, long live explainable ai! hypothesis-driven decision support using evaluative ai

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.993238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.208976Z digest=sha256:fdcd0be2721c03ab42e6fbc63c42a382a0a6b20ea7f20d8b4b64766bb7f1adb3

Observation 64981239-ccbe-4341-b08d-2e7cd7eb267c · outbound

This paper cites Evaluating the Effectiveness of Retrieval-Augmented Large Language Models in Scientific Document Reasoning.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Evaluating the Effectiveness of Retrieval-Augmented Large Language Models in Scientific Document Reasoning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:35:44.566829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.212531Z digest=sha256:4db0670dd0369e4429496da1857d1a1686fb04aac49d31ebad23ac1fc05b1e0c

Observation 834501c2-8c27-4588-ac62-e064f7f20e9a · outbound

This paper cites A review of relational machine learning for knowledge graphs.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models A review of relational machine learning for knowledge graphs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.981066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.216249Z digest=sha256:1621cdc21bb2de4abe4922659f23b8162e547016ab283fd627368440baf63e11

Observation 57c4f35d-7d46-41a5-9edf-ae47dc2b0b8c · outbound

This paper cites Can chatgpt be used to generate scientific hypotheses? Journal of Materiomics , 10(3):578--584, 2024.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Can chatgpt be used to generate scientific hypotheses? Journal of Materiomics , 10(3):578--584, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.967649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.219755Z digest=sha256:e155bc160dfb6ab249798cc09593df77dd23c6154a7e516c89c2f5f901d0a59a

Observation 060dc8bf-12ba-44c0-b17a-46f5cf51c4b6 · outbound

This paper cites Graph Retrieval-Augmented Generation: A Survey.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Graph Retrieval-Augmented Generation: A Survey

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.223554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.223554Z digest=sha256:3add8a8f153f97fa039553cb7c011281859946db9ead063c2f3d9fcb550649d3

Observation 78f50942-e345-4819-8346-23cc13c1cc7c · outbound

This paper cites Large Language Models are Zero Shot Hypothesis Proposers.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Large Language Models are Zero Shot Hypothesis Proposers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.227331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.227331Z digest=sha256:70e887df812e8efe46e0ba824839664299a042a4bde79f962eb02bcc6a57b974

Observation 81fd97d5-92ed-4bec-94cb-3c4fae8e3f0f · outbound

This paper cites Large language models as biomedical hypothesis generators: A comprehensive evaluation.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Large language models as biomedical hypothesis generators: A comprehensive evaluation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.953922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.230991Z digest=sha256:97a3e95271418d1b62cabc66b03fc932edfb791e2aa0fcc8cec0b00ac1c852c6

Observation 1534be72-5362-4fda-be8a-38b4207115f9 · outbound

This paper cites Human-LLM Compound System for Scientific Ideation through Facet Recombination and Novelty Evaluation.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Human-LLM Compound System for Scientific Ideation through Facet Recombination and Novelty Evaluation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.234504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.234504Z digest=sha256:d50ab9bdb2d407da48cf6b98d42861e92c5e1dfe70a3b32928d65a7c67725f0f

Observation 33beb8a4-b280-4b9e-8697-6ddda6f0e9d1 · outbound

This paper cites A review on large language models: Architectures, applications, taxonomies, open issues and challenges.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models A review on large language models: Architectures, applications, taxonomies, open issues and challenges

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.941574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.238181Z digest=sha256:d7bb6bad576c9dfa3161e4c56cf4ff44355775ee9e6cf90654c8ce1b27242de2

Observation a61408ac-1252-401d-9686-7f6bad901554 · outbound

This paper cites The probabilistic relevance framework: Bm25 and beyond.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models The probabilistic relevance framework: Bm25 and beyond

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.928045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.241541Z digest=sha256:7f1541ec632acbb49c2e8ed4624f32362786da3ad4f7c8435678f5b85c14c125

Observation e67400fa-6fdd-4b6c-929c-c5fbb44557c1 · outbound

This paper cites Knowledge Graph Large Language Model (KG-LLM) for Link Prediction.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Knowledge Graph Large Language Model (KG-LLM) for Link Prediction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.244864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.244864Z digest=sha256:7c2682a06d1737552b3867b7d0495bbd79ee19f7a649d60cbb23fdd32d066c2e

Observation f52cadca-fd1e-43a3-9402-851ebec90d94 · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.248638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.248638Z digest=sha256:45dad9b33320ae7d4dea15e24c7bffa49102277eecc73afdffdc759d9268ef33

Observation b9295f60-5ccf-4673-a227-772be0e047a6 · outbound

This paper cites Colidr: Concept learning using aggregated disentangled representations.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Colidr: Concept learning using aggregated disentangled representations

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.915303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.252434Z digest=sha256:2331b24211dfc24248290f2a1f61bc5d1b3640af77e713e1f24e208085f8c1a3

Observation fa061371-de2f-413d-83fa-dfa00be4ecdd · outbound

This paper cites A self-explaining neural architecture for generalizable concept learning.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models A self-explaining neural architecture for generalizable concept learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.902142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.255801Z digest=sha256:ad718bad8523c77d40d8c360c1f4621fa74954b8049d4b5fa61a33ed42f050d6

Observation f8496503-a5ca-408d-b310-6e7bcb19d0a2 · outbound

This paper cites Language agents achieve superhuman synthesis of scientific knowledge.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Language agents achieve superhuman synthesis of scientific knowledge

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.259538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.259538Z digest=sha256:64d5fc06679fda5624ff3e0a5ac03162c0e04dd771a91bc111a58b7519af2b58

Observation 7663d7d0-ddcd-4e91-921a-6d0bdf53a7c0 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.263315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.263315Z digest=sha256:81ca4039155de1fe188a4b2a8285a469304dd03f2d9fd2d1fea164349fc9ebe3

Observation 557ce5cc-e533-4404-af30-2349996f2b3f · outbound

This paper cites SciMON: Scientific Inspiration Machines Optimized for Novelty.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models SciMON: Scientific Inspiration Machines Optimized for Novelty

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.267051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.267051Z digest=sha256:957720b192bd5c70a2f36cd68f262feec9010af2877dbba474809b41f0ff5f85

Observation 2d3f3c92-ffb8-4b40-9e99-966639debedd · outbound

This paper cites Knowledge Graph Retrieval-Augmented Generation for LLM-based Recommendation.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Knowledge Graph Retrieval-Augmented Generation for LLM-based Recommendation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.271010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.271010Z digest=sha256:1143ebf6fee3255b546244419875578865b1d71677d7419f67a213a81c9e757d

Observation d8168298-e13f-4af9-ba2a-01d69ba1dee6 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.274755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.274755Z digest=sha256:741775425b2d0b46d25139f5c8970540ab036211cfbb0204248513ef241105c8

Observation d2646951-98f5-4df8-be59-cc0c32d6d9c2 · outbound

This paper cites Pubtator 3.0: an ai-powered literature resource for unlocking biomedical knowledge.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Pubtator 3.0: an ai-powered literature resource for unlocking biomedical knowledge

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.881358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.278183Z digest=sha256:5532b0fd8384a4b370cecb2e814f3cbefc7c7b014be5869a92f3929b98150dbf

Observation 3910ec8a-be06-49eb-bad5-4df7c989861e · outbound

This paper cites Generating Scientific Claims for Zero-Shot Scientific Fact Checking.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Generating Scientific Claims for Zero-Shot Scientific Fact Checking

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:35:44.425129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.281944Z digest=sha256:455b6061605fcc449feef935aabf52e7e0cf81477f0f58d647f89954e2dafcd8

Observation df39b9bf-4bd9-433d-b2d7-714b21ad6694 · outbound

This paper cites Dynamic link prediction using graph representation learning with enhanced structure and temporal information.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Dynamic link prediction using graph representation learning with enhanced structure and temporal information

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.868411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.285609Z digest=sha256:9142742b8c7d1d259f68ec972a2a8aca95142b6af9191654d55e2ccc50cdee3a

Observation 5cc55a5d-4f8e-41f6-9ed5-a8f1c8cf6990 · outbound

This paper cites Benchmarking retrieval-augmented generation for medicine.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Benchmarking retrieval-augmented generation for medicine

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.854624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.289178Z digest=sha256:0aaac1a7ef4ad2fe36c974bf01425103b98a0f1c1ab9acdc541908b0f0c32b93

Observation 09e65ccd-a000-4e53-82a3-0cf2a16c0f9f · outbound

This paper cites Improving retrieval-augmented generation in medicine with iterative follow-up questions.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Improving retrieval-augmented generation in medicine with iterative follow-up questions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.841311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.293238Z digest=sha256:96d4121e4fa3bebbc7913c861fe4d89f78fc20c69f12c214444a476b458ee2b3

Observation 4567ab4b-5647-475f-9c88-b6b85dbfc96c · outbound

This paper cites Improving Scientific Hypothesis Generation with Knowledge Grounded Large Language Models.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Improving Scientific Hypothesis Generation with Knowledge Grounded Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.296888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.296888Z digest=sha256:75cec4e6db82b544fae656d5b51ad6a263ad1f915e9c8d7499b5f86196404a88

Observation ac650f2b-899c-4d49-a0ba-9dcbdf241ec6 · outbound

This paper cites Large Language Models for Automated Open-domain Scientific Hypotheses Discovery.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Large Language Models for Automated Open-domain Scientific Hypotheses Discovery

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.300728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.300728Z digest=sha256:396cd9e529a32953cf25e8765e8ad2780c0ca2a68530037cb021c6f1284737de

Observation b892c00a-f5e6-4014-a4af-1942b5aa8c31 · outbound

This paper cites Large language models for rediscovering unseen chemistry scientific hypotheses.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Large language models for rediscovering unseen chemistry scientific hypotheses

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.828609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.304539Z digest=sha256:2a3a0be58b540e574e95f7a9699308af1a8f8a458b846764d0d2a0360e890000

Observation c8c4d3e3-d8df-46f4-8bd0-0db061e0ce4c · outbound

This paper cites Scientific Opinion Summarization: Paper Meta-review Generation Dataset, Methods, and Evaluation.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Scientific Opinion Summarization: Paper Meta-review Generation Dataset, Methods, and Evaluation

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:35:44.379404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.308287Z digest=sha256:f12d9362f1ee5f118f5ff2841458b1b70801063bfd1170d3b69da3540b07316e

Observation 55e1809b-e6eb-4843-b4e9-b048cf3c4d79 · outbound

This paper cites Link prediction based on graph neural networks.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Link prediction based on graph neural networks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.816240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.312383Z digest=sha256:fc372f7f8759ff7a477c4aa5f9430b814aaa37903be7e780d5ee08cae7e56078

Observation 342c9668-7aaf-4f82-8ce1-61aa02e612b4 · outbound

This paper cites Goal driven discovery of distributional differences via language descriptions.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Goal driven discovery of distributional differences via language descriptions

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:35:44.803688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:35:44.316068Z digest=sha256:103f7a946f2e76c4bd96a649ae6a102cafdcb5e7e9e95d4cee3e287410c81d70

Observation dcecd143-52f7-4e3a-8f13-682e8c310db2 · outbound

This paper cites Hypothesis Generation with Large Language Models.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models Hypothesis Generation with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.319549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.319549Z digest=sha256:8f5ab2b7a2771b223c4feffa7e243effa968a064e395f0ec2a25177c6144e3c3

Observation 92ce117e-cd9d-49aa-9c94-6730d710cb2e · outbound

This paper cites write newline.

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:44.323361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:44.323361Z digest=sha256:8be04e405a81779c1be1fcb6db635ae8a305e3ce38536bb0ccb07f401225bf07

Pith citing papers

Observation 6b7b2aa5-0e2d-4124-8482-8d7427b61194 · inbound

Interestingness First Classifiers cites this paper.

Interestingness First Classifiers Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T15:33:46.563630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:33:46.563630Z digest=sha256:9b8dbef1d9edd4f2c10278a02980916a294d092ea429d3ae0ddb68b9786bbdf6

Observation 44bf4178-6d59-45d0-879b-bc8a3ba14c66 · inbound

BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation cites this paper.

BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:44.598968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T04:07:26.224919Z digest=sha256:9a6b7d17b8d755ba16edf69973fb1f9b0301d62e64f8914e4284f164edfdc5c5