Pith. sign in

Paper Citation Record · LEDGER

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes

As of 20 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 2 inbound Pith citation observations for arXiv:2507.22940.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22940 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:03:15.426142Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T12:08:10.552789Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved56
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a15e3588-c29a-4b9e-a209-80d92dd4745d · outbound

This paper cites Language models are few-shot learners,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Language models are few-shot learners,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.034973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.034973Z digest=sha256:12198099dba83a1e9fe8f955cb5fcbb6a350133271e99bc8b5970d8db412b9f0

Observation a703d20e-27af-49b3-890e-7be3277f1b41 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Constitutional AI: Harmlessness from AI Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.039775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.039775Z digest=sha256:119c691dbc0974dad51f43ac9d9fe7578544936c37148cde3f39fc17befb1cec

Observation ddd8c491-abf3-4d90-8371-5138374303c5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.045001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.045001Z digest=sha256:59699525f78d29379684f2f4a6b5b01f988b9556aac09d917250caa0f049032b

Observation 1281603e-5df3-4ead-8fad-1d3b24c5bf5b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Gemini: A Family of Highly Capable Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.049765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.049765Z digest=sha256:b523f4050b3f4b0d3a04346386a215e836e89560803f68df7f79b572f9e486bc

Observation 0bcb41a8-d4a3-4cd1-b4bd-d24ef379f2bf · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.054477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.054477Z digest=sha256:00e07ccb77cc0023828c0cb901ef65cd530abe1b240b81e02b511167c3a4d79c

Observation 2c7dee8c-23be-4991-b986-28beaf76540e · outbound

This paper cites Qwen Technical Report.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Qwen Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.059198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.059198Z digest=sha256:6c76a5dd2e04af3f4b41a75085ba7975542311ed1a198df944183fcd62c1b37d

Observation 670d17e9-ea99-4c80-8b44-5ffbab56cb8b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.064263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.064263Z digest=sha256:b8e2b6ce3f325301f4d0f878bb4567230372dbd37cd5dfc2fe361cff7446855e

Observation 45787ff0-20a7-45c5-99a3-97e623309e16 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Qwq-32b: Embracing the power of reinforcement learning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.073137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.073137Z digest=sha256:576a3088377de6b9b655ed9cd622d1ee061e582dc2db555a7fbefbb92298b297

Observation 35d7ce23-942b-49c5-ae84-0aff7cbff47b · outbound

This paper cites Learning to reason with llms,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Learning to reason with llms,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.077594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.077594Z digest=sha256:d8f3d5134ebfb03d3f0a9ef9218c2738a5693b1a0fea481e986845863020266a

Observation 7c8e32c1-5a10-46cd-ab4f-db81477f63e4 · outbound

This paper cites Claude 3.7 sonnet and claude code,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Claude 3.7 sonnet and claude code,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.082334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.082334Z digest=sha256:10f2e0512cc4a760669888a017142c43c390ee7989ca1a910a87c6f45dfa8ffa

Observation f3ddf8ae-848d-484b-a1e6-3c2555323dd2 · outbound

This paper cites Large Language Models for Disease Diagnosis: A Scoping Review.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Large Language Models for Disease Diagnosis: A Scoping Review

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.086673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.086673Z digest=sha256:eaa25bbffddcd9e8f019a7d03b1a518b56f86269b96c461bf132214e0bcb3273

Observation 325289f8-ae28-4f7c-bacf-36aeac9ad512 · outbound

This paper cites Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.091368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.091368Z digest=sha256:00d524b30230220b5cf1a26cd7700c04fbac6940a1ee571b59f5a286b8424f7e

Observation d2df7d58-7eb6-465e-a51a-7b607ffab679 · outbound

This paper cites Towards Robust Legal Reasoning: Harnessing Logical LLMs in Law.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Towards Robust Legal Reasoning: Harnessing Logical LLMs in Law

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.095819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.095819Z digest=sha256:d50728754a7f5c6fa882aaead9c6991c4fa5340ce283483f7f9226091159080b

Observation 5504c403-47e9-4d4a-a0ca-2b3860540945 · outbound

This paper cites Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.100738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.100738Z digest=sha256:a78e5293eddbb589186c10883699db1be53785cff836e475d047a5005e02c36f

Observation 8d189ea0-b83d-4ba7-8d00-7ece3006bd71 · outbound

This paper cites INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.105798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.105798Z digest=sha256:7b3e879bdd63ee652f2829ebbb5d073e548a3a3b9bd65348082c9bb0eac97e5e

Observation 7006b9f1-b8a9-4c53-83a3-6dd0ff2fdb86 · outbound

This paper cites Finqapt: Empowering financial decisions with end-to-end llm-driven question answering pipeline,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Finqapt: Empowering financial decisions with end-to-end llm-driven question answering pipeline,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.110881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.110881Z digest=sha256:5e063b8abb9ce645046d6a400903ca415bd73e785653986b3c7620688a96d06b

Observation c604c2ed-a02e-4b43-9574-7c5cdbdb4d51 · outbound

This paper cites Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language-Models.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language-Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.115336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.115336Z digest=sha256:43e8bcabe16ba0b7d8fd13141880a2da7a12de95a0b42da05f0db0874d1323a9

Observation b2deed91-ab7e-489b-93ba-505d2a67a9f3 · outbound

This paper cites CMMLU: measuring massive multitask language understanding in chinese,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes CMMLU: measuring massive multitask language understanding in chinese,

Reference 18

Resolution
malformed identifier
no resolver link, observed 2026-08-15T18:03:15.120245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.120245Z digest=sha256:7dc295cfd070e44440379e80ca87c9bb7673eb2bd61a856f9de256da63ce7f94

Observation 737419bd-99e6-401f-956e-3993d1de2778 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Truthfulqa: Measuring how models mimic human falsehoods,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.124874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.124874Z digest=sha256:6d7300e6baa17944e27a98f4f5452aed614d1c14aec0e59151027b1dfa7071f6

Observation eba5314c-e4e5-4c89-be09-c2bc3b52ce52 · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.772475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.129559Z digest=sha256:0035ce9888299cc16571ce68b582c35a91fc565a740ab6b649c421f91d2eca01

Observation ac486191-a068-4dfb-a98a-4e058a533abd · outbound

This paper cites Factuality enhanced language models for open-ended text generation,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Factuality enhanced language models for open-ended text generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.758204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.134215Z digest=sha256:21459c13b4a809cb8744150e86889ff52baea4e4e8cc305615a84aa8e1e8574d

Observation 0faf971c-8388-4047-b042-60b13998e140 · outbound

This paper cites SKILL: structured knowledge infusion for large language models,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes SKILL: structured knowledge infusion for large language models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.138673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.138673Z digest=sha256:a5846d9a79c9ff42fca6e980fbbe2534f872d63991573476072a965fb3f5514f

Observation 4cb5748c-c181-4272-9707-55de5f7821af · outbound

This paper cites Contrastive learning reduces hallucination in conversations,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Contrastive learning reduces hallucination in conversations,

Reference 23

Resolution
verified exact
doi, observed 2026-08-15T18:03:16.742961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.143255Z digest=sha256:318a10fc2f0f0a8291868438f1e9333d7b3314a01961ede759ea72ce7ef663e2

Observation 8b54acc6-1798-464c-91bf-c21693d8f3b5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.148064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.148064Z digest=sha256:ef5735d35055083df42c08f6a02127cac77746d37c300c0fb1dbd2c5e67f1231

Observation 1461a0bd-247b-476f-81a4-fffb29c0e41e · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Mechanistic Interpretability for AI Safety -- A Review

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.157880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.157880Z digest=sha256:a7c11f5a8c4d7deff773519c82ef5ad3ff9c1d460de382acc1e60abf5f8eb0b6

Observation 3e875050-bdb9-4f1b-9283-55b503193cb2 · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big?.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes On the dangers of stochastic parrots: Can language models be too big?

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.727790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.162661Z digest=sha256:a3e09d0933b55eec249d524e5fda047404a26c7886ed9c62c67f83ba08a17a62

Observation e2439061-1e03-432d-9905-3a696d96c87d · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.171506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.171506Z digest=sha256:39aa7c03f323d15eda4c2019eba59ca98fe7f56f9f77d418409db26e9e2a2ee2

Observation 26aa5c46-90c4-41e5-86b4-be766c1a27e7 · outbound

This paper cites GPT-4 Technical Report.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes GPT-4 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.176522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.176522Z digest=sha256:dd3f2ca1baad993d28ef285cba9706f24f6ef0c65a47daefd7905359e6826552

Observation b589ed74-e21d-4091-87d5-2473db998aca · outbound

This paper cites Microsoft Bing: Get to know Bing,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Microsoft Bing: Get to know Bing,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.712392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.185968Z digest=sha256:e31f60e4eac7975048203418d91e937f3ccf3e928d54a5395f0f82675edd7566

Observation f5384afd-c889-400f-ba3c-8c6333b7c8d6 · outbound

This paper cites [Online].

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes [Online]

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.680863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.195095Z digest=sha256:d0fbcb551ad53919403321064f6920175ed3acf6300fc4307a7d244455f3c4a6

Observation 14189d15-bb15-4b9c-b4af-dba1b88fafe2 · outbound

This paper cites an unresolved cited work.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:03:16.665886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.199452Z digest=sha256:b5de482b38b49acb8ab699f0d3ffbfa9521fdb9a72d52fbfcc4c4e68e13c32fa

Observation ca075c6c-b6f4-482d-962e-e4e02412d3e5 · outbound

This paper cites an unresolved cited work.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:03:16.651350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.203943Z digest=sha256:404cd3c3caf952c2916759bca1a5af25c6536f7589ff2275d5b8ccd69b541d93

Observation 0386fc58-86dd-4e69-82b9-c1bd0fa82b1b · outbound

This paper cites Chatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts Architecture.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Chatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts Architecture

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.208124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.208124Z digest=sha256:f641705fd4c0c39e54d015eec741ab5a92fefd079c26559f63f94680ae9bc4b2

Observation b27f6346-904f-4d31-b57d-483dcd3c058b · outbound

This paper cites Available: https://www.microsoft.com/en-us/bing.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Available: https://www.microsoft.com/en-us/bing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.696479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.190506Z digest=sha256:2603359f21732cc04d78bb8f35a2e326a62fcc363c996cb24d6f168345277b49

Observation 76e5f2bb-69a3-4141-9ef1-914304b0c6d1 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.621437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.221491Z digest=sha256:abd33fcff1627cbd06797837f0d3e93114f916d1c726e7a82052299f87576dba

Observation 715d0362-a11b-491c-8cd7-4fe1e6eafcbe · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes WebGPT: Browser-assisted question-answering with human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.225818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.225818Z digest=sha256:4e7f617edb364a5e39f2ab0d5f4c2cb826bbefcd8f00da88dbeb4420bd3f7a8f

Observation 559b081b-0c80-4b74-915f-6345a752d75e · outbound

This paper cites Webcpm: Interactive web search for chinese long- form question answering,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Webcpm: Interactive web search for chinese long- form question answering,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.230528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.230528Z digest=sha256:e5f8120094e4d794b1eb39b6f859fb2daa69ab143f83a20790a2a019bc7cc275

Observation f7329238-d0d6-44c3-a748-366e42d93b7c · outbound

This paper cites The reversal curse: Llms trained on.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes The reversal curse: Llms trained on

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.596007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.235013Z digest=sha256:fdd620d7662217494d116e85c161d772b1791061f8da9875ca36a9834c9c05ee

Observation 0bda6a5d-60e8-4ce0-b4f4-756fa9a48ec9 · outbound

This paper cites Survey of hallucination in natural language generation,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Survey of hallucination in natural language generation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.636748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.212747Z digest=sha256:f1dcd4441a8a73648fd59fe2fb96932f6369765e2395b00b73a329cc19533d2a

Observation 928116c9-de0f-43ce-8e46-b357386115d7 · outbound

This paper cites Available: https://doi.org/10.1145/3571730.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Available: https://doi.org/10.1145/3571730

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.217039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.217039Z digest=sha256:975766faaad03285476a099199ac7506499c024b157e18fd178b308c8bfdbce8

Observation a8b1339a-ba1c-40ef-b4db-252f153fd464 · outbound

This paper cites MISGENDERED: limits of large language models in understanding pronouns,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes MISGENDERED: limits of large language models in understanding pronouns,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.252829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.252829Z digest=sha256:460e94d8f7ef1c63b3872ee6bb7d325bfd0b33ec694df25d87845c35165cbeef

Observation 9ce2cc9b-185e-490f-a00a-e11137e6246f · outbound

This paper cites Knowledge neurons in pretrained transformers,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Knowledge neurons in pretrained transformers,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.257431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.257431Z digest=sha256:ec242becad611e1a85b51b2bd6a92c936a7fc6c272497a21bcd65b0522c8a867

Observation 46cf68bc-395a-4ef4-bde2-cfcc16163b02 · outbound

This paper cites Locating and editing factual associations in GPT,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Locating and editing factual associations in GPT,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.566089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.261672Z digest=sha256:8486ca1dae9f3b60029c724e3a759db4d9c2e20ba9e972dede3410d3d3ffc5e9

Observation 7a1cf624-05dc-4f02-a156-d9f08fc58eff · outbound

This paper cites Improving factuality and reasoning in language models through multiagent debate,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Improving factuality and reasoning in language models through multiagent debate,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.550476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.266237Z digest=sha256:88ff50e9f4aeadccea2ed5e03f42a30923b86d3f544f81ea63c2c7bd1faf713c

Observation b1192ac6-21eb-4fed-8d84-d5ada10471d4 · outbound

This paper cites Understanding catastrophic forgetting in language models via implicit inference,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Understanding catastrophic forgetting in language models via implicit inference,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.581022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.239779Z digest=sha256:dd849210fa74525bf8a4da4222a8a3b64ffa04c3ca74c09b714f247b556a1768

Observation 8a8a640f-38fc-4591-935f-35ff28a86700 · outbound

This paper cites Bias and Fairness in Large Language Models: A Survey.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Bias and Fairness in Large Language Models: A Survey

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.244043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.244043Z digest=sha256:06b1f3d177ae37661e62926e7d594d39ea7e912ec29e7f44b747d8c2dd89f687

Observation 7e2dfbe5-f929-45c6-ad0a-786c7a9b4a94 · outbound

This paper cites Bias and Fairness in Large Language Models: A Survey.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Bias and Fairness in Large Language Models: A Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.248522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.248522Z digest=sha256:fc9adb35cfd6faf0c32432b4b5a1216fc52eeba3b021e7ccbeb186f7eec6fcf0

Observation 13d67ff1-81f3-40fe-a878-a78b5149c074 · outbound

This paper cites Dola: Decoding by contrasting layers improves factuality in large language models,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Dola: Decoding by contrasting layers improves factuality in large language models,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.505457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.283636Z digest=sha256:3f25f6bcbed557ce33cf1fee136fb491ae3d234b815b12a6eb8bcc4c8a9fa986

Observation 2f2eede1-b76a-497b-a974-34a5e89c215a · outbound

This paper cites Improving language models by retrieving from trillions of tokens,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Improving language models by retrieving from trillions of tokens,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.490560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.288064Z digest=sha256:ef6839b617ecdbda5964170ff7b1d81fdeb292077e23558202037169f5a4b60d

Observation 04a1613a-439a-4e90-9fd8-ad2f8450207e · outbound

This paper cites Internet-augmented language models through few-shot prompting for open-domain question answering.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Internet-augmented language models through few-shot prompting for open-domain question answering

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.292797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.292797Z digest=sha256:0d9ec33c0aa3e77e7dd2784d8463441801a1a2829ba2befac62e1d4647ba4cea

Observation 184e045a-675c-4fa9-a5fd-cd3289174221 · outbound

This paper cites Rethinking with Retrieval: Faithful Large Language Model Inference.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Rethinking with Retrieval: Faithful Large Language Model Inference

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.302072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.302072Z digest=sha256:0128b3e4c36f0f95760b4f85f66b3d5f31707da2e3ffbd31fac55dfe6085fb73

Observation 456833a8-8fea-4aa8-b915-ef5809a1e2eb · outbound

This paper cites LM vs LM: detecting factual errors via cross examination,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes LM vs LM: detecting factual errors via cross examination,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.270542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.270542Z digest=sha256:1f18f8b0f95450d2aa96590d71d7f9c9485e09292e0b58ee71fc6bb14bc58041

Observation b0c610ac-83b6-4e09-8fff-61838567a76e · outbound

This paper cites Generate rather than retrieve: Large language models are strong context generators,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Generate rather than retrieve: Large language models are strong context generators,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.534423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.274844Z digest=sha256:9935295d26af49ab64af7dc142cf7ae18b3454f9d7a0631bf8a247a452d6548e

Observation 7c03668b-0cf7-4ec4-bef5-cbec39ad4d73 · outbound

This paper cites ”according to.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes ”according to

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.520011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.279249Z digest=sha256:c96ebd31a926ae25f5d4ab4ea9ad5e26372ff4569d3df36d174974b49ec76c5e

Observation 0053207f-b76d-465c-9c10-021ffe283d98 · outbound

This paper cites SAIL: Search-Augmented Instruction Learning.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes SAIL: Search-Augmented Instruction Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.320474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.320474Z digest=sha256:31b1e843ca9518690cc7c48bcc6beacce7dd02a8c6fa2a8dcdae70831e290375

Observation afc209e6-ea1d-447c-ba99-250469990d92 · outbound

This paper cites Decoupled context processing for context augmented language modeling,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Decoupled context processing for context augmented language modeling,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.449676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.325340Z digest=sha256:f1bf8ce00948a575767e9bfdf9cd2aa40dce73497e0043cfe625f1956f8a176c

Observation 54223f28-3cf7-4c84-8fc9-35be5a160c0a · outbound

This paper cites G-MAP: general memory-augmented pre-trained language model for domain tasks,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes G-MAP: general memory-augmented pre-trained language model for domain tasks,

Reference 57

Resolution
verified exact
doi, observed 2026-08-15T18:03:15.633949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.329994Z digest=sha256:5a2206bf87a69dffd2002955671252b59e40be7120bbc79940d3c28a00654660

Observation 1f20ef7d-40a0-4bec-ba68-9acc7787aa29 · outbound

This paper cites The Knowledge Alignment Problem: Bridging Human and External Knowledge for Large Language Models.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes The Knowledge Alignment Problem: Bridging Human and External Knowledge for Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.334546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.334546Z digest=sha256:edf1c21294c8033757b92b8459748720b458f87625cde8177e47665722605cab

Observation 7a37df6f-29b3-4ca4-97cf-413fb9b8c2f9 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Chain-of-thought prompting elicits reasoning in large language models,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.339758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.339758Z digest=sha256:44836055acd6dea570928e3e4560e819b43fc8209e74d56c9798896305c4156e

Observation 5c0bb3d0-5fd7-40b9-bd31-9a0f4c362cbc · outbound

This paper cites Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions,

Reference 60

Resolution
malformed identifier
no resolver link, observed 2026-08-15T18:03:15.306845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.306845Z digest=sha256:0c74d83e450d23a3403e22047c9acfdb4d9003b8fe1964558d7e7feb1b492568

Observation 6d086de5-e5b8-4a23-9386-6c76f13245af · outbound

This paper cites Atlas: Few-shot learning with retrieval augmented language models,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Atlas: Few-shot learning with retrieval augmented language models,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.475090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.311660Z digest=sha256:dc5dd8cb67b0f594d3d22723a599a0f9909715138fdfa55bb04fb175d02ca4f9

Observation 5d7cab6a-d398-45cb-9c58-1d26511bb7f7 · outbound

This paper cites REPLUG: retrieval-augmented black- box language models,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes REPLUG: retrieval-augmented black- box language models,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.316124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.316124Z digest=sha256:474da05f8bd309437130f7d55414d07935d6633b2fbceddb1b4dd5ec921263f6

Observation 3cfadc84-5c2e-4228-8832-5abc2a0321fc · outbound

This paper cites LoRA Learns Less and Forgets Less.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes LoRA Learns Less and Forgets Less

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.357576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.357576Z digest=sha256:715066da7035a3f830ef07f49f14c9b325dcc9cc4bf47c0e0a3e0fb34278de20

Observation 2239fb46-80d1-49ce-8f4b-9211b30804ca · outbound

This paper cites Decoupled weight decay regularization,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Decoupled weight decay regularization,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.391642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.362191Z digest=sha256:63021c39a956740fa004f430b85e42ad582b9ad12182a174b1885cfd7ca673ac

Observation de75e81d-39e3-4939-8d72-b9703538d3ad · outbound

This paper cites Proximal Policy Optimization Algorithms.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Proximal Policy Optimization Algorithms

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.366675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.366675Z digest=sha256:07c3f641a857dd03e1f207e45ea0100cf36e67d1f279952006916736ad07e530

Observation 9af4a418-ec9d-42f7-a66f-ddc5ba04a6a4 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.371184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.371184Z digest=sha256:99905c80bffb5f913c1da2ad19df216c2decea5d3f177a2bbbae1f8dd5117117

Observation 291a3523-2418-4374-8a24-16168096b4c8 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes A Survey on LLM-as-a-Judge

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.380576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.380576Z digest=sha256:0f1e6589bee3934a0820c6b9f095395b8bef7d83a79d9708aa2bc84502095090

Observation acc37e6c-6a6b-48ff-a366-79824e0a04c8 · outbound

This paper cites FLAIR: An easy-to-use framework for state-of-the- art NLP,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes FLAIR: An easy-to-use framework for state-of-the- art NLP,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.424176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.344107Z digest=sha256:c3b139cd2dd4595933105be539f045be965e47aa5796b4cf4fc74dd8f54e3766

Observation 450acbdb-be2d-4714-90ca-a5bfaec99d70 · outbound

This paper cites Instruction tuning for large language models: A survey,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Instruction tuning for large language models: A survey,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.348513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.348513Z digest=sha256:b313e6e018c599bae789e7fa2743f7c865306f53cfc8c3b2b76d37b17d1c2431

Observation 03fd2163-99c4-4688-8658-abbb84329f4b · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Lora: Low-rank adaptation of large language models,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.408784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.353075Z digest=sha256:5f05ef77f500884c0f884cd2b5c720ed18670308dfbd07f29fb53347eda33f60

Observation a8ee7b24-83d0-4bd9-b38f-c879b987989a · outbound

This paper cites Lighteval: A lightweight framework for llm evaluation,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Lighteval: A lightweight framework for llm evaluation,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.345793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.407860Z digest=sha256:350bd6725cee065f300efe25f9c94612eec86cfd1a235cd53d4d608b6a03d9a6

Observation f3f3a85f-81f4-485b-b111-13ea0cf93655 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Efficient memory management for large language model serving with pagedattention,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.330377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.412408Z digest=sha256:dbd4745466a07ff73cda5ffd82a3c412e470efa7470c6d56277b63d4c70fc011

Observation ef580638-c827-423f-b193-3e9c54128cf6 · outbound

This paper cites DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.417052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.417052Z digest=sha256:e6f8ed13576f1d651cf11c6efedbbac66259194754530bf58ffa1139f657c91c

Observation f684f81b-bfc8-496d-87b7-22b60993e257 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.376088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.376088Z digest=sha256:66d80249c9b1b87c76c9116579906eb337f65352e866b4c6ef32cb9301e04004

Observation 24df5dcb-a85c-4355-81b9-2ec94314926a · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Open r1: A fully open reproduction of deepseek-r1,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.376272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.385074Z digest=sha256:c9a3b36e34b54cf284a810f7969bb28f60e1d204ec2a48cf0de9794b753f342f

Observation 80afaed0-6d58-4f34-8660-bd58f1d3a721 · outbound

This paper cites Gemini 2.0 flash,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Gemini 2.0 flash,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:03:16.361249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:03:15.389343Z digest=sha256:5f5145508e0cd39bf9783dd4e6480d87c0d4d4134282cb024fbbb08d6761c795

Observation 36c1c478-8bdb-49fd-8a1c-a4f9eae35fdd · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.393616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.393616Z digest=sha256:065b93916c9550b71f2c72fd488e1018cce2660f9334493e13f65355ac71595d

Observation c331b2c8-9376-400f-81ec-44c44298a02d · outbound

This paper cites DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.421740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.421740Z digest=sha256:5b88e1591f8c66e9af6d1e0b9b1f6de187ac8487194bb2bba818134bebefcfb4

Observation d2ee8123-3fea-457a-9d42-8e0d9ecfbdd8 · outbound

This paper cites An analysis of encoder representations in transformer-based machine translation,.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes An analysis of encoder representations in transformer-based machine translation,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.426142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.426142Z digest=sha256:5046fcba8ee051448326a6f6caadb5e9f2730bd074ed9d7ea1af527f5c69bb36

Observation a3c5881c-adba-43e6-ae71-bf288a10697e · outbound

This paper cites Available: https://doi.org/10.1145/3442188.3445922.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Available: https://doi.org/10.1145/3442188.3445922

Reference 623

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.167069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.167069Z digest=sha256:b40f160c3b581db9e081c1965dc61df1d4e9e235d69a88ee92510d08483b1172

Observation 9227a4b4-1a4b-4cc3-8a05-906517eae6a1 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.403479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.403479Z digest=sha256:2b771dfcc07d72a413a3b5ffac37d6128d03895926ee0fae32b660d178c14868

Observation 518519e3-ca66-492f-af2a-c65d80a00a9e · outbound

This paper cites Internet-augmented language models through few-shot prompting for open-domain question answering.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes Internet-augmented language models through few-shot prompting for open-domain question answering

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.297259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.297259Z digest=sha256:3dec45efec661daebd614c7ee8e2b6fd52063ffd7da4f652420ec633129131a5

Observation 21e9eb86-ad22-4b40-aba5-dc4d67e50d0b · outbound

This paper cites GPT-4 Technical Report.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.181466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.181466Z digest=sha256:0e546470e3aa71011d4245c3516999bfadd035d6f419e740874e79d1259b6077

Observation 9e7d88cc-daad-4bb5-8de9-6eaaeef674b6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.152972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.152972Z digest=sha256:d2f1478aaea3480f5570df4eca3e6bd63ebfe8f97955cee3be2e15c310de21d0

Observation c0924e61-864d-4718-93ad-d217650a8504 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.068746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.068746Z digest=sha256:8dd5ac5a0e1f122c84dff4050203c9de641f111975bc7ab8efb936c67a7c6edd

Pith citing papers

Observation 104cf56b-0a3d-410a-88c1-2ceda34d5c0b · inbound

Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution cites this paper.

Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:48:05.981542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T06:43:23.870127Z digest=sha256:993ebf2a3272d7498415bc6238cd7e43aa5ce7c40970199d5241209797cc9d8a

Observation 294bc273-77c1-43e4-b4c7-f686bdddfeab · inbound

Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research cites this paper.

Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T12:12:07.831497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T12:08:10.552789Z digest=sha256:445c59cb8aed0c1c012fb1bb31b380a2ace39423e6ee3c195bfeb2505b8f1432