Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:39:05.842706Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 100 of 299 outbound references and 8 inbound Pith citation observations for arXiv:2509.03871.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:39:05.842706Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:40:57.907170Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:49:51.497057Z
100 of 299 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7d684ae2-6ab1-4786-8186-c77c6ece9bcd · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d17092a0-1d6f-47c6-a1fa-d05acfb57e6b · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae816db-1f4e-48cd-a66a-b5bc49ffa739 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8360dd42-bfa8-4b80-8ac4-5808aa6baa3c · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2be066f5-ef87-4f50-b954-115819e36f6e · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Attacks, defenses and evaluations for llm conversation safety: A survey
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c254488c-c2f7-46ee-beaf-88b8bd574a9e · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Large Language Model Safety: A Holistic Survey
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2c15ced-bc72-40a5-a4bc-5d5456d2d60b · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 934e544c-2bb9-4c60-bb5f-57e5d2618c37 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26d5bf3-a448-4ae3-a547-ec663a5f02c0 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 170bab49-91d8-4810-8e0d-9438cdcd2257 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 610366fc-8b1b-4b47-a96a-2bbf9621125a · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Efficient reasoning models: A survey
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ea0c3d-8c8a-41f9-9473-58e19b5eb048 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Safety Reasoning with Guidelines
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0914dd50-21dd-4180-8a74-a233e801d714 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models GPT-4 Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f79476f6-7db3-4486-9aad-0a71a211d833 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bd43721-3fdd-45b1-ac90-cd091f384b50 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7e80236-c9dd-4894-8307-84af749380bc · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Chain-of-thought prompting elicits reasoning in large language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68fd7f7e-d527-4b46-8a6d-c03af03e61be · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Large language models are zero-shot reasoners
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a301d2b-b3d1-4553-8f3e-5faf5958662f · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Language models are few-shot learners
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d49a3beb-5b65-472a-83fe-2c298996ec15 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Chain-of-scrutiny: Detecting backdoor attacks for large language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb86ee6d-174b-49a4-a54f-8791fafab37c · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models OpenAI o1 System Card
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d44e522-bf83-45c4-9b9f-51c8f7fd725c · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54025218-799b-439d-b3b8-6be341fb747c · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models O1 Replication Journey: A Strategic Progress Report -- Part 1
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24ce318b-8b3f-498d-af59-d535234a24c3 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c18fd69-1946-4384-aee6-64b2651c3071 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2bd868a-d220-45f5-ab20-d92fb7837cc5 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6732203c-4bca-45ba-af03-c40b798a731c · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models A survey of monte carlo tree search methods
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9811b662-a57b-4ef7-9c3b-ec3c8655bd16 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Training Verifiers to Solve Math Word Problems
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 631cfcfb-ec8a-4b8a-9621-df358ffbb683 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Measuring mathematical problem solving with the math dataset
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3396605-0c22-44c6-9873-37eb238d2acf · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models MARIO: MAth Reasoning with code Interpreter Output–A Reproducible Pipeline
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4b98ad7-d3b4-404e-a961-4c5e6a7565a3 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd470b81-3047-4507-a0f5-029c2442a9f9 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Proximal Policy Optimization Algorithms
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c25a6645-d1df-4227-8f54-377d6ead8131 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19afe0dc-4eeb-40c5-9c4e-731a3a6d08c3 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fefff651-33ef-465c-95da-8d89f584b8f0 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Training large language models to reason in a continuous latent space
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c05d5f40-b122-4a55-9c32-ebb112432f60 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models DeepSeek-V3 Technical Report
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 513f4a01-cbad-4bf3-a7cb-652a27b05a98 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Qwen2.5 technical report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f04f7e6-5f31-4fde-82e8-4e23f68166ea · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The llama 3 herd of models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa71491-3a98-4766-bfa4-479123a9a6b0 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Solving math word problems with process- and outcome-based feedback
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57befc21-8b38-4b9b-af18-ca9b8ff88b48 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Star: Self-taught reasoner bootstrapping reasoning with reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d64622b8-f447-4ee9-b1de-0bc8c97a6ea1 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Let’s verify step by step
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e487db32-9fc0-4e2f-95d8-555ccdea0028 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fafe2158-ea6e-4f58-b8e3-22c6740d5d33 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Reinforcement Learning with Verifiable Rewards: GRPO’s Effective Loss, Dynamics, and Success Amplifi- cation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 194346aa-9f27-42e6-b08f-ae432e5de4b5 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d121e73-d805-4eef-a774-ff999f4b0b1d · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d24712ef-2e8f-464e-9776-4b8359319493 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Multimodal chain-of-thought reasoning in language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe95329-852b-4cec-bd1e-6ea494c9cff8 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Video-of-thought: Step-by-step video reasoning from perception to cognition
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 140f773a-c2e3-468a-b8c4-9dc3f4a975b4 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d8e403c-73af-42f9-8e7c-e3fe17d32bd8 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c804ad-0e2b-49a0-b4a0-5a494adfda72 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 383c04bc-624c-4b91-a06b-d9a1253be304 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22b6c7f7-abb2-4156-85da-64c777a29856 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67f349f2-fb93-440c-b0d2-ac3572e510f8 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Improve Vision Language Model Chain-of-thought Reasoning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51e2896e-645a-4700-adc9-049ae7917dff · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Insight-v: Exploring long-chain visual reasoning with multimodal large language models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b685855f-c7de-430b-a29e-3eb89d776736 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2748fbf2-ba15-4b40-92e5-fafea5ad881d · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d1ef5d2-b76f-44ec-a949-dd3d6e8c89ba · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a383f797-e3bd-4b2b-bfb0-72f46d5084d5 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models HalluMeasure: Fine-grained hallucination measurement using chain-of-thought reasoning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0254abf4-680a-4735-85bc-717af5572cbd · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models CLATTER: Comprehensive Entailment Reasoning for Hallucination Detection
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 679fcc8f-0ad7-47aa-ab15-37384e54570d · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language-Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51891bc4-500e-4224-9a61-e537c1f404e1 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Grounded Chain-of-Thought for Multimodal Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b7cb315-8224-44e6-b02d-f9c0e3e8ec5e · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models CoMT: Chain-of- Medical-Thought Reduces Hallucination in Medical Report Generation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f61cce6-5c2d-4d17-8b3f-2c8cc3ad9045 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1ed60c3-31ad-4d33-8b55-7c3267e55276 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The Hallucination Tax of Reinforcement Finetuning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aec353aa-b126-4c32-a2ad-2a8376ce7f35 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0fcb64b-11b0-4b10-b200-be3b69dc526a · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Are Reasoning Models More Prone to Hallucination?
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 069a4431-2cd7-4e46-9e11-26cf1ab698aa · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 600848f0-cea2-479a-a9dd-11a0545b5a32 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1717b948-8b70-4e97-ae2c-efaf562f2d89 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The Hallucination Dilemma: Factuality-Aware Reinforcement Learning for Large Reasoning Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd1a26e2-50c8-46a3-89b6-67a1680db91f · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Analyzing Logical Fallacies in Large Language Models: A Study on Hallucination in Mathematical Reasoning
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e60f5e06-b4b8-41eb-95bb-75a6fb4b37d1 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cab1ef21-b8ef-4cd2-bb63-4d76cee974f3 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36b76c94-a499-49a2-bfd9-efb6d447c250 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c49d9c9-e8d3-4485-9753-cb07cd765495 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c025b55a-1598-4e64-94b7-88672eb309d0 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea3174b3-7e63-4164-af9b-2a26598b591f · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Measuring Faithfulness in Chain-of-Thought Reasoning
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2dbe24d-52f7-47e1-9892-fc18aa3ff9cd · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b1c5bcd-3e3b-4a0e-b0fe-74cb0090ec0c · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Measuring faithfulness of chains of thought by unlearning reasoning steps
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df660e94-2abd-4754-be99-949a2751cd90 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f379e7e-4881-4ecd-bdfe-858513bb4c92 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Chain-of-Thought Unfaithfulness as Disguised Accuracy
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dd18d34-1be3-434d-8793-e115c674b165 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fe62b89-f21d-4614-96a3-6c8c9b36cc8f · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Are DeepSeek R1 And Other Reasoning Models More Faithful? In ICLR 2025 Workshop on Foundation Models in the Wild, 2025
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdf6b7b3-ba69-480a-afc1-3816552f5932 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Reasoning Models Don’t Always Say What They Think
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 349ac01d-492b-4973-b758-b9a865f9f38b · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f73ef14-9207-4f9d-bf08-4e403d033815 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc2adb70-ea5d-451b-8821-ec6633379986 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models How Likely Do LLMs with CoT Mimic Human Reasoning? In Proc
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b318b7-55c1-473d-b04d-b1a9add6b090 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models On the difficulty of faithful chain-of-thought reasoning in large language models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9e417f7-0f61-4660-a805-f917c0d1acf1 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models On the impact of fine-tuning on chain-of-thought reasoning
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 833bd429-3cfe-479f-a311-5cdf9586a51a · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80c8271e-838c-43a1-a83b-a94817eb97a6 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Faithful logical reasoning via symbolic chain-of-thought
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b34a989e-e8b8-460a-b41d-e1f0273933fc · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Question Decomposition Improves the Faithfulness of Model-Generated Reasoning
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eb7d797-5283-4e3c-b0bc-2597c266defe · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Faithful chain-of-thought reasoning
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36edc327-6624-444e-9786-0bce5bdc0433 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe67efb8-5a17-4ce7-bdc6-1cdef025df79 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models FLARE: Faithful Logic-Aided Reasoning and Exploration
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf457c86-83c9-4be5-b4eb-0d1c3fa4e3d5 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models CoMAT: Chain of mathematically annotated thought improves mathematical reasoning
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff9a1246-84ea-4b1f-ba08-851c3bfa4e01 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Causal-driven Large Language Models with Faithful Reasoning for Knowledge Question Answering
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f4e3bc-0fbc-487f-ba06-bafe7f586ae8 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Fact: Teaching mllms with faithful, concise and transferable rationales
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edf8e7fb-fd5c-4290-bbcc-c0382816e13b · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Markovian Transformers for Informative Language Modeling
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6273c2d2-6584-4c26-a111-7a8d6cce3823 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19cfad14-82b7-4d3a-a0ed-3d63e836e747 · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e940c272-3b52-427d-9301-24d1d299e2ec · outbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The hidden risks of large reasoning models: A safety assessment of r1
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6529f910-4f45-4646-b22d-8b94c3ae006a · inbound
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dfa30600-790c-40a8-87a8-3eddea1fb926 · inbound
Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b32d8cb4-6b38-4187-9c65-dddf41cf8eb7 · inbound
From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 10f6d2b5-7061-4305-af0d-f0e670375250 · inbound
Pause or Fabricate? Training Language Models for Grounded Reasoning A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9f7f2ab7-39d7-41c1-83be-508b2090a737 · inbound
Auditing Reasoning-Trace Memorization Claims after Unlearning with Head-Conditioned Canaries A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3199632a-16f2-449e-bb81-e1a2c6d87600 · inbound
Where Do CoT Training Gains Land in LLM based Agents? A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a1e99ab0-aaf5-492f-ad01-23543337c690 · inbound
Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4ef013ca-8e43-4489-a143-943009f877fe · inbound
Risky Business: Measuring The Faithfulness-Safety Tension A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.