Pith. sign in

Paper Citation Record · LEDGER

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

As of 23 August 2026, this Paper Citation Record lists 100 of 299 outbound references and 8 inbound Pith citation observations for arXiv:2509.03871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03871 v1

Coverage vector

measured 100 of 299 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:39:05.842706Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:40:57.907170Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:49:51.497057Z

Reference resolution

100 of 299 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d684ae2-6ab1-4786-8186-c77c6ece9bcd · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.142775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:58.142775Z digest=sha256:4de1367973d811b7e678aea533d7ead154ef5e88aaddc28185ceecdb8fe397f6

Observation d17092a0-1d6f-47c6-a1fa-d05acfb57e6b · outbound

This paper cites Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.228860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:58.228860Z digest=sha256:91a7259a0cef792548d15117bb7780ae34ba0e081c9cf7f9572fba66fa996057

Observation 2ae816db-1f4e-48cd-a66a-b5bc49ffa739 · outbound

This paper cites Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.311706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:58.311706Z digest=sha256:39e7aae3cefdb2048b7f2f79601c1e9982b5c1967c6015b9e028e025e4c2415c

Observation 8360dd42-bfa8-4b80-8ac4-5808aa6baa3c · outbound

This paper cites A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.402887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:58.402887Z digest=sha256:6bbaa1addfdbba91ffff3e3d1f62317557f08ddc5e7312bbcef3065023877fb7

Observation 2be066f5-ef87-4f50-b954-115819e36f6e · outbound

This paper cites Attacks, defenses and evaluations for llm conversation safety: A survey.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Attacks, defenses and evaluations for llm conversation safety: A survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.485426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:58.485426Z digest=sha256:7d467839a7a5947eb17446c71a4e1681ec0397033d40fe629253578f2506476a

Observation c254488c-c2f7-46ee-beaf-88b8bd574a9e · outbound

This paper cites Large Language Model Safety: A Holistic Survey.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Large Language Model Safety: A Holistic Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.561446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:58.561446Z digest=sha256:05fea2e3ad1a4bf8f1daa30a79e640484c38d9fc87a03d10c3f09b1c26fa33fa

Observation e2c15ced-bc72-40a5-a4bc-5d5456d2d60b · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.646795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:58.646795Z digest=sha256:aa2e462e255ca3725cfd1788efbf0677257d0369ef93663de93f3adbc22aff57

Observation 934e544c-2bb9-4c60-bb5f-57e5d2618c37 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.714356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:58.714356Z digest=sha256:f5d32119b46210b360bb8b8a721b42584598d7605f4a8cce8a1cc6da135640ab

Observation b26d5bf3-a448-4ae3-a547-ec663a5f02c0 · outbound

This paper cites A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.806352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:58.806352Z digest=sha256:5674a742f0d1e6c829be54b99e0fbe3e64a5cc82e0a9efa930dfc6fa94a8e70f

Observation 170bab49-91d8-4810-8e0d-9438cdcd2257 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.871065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:58.871065Z digest=sha256:19e034b770e9573017f9281b5cb2add003774805a66a54ca919b04627c75b3a3

Observation 610366fc-8b1b-4b47-a96a-2bbf9621125a · outbound

This paper cites Efficient reasoning models: A survey.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Efficient reasoning models: A survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.960029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:58.960029Z digest=sha256:8b677a80bb418ea44416030fb8871a64bdce5a7fdc05fd7057714f851b6b01dc

Observation d9ea0c3d-8c8a-41f9-9473-58e19b5eb048 · outbound

This paper cites Safety Reasoning with Guidelines.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Safety Reasoning with Guidelines

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:59.035605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:59.035605Z digest=sha256:8e08312ade9a27b72925f85e6521338e2ad4184cf577534101bb3867fc886974

Observation 0914dd50-21dd-4180-8a74-a233e801d714 · outbound

This paper cites GPT-4 Technical Report.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models GPT-4 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:59.115959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:59.115959Z digest=sha256:ae3746de3a19a804a56b8da40b4fd715abc1f88260888d0e4ce9a602eb162b9e

Observation f79476f6-7db3-4486-9aad-0a71a211d833 · outbound

This paper cites The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:59.182083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:59.182083Z digest=sha256:a2814b8280d42ec4edac11fb9c772b1a0be0a78d42761cd76dae35d1adf1e562

Observation 5bd43721-3fdd-45b1-ac90-cd091f384b50 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:59.242813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:59.242813Z digest=sha256:f277711b837ed6765024b983ae4fa2f005f85627a9c22f809c69a22abb48bac5

Observation c7e80236-c9dd-4894-8307-84af749380bc · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:59.344978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:59.344978Z digest=sha256:8444dcbd5a3e4bc9c660ad8afd5797a10b013888ed1290b542f01a4e49faf10f

Observation 68fd7f7e-d527-4b46-8a6d-c03af03e61be · outbound

This paper cites Large language models are zero-shot reasoners.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Large language models are zero-shot reasoners

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:59.419222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:59.419222Z digest=sha256:998fe105dda55d19dd8e41135d5cdef97ba2868774b316432197ecd7149588dd

Observation 6a301d2b-b3d1-4553-8f3e-5faf5958662f · outbound

This paper cites Language models are few-shot learners.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Language models are few-shot learners

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:59.510055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:59.510055Z digest=sha256:fc77da98fccdb00e8d4b092a63770fe2b52e7ed10299435c029a4aa4980c91f0

Observation d49a3beb-5b65-472a-83fe-2c298996ec15 · outbound

This paper cites Chain-of-scrutiny: Detecting backdoor attacks for large language models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Chain-of-scrutiny: Detecting backdoor attacks for large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:59.584929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:59.584929Z digest=sha256:f22389fa8f46b00030b7ae02b5ca7152eb82d9e3ec5e528bda61f13774601e36

Observation cb86ee6d-174b-49a4-a54f-8791fafab37c · outbound

This paper cites OpenAI o1 System Card.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models OpenAI o1 System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:59.680716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:59.680716Z digest=sha256:c1a6aa013fd653fa276eef05a2b2eb9e5ba39cd58080316f639a9dddea6c209b

Observation 8d44e522-bf83-45c4-9b9f-51c8f7fd725c · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:59.785250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:59.785250Z digest=sha256:1aef80539cf0522a73d3c830d417dcc37b34491aeb680a7a9ede1777f304d365

Observation 54025218-799b-439d-b3b8-6be341fb747c · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:59.862669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:59.862669Z digest=sha256:550357e0814749ee2b88cc3491ef9c5ae751a129134d78c092175840d509d746

Observation 24ce318b-8b3f-498d-af59-d535234a24c3 · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:59.955343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:59.955343Z digest=sha256:8a123a17cdcbebdaf32102c97f9e781c8d65a68f07e470f90e0bee4ffeff959c

Observation 4c18fd69-1946-4384-aee6-64b2651c3071 · outbound

This paper cites O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.021034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.021034Z digest=sha256:cfd6e29a451b75fbf8d189d7a8fb664160eea20c0f1e5f06e58237dbfd49ee8a

Observation f2bd868a-d220-45f5-ab20-d92fb7837cc5 · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.088338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.088338Z digest=sha256:78f2698ba61798e8d2b909de8d1f2da2f85f81d3b80fb3977fb5d86852157470

Observation 6732203c-4bca-45ba-af03-c40b798a731c · outbound

This paper cites A survey of monte carlo tree search methods.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models A survey of monte carlo tree search methods

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.151852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.151852Z digest=sha256:c87f094e8f6ee1309ea08ae303644af214b0f01e0d40ad6d08bbc46b6efea384

Observation 9811b662-a57b-4ef7-9c3b-ec3c8655bd16 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Training Verifiers to Solve Math Word Problems

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.239660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.239660Z digest=sha256:728afa68af31a3ea9a444029a51e3fb91b4871243efe90ba6dd857ac1848a618

Observation 631cfcfb-ec8a-4b8a-9621-df358ffbb683 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Measuring mathematical problem solving with the math dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.328322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.328322Z digest=sha256:ff6bc5d1690a257565bd2e2f1408331aeca6998fdbafe98de4d2ee6e053443da

Observation b3396605-0c22-44c6-9873-37eb238d2acf · outbound

This paper cites MARIO: MAth Reasoning with code Interpreter Output–A Reproducible Pipeline.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models MARIO: MAth Reasoning with code Interpreter Output–A Reproducible Pipeline

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.404102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.404102Z digest=sha256:2a65725f60956c01579bd8a9e9fc5ce2249e4b0ac7f06312c63b0f2f3e0cd1ef

Observation e4b98ad7-d3b4-404e-a961-4c5e6a7565a3 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.496497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.496497Z digest=sha256:a7316524c1592db286427a4365fbde12b77efe030e4fb27def3cd6872a392c1f

Observation dd470b81-3047-4507-a0f5-029c2442a9f9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Proximal Policy Optimization Algorithms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.582402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.582402Z digest=sha256:a392d93ba906787888f15dacc34ebb63c922d4dab3130da576e3cc3e5fb95f4a

Observation c25a6645-d1df-4227-8f54-377d6ead8131 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.647941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.647941Z digest=sha256:21d6ec77380065a454c080a42b2b116db9870c1ef961c6c9adb0df2df54cc4db

Observation 19afe0dc-4eeb-40c5-9c4e-731a3a6d08c3 · outbound

This paper cites Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.744945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.744945Z digest=sha256:46b8de101f2598a294e2e230a09ffa1a0a805bae071fe50ecd04ef119573c3ef

Observation fefff651-33ef-465c-95da-8d89f584b8f0 · outbound

This paper cites Training large language models to reason in a continuous latent space.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Training large language models to reason in a continuous latent space

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.819156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.819156Z digest=sha256:487d5c94222390fd409609a00cbb24507f613db36c7c27c63d961860c4ae25bb

Observation c05d5f40-b122-4a55-9c32-ebb112432f60 · outbound

This paper cites DeepSeek-V3 Technical Report.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models DeepSeek-V3 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.885552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.885552Z digest=sha256:efede426d58fec3819bfce9ed06b2961cfa7cca46528bd33bb5c834562059055

Observation 513f4a01-cbad-4bf3-a7cb-652a27b05a98 · outbound

This paper cites Qwen2.5 technical report.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Qwen2.5 technical report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:00.981498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:00.981498Z digest=sha256:9d0c67e0102c5492d3b16434ebc74dbf976b83493e389d9e2ccf72041ce9f97a

Observation 1f04f7e6-5f31-4fde-82e8-4e23f68166ea · outbound

This paper cites The llama 3 herd of models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The llama 3 herd of models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.040517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.040517Z digest=sha256:1b12870b2cada846bf8e14e55fe5652c352b4b6d38f177d16d4aaec21acc81f4

Observation 5aa71491-3a98-4766-bfa4-479123a9a6b0 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Solving math word problems with process- and outcome-based feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.109679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.109679Z digest=sha256:80f08b66f4347591ebe7e3aea7582b06dac3d97239e3bcd2459a5ba9927d9372

Observation 57befc21-8b38-4b9b-af18-ca9b8ff88b48 · outbound

This paper cites Star: Self-taught reasoner bootstrapping reasoning with reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Star: Self-taught reasoner bootstrapping reasoning with reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.199920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.199920Z digest=sha256:0763e9070a45982158cf96ba1bf3793af68bd0f30709142464b446597a2089fd

Observation d64622b8-f447-4ee9-b1de-0bc8c97a6ea1 · outbound

This paper cites Let’s verify step by step.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Let’s verify step by step

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.252139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.252139Z digest=sha256:5e07aacda4443a63978e1adb6404cdd876a12ffc423e97f6d359632deecb0686

Observation e487db32-9fc0-4e2f-95d8-555ccdea0028 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.338262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.338262Z digest=sha256:f40d64e18aa16f9eced476c58cd93f4adb5cd47318788425241369d2e3dc9b13

Observation fafe2158-ea6e-4f58-b8e3-22c6740d5d33 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards: GRPO’s Effective Loss, Dynamics, and Success Amplifi- cation.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Reinforcement Learning with Verifiable Rewards: GRPO’s Effective Loss, Dynamics, and Success Amplifi- cation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.417577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.417577Z digest=sha256:e77afc31a0314d2855e9c4117e1e19c7b829d4578a73218b6859c8e5de460037

Observation 194346aa-9f27-42e6-b08f-ae432e5de4b5 · outbound

This paper cites Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.493981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.493981Z digest=sha256:e9075bf9fefdf4f1ee97c7bf7a281bacaaddb78358d459ddebc322e63c0ce1cc

Observation 6d121e73-d805-4eef-a774-ff999f4b0b1d · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.576595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.576595Z digest=sha256:e04201123f4a94ce26f275c72c84c085bac20dd38d46238b3fe8b7a908b99985

Observation d24712ef-2e8f-464e-9776-4b8359319493 · outbound

This paper cites Multimodal chain-of-thought reasoning in language models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Multimodal chain-of-thought reasoning in language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.651618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.651618Z digest=sha256:d39e05326d3c3963e3751249f924e3c8d628cfdfbd4240798b800a2dee6804a8

Observation cbe95329-852b-4cec-bd1e-6ea494c9cff8 · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.740213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.740213Z digest=sha256:22bbf82b5c5a206c839b2b1c24205b57a75fd5e495f2df73a387361469d6bc18

Observation 140f773a-c2e3-468a-b8c4-9dc3f4a975b4 · outbound

This paper cites Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.812797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.812797Z digest=sha256:288cb3b6429dab1e61f025b172988863c0c464e11ce70a6cecff36e71e3fa155

Observation 5d8e403c-73af-42f9-8e7c-e3fe17d32bd8 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.904314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.904314Z digest=sha256:2bec51265b29c4a0ab8507555ee26999c975d4fd675933dcd849a57745668d8f

Observation 85c804ad-0e2b-49a0-b4a0-5a494adfda72 · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.002966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.002966Z digest=sha256:8440456e5ae1735408396412c73525150fb358766e3d80d86265ef114e728f9c

Observation 383c04bc-624c-4b91-a06b-d9a1253be304 · outbound

This paper cites RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.106102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.106102Z digest=sha256:a3ef001e1b2e2c36d22a067dda5ef95fad39388febcf50f4257ff44952aba497

Observation 22b6c7f7-abb2-4156-85da-64c777a29856 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.195906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.195906Z digest=sha256:9bdcc995196e2389903def4db9cd3a7403a29c2dfe5f2727f9be896788dadd93

Observation 67f349f2-fb93-440c-b0d2-ac3572e510f8 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Improve Vision Language Model Chain-of-thought Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.269696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.269696Z digest=sha256:01de670a65101ecc1459e3350c14da1d50b56241df52ba6cac63e1407c054e13

Observation 51e2896e-645a-4700-adc9-049ae7917dff · outbound

This paper cites Insight-v: Exploring long-chain visual reasoning with multimodal large language models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Insight-v: Exploring long-chain visual reasoning with multimodal large language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.347544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.347544Z digest=sha256:d05e6ae97a06037cabad48d6b808a5663960e278a4d7d79395224097f1f3821d

Observation b685855f-c7de-430b-a29e-3eb89d776736 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.399950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.399950Z digest=sha256:8a649a2321a831e8b0789c49775f27e2111a1d9ff279b8ddd33e430540043bd0

Observation 2748fbf2-ba15-4b40-92e5-fafea5ad881d · outbound

This paper cites MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.499587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.499587Z digest=sha256:fc084b8706452ab0ec94a3a628723f382ecf6d3f3d5bb34aa415984e78fb172b

Observation 4d1ef5d2-b76f-44ec-a949-dd3d6e8c89ba · outbound

This paper cites Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.581234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.581234Z digest=sha256:dfa92eb9733ea2324902d49715afa7682b5dbac7ba86c2b660406e3894626087

Observation a383f797-e3bd-4b2b-bfb0-72f46d5084d5 · outbound

This paper cites HalluMeasure: Fine-grained hallucination measurement using chain-of-thought reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models HalluMeasure: Fine-grained hallucination measurement using chain-of-thought reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.666630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.666630Z digest=sha256:5ca212407dea6e7b829236dfce86ca88f8ae19204abf8b3c8829a57aa62572fc

Observation 0254abf4-680a-4735-85bc-717af5572cbd · outbound

This paper cites CLATTER: Comprehensive Entailment Reasoning for Hallucination Detection.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models CLATTER: Comprehensive Entailment Reasoning for Hallucination Detection

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.718004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.718004Z digest=sha256:b42cf11e24a2d21739fe537a02b656e22bcb884ec1cd7eb5d73c7c02fc84259b

Observation 679fcc8f-0ad7-47aa-ab15-37384e54570d · outbound

This paper cites Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language-Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language-Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.811140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.811140Z digest=sha256:5c389da6bd1e42df573f396a024cac56e4c4b71fd14353dcfb831025b12c28c6

Observation 51891bc4-500e-4224-9a61-e537c1f404e1 · outbound

This paper cites Grounded Chain-of-Thought for Multimodal Large Language Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.881423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.881423Z digest=sha256:0b7de5404e68bbaf05f465f679b89bc851bc64a6b1c340de3a6646c6d28e0d17

Observation 3b7cb315-8224-44e6-b02d-f9c0e3e8ec5e · outbound

This paper cites CoMT: Chain-of- Medical-Thought Reduces Hallucination in Medical Report Generation.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models CoMT: Chain-of- Medical-Thought Reduces Hallucination in Medical Report Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.968136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.968136Z digest=sha256:ad319b220436972230e5b7c1ee83fee43b5e4ec790aa5a28d68098a6a731654d

Observation 8f61cce6-5c2d-4d17-8b3f-2c8cc3ad9045 · outbound

This paper cites MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.028373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.028373Z digest=sha256:bcb0a91ce799688cca185a46286e0a1c83d35baa864249ec028ab49ca9bb50f6

Observation d1ed60c3-31ad-4d33-8b55-7c3267e55276 · outbound

This paper cites The Hallucination Tax of Reinforcement Finetuning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.113610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.113610Z digest=sha256:6f09be1eba562d08568ecaf189f8c1621039bc70c61ede59e8e3a90c066d7697

Observation aec353aa-b126-4c32-a2ad-2a8376ce7f35 · outbound

This paper cites More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.190809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.190809Z digest=sha256:1b05af5868643cb59bc8aa8ef5142377371de11eb3c320f192ba2f7d4091d80b

Observation b0fcb64b-11b0-4b10-b200-be3b69dc526a · outbound

This paper cites Are Reasoning Models More Prone to Hallucination?.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Are Reasoning Models More Prone to Hallucination?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.278036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.278036Z digest=sha256:70bc9c355c87db95d0ed38102cb141c34a1e3bd0344318eb9f39491592413dcf

Observation 069a4431-2cd7-4e46-9e11-26cf1ab698aa · outbound

This paper cites AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.364570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.364570Z digest=sha256:0fdab4e3a024e896ce0c6bb1b637df17da857d613cb783ace4596a2193415bfa

Observation 600848f0-cea2-479a-a9dd-11a0545b5a32 · outbound

This paper cites Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.438085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.438085Z digest=sha256:45b0b5553ea2c57caf476b32b4c8975bcb241fa7e66470ff0e4be7255e590c9e

Observation 1717b948-8b70-4e97-ae2c-efaf562f2d89 · outbound

This paper cites The Hallucination Dilemma: Factuality-Aware Reinforcement Learning for Large Reasoning Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The Hallucination Dilemma: Factuality-Aware Reinforcement Learning for Large Reasoning Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.522217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.522217Z digest=sha256:edc8b5f3f8df88fb0a3d71141658ab8d4243771cb5cb8fbb28130b5b0976fd06

Observation fd1a26e2-50c8-46a3-89b6-67a1680db91f · outbound

This paper cites Analyzing Logical Fallacies in Large Language Models: A Study on Hallucination in Mathematical Reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Analyzing Logical Fallacies in Large Language Models: A Study on Hallucination in Mathematical Reasoning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.602018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.602018Z digest=sha256:e7f00c32acd9c165810c178d4307f9edce528ee057f1e26c8743f9ea3a06066c

Observation e60f5e06-b4b8-41eb-95bb-75a6fb4b37d1 · outbound

This paper cites Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.674536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.674536Z digest=sha256:b35926857e647cf92b3b6e60d983f56710d6baae830188216eda29b283350f8d

Observation cab1ef21-b8ef-4cd2-bb63-4d76cee974f3 · outbound

This paper cites Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.728877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.728877Z digest=sha256:f496f82ef33aad3f6b6939c2c0e6a079e660f13f8f29b63d741002c20ecac70a

Observation 36b76c94-a499-49a2-bfd9-efb6d447c250 · outbound

This paper cites Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.777748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.777748Z digest=sha256:7eb7f84df99eb08dc1fd5eeca2171e47a849ae1291b212e65c47cfecc286229e

Observation 0c49d9c9-e8d3-4485-9753-cb07cd765495 · outbound

This paper cites Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.846399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.846399Z digest=sha256:5b6cbf1ab4afb21de6c4976bfffdf729ebc91ff29c6f978902f4742dd22553c9

Observation c025b55a-1598-4e64-94b7-88672eb309d0 · outbound

This paper cites Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.933007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.933007Z digest=sha256:82f01229c8b0a25973c38516eec86f7c235d1614d63d5e0a392f9a114cca1140

Observation ea3174b3-7e63-4164-af9b-2a26598b591f · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.017581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.017581Z digest=sha256:3c6a39f583df43df5f321aa5ed4ef435a45acfd67c637dd7297ce9cb82660e00

Observation b2dbe24d-52f7-47e1-9892-fc18aa3ff9cd · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.122351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.122351Z digest=sha256:b7e9c25bb3d619f77766194ea09a5b0876731efe0bac725d36b75290af264eb8

Observation 1b1c5bcd-3e3b-4a0e-b0fe-74cb0090ec0c · outbound

This paper cites Measuring faithfulness of chains of thought by unlearning reasoning steps.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Measuring faithfulness of chains of thought by unlearning reasoning steps

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.202832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.202832Z digest=sha256:3f1c15fc5a73b3dcd397330633f2566163b6b739c4c568410f7844eacdf19bfa

Observation df660e94-2abd-4754-be99-949a2751cd90 · outbound

This paper cites Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.273817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.273817Z digest=sha256:9315b64eee7ca5293a6824a56b99d7e16d45279df1f5de818479a6a0f7607f09

Observation 8f379e7e-4881-4ecd-bdfe-858513bb4c92 · outbound

This paper cites Chain-of-Thought Unfaithfulness as Disguised Accuracy.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Chain-of-Thought Unfaithfulness as Disguised Accuracy

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.365681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.365681Z digest=sha256:a41ba63bd077a7b11ba21d9b7c1191ccd60738208e79a3b05c82672d3a0948c4

Observation 2dd18d34-1be3-434d-8793-e115c674b165 · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.414376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.414376Z digest=sha256:31d8ec84f5bd703b32b2ee9bdb49705cd7c745e3ccc446635adef12f40b44ae3

Observation 8fe62b89-f21d-4614-96a3-6c8c9b36cc8f · outbound

This paper cites Are DeepSeek R1 And Other Reasoning Models More Faithful? In ICLR 2025 Workshop on Foundation Models in the Wild, 2025.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Are DeepSeek R1 And Other Reasoning Models More Faithful? In ICLR 2025 Workshop on Foundation Models in the Wild, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.465800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.465800Z digest=sha256:2c754158f8029ed8f286a8d89c0dc8f5146d0a55262a06aa9754910d161f51a6

Observation bdf6b7b3-ba69-480a-afc1-3816552f5932 · outbound

This paper cites Reasoning Models Don’t Always Say What They Think.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Reasoning Models Don’t Always Say What They Think

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.520817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.520817Z digest=sha256:28ad6dc8a9b9cb5b4d599ac4c78403b028a159a2bfcce44a7831b5cef07db25d

Observation 349ac01d-492b-4973-b758-b9a865f9f38b · outbound

This paper cites Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.568023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.568023Z digest=sha256:eeee2381b3dc2760f4fd8d2bf8331794b25bbd4f087a64609bcfa98f0f7b4055

Observation 3f73ef14-9207-4f9d-bf08-4e403d033815 · outbound

This paper cites Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.627878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.627878Z digest=sha256:795dc1c28e9cc4dfae43a3e7fc5468810f9ad9ff9e60bdb9844579ac11217bb7

Observation bc2adb70-ea5d-451b-8821-ec6633379986 · outbound

This paper cites How Likely Do LLMs with CoT Mimic Human Reasoning? In Proc.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models How Likely Do LLMs with CoT Mimic Human Reasoning? In Proc

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.683690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.683690Z digest=sha256:74e6ded23dd8763d99299ab37db82a8f6bd77f55df5208e263b6ac7967563974

Observation f1b318b7-55c1-473d-b04d-b1a9add6b090 · outbound

This paper cites On the difficulty of faithful chain-of-thought reasoning in large language models.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models On the difficulty of faithful chain-of-thought reasoning in large language models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.736151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.736151Z digest=sha256:7c30d5156933a9806d17386210fdd7573fcd481688bb275c8b50ddcd5bfc57a7

Observation f9e417f7-0f61-4660-a805-f917c0d1acf1 · outbound

This paper cites On the impact of fine-tuning on chain-of-thought reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models On the impact of fine-tuning on chain-of-thought reasoning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.787173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.787173Z digest=sha256:0f43eba02fe4519f23d67bf37fa373c0fcd4003ecebb49cbc7d19b1af7b1fee2

Observation 833bd429-3cfe-479f-a311-5cdf9586a51a · outbound

This paper cites Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.834930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.834930Z digest=sha256:318586dc4e81b044ab1b40b9bf7f260566fd2468d6c4463b70efa5fa5bfd1979

Observation 80c8271e-838c-43a1-a83b-a94817eb97a6 · outbound

This paper cites Faithful logical reasoning via symbolic chain-of-thought.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Faithful logical reasoning via symbolic chain-of-thought

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.920168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.920168Z digest=sha256:ae792018ba34bdae0adf29424abbc218bab369ee83cfd5e28dd0f1ec089feb69

Observation b34a989e-e8b8-460a-b41d-e1f0273933fc · outbound

This paper cites Question Decomposition Improves the Faithfulness of Model-Generated Reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Question Decomposition Improves the Faithfulness of Model-Generated Reasoning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.996814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.996814Z digest=sha256:2f457dbc7dc6dba11e51a7f537686ab0f4dd4e724f1eaa7dc4fabe62dc702eae

Observation 7eb7d797-5283-4e3c-b0bc-2597c266defe · outbound

This paper cites Faithful chain-of-thought reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Faithful chain-of-thought reasoning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:05.049781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:05.049781Z digest=sha256:e5345c71ca60872c7f01c6893e32610abd5ad38220a24f7096001943513ac1a8

Observation 36edc327-6624-444e-9786-0bce5bdc0433 · outbound

This paper cites Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:05.118378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:05.118378Z digest=sha256:38a55fbf23f6d88cd0f7f753ab5d4e70d1ff41931b4ca9a4e8b1190fdd2527a2

Observation fe67efb8-5a17-4ce7-bdc6-1cdef025df79 · outbound

This paper cites FLARE: Faithful Logic-Aided Reasoning and Exploration.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models FLARE: Faithful Logic-Aided Reasoning and Exploration

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:05.186460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:05.186460Z digest=sha256:3eb331dd83e4eec516729e6bd11cfbb6ee752b2be16f6001958d19d3e2c06067

Observation bf457c86-83c9-4be5-b4eb-0d1c3fa4e3d5 · outbound

This paper cites CoMAT: Chain of mathematically annotated thought improves mathematical reasoning.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models CoMAT: Chain of mathematically annotated thought improves mathematical reasoning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:05.189613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:05.189613Z digest=sha256:c6bfde1b6121620720596ff48e1e9b34b217ce7189d8ead51c478f452e6266ea

Observation ff9a1246-84ea-4b1f-ba08-851c3bfa4e01 · outbound

This paper cites Causal-driven Large Language Models with Faithful Reasoning for Knowledge Question Answering.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Causal-driven Large Language Models with Faithful Reasoning for Knowledge Question Answering

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:05.267606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:05.267606Z digest=sha256:2b93777d7b619ece1df9383711ce84becad715603c90a22f512fd0eb310a114c

Observation 77f4e3bc-0fbc-487f-ba06-bafe7f586ae8 · outbound

This paper cites Fact: Teaching mllms with faithful, concise and transferable rationales.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Fact: Teaching mllms with faithful, concise and transferable rationales

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:05.365746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:05.365746Z digest=sha256:e6e2565a9c5759fe34f4c98552a23c0db533c2bcc57dd12b3cb4e4876a92bfbc

Observation edf8e7fb-fd5c-4290-bbcc-c0382816e13b · outbound

This paper cites Markovian Transformers for Informative Language Modeling.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Markovian Transformers for Informative Language Modeling

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:05.468396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:05.468396Z digest=sha256:c4e881df9db29c6058bb056f997cb7216db01906eeffd972e30679a55f465767

Observation 6273c2d2-6584-4c26-a111-7a8d6cce3823 · outbound

This paper cites Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:05.612352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:05.612352Z digest=sha256:7602bc6c56851e4e1d2e263d9fb87a96b119962d602ff5596abecad8a3df249c

Observation 19cfad14-82b7-4d3a-a0ed-3d63e836e747 · outbound

This paper cites Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:05.730113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:05.730113Z digest=sha256:a3114e0a2b6acde279d5f6ede6315722faa3e3daf2ffdc89e5d247ee0b502fde

Observation e940c272-3b52-427d-9301-24d1d299e2ec · outbound

This paper cites The hidden risks of large reasoning models: A safety assessment of r1.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The hidden risks of large reasoning models: A safety assessment of r1

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:05.842706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:05.842706Z digest=sha256:8ce21d8dafb4918142f10bb837f8eb82e55042b8558e7b1c58855b69d3e6adf1

Pith citing papers

Observation 6529f910-4f45-4646-b22d-8b94c3ae006a · inbound

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs cites this paper.

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:27:23.292121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T05:22:59.050348Z digest=sha256:a8103869c7f8f450901bf63d2bb10cd78727aaceb698771513ea6c53e10abb9d

Observation dfa30600-790c-40a8-87a8-3eddea1fb926 · inbound

Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework cites this paper.

Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

Reference 25

Resolution
malformed identifier
arxiv_id, observed 2026-05-10T23:45:52.713221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:53:12.473010Z digest=sha256:fea5831f48846903df26d4f528b700260ce9dec7d980c9de1c0e1870266ba756

Observation b32d8cb4-6b38-4187-9c65-dddf41cf8eb7 · inbound

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space cites this paper.

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:41:04.152717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T12:50:57.603403Z digest=sha256:9fcdf24830d8a99de74bc2bcc208aafb717a4a3980e719503c9c5511aefd1a97

Observation 10f6d2b5-7061-4305-af0d-f0e670375250 · inbound

Pause or Fabricate? Training Language Models for Grounded Reasoning cites this paper.

Pause or Fabricate? Training Language Models for Grounded Reasoning A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:46:05.200951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T03:01:58.366028Z digest=sha256:2ffdc2835952f60544ddc4374cad786b43d51a10065b950f4127e61496267251

Observation 9f7f2ab7-39d7-41c1-83be-508b2090a737 · inbound

Auditing Reasoning-Trace Memorization Claims after Unlearning with Head-Conditioned Canaries cites this paper.

Auditing Reasoning-Trace Memorization Claims after Unlearning with Head-Conditioned Canaries A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.269444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T14:10:50.885027Z digest=sha256:4ebd46b4869280e543b9e9becae7ca9dbb81d93df91b56b10b48540e5d7b925e

Observation 3199632a-16f2-449e-bb81-e1a2c6d87600 · inbound

Where Do CoT Training Gains Land in LLM based Agents? cites this paper.

Where Do CoT Training Gains Land in LLM based Agents? A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:51.498548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T04:55:14.293452Z digest=sha256:a34252f0ea946112007abea7983d7b5612025632970a320a4b6157b91b48fe53

Observation a1e99ab0-aaf5-492f-ad01-23543337c690 · inbound

Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning cites this paper.

Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:14:18.884505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T06:12:49.839459Z digest=sha256:e4219a62cc48898a86356a356c2339601e39b8a594acd88aecc9eb50edd93042

Observation 4ef013ca-8e43-4489-a143-943009f877fe · inbound

Risky Business: Measuring The Faithfulness-Safety Tension cites this paper.

Risky Business: Measuring The Faithfulness-Safety Tension A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T13:40:57.907170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:40:57.907170Z digest=sha256:2dca96ac10d803649595fcf0c8cb9a756d5e29eb6580a22039b624efad9b9b8e