Pith. sign in

Paper Citation Record · LEDGER

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models

As of 9 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2505.19690.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19690 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:13:57.131554Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:39:06.815939Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T01:02:54.838590Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5535be79-d61c-4798-b985-68a69e82c44a · outbound

This paper cites OpenAI o1 System Card.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models OpenAI o1 System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.333820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.333820Z digest=sha256:d2529211a0b605cc7e2713b200452a2716192539a3a995803e84d9d1174057fc

Observation 55a0d2a8-2ed7-4151-a0e4-f11d21e65da0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.444148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.444148Z digest=sha256:bd8c8d0c851807b1670042c749a7f0f0ee266ad4e9dc3d01203ac0c105f09b87

Observation 31cf1929-a18f-441d-83e2-5ceb75a7b59c · outbound

This paper cites Qwen2.5 Technical Report.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Qwen2.5 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.568128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.568128Z digest=sha256:4718501f86d24d14fd9a29a1df5b0dc72a02ef0fcc4b4c5519b5d4c637f077ab

Observation 78c0d864-884f-4219-ae30-0e4c50eb0bc5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.688886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.688886Z digest=sha256:af7a75b62b4cf1421fe3c5beb97a3bb9cbad822a12a68031418226e4d6893499

Observation cd40c532-af5d-4742-ad30-1ddb0e71200c · outbound

This paper cites Safety in Large Reasoning Models: A Survey.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Safety in Large Reasoning Models: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.771889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.771889Z digest=sha256:a751dd22286edcc664035fe74a8e7f880b6e6af1682a42f913c2b2c963560ed2

Observation cb9dbb89-0ad2-42af-a47f-2198161e37b9 · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:49.893258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:49.893258Z digest=sha256:3be7dcecc43d707ddd545080880741ced734842fdc5b73204abc28563b0d67f2

Observation b4dc0665-1d70-4883-9ff1-5f83f3f7fcdf · outbound

This paper cites Trading Inference-Time Compute for Adversarial Robustness.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Trading Inference-Time Compute for Adversarial Robustness

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.014107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.014107Z digest=sha256:8d2d83146256abe29d32b0b47f3de70c4593446449641cb5dba648e9334ce48f

Observation be72498e-6cf4-4317-b806-236374894a5f · outbound

This paper cites ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.101785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.101785Z digest=sha256:27ad8383692e9977c92a8ff221ab3b8526aa6cad5fbde27bdb1064b0e9c7dddc

Observation b86248c1-ba00-429c-bbab-e57b4b4fae77 · outbound

This paper cites Overthinking: Slowdown attacks on reasoning llms,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Overthinking: Slowdown attacks on reasoning llms,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.189517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.189517Z digest=sha256:e8546b753df466a9d113f7cabf7498d9b8ed628a9d5630363da860296c56e739

Observation 26fc7304-851a-4c7a-900c-837a097bee99 · outbound

This paper cites Alignment faking in large language models.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Alignment faking in large language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.292007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.292007Z digest=sha256:639aa23ea320f18604475c8968c1d115e67a443e43c667e945fdf156f55d2ea5

Observation 308be687-ef63-4519-9927-b974cc7d1509 · outbound

This paper cites Deepseek-r1 thoughtology: Let’s< think> about llm reasoning,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Deepseek-r1 thoughtology: Let’s< think> about llm reasoning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.391944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.391944Z digest=sha256:6e084a76bbccc9bc14537addb5d49b186a97b303964bfcbe79cb5108f1e95591

Observation 15ecdde5-f430-4616-8d8d-7541c6213c5d · outbound

This paper cites Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.508020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.508020Z digest=sha256:f01789128557e26ed5f6e7f19b940187544019fd4c92d95b638386613c1c43aa

Observation f0eb9e95-cf39-4324-9b69-27947e6cce1e · outbound

This paper cites Safety Evaluation of DeepSeek Models in Chinese Contexts.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Safety Evaluation of DeepSeek Models in Chinese Contexts

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:13:57.994867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:50.590848Z digest=sha256:d7bfaf188bccf1b5314968819dc7d3f18027c470b6ae641e8f0c3d99013428b5

Observation c74ab0e6-58f7-4c82-b0fd-0ccdce93db08 · outbound

This paper cites Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.720640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.720640Z digest=sha256:2f1103c191a74fd274d41e9202193cdafd6544deca29230c7bab1c8f2b2dd4d0

Observation 01b9a4ff-b677-440c-aa06-5bcd0b721490 · outbound

This paper cites HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.813253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.813253Z digest=sha256:f6ab8199a8cd50289c2d8460ff651b985646d53923c80c3cd16bda05752985d0

Observation 1bfad11f-c468-461a-b52f-0a1b8827af05 · outbound

This paper cites Star-1: Safer alignment of reasoning llms with 1k data,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Star-1: Safer alignment of reasoning llms with 1k data,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.907629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.907629Z digest=sha256:1fcffe7f0f184720f4ccb484acdfc3b04935bf3fbc62b298b5b15554b921c77b

Observation ab77b23f-1485-4986-a3f9-675a575d0896 · outbound

This paper cites RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.995870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.995870Z digest=sha256:2f4c3251c15880f9389e73cf37e8ac94805b38b5bc55158a33c085eecbc03de5

Observation 1e882a4c-391c-45d6-94c6-375df79ea19c · outbound

This paper cites The Llama 3 Herd of Models.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.098100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.098100Z digest=sha256:99d6f5b39e5722855a04407056e2ead9474388fed8dca25820b583a5fc1e00ea

Observation 62328b63-bc7c-4b15-8e87-7bf261f73bfa · outbound

This paper cites Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.212632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.212632Z digest=sha256:6f32060786f41d01e9e3f4892dc8edb1791f30c35a45c966a6aeb1077f18b110

Observation dddadb82-3c27-4d83-b6a4-a508bcd7b558 · outbound

This paper cites Guardrea- soner: Towards reasoning-based llm safeguards,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Guardrea- soner: Towards reasoning-based llm safeguards,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.306406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.306406Z digest=sha256:8a3e1b914ae78cfd5852b9efb9cab4deb30ba789cea625e60f6f5c032e0ff112

Observation 6562588d-e7fb-4669-88ad-927cea5e4a1c · outbound

This paper cites ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.424723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.424723Z digest=sha256:eb47cb83c04a79adb4217ec68a02bf9d4d2c2b521e3e4d0f3df587fd1ea3744f

Observation 7277bf91-625e-450a-a6b9-dc573a7ecbe7 · outbound

This paper cites RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.539892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.539892Z digest=sha256:33b9b6c3808e1d021215efae107c89749c1de01b87fa912340ffa2861e09d0ac

Observation 7037e839-8af1-4b57-a9e0-591660cdd27b · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.647994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.647994Z digest=sha256:0ae39d37b59eda90aa431c832f37e1092d7d16a67cf4dc841c3a0111e62ae8df

Observation 4e6727da-bd81-48d1-af2f-45f0316204dd · outbound

This paper cites Ai deception: A survey of examples, risks, and potential solutions,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Ai deception: A survey of examples, risks, and potential solutions,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:03.895924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:51.733525Z digest=sha256:cbed2e8ea6628f9bad66e5cba316eb04c936b52da7e75bfee92e1344d16fde66

Observation 99598a03-01f7-4abe-8173-f38f7f05f47e · outbound

This paper cites Ai sandbagging: Language models can selectively underperform on evaluations,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Ai sandbagging: Language models can selectively underperform on evaluations,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:03.733841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:51.843560Z digest=sha256:f3ce455e67721f4beebff38ba778845b1a3429f42d0c3ee80c3d898724134af2

Observation aa202458-0103-447c-b666-20ea3ed08adb · outbound

This paper cites Auditing language models for hidden objectives.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Auditing language models for hidden objectives

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.942324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.942324Z digest=sha256:0dc783881c79f14623d68fb3b7f47afdcca8c4a05d5754fe277cb8b27508d610

Observation e2fcd080-e72a-4399-9561-abacfd1c8625 · outbound

This paper cites Large language models can strategically deceive their users when put under pressure,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Large language models can strategically deceive their users when put under pressure,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:03.569055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:52.044693Z digest=sha256:5c800a97f654e176846ad3167f02492bb72a6ad69264a24300fcfa33ad0f1983

Observation 4d3872db-06e7-49e6-8c42-ac29188ba1f1 · outbound

This paper cites Towards understanding sycophancy in language models,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Towards understanding sycophancy in language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:03.408967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:52.156517Z digest=sha256:12aae2efb43c2aef0ec77e08f29a36c6faa38f237e94f3dd171377cce90d7460

Observation e31f9482-b5a0-44ba-ab4e-7991955b234c · outbound

This paper cites Fake Alignment: Are LLMs Really Aligned Well?.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Fake Alignment: Are LLMs Really Aligned Well?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.256199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.256199Z digest=sha256:092daeb670d8d81428b05440a0307227ec2618ee1f905307b94b8af8be29c8d5

Observation df34b300-69f0-4081-a93f-1d8bac326022 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Beavertails: Towards improved safety alignment of llm via a human-preference dataset,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.338828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.338828Z digest=sha256:4051a67279528f87a33c17d0c3fb3d7eccdd78769f9a11f3303457d1d13c9987

Observation 62fb4ef6-d5c8-4425-9bab-d36d4cfe9f14 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.457692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.457692Z digest=sha256:dd64997ca36aa8435d166cc57e0a85e3e21024b7a0ec9dbce44238bdd24c85ee

Observation 3ada40e2-6b9b-427a-82a7-f2956c25f27c · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.566860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.566860Z digest=sha256:785e937c63bc28b0366ae02c8a114475d09a700865f89c9eb71d3e8292675402

Observation 6467970d-c79f-4984-ac17-5e3ba0ceb46f · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.672411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.672411Z digest=sha256:56e5d867e7ab0efed9b37df3b1c2410514ced33ce69efc9101a36e340ba7a3b8

Observation 18640e6f-a309-42af-930a-89d013c6b357 · outbound

This paper cites H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.872208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.872208Z digest=sha256:6b7607db1d93f26953f06d6a18d2e263f398e47095a6aee4a350fdf9c33f9465

Observation f8819dd7-2025-42da-b065-8d2e3f16f4b6 · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:53.012947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:53.012947Z digest=sha256:3102fb9f1ad911732d93fbc3943d65299648035d9060e8761d67d2118375ce4a

Observation 8d3e351c-1344-4fec-9943-9b21c1f90314 · outbound

This paper cites Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:53.204922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:53.204922Z digest=sha256:043ed8526ce5956719e217474547977db46677482aade01e8cf1dff85baa009d

Observation aee950f9-1eaa-4ed8-a472-454f7a619694 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models A Survey on LLM-as-a-Judge

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:53.369850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:53.369850Z digest=sha256:baf05bed328f99941abb61983f2969e9b6c8cc9c2ff0945c1f8ac126b8a9b7e8

Observation 9c76d516-b3af-4509-b14f-fa57bb51077b · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:53.561043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:53.561043Z digest=sha256:3b89d383e13731c11eb3faabcbd7602766e1935193278ade7aa0a4747e16e5dc

Observation 15a050fd-41e9-4aea-9f21-c21a4429791f · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:53.757909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:53.757909Z digest=sha256:f4be56f30596d29e8d3b5065da5ba9dcae61a40031cfce467e410e7dc2aba553

Observation c3f4d0ad-0bba-4d54-aff5-0d993aa0f1d8 · outbound

This paper cites Llm-as-a-judge: a complete guide to using llms for evaluations,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Llm-as-a-judge: a complete guide to using llms for evaluations,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:03.259431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:53.849787Z digest=sha256:505143b3fbd8ccbaf729df6e1c4891d23cd103efb1ce77e3f68b67647d5d435e

Observation 8bdc18ab-355a-493e-b4ee-b2f0734a296f · outbound

This paper cites Llm-as-a-judge simply explained: A complete guide to run llm evals at scale,.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Llm-as-a-judge simply explained: A complete guide to run llm evals at scale,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:03.079495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:53.976647Z digest=sha256:c58213c2950214c41ac4ae6661136bc2631453a3f9e941e040bac7f68263af26

Observation 7c674e3d-7334-4773-91d8-1318731fe9a6 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:54.133916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:54.133916Z digest=sha256:bbe148e60e8c7814fa204c634eea04f606869a76281a22eb2353426ad372df72

Observation 104f63dc-17ea-41e7-bf3c-5c73d04616ad · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Constitutional AI: Harmlessness from AI Feedback

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:54.323505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:54.323505Z digest=sha256:fb794cd9d8db6107b1256a07452c56ade0f99b6125012eabe4c2fc3be459ca02

Observation 843de4f0-38dc-4827-a444-dc5a8d97eef3 · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:02.925816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:54.480196Z digest=sha256:972da55ce9f1dcb1fa434ec5f9d1c996d5577c5f4d61abea409ebbd766621a85

Observation f01efde8-bfac-4c25-98bc-007713015ac8 · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:02.747764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:54.667429Z digest=sha256:6a7e622f380ae27d4d01ad4db99a98bb5fd6d77597f7cdac38be1f150466703d

Observation bace8e75-82cb-4877-ab60-247932cfbb00 · outbound

This paper cites # Evaluation Guidelines.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models # Evaluation Guidelines

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:02.627227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:54.885944Z digest=sha256:cf78b17ffa657f59432635f3637552160ed923e643cb5b712f1a87adedb42f65

Observation 82dce026-34f9-413c-854f-9b8f3312384d · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:02.458444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:55.031756Z digest=sha256:5043394449dc0069454441a57a39939123a783ac786629c700b732ba8c81becf

Observation 826cb8a3-daa5-47cd-9c16-de6d5a72668f · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:02.286632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:55.254663Z digest=sha256:01fc1fbdd53dc55c442358a4967fe663ca2111ed236f8d8663d8e07572209a67

Observation 40bf18e0-938d-42ef-b79c-356adec2d824 · outbound

This paper cites Reasoning Quality Evaluation.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Reasoning Quality Evaluation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:02.130403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:55.338077Z digest=sha256:d74ae05066236faddb415017132e0750372a9126f2d6c2b718a20f4bcbfb7a3d

Observation 355bdf47-f487-4948-b442-9f13ef2f751f · outbound

This paper cites (Usually includes two main risks, e.g., risks of insulting others and privacy violations.) 15.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models (Usually includes two main risks, e.g., risks of insulting others and privacy violations.) 15

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:01.950266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:55.449364Z digest=sha256:e4f97e0b48ab65e8027f6f5cbabbe942e71dfe120a95267053b07d9e43d2d692

Observation d8d671c0-3f5f-43cf-81a9-734cef4c3b9a · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:01.850665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:55.621282Z digest=sha256:7bb9e8374675bbf3f2615942cdcf59c10fe17517c0a46ec0bb211352bd88713e

Observation 359b3b1e-7b58-4d6b-84d1-13683bc3dba9 · outbound

This paper cites # Evaluation Guidelines.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models # Evaluation Guidelines

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:01.685722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:55.715943Z digest=sha256:a14244d00afd4f52211a1b0bb639bd983052543f3efffe3b3d9d74e7bfbdbcc4

Observation f0a35677-0886-410d-afd9-d11300c7ecf2 · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:01.463138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:55.828696Z digest=sha256:5c2fc30cfdb06d9c9573e600bfa67b351e85ca58ca5a5f62794b3fa4ae09f6d2

Observation 3650a463-b1a9-4a70-a63e-72322a235a65 · outbound

This paper cites Reasoning Quality Evaluation.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Reasoning Quality Evaluation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:01.207257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:55.929030Z digest=sha256:f20e81f219a74f0359d5687608682bd8f07f37a38371a2a1f4bed77bf07af5d3

Observation 54e0834a-3ca0-4756-8235-c2098859ea3d · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:01.017613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:56.031612Z digest=sha256:bf2e7f80a4828039c2fcde05de5f71f7a35bd0c436a66ee12d922a3ba399f7b7

Observation 9d172e54-b009-4572-ad51-d267a2cf3e9c · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:00.804263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:56.106352Z digest=sha256:126acc5954dd910e5d08201b99036da7eefaf34426194435dba7258366a7aa9f

Observation 7898a7dc-63ce-4948-9441-5617486baea4 · outbound

This paper cites # Evaluation Guidelines.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models # Evaluation Guidelines

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:00.542408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:56.212901Z digest=sha256:1a1a61f590daa39616667ae5787a1f8670be71a47e7607381b597672925b14a9

Observation ce03aac8-dfdc-405b-b9b9-3bae91286b2c · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:00.349370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:56.326896Z digest=sha256:84a963fc6c17442b588ffed2f4f34ebc42e04c8c28a380ca0c773bc2eb28bde8

Observation 03040abd-0260-44df-ba73-0cc4c5581b7c · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:00.146106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:56.427560Z digest=sha256:8252251fcfc92f9a158279bcaede8db27ebeefd695d64a70242b690bbb422fee

Observation ca921e1e-4bba-4c8f-8904-358688b89094 · outbound

This paper cites Reasoning Quality Evaluation.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Reasoning Quality Evaluation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:13:59.874821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:56.543372Z digest=sha256:a7a8f5bc03756c7a4ea189423d80848de0e0acc6778034ef8e3bec95804e9d3d

Observation 912e62e4-1d45-45b0-a113-c96b0eb3b7c6 · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:13:59.646339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:56.670741Z digest=sha256:f90bed67b78e7a4f40cfa8e0fd7a78442d81e100fc62e49c0ed2519e3e878a08

Observation 95c3a5a2-aeb7-447e-b00b-fd54562b9199 · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:13:59.406710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:56.779273Z digest=sha256:7a97b1ec16b941b54aff34a124b7fa7ecb83a9de94588e0b16de89822bf83194

Observation 4edb71d9-1afd-4285-b5c3-08694e316455 · outbound

This paper cites **Evaluation Guidelines:**.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models **Evaluation Guidelines:**

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:13:59.198194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:56.866657Z digest=sha256:69f5da0bba5bc5ba81b1d8fc8dd3a0227a483181da8d6b5e6790dbead72e9cf2

Observation a8a62d21-5e73-44f6-ab92-6deeec9b3a54 · outbound

This paper cites A safe answer must refuse harmful requests or appropriately redirect, without producing problematic content.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models A safe answer must refuse harmful requests or appropriately redirect, without producing problematic content

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:13:58.988709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:56.972144Z digest=sha256:2b862936e9984be3dffdeb5ac562df0747ec80ede0453f60ac6085b8815bd2ae

Observation d0580437-d369-4bc6-9da0-9bd3c761afe2 · outbound

This paper cites an unresolved cited work.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:13:58.863870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:57.037469Z digest=sha256:5f74d5d18c37e66ef8d05757ed08ec7479c89354080de873190b3958ee871821

Observation a736e2f0-fb16-4d07-81e9-83e3bbfd6e19 · outbound

This paper cites User Question.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models User Question

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:13:58.565598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:13:57.131554Z digest=sha256:c90722b2ab413be8f44538a6a0bd4746acea1f32e6927c2e367513b15811eed1

Pith citing papers

Observation 8a205ab1-4ae9-4570-93e8-02c8db525cb1 · inbound

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments cites this paper.

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:02:54.841578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T01:02:07.088724Z digest=sha256:ccf4a1ac84ebb5bbc14e2d3601601f303de737b851bf54455b68a6f7f2e8c51c

Observation 75343ba8-4f88-49aa-8c3b-a262141cdd23 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:06.815939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:06.815939Z digest=sha256:8903fa4bddee107d558b7ec1868659167c34e26117287c19b0e28ee3a4d94855