Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:13:57.131554Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2505.19690.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:13:57.131554Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:39:06.815939Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T01:02:54.838590Z
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5535be79-d61c-4798-b985-68a69e82c44a · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models OpenAI o1 System Card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55a0d2a8-2ed7-4151-a0e4-f11d21e65da0 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31cf1929-a18f-441d-83e2-5ceb75a7b59c · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Qwen2.5 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78c0d864-884f-4219-ae30-0e4c50eb0bc5 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Chain-of-thought prompting elicits reasoning in large language models,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd40c532-af5d-4742-ad30-1ddb0e71200c · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Safety in Large Reasoning Models: A Survey
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb9dbb89-0ad2-42af-a47f-2198161e37b9 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4dc0665-1d70-4883-9ff1-5f83f3f7fcdf · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Trading Inference-Time Compute for Adversarial Robustness
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be72498e-6cf4-4317-b806-236374894a5f · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86248c1-ba00-429c-bbab-e57b4b4fae77 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Overthinking: Slowdown attacks on reasoning llms,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26fc7304-851a-4c7a-900c-837a097bee99 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Alignment faking in large language models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 308be687-ef63-4519-9927-b974cc7d1509 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Deepseek-r1 thoughtology: Let’s< think> about llm reasoning,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15ecdde5-f430-4616-8d8d-7541c6213c5d · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0eb9e95-cf39-4324-9b69-27947e6cce1e · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Safety Evaluation of DeepSeek Models in Chinese Contexts
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c74ab0e6-58f7-4c82-b0fd-0ccdce93db08 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01b9a4ff-b677-440c-aa06-5bcd0b721490 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bfad11f-c468-461a-b52f-0a1b8827af05 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Star-1: Safer alignment of reasoning llms with 1k data,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab77b23f-1485-4986-a3f9-675a575d0896 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e882a4c-391c-45d6-94c6-375df79ea19c · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62328b63-bc7c-4b15-8e87-7bf261f73bfa · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dddadb82-3c27-4d83-b6a4-a508bcd7b558 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Guardrea- soner: Towards reasoning-based llm safeguards,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6562588d-e7fb-4669-88ad-927cea5e4a1c · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7277bf91-625e-450a-a6b9-dc573a7ecbe7 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7037e839-8af1-4b57-a9e0-591660cdd27b · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e6727da-bd81-48d1-af2f-45f0316204dd · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Ai deception: A survey of examples, risks, and potential solutions,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 99598a03-01f7-4abe-8173-f38f7f05f47e · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Ai sandbagging: Language models can selectively underperform on evaluations,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation aa202458-0103-447c-b666-20ea3ed08adb · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Auditing language models for hidden objectives
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2fcd080-e72a-4399-9561-abacfd1c8625 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Large language models can strategically deceive their users when put under pressure,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4d3872db-06e7-49e6-8c42-ac29188ba1f1 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Towards understanding sycophancy in language models,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e31f9482-b5a0-44ba-ab4e-7991955b234c · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Fake Alignment: Are LLMs Really Aligned Well?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df34b300-69f0-4081-a93f-1d8bac326022 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Beavertails: Towards improved safety alignment of llm via a human-preference dataset,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62fb4ef6-d5c8-4425-9bab-d36d4cfe9f14 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ada40e2-6b9b-427a-82a7-f2956c25f27c · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models OR-Bench: An Over-Refusal Benchmark for Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6467970d-c79f-4984-ac17-5e3ba0ceb46f · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18640e6f-a309-42af-930a-89d013c6b357 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8819dd7-2025-42da-b065-8d2e3f16f4b6 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d3e351c-1344-4fec-9943-9b21c1f90314 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aee950f9-1eaa-4ed8-a472-454f7a619694 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models A Survey on LLM-as-a-Judge
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c76d516-b3af-4509-b14f-fa57bb51077b · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a050fd-41e9-4aea-9f21-c21a4429791f · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3f4d0ad-0bba-4d54-aff5-0d993aa0f1d8 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Llm-as-a-judge: a complete guide to using llms for evaluations,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8bdc18ab-355a-493e-b4ee-b2f0734a296f · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Llm-as-a-judge simply explained: A complete guide to run llm evals at scale,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7c674e3d-7334-4773-91d8-1318731fe9a6 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 104f63dc-17ea-41e7-bf3c-5c73d04616ad · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Constitutional AI: Harmlessness from AI Feedback
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 843de4f0-38dc-4827-a444-dc5a8d97eef3 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f01efde8-bfac-4c25-98bc-007713015ac8 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bace8e75-82cb-4877-ab60-247932cfbb00 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models # Evaluation Guidelines
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 82dce026-34f9-413c-854f-9b8f3312384d · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 826cb8a3-daa5-47cd-9c16-de6d5a72668f · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 40bf18e0-938d-42ef-b79c-356adec2d824 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Reasoning Quality Evaluation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 355bdf47-f487-4948-b442-9f13ef2f751f · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models (Usually includes two main risks, e.g., risks of insulting others and privacy violations.) 15
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d8d671c0-3f5f-43cf-81a9-734cef4c3b9a · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 359b3b1e-7b58-4d6b-84d1-13683bc3dba9 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models # Evaluation Guidelines
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f0a35677-0886-410d-afd9-d11300c7ecf2 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3650a463-b1a9-4a70-a63e-72322a235a65 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Reasoning Quality Evaluation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 54e0834a-3ca0-4756-8235-c2098859ea3d · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9d172e54-b009-4572-ad51-d267a2cf3e9c · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7898a7dc-63ce-4948-9441-5617486baea4 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models # Evaluation Guidelines
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ce03aac8-dfdc-405b-b9b9-3bae91286b2c · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 03040abd-0260-44df-ba73-0cc4c5581b7c · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ca921e1e-4bba-4c8f-8904-358688b89094 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Reasoning Quality Evaluation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 912e62e4-1d45-45b0-a113-c96b0eb3b7c6 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 95c3a5a2-aeb7-447e-b00b-fd54562b9199 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4edb71d9-1afd-4285-b5c3-08694e316455 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models **Evaluation Guidelines:**
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a8a62d21-5e73-44f6-ab92-6deeec9b3a54 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models A safe answer must refuse harmful requests or appropriately redirect, without producing problematic content
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d0580437-d369-4bc6-9da0-9bd3c761afe2 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a736e2f0-fb16-4d07-81e9-83e3bbfd6e19 · outbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models User Question
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8a205ab1-4ae9-4570-93e8-02c8db525cb1 · inbound
ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 75343ba8-4f88-49aa-8c3b-a262141cdd23 · inbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.