Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T05:05:01.269473Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2501.17030.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T05:05:01.269473Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:50:12.427431Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T15:59:41.195117Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 31c8ab8e-1d47-43db-8686-eb16c03313d6 · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ab578b07-dcaf-4e0b-936a-3bcfc330ff2f · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 116f0772-e6d7-492b-acb3-9975bd43de23 · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies A survey of reinforcement learning from human feedback, 2024
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1d97cb7f-1ed0-49fa-ac70-b7e5926ed309 · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bf6030c7-23ab-4f9b-9378-6967f3319124 · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies A survey on knowledge distillation of large language m odels, 2024
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fe8a78d5-a66e-4ec3-8031-b8df746f6cd5 · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b8072c70-2e3f-4e8c-ba4c-d045863e50aa · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Ai alignment through r einforcement learning from human feedback? contradictions and limitations, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 04f5af42-5c1d-4a86-ac80-0446af801e1a · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Making harmful behaviors unlearnable for large language models, 2023
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 01e4d683-b0ca-4d4e-a9cf-05f371b640b4 · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Reinforcement learning enhanced llms: A su rvey, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7d49c436-a5f7-4f2a-a103-691e2bca5197 · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Harmful fine-tuning attacks and defenses for large language models: A survey, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 194398b7-781d-44ec-b522-a9f1e9467b56 · outbound
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bebbc818-4e1b-4dcf-95a9-82f42aebfe2c · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Safety-awar e fine-tuning of large language models, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d9ea1634-02e3-4156-b906-28df3de31eaa · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Reinfo rcing thinking through reasoning-enhanced reward models, 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c2ab986f-092b-4b24-b03e-847b9bbc8efb · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Learning and forgetting unsafe examples in large language models, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a479a1f7-77b0-4096-9c38-96cb9eaab9ee · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies U nlock the correlation between supervised fine- tuning and reinforcement learning in training code large la nguage models, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 384eaee3-e832-488d-b94d-28ca5a5ec35a · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Q-sft: Q-lea rning for language models via supervised fine-tuning, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bbb3a527-114c-458b-9c92-b8f88b161636 · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Superhf: Supervised iterative learning from human feedbac k, 2023
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 60786ee8-b584-423d-88cb-e60596daf280 · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Reinforcement learning fine-tuning of language models is biased towards more extractable features, 2023
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1950a342-4a06-495c-ba88-e7a994f7c3af · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies A closer look at the limitat ions of instruction tuning, 2024
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 19b95f18-8d05-46cb-bc42-d5cfb8818056 · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Mitigating forgetting in llm supervised fine-tuning and pre ference learning, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bcaec5dd-67fe-445b-83f8-6a9c48200214 · outbound
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies Supervised fine-tuning as inverse reinforceme nt learning, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 38c0dfef-a61f-4d89-9103-1c632de164db · inbound
DeepSeek in Healthcare: A Survey of Capabilities, Risks, and Clinical Applications of Open-Source Large Language Models Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c8e444-b67c-4f7e-a612-4da9aa17c510 · inbound
Reasoning LLMs in the Medical Domain: A Literature Survey Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.