Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T01:16:05.392554Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.00133.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T01:16:05.392554Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1cf5dcd9-b8a1-419b-aa3c-91e86c17e31a · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Genie: Generative Interactive Environments
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d729623-6732-4a19-b31b-0a77d438c5f2 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Soft Actor-Critic for Discrete Action Settings
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19a6c79a-b59f-43f6-8274-d9d792531765 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31225c62-a18c-46fb-88a0-af4f988b016a · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Improving alignment of dialogue agents via targeted human judgements
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a08dccf8-c088-4d9e-b847-916002338995 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Learning Control Barrier Functions and their application in Reinforcement Learning: A Survey
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de2b33a2-0b57-466f-8614-b71b6d803c74 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models World Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 660c4eec-04d9-462b-9a86-a13ec9e8fdda · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Soft Actor-Critic Algorithms and Applications
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 939e08b6-7e87-40a3-b90f-c4bbb335b98a · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Mastering Diverse Domains through World Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac4b6f60-c66a-4eaf-96f4-0e00bcbea1f6 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Nicklas Hansen, Xiaolong Wang, and Hao Su
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8ce6f9a-7e3d-42fd-89cc-5cd12670f129 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Dropout Q-Functions for Doubly Efficient Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0381c50-8cf6-4358-b515-fd7660eb2a3e · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models ORPO: Monolithic Preference Optimization without Reference Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b90705c3-4260-4570-8d66-8578008d5b04 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models A Control Barrier Function-Constrained Model Predictive Control Framework for Safe Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 003ce360-d3df-4b6e-a8b5-bfc5c88f9d4b · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c36079bd-0e77-47aa-8dcc-6f668a865610 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models A Review On Safe Reinforcement Learning Using Lyapunov and Barrier Functions
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f032c323-925b-4c94-a2a7-1ac27b426398 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2447f6b6-a2e2-4152-bfa4-e84abf25a818 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Continuous control with deep reinforcement learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e71f77-bbbb-44d0-91f7-57e680ea3b4e · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Online finetuning decision transformers with pure reinforcement learning gradients.arXiv preprint arXiv:2601.00167,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3ccb85a-90fc-470d-be34-1ff702fb6851 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Efficient soft actor-critic with LLM-based action-level guidance for continuous control.arXiv preprint arXiv:2603.17468,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1982bf56-417a-4cf6-96dd-b3a3fcb8088c · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be68d9f-2ae8-4ba6-ae08-d67bf38cf296 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fb2f258-a3a1-47d9-bf4b-d0edba7f3b23 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models How to train your latent control barrier function.arXiv preprint arXiv:2511.18606,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eba33d7a-9c7d-4e48-911d-9ecc8dcb641b · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20b1dd58-bc4b-4dfb-95d6-1c4d4426f72d · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2111a3b-d4cc-4667-9bf6-50479691f8eb · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d9d1167-8c6e-4b9e-ad31-8ac7a273542f · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Towards Understanding Sycophancy in Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b3cc78-0585-4d7a-b578-47d9d18788a3 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Defining and Characterizing Reward Hacking
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f028ebc4-b306-4de3-b1a3-e632691dbc5f · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Target Return Optimizer for Multi-Game Decision Transformer
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd40a34f-c338-47fc-83fb-c2ffc2a069b5 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Gymnasium: A Standard Interface for Reinforcement Learning Environments
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86f53e1a-e55a-4873-8bf7-d53c20bd42a8 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Deep Reinforcement Learning and the Deadly Triad
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32df2db8-2384-4a95-ad7f-c374e6cba6e4 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Behavior Regularized Offline Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f70470f-965c-4888-9572-6e5548bbf852 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c55b42e7-558f-4eba-b3d0-6d259ae5a02c · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models CBF-RL: Safety Filtering Reinforcement Learning in Training with Control Barrier Functions
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 143bfbb0-2a4a-4c98-989e-e08eb895f6c5 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48e765bb-755c-4a4d-b144-67fca8c8c5da · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c85bfbf-931a-4c06-855b-513c0439af37 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models A Deeper Look at Experience Replay
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24175e85-901a-43af-9389-604fd30d8ee7 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Revisiting Discrete Soft Actor-Critic
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c24103c3-e6bd-4024-95f5-f708b29fdf6f · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Fine-Tuning Language Models from Human Preferences
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e11276-cd1e-490d-8e1a-aedfd8134572 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1993
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c61d01a7-d41e-460e-8973-022a07b7e244 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 1998
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aadab0d4-40bf-47b7-b095-e1c44389b842 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Delgrange et al
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2988e363-41f8-479f-90ad-9503baae70d0 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fea863f-6f42-45cc-b4a4-e1fa5487bc6d · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Concrete Problems in AI Safety
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99bc4337-08ae-461a-9bd6-1a2396b92e0c · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Constitutional AI: Harmlessness from AI Feedback
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd6d564e-4faf-4ee3-81ba-5f8d7d12e922 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models 2024 ACM a.m
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1888d178-3b4b-42cf-ad81-6866b1b56b89 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Training Verifiers to Solve Math Word Problems
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a6e3e87-58e6-4cb1-a27d-90494072aec9 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Process Reinforcement through Implicit Rewards
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e73967b3-8b6f-4d3a-a30b-9f3235461112 · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models OpenAI Gym
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65176c4b-f0e2-447f-8ebf-363b34b5ffdf · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb622e9c-e10a-4359-bdc1-e8dd456c3d5e · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 009c66c3-36ad-4803-b635-9cdf1538379b · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models arXiv preprint
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b13b507c-def0-489c-8653-a2b950219aed · outbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Playing Atari with Deep Reinforcement Learning
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.