Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:02:38.886258Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2507.01327.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:02:38.886258Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 31a208ec-737b-4732-823d-902ac8dc57a8 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f9e17e1-3cc5-483a-b922-d22bd53228f7 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy o3-mini vs DeepSeek-R1: Which One is Safer?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14b9f2ce-14ee-4cf7-a291-fb1d8309115d · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d5c8270-2ff8-49f1-b33a-2dbbc9d82af1 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 023fce87-47f5-49f7-8100-39b5698eaf4c · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aa882ab-7b75-4db8-8aef-e46eea1ddc3a · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86e7978e-5dd8-4787-9961-5830e10eec17 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce74fb56-f1ee-4738-a1c1-2e2a673cc6c3 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c657feba-84a2-4a22-8038-9a9c4d28439b · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4835359-7fc4-40df-b9fd-6224fa5dd951 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c2ad01c-1d6e-4ccf-abb3-9329f1992d0e · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1327e2c-c0e8-4beb-85d4-dfe89e595765 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e57f1145-b8fc-42c4-bed9-231ca0008c03 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8878cbaa-47b4-4a76-80a2-7262b7f49da0 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd28b2ec-b78f-4368-bd84-ba7e374f707d · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c259246-689b-4490-928f-5ad5de36cd71 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy OpenAI o1 System Card
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7e29413-02af-4dc8-8d42-a5e0a05c3671 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Scaling Laws for Neural Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f43a6cd5-3164-47a8-a521-720e6188e685 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Gonzalez, Hao Zhang, and Ion Stoica
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12f06d24-1b9c-4673-bf13-d373f5c1b33a · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b8e6642-fd7f-4db2-babd-ce61d14f1753 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Towards General Text Embeddings with Multi-stage Contrastive Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cb29d71-2aa2-4135-8315-01a3eedbb761 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3d2dc45-2930-4ced-9b96-f57ef3b601f6 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 806618c4-0cc6-4a35-9b9a-34319bb0c842 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Understanding R1-Zero-Like Training: A Critical Perspective
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7586d958-0dc2-4586-9fd6-44af33d8ac51 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07392453-e916-4d8d-8ef8-bc9b199f4f63 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aeb3252-57f0-4b23-9cf4-20a64e488caa · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b96b7442-c81b-4c59-80d4-041fdf07e4dd · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d6d9956-b807-470b-ad2e-2f02d026acf9 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ceacf498-c68d-439d-9547-2842c0e27f5d · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy GPT-4 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1d61937-ff6d-4540-bd52-7cfa4ed879ee · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04ea2fdc-426e-4521-8c9c-92d6778a9918 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a00037ce-f4da-4fb2-ba28-2a7ffbd0fbe1 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Multi-class Text Classification using BERT-based Active Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32ffb3d0-ec8e-4c7b-b520-c4910cf7dc8b · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0591341e-520c-4e6c-af89-90d0f5dffecf · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Qwen2.5 Technical Report
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7ace076-2e60-4f14-8da7-2890c9818394 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9babff3-870f-4bcd-b9bc-6ed2395b8610 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2715de2e-6c1d-45f2-ac73-649341fc1bea · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Ray Interference: a Source of Plateaus in Deep Reinforcement Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b522326-3a75-4857-915a-6ccf1c508f45 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Proximal Policy Optimization Algorithms
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb864e8-2e6b-4ad6-92e9-26720e25ec49 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4571d06d-73c8-4ff6-83e7-161d0e3e36d0 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy HybridFlow: A Flexible and Efficient RLHF Framework
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 042cd7d2-6082-4890-bce1-bb2480a4d2e9 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e05d6af9-440d-4d16-bb6d-5ce7cc0b3662 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11905553-2ea7-4744-ad24-d71fc44b1561 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e72aea4-7d5f-474f-be6f-67a07e0a119e · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Fine-tuned vs. Prompt-tuned Supervised Representations: Which Better Account for Brain Language Representations?
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a02cf4b-ebfd-4335-aac8-44d35fa87c26 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df9f09b6-8183-466e-98d8-68bd965f2c22 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy LLaMA: Open and Efficient Foundation Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33e7ef3a-5699-49e0-8229-fa15229b2143 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy ChatGPT Empowered Long-Step Robot Control in Various Environments: A Case Application
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5765637a-d77e-4077-9058-a0f4022bc6d1 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Finetuned Language Models Are Zero-Shot Learners
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e8d32ef-5428-4420-81a6-7d0ec150fb46 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd18e72e-1c9a-4bbb-92b7-fee9f47c2141 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20842445-d568-402d-b50f-703292704157 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dc20609-ed91-4ef5-a216-9aa1bdba487c · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 777fff92-f2f1-415f-9946-188530af05c8 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1c42d31-9ed7-48d7-8c42-8cefef7bf7ad · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bb0e20d-945b-41d7-a4a4-30fc039d6810 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48d27688-5e9e-45c2-a2d9-51bd55bfa345 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71613651-96da-44e0-bf60-01d6a1a7ebf9 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0df93c6-3af7-4ce3-ad93-beb0d070b678 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c168912-83aa-46d6-94cc-af00a997d0dd · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0586836-fa80-412a-920f-13e07fa9f098 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 656c26ef-ca14-4248-8233-0421db02e172 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d90440a-2dd7-43ce-8ece-082d594f8f6d · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy online" 'onlinestring :=
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16f3e57d-a978-4b96-9d35-2180688f1449 · outbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy write newline
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.