Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:33.995648Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 100 of 263 outbound references and 11 inbound Pith citation observations for arXiv:2509.16679.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:33.995648Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T20:14:04.060992Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 263 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 75da94fe-2a7c-4b6d-8726-f40e5b02ced5 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle NeMo RL: A Scalable and Efficient Post-Training Library
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b0aae59-024d-4ad4-bded-1c6c3f5b4e68 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc88b0d3-dfa3-46bd-8d37-5731a868c0b8 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae598da4-298a-45ac-841d-6fc736736556 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle A General Language Assistant as a Laboratory for Alignment
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59a51287-89e9-49b0-8d36-c8fa5e843dfa · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b3e214e-4027-44e3-ad3e-539f0a8410da · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3bdfe05-bfcc-4e36-90e5-fe91b5390176 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Constitutional AI: Harmlessness from AI Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75c084d6-33c5-44d8-81c1-3e1257b7ef9d · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Thinking Machines: A Survey of LLM based Reasoning Strategies
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f7f2e04-4a7c-4d6a-8a75-32e8e1cf9336 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Do, Yan Xu, and Pascale Fung
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd246f83-02c2-4b8d-b500-eaf06e80367f · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d82380ff-a5ea-4668-81f8-f646c6492e78 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1916877-41dd-489e-98f8-b63f121d0c27 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reasoning Language Models: A Blueprint
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9561ea-e3fb-4436-a269-e5d4cf91808e · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d8175f-0578-4fcc-a3dd-97a8dab4cdfe · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle On the Opportunities and Risks of Foundation Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04f9b39a-d183-48f0-97f9-8ea6c5e4c965 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07a5c395-e85b-47e6-8944-e6f28aab8d59 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f11dde3-5b1c-4a67-a18c-3dace7c545f6 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64a5b613-97f8-40e5-bd7e-3f5e9678c71d · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe009cb9-2b50-4598-83d3-acbabef4f6b3 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b56924-bdac-44a4-8057-504c782c8133 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5722edcc-a5bd-4e16-9781-3eceeb8c3177 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7273204-5657-4801-966e-990054462abe · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c08c5d-0f4e-4c78-a89e-f2ee7125aaf4 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 204d6904-3f0c-4a47-b150-6d88ee4dda7f · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Compile Scene Graphs with Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0336c218-8eeb-441a-85a1-478bcc30418e · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e3e9ae-e1a9-492a-8e82-60b49c83205e · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Training Verifiers to Solve Math Word Problems
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a572717-c055-49e3-927a-3930807cca81 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6506b4ce-8fa9-4950-8f8a-c64a4fefe82d · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994cb6c9-e3dc-4725-8a93-657b246f0fd7 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c39a617e-8a38-456c-a1b9-3b25bf595665 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 481bd3f4-b9cc-419d-be04-f5ae99ef2ace · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcement Pre-Training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8be9316-60c5-4dd4-8320-403e4905d74c · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 303355bc-dd3c-43de-b77b-526cd55a09f5 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf3cca11-4862-4e85-a765-c753ea3879df · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db69603-4fa2-49dd-86c8-5e16e100152c · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b052fbbe-70e8-48c3-a101-990c49522374 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Thinkless: LLM Learns When to Think
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e14da2-97a7-4a3f-b6da-1696f3588219 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c001a136-ba18-4657-9af6-862cfa25bf17 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b4b132-0618-47eb-a124-a9f0de000942 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Group-in-Group Policy Optimization for LLM Agent Training
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d89261b-1454-43a4-a0b5-9ea6d18d6b1d · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3d2f39a-8b98-457d-9cc5-1bf124dc3f08 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8d6583d-a411-4896-ac62-6ef44ab2e356 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a302257-1715-428a-ad23-c15f21868d2b · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74598496-fe99-4539-bf71-53871d05cb86 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3770ca00-2666-4ab3-8baf-d1275f82cd24 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Visual Pre-Training on Unlabeled Images using Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eb84988-09be-4dc8-9337-a1b723f9e5a5 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d72453d-d5e8-45fc-a2a1-35a9ac617340 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 319648f6-4b28-432e-a38d-9b768458a376 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3900aa1-9feb-460d-9ceb-9b19a2dc5c8d · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reward Reasoning Model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9aff45a-0e39-4acf-9076-ba84fa7fa73c · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Synthetic Data RL: Task Definition Is All You Need
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb229a9-44a0-47da-acb9-7bc50e4de767 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd906e2e-ba42-45cf-b81b-8cf72f147eca · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9442083a-9dd8-4aa2-9f9f-c10435c2f283 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2e83bd1-7106-43b0-9502-5f17c29e3246 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cff38d40-0384-4dfd-b50f-4b7a1be48da5 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 192b405b-a9bc-482a-98b8-9088abfc0528 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e89d2e0-c61a-481a-acad-d83d5fdb580e · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe72ef61-5397-4f90-b97a-143a3c5125c4 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73fda0c8-0b55-468b-9458-326729616f46 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51a23c9c-8d3c-461a-beaa-16414fcc4226 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31792538-d656-4233-b016-9db69d3c341a · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00ae4d2d-d592-4a4c-ab4d-5ea56b6be92b · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee7e32b-36b7-4bf1-9db7-2c0e8e53d450 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fb8a73f-287e-4ea6-a7d3-064619ab26af · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SLOT: Sample-specific Language Model Optimization at Test-time
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe90b436-1748-42c1-a9ec-edf95282059b · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3966b4c6-aee4-4b0f-a038-e38f848f33ca · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 460d4a8a-d80b-43fe-bcf5-c16251f52e35 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle 3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 423caa36-317e-440e-bc6b-5f65775b5463 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5285291a-8157-4889-a271-cdd9944ca92b · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle GPT-4o System Card
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc91860a-5043-4c37-a387-c2d39244a6e2 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle OpenAI o1 System Card
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04086387-639e-4a0b-b1bf-1820ffbbd2a5 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f62552a-42ae-461a-a31e-025c52a5da8e · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle A Survey on Progress in LLM Alignment from the Perspective of Reward Design
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2497ffc9-72aa-4900-8f5f-06a21319f463 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11a487ad-e7a2-4edf-88fa-750f4e1bae50 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Think Only When You Need with Large Hybrid-Reasoning Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c595a2f4-a866-47f8-9ec6-70a2fc92d731 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle A Survey on Human Preference Learning for Large Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff833ea8-a32c-4dd3-9eb5-cbd452a1d7f4 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d877f3d-03f5-4ab8-b772-c70b75a319ca · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02979552-fac8-4515-b377-b34adf2a43c8 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff89a9f-4a92-4457-9983-5683a446d207 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle BIG-Bench Extra Hard
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fca4bac-f0e6-4cbd-8166-6baed38a03e1 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06574205-1244-4c8b-b693-d2455cf9930e · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Alignment of Language Agents
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a69d7b46-fc1f-49b9-8fc5-fe3c25a26548 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c8251bc-4df4-40ef-98e5-96a26da71e88 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7cbecd9-fc24-49bb-99c5-b6f5dd58c1ed · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1255068-badb-4245-80ee-4c78eba503a5 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle LLM Post-Training: A Deep Dive into Reasoning Large Language Models
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 648ae64a-db90-4202-b7ff-5b8212cf3953 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d7a65e5-8847-4f1d-9a0e-78372b7d687d · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84e29578-dfb1-4258-aeb6-5a4d37540d91 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08dbc974-e252-4e5d-ad6b-9a2077e914ce · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfe83f2a-4ba7-47b6-aaac-e62ec57aaf9d · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a0e690e-1d98-4c45-b51b-c1a304d95dfb · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8cfb525-bb79-4e86-b03a-f8abf63d5b85 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8a6185a-fc02-4d71-8a46-355ab7680a42 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac6d0e7b-4034-494d-b8cf-bf0226db0283 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle ToRL: Scaling Tool-Integrated RL
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0657594b-20a4-42e4-84e5-f1d2760997d2 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cb9acb6-6d08-4dee-9e45-d2c4442bcc98 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb5a63db-4682-4902-abd3-5b5346672270 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Unresolved cited work
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f1922a0-0415-4eae-b74f-083630c02ee7 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Generalist Reward Models: Found Inside Large Language Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6776643c-97a0-4199-9eb2-0fd15a3547b6 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a7ed2cc-0db1-417c-8b68-aae1f1a6ccf7 · outbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Improved Visual-Spatial Reasoning via R1-Zero-Like Training
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39461bab-9f44-4b56-9a4f-484a762757cb · inbound
Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0e13ce9d-1ac6-4f46-ba5e-ece760393791 · inbound
EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f5b7743-6b98-41d1-b985-12754fba7734 · inbound
StaRPO: Stability-Augmented Reinforcement Policy Optimization Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation abc577b2-6a79-417c-8b8f-d5ddc2dc6734 · inbound
Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4a6c27ff-3c2d-4266-a81b-35f03d40a2b4 · inbound
Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5ef49bc2-7813-4ba6-8c46-9dddd3f78362 · inbound
TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-LLM Reasoning Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bb33b206-c858-4eb3-88ed-e13ab3645cda · inbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 541498bc-8b4c-49a6-8fac-e794c3e3ff6b · inbound
ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bda8178b-d2b3-48f7-adc5-2f88c545bbff · inbound
Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3a2061e6-be90-4961-b661-0e466d79821e · inbound
Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 725f57e9-41a6-4f37-b26a-26ef05d0775f · inbound
Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.