Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2506.21495.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:50:46.549563Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:40:08.218379Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 3ab7ef80-1dbb-46a4-821c-b519340abb04 · inbound
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework Bridging Offline and Online Reinforcement Learning for LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d3c3b757-9212-422a-b8f0-97949b079396 · inbound
Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation Bridging Offline and Online Reinforcement Learning for LLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202772d5-4b27-4529-a7b7-41fb2b972c45 · inbound
Safety Alignment of LMs via Non-cooperative Games Bridging Offline and Online Reinforcement Learning for LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089ad325-e523-45ab-a2c4-a1e5469647d6 · inbound
OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Bridging Offline and Online Reinforcement Learning for LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7465628b-7cca-4c32-8f66-7c8684b401f7 · inbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Bridging Offline and Online Reinforcement Learning for LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b7230ccc-d595-499e-98fc-4d0eb1787025 · inbound
Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments Bridging Offline and Online Reinforcement Learning for LLMs
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 507f7013-5128-4117-8a77-4d64895355cc · inbound
Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments Bridging Offline and Online Reinforcement Learning for LLMs
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc445246-eaec-4226-b5fc-f4fefe4861de · inbound
Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces Bridging Offline and Online Reinforcement Learning for LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7ee3a065-8118-4b92-9708-1f8e7d080591 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Bridging Offline and Online Reinforcement Learning for LLMs
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7e0c52fc-5ece-4b58-9779-38d28c66da07 · inbound
Autodata: An agentic data scientist to create high quality synthetic data Bridging Offline and Online Reinforcement Learning for LLMs
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f177b23e-6f8c-4e22-b30a-ef7b1335dae5 · inbound
Autodata: An agentic data scientist to create high quality synthetic data Bridging Offline and Online Reinforcement Learning for LLMs
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9cbe79a4-8756-444d-a1b3-b98c5f4f1484 · inbound
LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training Bridging Offline and Online Reinforcement Learning for LLMs
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.