Pith. sign in

Paper Citation Record · LEDGER

Bridging Offline and Online Reinforcement Learning for LLMs

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2506.21495.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21495 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:50:46.549563Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:40:08.218379Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3ab7ef80-1dbb-46a4-821c-b519340abb04 · inbound

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework cites this paper.

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework Bridging Offline and Online Reinforcement Learning for LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:36:24.983666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T13:34:22.790199Z digest=sha256:bc611bf65211ff7b51aa0bb2ba94ad4332df8848d982cafc612e142512787e4b

Observation d3c3b757-9212-422a-b8f0-97949b079396 · inbound

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation cites this paper.

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation Bridging Offline and Online Reinforcement Learning for LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T09:50:46.549563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:50:46.549563Z digest=sha256:056a26d1ec68723a50d21df2c8555307435fa48428b9950ba0d9719bd020a8fe

Observation 202772d5-4b27-4529-a7b7-41fb2b972c45 · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Bridging Offline and Online Reinforcement Learning for LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:21.679930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:21.679930Z digest=sha256:95c8147b7a7da76c5f76290a1b62302507beb763fed6d086b1ed761ac88c3f09

Observation 089ad325-e523-45ab-a2c4-a1e5469647d6 · inbound

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning cites this paper.

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Bridging Offline and Online Reinforcement Learning for LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:15.425803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T04:29:21.897215Z digest=sha256:fa0c52d69304a2d86eff2e377f7f2b16b8c6ef761e551afc9b940c432bd949cc

Observation 7465628b-7cca-4c32-8f66-7c8684b401f7 · inbound

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO cites this paper.

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Bridging Offline and Online Reinforcement Learning for LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.010858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T23:54:32.621093Z digest=sha256:042d967bcbb5b21c34266fc8bb4a5f68f5b6c18287d0608c005c29db2caf1824

Observation b7230ccc-d595-499e-98fc-4d0eb1787025 · inbound

Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments cites this paper.

Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments Bridging Offline and Online Reinforcement Learning for LLMs

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:27.027263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:29:46.554292Z digest=sha256:b4c08a9325a3f5d4aa20f99a666992b5a82677892b2dee8d52e974f2e333b9af

Observation 507f7013-5128-4117-8a77-4d64895355cc · inbound

Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments cites this paper.

Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments Bridging Offline and Online Reinforcement Learning for LLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T12:30:46.217716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:30:46.217716Z digest=sha256:749356885b07592f3be7cecaebde9f4208cb3473e92595b54389d7c96ee1a5ee

Observation cc445246-eaec-4226-b5fc-f4fefe4861de · inbound

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces cites this paper.

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces Bridging Offline and Online Reinforcement Learning for LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.927436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T05:46:26.938277Z digest=sha256:923790f98ec9b600db9b45bd2872d1412c3ab2edecfeacd8e0e8a0fdb5335d29

Observation 7ee3a065-8118-4b92-9708-1f8e7d080591 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Bridging Offline and Online Reinforcement Learning for LLMs

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.509771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:3d2c4e44f0567eeb7b75205ebe88384778dc15a1c56d5e6372a62d1839c6302f

Observation 7e0c52fc-5ece-4b58-9779-38d28c66da07 · inbound

Autodata: An agentic data scientist to create high quality synthetic data cites this paper.

Autodata: An agentic data scientist to create high quality synthetic data Bridging Offline and Online Reinforcement Learning for LLMs

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:40:08.219887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-25T19:50:35.574454Z digest=sha256:6755ca561ef3d3d58224a8df126deb8e8e1080982d4b42b3c2a7bcb8a1588f87

Observation f177b23e-6f8c-4e22-b30a-ef7b1335dae5 · inbound

Autodata: An agentic data scientist to create high quality synthetic data cites this paper.

Autodata: An agentic data scientist to create high quality synthetic data Bridging Offline and Online Reinforcement Learning for LLMs

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:51.287491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T05:16:12.361470Z digest=sha256:cb597104b9c1a551476ef0a6cc343ed67ee43893c409c152309dc5b2a957dc6a

Observation 9cbe79a4-8756-444d-a1b3-b98c5f4f1484 · inbound

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training cites this paper.

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training Bridging Offline and Online Reinforcement Learning for LLMs

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:15:47.555967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T04:21:38.825926Z digest=sha256:1af2504d24f34101a0b3182c0d92952bb6cfc26c5832c64c45cdac5ffceb7252