Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T18:38:09.609972Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2605.20256.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T18:38:09.609972Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8158e572-a45c-4f6b-98fe-b5f5991614b9 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Minif2f: a cross-system benchmark for formal olympiad-level mathemat- ics
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 46bbace4-1160-4249-9603-213313f6a2a4 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning TravelPlanner: A Benchmark for Real-World Planning with Language Agents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dbbc1452-4ab8-4b96-abff-b80eb66459a7 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Learning to Reason under Off-Policy Guidance
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 59e770ff-1a96-4679-a5d8-cafb82a1ea4c · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a6151876-ba61-428d-b574-e0e569814d44 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fba08b0a-187b-4564-95fb-8c87d673f41c · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Training language models to follow instructions with human feedback
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d246294-27db-41e4-8094-dcd04e7434af · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Constitutional AI: Harmlessness from AI Feedback
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d5e0a45-f96f-456e-aaea-1d57fd339105 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 111c27f2-ad05-4fdc-ab6a-c28212bd77aa · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Self-refine: Iterative refinement with self-feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31389433-791c-4249-9512-441e860fe66f · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Reflexion: Language agents with verbal reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 75573ef9-c1df-4fa6-9239-8102996716bc · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Training language models to self-correct via reinforcement learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a88a53a4-6b40-4b3d-999b-3550d06a8dd9 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 468692c0-e142-4be7-8418-8c0a3f8749da · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Let's Verify Step by Step
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 053351d8-a1f8-4de5-8a8b-9de39a0419ff · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Large Language Models Cannot Self-Correct Reasoning Yet
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e62c7fa-3064-4aa4-ba16-018c874f5038 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b0dfe489-aae0-40a6-804c-cd4b4559b40f · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Math-shepherd: Verify and rein- force llms step-by-step without human annotations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e90d92ab-85e5-4b7e-915a-3c958f7e8585 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Critic: Large language models can self-correct with tool-interactive critiquing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3cc9602d-eacd-45e8-80d9-d632ab112720 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Generating sequences by learning to self-correct
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5c49ffdd-5c4a-433c-a872-1b3b56046084 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Self-critiquing models for assisting human evaluators
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 62cfb1a4-fa64-47c8-be10-747a8cd75bf1 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Self-rewarding language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2187babb-5a66-4b0e-9bc3-f61a0dd94ef3 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05bf0cfb-0094-4d53-977e-58f0dab65e35 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b52c1f3-20d6-414c-a0f4-58cce9c5c2f5 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d49d42a1-cc67-4d85-ab7d-241102b51bb8 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b90a97b5-82a6-4100-81b2-cd3ab7c3b356 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ca022e7d-b208-48d7-a81b-30f6cfedad54 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Self-play fine-tuning converts weak language models to strong language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d7f93e4-ae92-4d01-809f-91873f02c73f · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Constraining Statistical Isotropy using 21cm Power Spectrum and Bispectrum
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 736b2028-1a9e-4653-842f-5086b3d901ee · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Exploration–exploitation trade-off in reinforcement learning for large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a321ba01-88cc-4b99-8406-10ed859c32b0 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Recursive introspection: Teaching language model agents how to self-improve
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bb098042-05dc-4cbe-8b8c-d0dffae9c050 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cf5e7e3c-9472-4968-8ee3-73a547e12a96 · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Training large language models for reasoning through reverse curriculum reinforcement learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e8c59bb4-f092-4d46-a3cb-c031ef797bff · outbound
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Process Reinforcement through Implicit Rewards
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.