Pith. sign in

Paper Citation Record · LEDGER

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2605.20256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.20256 v3

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T18:38:09.609972Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact18
  • verified fuzzy14
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8158e572-a45c-4f6b-98fe-b5f5991614b9 · outbound

This paper cites Minif2f: a cross-system benchmark for formal olympiad-level mathemat- ics.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Minif2f: a cross-system benchmark for formal olympiad-level mathemat- ics

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.906125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:9e7cb99ff710f4cbc1cd79323f353950b7879b140d3d29a39a198e8543ddb3a1

Observation 46bbace4-1160-4249-9603-213313f6a2a4 · outbound

This paper cites TravelPlanner: A Benchmark for Real-World Planning with Language Agents.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning TravelPlanner: A Benchmark for Real-World Planning with Language Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:55:48.262269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:a7d18d19ca5fe5ab7b89df82a42cfee00e1c44a95f594af50f4b1c68dbfa563f

Observation dbbc1452-4ab8-4b96-abff-b80eb66459a7 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Learning to Reason under Off-Policy Guidance

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.256309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:8257dc8b4995047451674bd968c8f5537a07b067f4dc303b870864df8817159f

Observation 59e770ff-1a96-4679-a5d8-cafb82a1ea4c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.283447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:368db1e131f0354c9692c468fb686f7524cbeb9a0d4f171b6d5e4f6eebd18e26

Observation a6151876-ba61-428d-b574-e0e569814d44 · outbound

This paper cites Proximal Policy Optimization Algorithms.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.276567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:a56e71ede8073c8e6a3259c36f042f90f541cb6e9cc206cd25f8ee97cdea6cb3

Observation fba08b0a-187b-4564-95fb-8c87d673f41c · outbound

This paper cites Training language models to follow instructions with human feedback.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Training language models to follow instructions with human feedback

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.908097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:3c038a0863ac1832775188a7bc402be9a8f943119ff2cf593772bd742deefb83

Observation 2d246294-27db-41e4-8094-dcd04e7434af · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Constitutional AI: Harmlessness from AI Feedback

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.279183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:57974842fd53f7442ec70d521cc665fedb7ea978d017a21b5525bda10b281c19

Observation 2d5e0a45-f96f-456e-aaea-1d57fd339105 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.910177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:7c6b8217cff826a2611dd6f1bb769472c03483e21bec6c16f63191cfe49233f0

Observation 111c27f2-ad05-4fdc-ab6a-c28212bd77aa · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Self-refine: Iterative refinement with self-feedback

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.904279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:f710602bbecfe5358343bcf449e7645255f1cbae3c86742faec346782ae734bf

Observation 31389433-791c-4249-9512-441e860fe66f · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Reflexion: Language agents with verbal reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.900407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:be32f90dfb0fe098936e88d68d4f67e51b8e9c9548a0c4c9c1c07ddc93454e0c

Observation 75573ef9-c1df-4fa6-9239-8102996716bc · outbound

This paper cites Training language models to self-correct via reinforcement learning.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Training language models to self-correct via reinforcement learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.902508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:b177933db241ed797f57420fd9621959277e4a739ed7fffca64ca796d865cab3

Observation a88a53a4-6b40-4b3d-999b-3550d06a8dd9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.287734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:5ef6eb5b17b53b4b1bb01eebc53fe49de2db48c3cc8adfec179876f473638169

Observation 468692c0-e142-4be7-8418-8c0a3f8749da · outbound

This paper cites Let's Verify Step by Step.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Let's Verify Step by Step

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.262531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:1f1640a3a36b331e3fe63001a899a6b5721eb184536192b5539df0cc2f24c932

Observation 053351d8-a1f8-4de5-8a8b-9de39a0419ff · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Large Language Models Cannot Self-Correct Reasoning Yet

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.273946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:e71beb2deb06d18d7f9a5feb6136e1b3cd84aa1fbc2f09affbe7c79f38f008d8

Observation 8e62c7fa-3064-4aa4-ba16-018c874f5038 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.259700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:39ddf1968d83314c01d56b3f51204c99be28df9239b5990fb8098dc87db0d660

Observation b0dfe489-aae0-40a6-804c-cd4b4559b40f · outbound

This paper cites Math-shepherd: Verify and rein- force llms step-by-step without human annotations.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Math-shepherd: Verify and rein- force llms step-by-step without human annotations

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.898384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:024b455f50020000ddfd825bcdb086106a551cd55dd7adac2ef5eeba15d1ee08

Observation e90d92ab-85e5-4b7e-915a-3c958f7e8585 · outbound

This paper cites Critic: Large language models can self-correct with tool-interactive critiquing.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Critic: Large language models can self-correct with tool-interactive critiquing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.896197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:a8529d00b3d57547aa940c58f574eda3c9ecdaf91c263a8dfd8dedcb0400e2a3

Observation 3cc9602d-eacd-45e8-80d9-d632ab112720 · outbound

This paper cites Generating sequences by learning to self-correct.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Generating sequences by learning to self-correct

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.892292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:656da9cc6924da51751341dee87732fe4d4a5a66e29fffdcd75251eadb2e9f49

Observation 5c49ffdd-5c4a-433c-a872-1b3b56046084 · outbound

This paper cites Self-critiquing models for assisting human evaluators.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Self-critiquing models for assisting human evaluators

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.280312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:8d8ad1f677060af3dfe22a3ee0e4e615022e6b715ca98025a79de6c5c1f100e3

Observation 62cfb1a4-fa64-47c8-be10-747a8cd75bf1 · outbound

This paper cites Self-rewarding language models.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Self-rewarding language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.883887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:a16588d43806b2ac7a38c74e1f6c25e73434e49b05263ee3c46ec95a4232b1d0

Observation 2187babb-5a66-4b0e-9bc3-f61a0dd94ef3 · outbound

This paper cites Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.890515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:bf8732fd87bc78426e231a00e7e7fc32f925669ca855717cbc2cc7d319b4e991

Observation 05bf0cfb-0094-4d53-977e-58f0dab65e35 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.267554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:31fb9730e3cf0114f69dd429a2a4341206947304f66519955309a547c9a42c22

Observation 0b52c1f3-20d6-414c-a0f4-58cce9c5c2f5 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.277325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:59539e26acd0b5462527c3cfb71cb68db4f408302c959bc4630be928d643e303

Observation d49d42a1-cc67-4d85-ab7d-241102b51bb8 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.264942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:93f1fe5e3e0850e87ef3517ad80347cd46135bdcae3866ab6490289a89ec9bcc

Observation b90a97b5-82a6-4100-81b2-cd3ab7c3b356 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.284952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:fd2ebdb338e7b261ea8f579ae5727e233f77c00693eff7b05cc8f506fb11462f

Observation ca022e7d-b208-48d7-a81b-30f6cfedad54 · outbound

This paper cites Self-play fine-tuning converts weak language models to strong language models.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Self-play fine-tuning converts weak language models to strong language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.885989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:da948112b34e40d12c7ebfac7928ed526ecff646912867c56eccae5440ace3e8

Observation 0d7f93e4-ae92-4d01-809f-91873f02c73f · outbound

This paper cites Constraining Statistical Isotropy using 21cm Power Spectrum and Bispectrum.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Constraining Statistical Isotropy using 21cm Power Spectrum and Bispectrum

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:55:48.271182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:ddc422def94ffa4f5f124a54622c921ab5ded20930ef69279bd12a3d500b6735

Observation 736b2028-1a9e-4653-842f-5086b3d901ee · outbound

This paper cites Exploration–exploitation trade-off in reinforcement learning for large language models.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Exploration–exploitation trade-off in reinforcement learning for large language models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:55:48.290623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:f24ff0b6f1e07420b7fea04940e8cc63c3832c08284387ade46f0b24e0637860

Observation a321ba01-88cc-4b99-8406-10ed859c32b0 · outbound

This paper cites Recursive introspection: Teaching language model agents how to self-improve.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Recursive introspection: Teaching language model agents how to self-improve

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.888447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:6abd7aae14c6ca1ca52231f8d78d9750924d3d50984e531974959e976fd21d1a

Observation bb098042-05dc-4cbe-8b8c-d0dffae9c050 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.290536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:ce1543979a54ae66506788606c2a3ebbeccfd9cec1ccb1e60a610773594dc022

Observation cf5e7e3c-9472-4968-8ee3-73a547e12a96 · outbound

This paper cites Training large language models for reasoning through reverse curriculum reinforcement learning.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Training large language models for reasoning through reverse curriculum reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:44:29.894355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:984f2dd4f05046cf17b8b399005004717d99e9577372b1aec5c6463e34a360b1

Observation e8c59bb4-f092-4d46-a3cb-c031ef797bff · outbound

This paper cites Process Reinforcement through Implicit Rewards.

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.265222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:38:09.609972Z digest=sha256:9eccdc974a0dff9b5db70b244fb2b1f368461096ee635b41da0efc453f62e760

Pith citing papers

No inbound Pith citation observations are available.