Pith. sign in

Paper Citation Record · LEDGER

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance

As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2506.05748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05748 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:36.208057Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfc05f2a-2cd1-423e-b517-028202c5982b · outbound

This paper cites Each category of algorithms presents unique benefits and constraints, rendering their integrated application beneficial in real -world scenarios.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Each category of algorithms presents unique benefits and constraints, rendering their integrated application beneficial in real -world scenarios

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:41.116682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:32.904699Z digest=sha256:554e34a946a937a96969ca832fe843170838c0d805eea47c9661f4eff73a202a

Observation 4bf31566-4f24-4634-8570-42b5186bb4d4 · outbound

This paper cites LLM-as- a-Judge.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance LLM-as- a-Judge

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:41.070207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:32.952718Z digest=sha256:b12b449c0563ffbfbc458f1a03a0ddb34f83404a8c699609ccbe1aa3f2044cfe

Observation c4e8b6e8-0188-4312-a89d-4c30f929e8c7 · outbound

This paper cites score" field in [-1, 1] and a short.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance score" field in [-1, 1] and a short

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.876633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:33.128347Z digest=sha256:08ca8c5c876ad53597acee76234c6aeee4bf698f893f7f0da73332338e010dbf

Observation 9c517a11-173d-4123-849e-a78414a33087 · outbound

This paper cites 𝑏𝑒𝑡𝑡𝑒𝑟":.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance 𝑏𝑒𝑡𝑡𝑒𝑟":

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.636409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:33.304988Z digest=sha256:9fa5bdd84055a6c12182c662403a28845a6f3cf7d3ab6a6948012f112f14acac

Observation 0394a0ad-67b0-4824-8573-f5df187a27c4 · outbound

This paper cites be funnier.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance be funnier

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.765243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:33.229860Z digest=sha256:8351f49e0a2f4e8134c2d5fa7e2d00220ee34e9e7d8c1036c09cc43eb3982edc

Observation 21804ec0-a5c9-4bad-a29b-1e5668030674 · outbound

This paper cites plug -and-play.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance plug -and-play

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.366764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:33.434466Z digest=sha256:a60de91976350fd60ca275aec6925134e9c3f1f8da1961ca3f1bb7bb3d3e8714

Observation ed1ee598-6ad1-4a86-b403-bab2b8c2cefc · outbound

This paper cites Which answer is better? Return ‘A’ or ‘B’ only.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Which answer is better? Return ‘A’ or ‘B’ only

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.499258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:33.363323Z digest=sha256:9c1219cd713e290ce0a1bfc49c26adb99ae936859885c8156a09f41973b1651a

Observation 2966fd8d-7156-4a75-887e-f9bfc63802e3 · outbound

This paper cites an unresolved cited work.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:39.929275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:33.586566Z digest=sha256:b73371c1d9e8dab716898edae494da29372dae92a6711843ddd15e832e228fca

Observation 07f5747b-7a14-49e1-9a5b-c3a911624fd7 · outbound

This paper cites A” or “B.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance A” or “B

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.160618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:33.507669Z digest=sha256:d0d0742f4d8c59434474fc5521d2f7c60dc906b0500f9ce8848597c209b7f833

Observation 008d4bbd-9253-4423-838d-373d19c01c03 · outbound

This paper cites Online and Offline Reinforcement Learning by Planning with a Learned Model,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Online and Offline Reinforcement Learning by Planning with a Learned Model,

Reference 10

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T10:19:37.491599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:34.462170Z digest=sha256:60e7f907f52eda6c3fb0ba33155a49c381f25e0947742b4a48ab4cacbfc44630

Observation 8bb93105-9cad-4e16-aae9-7de035cd3255 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model Oral,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Direct Preference Optimization: Your Language Model is Secretly a Reward Model Oral,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:39.631521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:33.659848Z digest=sha256:9928b0cfe6f0f407f60355b35bdbf1bb6337edadf429d1d61a80936c8f9432cd

Observation 4230bc2a-447e-413d-b868-637db43f9f7d · outbound

This paper cites A Survey of Reinforcement Learning from Human Feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance A Survey of Reinforcement Learning from Human Feedback,

Reference 12

Resolution
verified exact
doi, observed 2026-08-07T10:19:36.981902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:33.727935Z digest=sha256:4b51d607608293fce0e049a4ab520f136c1455e87e430c4fd98e7eab212f699e

Observation 52201413-e62d-4376-99d4-f540ac16dbc2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:33.848792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:33.848792Z digest=sha256:9d20232a0d71db21a7d602b2f5d68fc5d25abcc37a70511693dcf9f864fb9f58

Observation 650048c6-1110-4069-93d1-b38e3de8259a · outbound

This paper cites Security and Privacy Challenges of Large Language Models: A Survey,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Security and Privacy Challenges of Large Language Models: A Survey,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:33.926564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:33.926564Z digest=sha256:9f03f9618dab7ec79112219d260897d4837f600241991944242d46e25422e958

Observation 9f9e4aee-6743-4532-8a83-70b6c4ab6cb9 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.046186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.046186Z digest=sha256:6cb401f14addd999b395b8651c8d577da7b09d4e4447ea327928a4483e551e41

Observation 9a466f73-917c-48b2-9f60-e20479759f8d · outbound

This paper cites Qwen2.5 Technical Report.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Qwen2.5 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.135241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.135241Z digest=sha256:03cbb35d8b25b02e373e2b4705af8e8bb8c232632a6f0e978f8e239a52a84101

Observation 008311f0-f88e-4e7a-a34e-34c079fed166 · outbound

This paper cites Self-rewarding language models,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Self-rewarding language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:39.400464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:34.252515Z digest=sha256:b8b15c074ad5eddcec9116d84093bf6141b4696f8040318cc31ae6588dc991d8

Observation d4f62995-bf80-46ad-bb75-e2a1664aeb75 · outbound

This paper cites RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:39.116623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:34.294065Z digest=sha256:8adb27202ac4f0f3e886563b130827bedb93ab6789977ced5ed4aa50b16045db

Observation 3cdbe7f3-4c87-41ac-8be6-f70283a8170b · outbound

This paper cites Evaluating Text -to-Visual Generation with Image -to-Text Generation,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Evaluating Text -to-Visual Generation with Image -to-Text Generation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.415108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.415108Z digest=sha256:775fc366239ce9af6c16b9f3b92e08429c47860385479a41006e812d4d259a61

Observation 2e689453-8921-4ec3-85a3-8d38b9adadda · outbound

This paper cites more proficient.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance more proficient

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:41.034025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:33.043123Z digest=sha256:07cf5d2338a80ca74b1682423f719663a7aa007171ca4beacd511f0a6dbc189f

Observation 50ef29fd-e9e9-4a45-bdb5-a57243e42712 · outbound

This paper cites Training Language Models to Follow Instructions with Human Feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Training Language Models to Follow Instructions with Human Feedback,

Reference 21

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T10:19:37.289842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:34.523346Z digest=sha256:018db6d061ea6c9b2e7ec46fc147d950dfa9177236790c103c48ba731ab3281f

Observation b3f198b2-d25f-4a79-afe5-066bbddecdb4 · outbound

This paper cites Reinforcement Learning Enhanced LLMs: A Survey.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Reinforcement Learning Enhanced LLMs: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.611764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.611764Z digest=sha256:25919de7f132d49b5e2107a9acec98a015d7be81527315aa904a091280f2c3a3

Observation 27d2466d-933e-4b88-ad92-f2b3074f42d3 · outbound

This paper cites Survey on Large Language Model -Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Survey on Large Language Model -Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.727234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.727234Z digest=sha256:d66d1fea38bf160dbc4359aea3fef962f5bd77bbee8ca4006d0460561deeb1f7

Observation 74ae2aea-c188-4db3-8f8c-e0aa6aeca16a · outbound

This paper cites On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.793339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.793339Z digest=sha256:86e660606510b7ce20fbb57e59bfba833d07ebd15d488c67467f4c1bf4621e00

Observation 7ed344d8-cc99-4521-8a1f-236530c45dda · outbound

This paper cites Human-like Summarization Evaluation with ChatGPT.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Human-like Summarization Evaluation with ChatGPT

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.827221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.827221Z digest=sha256:22e8a181dccfc25bc5de04cc7c9063629c2c8786c09e19bcda79e04027155456

Observation 2d96a1e9-2f2a-4270-9eba-9578b9740e94 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.917509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.917509Z digest=sha256:15ea0bc017cc5fc2281fb9838b2ff5d8323269f3bb686869e26e99a6597328eb

Observation be93dd90-2e75-4460-9faa-018aaa2009e0 · outbound

This paper cites LLM-as-a-Judge & Reward Model: What They Can and Cannot Do.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance LLM-as-a-Judge & Reward Model: What They Can and Cannot Do

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.977810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.977810Z digest=sha256:b10c4817c567740f5786524705dfc2de064e56188598478a3ba22008bebdae27

Observation b0a10ede-dbbe-42c8-b22f-a4d359ab22be · outbound

This paper cites RLAIF vs. RLHF: scaling reinforcement learning from human feedback with AI feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RLAIF vs. RLHF: scaling reinforcement learning from human feedback with AI feedback,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.869784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:35.034720Z digest=sha256:5092195868914084c2b33b8463256882ae44e3aa7830fe8c435f806dc32ecede

Observation fd153a62-f658-4573-8f88-c7e1ac52c102 · outbound

This paper cites Large Language Models Can Self -Improve,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Large Language Models Can Self -Improve,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.096249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.096249Z digest=sha256:cdf40d9c96514674aed45c74c56b3667b54d42ab7a2a3978073eacc2054c2931

Observation 8bac4ca2-9414-4a58-9b45-926eb8f81762 · outbound

This paper cites Advancing Large Language Model Attribution through Self-Improving,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Advancing Large Language Model Attribution through Self-Improving,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.691907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:35.180761Z digest=sha256:e03444086a4652fe801615b77b7da3c4ac4d4edec53db0aa747f3876d934f421

Observation 7c9c9731-28a4-45d0-914a-91332507933e · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.386305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.386305Z digest=sha256:a36252436bcf307c04ed446585bc937d0a791156f7e7d4f726b8e243b4842515

Observation 3ed7257e-2d9d-42ea-b02f-56b633518d3b · outbound

This paper cites Can LLM be a Personalized Judge?,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Can LLM be a Personalized Judge?,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.466565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.466565Z digest=sha256:96756bd2c5dc811941ac8b2984fa3e4e994897e2b51f2fa0776dc372871b044b

Observation 485880b4-b338-414d-b015-92fa35054ec6 · outbound

This paper cites ReST-MCTS*: LLM Self- Training via Process Reward Guided Tree Search,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance ReST-MCTS*: LLM Self- Training via Process Reward Guided Tree Search,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.522228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:35.548712Z digest=sha256:c152b9828a1f86bf4387853e0b9bd3dce43399e0138abead930c25565f0395a6

Observation 56da1a36-06d1-4eee-a34c-2ce65173a5a5 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Self-Play Preference Optimization for Language Model Alignment,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.280897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:35.634935Z digest=sha256:5c782d34bae7dbcd8ca8ba95a776f627d75811c40c474cdd3f65028eb90076fa

Observation 599d70aa-640b-4ce5-836e-8eaff13c67e2 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Training language models to follow instructions with human feedback,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:37.785037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:35.826010Z digest=sha256:2a42673327beef44100f2b6e4d2497598d33fbafcc36aab18602e0c42a61cc60

Observation 3bcf55f9-03af-4a07-bc39-bd4c46f401cb · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Constitutional AI: Harmlessness from AI Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.928792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.928792Z digest=sha256:034125c5c0fb279c159850bd6d20feebd76138999fe88b8139a91665cbfb79eb

Observation b18f9567-20f9-477a-9a02-cfd1064c7b01 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.995780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.995780Z digest=sha256:ef0310e508555591860ff5b78f1b0cf0d53d5d991d7073ebf43f2e587fef619b

Observation 99d79685-81e9-4773-abbd-48ea36bbf58a · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:36.068399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:36.068399Z digest=sha256:c0f563c11d23d5be9f32b6ad6dbb4196f50d9acb9eb0d90854de39bcd025a9e9

Observation e97fd2f1-6d6e-44c7-9cc1-9fcf64545a28 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RewardBench: Evaluating Reward Models for Language Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:36.134562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:36.134562Z digest=sha256:392f32a75e1e0ba7c5659364bf6b2beaa6b13d197f6604a65f69fb0444cb7b56

Observation 3be384a4-a619-4978-9f80-881a09551d7f · outbound

This paper cites Iterative Reasoning Preference Optimization,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Iterative Reasoning Preference Optimization,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:37.604055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:36.208057Z digest=sha256:9c97df8068fd0580981e8edee0447a8abb97e16f5d73794ae01abb4d6fcd3028

Observation 16b69448-1360-4ae8-affb-5f829c56681e · outbound

This paper cites Available: https://neurips.cc/virtual/2024/108142.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Available: https://neurips.cc/virtual/2024/108142

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.000462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:19:35.719817Z digest=sha256:96032d2940ce292c8543c167b30f9d0fe4d1bb5c3cd432d1a748efba03487ad1

Observation b797fe01-8b85-4336-b0b0-c115757d42f8 · outbound

This paper cites an unresolved cited work.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Unresolved cited work

Reference 3836

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.287341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.287341Z digest=sha256:5600cb5ac7338fa3b7cb64b14992d23e1f1cee1335557228d588ba17dc4f5cef

Pith citing papers

No inbound Pith citation observations are available.