Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:36.208057Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2506.05748.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:36.208057Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bfc05f2a-2cd1-423e-b517-028202c5982b · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Each category of algorithms presents unique benefits and constraints, rendering their integrated application beneficial in real -world scenarios
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bf31566-4f24-4634-8570-42b5186bb4d4 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance LLM-as- a-Judge
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4e8b6e8-0188-4312-a89d-4c30f929e8c7 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance score" field in [-1, 1] and a short
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c517a11-173d-4123-849e-a78414a33087 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance 𝑏𝑒𝑡𝑡𝑒𝑟":
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0394a0ad-67b0-4824-8573-f5df187a27c4 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance be funnier
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21804ec0-a5c9-4bad-a29b-1e5668030674 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance plug -and-play
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed1ee598-6ad1-4a86-b403-bab2b8c2cefc · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Which answer is better? Return ‘A’ or ‘B’ only
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2966fd8d-7156-4a75-887e-f9bfc63802e3 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07f5747b-7a14-49e1-9a5b-c3a911624fd7 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance A” or “B
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 008d4bbd-9253-4423-838d-373d19c01c03 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Online and Offline Reinforcement Learning by Planning with a Learned Model,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8bb93105-9cad-4e16-aae9-7de035cd3255 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Direct Preference Optimization: Your Language Model is Secretly a Reward Model Oral,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4230bc2a-447e-413d-b868-637db43f9f7d · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance A Survey of Reinforcement Learning from Human Feedback,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52201413-e62d-4376-99d4-f540ac16dbc2 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 650048c6-1110-4069-93d1-b38e3de8259a · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Security and Privacy Challenges of Large Language Models: A Survey,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f9e4aee-6743-4532-8a83-70b6c4ab6cb9 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a466f73-917c-48b2-9f60-e20479759f8d · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Qwen2.5 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 008311f0-f88e-4e7a-a34e-34c079fed166 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Self-rewarding language models,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4f62995-bf80-46ad-bb75-e2a1664aeb75 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3cdbe7f3-4c87-41ac-8be6-f70283a8170b · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Evaluating Text -to-Visual Generation with Image -to-Text Generation,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e689453-8921-4ec3-85a3-8d38b9adadda · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance more proficient
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50ef29fd-e9e9-4a45-bdb5-a57243e42712 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Training Language Models to Follow Instructions with Human Feedback,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3f198b2-d25f-4a79-afe5-066bbddecdb4 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Reinforcement Learning Enhanced LLMs: A Survey
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27d2466d-933e-4b88-ad92-f2b3074f42d3 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Survey on Large Language Model -Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74ae2aea-c188-4db3-8f8c-e0aa6aeca16a · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ed344d8-cc99-4521-8a1f-236530c45dda · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Human-like Summarization Evaluation with ChatGPT
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d96a1e9-2f2a-4270-9eba-9578b9740e94 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be93dd90-2e75-4460-9faa-018aaa2009e0 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0a10ede-dbbe-42c8-b22f-a4d359ab22be · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RLAIF vs. RLHF: scaling reinforcement learning from human feedback with AI feedback,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd153a62-f658-4573-8f88-c7e1ac52c102 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Large Language Models Can Self -Improve,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bac4ca2-9414-4a58-9b45-926eb8f81762 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Advancing Large Language Model Attribution through Self-Improving,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c9c9731-28a4-45d0-914a-91332507933e · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ed7257e-2d9d-42ea-b02f-56b633518d3b · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Can LLM be a Personalized Judge?,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 485880b4-b338-414d-b015-92fa35054ec6 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance ReST-MCTS*: LLM Self- Training via Process Reward Guided Tree Search,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56da1a36-06d1-4eee-a34c-2ce65173a5a5 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Self-Play Preference Optimization for Language Model Alignment,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 599d70aa-640b-4ce5-836e-8eaff13c67e2 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Training language models to follow instructions with human feedback,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3bcf55f9-03af-4a07-bc39-bd4c46f401cb · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Constitutional AI: Harmlessness from AI Feedback
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b18f9567-20f9-477a-9a02-cfd1064c7b01 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d79685-81e9-4773-abbd-48ea36bbf58a · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e97fd2f1-6d6e-44c7-9cc1-9fcf64545a28 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RewardBench: Evaluating Reward Models for Language Modeling
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be384a4-a619-4978-9f80-881a09551d7f · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Iterative Reasoning Preference Optimization,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16b69448-1360-4ae8-affb-5f829c56681e · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Available: https://neurips.cc/virtual/2024/108142
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b797fe01-8b85-4336-b0b0-c115757d42f8 · outbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Unresolved cited work
Reference 3836
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.