Pith. sign in

Paper Citation Record · LEDGER

RRHF: Rank Responses to Align Language Models with Human Feedback without tears

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 47 inbound Pith citation observations for arXiv:2304.05302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.05302 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 47 of 47 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:48:04.370319Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T23:15:07.886084Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 72976909-ea44-46d9-88ba-833ff05566db · inbound

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment cites this paper.

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:46:56.897089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:46:56.664582Z digest=sha256:d6f959d721f7e0ae530854ba9bf0de94762c87e03b0ea9a428a1eae53b876cbc

Observation 9ec7fa5b-d481-4847-b890-79898516f46a · inbound

WizardLM: Empowering large pre-trained language models to follow complex instructions cites this paper.

WizardLM: Empowering large pre-trained language models to follow complex instructions RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:28:24.991784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T07:28:24.827546Z digest=sha256:413cff9db0c6a7e86bc4aede5f95053419314a22c458ea8fe154e6b32c93c76c

Observation a59b01a7-cea2-4b29-b9d2-895df4f38fb0 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.546046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:8f64205eb57b2c82addc3f44fcb52e39cbe2f463861534a355f5625bec7c438e

Observation e61a961f-06fa-44a7-8530-96ed0368ac11 · inbound

Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment cites this paper.

Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:30:44.734358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T22:30:44.520703Z digest=sha256:3d3fbb2096fd0ee3f88641ac358a800730edf5c7fe3208ae87d13acf53bbe4fc

Observation 64edbcdf-4674-4375-b44a-2eb0f48d4618 · inbound

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models cites this paper.

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:23:49.616672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T03:23:18.827351Z digest=sha256:2c60e10358d7610f9df1f3dd809833db1c75c517edf9bd1bd2779705a5e20584

Observation 9d66a7fd-3736-466f-abf7-3f142b5d1427 · inbound

A Survey on Knowledge Distillation of Large Language Models cites this paper.

A Survey on Knowledge Distillation of Large Language Models RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:31:11.531051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T23:31:11.213552Z digest=sha256:8220cabf1f8d4593491ff45800c9b62ace32d29df80f7a072862beaaa78ef46e

Observation 8fa9f457-f4b5-46d5-b3bf-c65a6b2866e0 · inbound

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types cites this paper.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T21:23:27.482146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:504963fd0b973b96d99d94319c047239db79472d6d183a1edeb2f0c73441d319

Observation b510fc61-533d-4dc4-b764-8299bbf4e1bc · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 175

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T20:58:25.928533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:bf6274b5cd2a5dff093962d267c4def068f2b2aedc01d4b31abd92421877b903

Observation baab4df5-0a6d-4668-9373-c572083d1de9 · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:53:38.963721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:0d75a2c4695273da85a8027232d1a4c079a801f2727ec1ff096b664104f85f35

Observation 1944b7f2-59b0-450d-b4d3-a13dc4e8c1eb · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 200

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:43.900965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:162ac7c027b24e0c4df1c287e276cfdfa1957026146a8aae3b8a65de3a4265a0

Observation eb4a2363-412a-44da-98fa-a8a3cae7000a · inbound

AnnoDPO: Protein Functional Annotation Learning with Direct Preference Optimization cites this paper.

AnnoDPO: Protein Functional Annotation Learning with Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:21.717895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:21.717895Z digest=sha256:53fe7c513c186bf52cd052330b1c4d380d3b0a83a05f2b1c6c4f089e2b75b564

Observation a6c9871d-7ba5-4d27-ba1d-24fa4ab96c53 · inbound

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models cites this paper.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:04.370319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:04.370319Z digest=sha256:a40eff3cff4dfad5503226ae681e28a4ec8c3f1b8491f9f46372c98d97bd11c5

Observation 0d3ab8c7-82c6-472a-85b1-493e848b096b · inbound

Value-Free Policy Optimization via Reward Partitioning cites this paper.

Value-Free Policy Optimization via Reward Partitioning RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.283692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.283692Z digest=sha256:f61ada2ee79a916f08948f3ef7ddda61a4f0993a36be71a62eeb71a26ef9efce

Observation a45831ae-3018-49f0-b5f9-0c5dc3b00652 · inbound

From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation cites this paper.

From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:48.654069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:48.654069Z digest=sha256:7d153d79c8cec4872cbe5fb364ff314f5f629b2e71e6ed4aef49c8c3ad6e09b9

Observation e8f5b9bb-35ec-4f6f-84c6-d5d8950363da · inbound

Activation Reward Models for Few-Shot Model Alignment cites this paper.

Activation Reward Models for Few-Shot Model Alignment RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.520388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.520388Z digest=sha256:867750087a50e82808a3481373727e95979988aae3409855a1cafacee8ae45dc

Observation de9c2c0e-ad7c-44cd-93c5-03e4f816e9f0 · inbound

Flipping Knowledge Distillation: Leveraging Small Models' Expertise to Enhance LLMs in Text Matching cites this paper.

Flipping Knowledge Distillation: Leveraging Small Models' Expertise to Enhance LLMs in Text Matching RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:29:10.475405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:29:10.475405Z digest=sha256:71aa5f65a99454987eee692ab3d4ff657870d9da79e3a69bf6a995d2375ad9e3

Observation f5e67739-421f-4012-a12f-e1cddf0cac87 · inbound

Enhancing Safe and Controllable Protein Generation via Knowledge Preference Optimization cites this paper.

Enhancing Safe and Controllable Protein Generation via Knowledge Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:26:45.506333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:26:45.506333Z digest=sha256:bb0ec424fc0c3bc7e980c09e7bde511c1682bfdc1a63a1cc5cb999e33acbfb00

Observation 26a849c7-7e62-4a1a-9237-5698fd209cb9 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.279111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.279111Z digest=sha256:259b55b66d537826cdff80a956591b127bab52622e9c9d6164d724f025aa9f65

Observation 48f014e5-7d23-4d3f-bc37-cd9703e6a585 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 260

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.226831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.226831Z digest=sha256:e3f59923af88bab7a81359ba751c83a17c586c79a595a0364fe0bd69c5afc422

Observation a8cc6708-05ee-4ece-a79a-1f80bff6b554 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 196

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:15.036632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:5a56f7ae422acf469084b3cc4e8b19ee9ea466229ec63e26fd35467b4cd3714e

Observation b5d488ce-ff01-4ef7-b4b2-c88ba28494c0 · inbound

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security cites this paper.

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:39.829280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:39.829280Z digest=sha256:c3ca38e9c2264299e18a64b6bc0e25bd3fa0e29c4c16c1e843d064f615a52c65

Observation 6735acff-92f5-4e47-a1e8-a09b988f2502 · inbound

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning cites this paper.

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:25.428063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:25.428063Z digest=sha256:b19f74104361f79e47647180620a9b0fffdb8e2c692b7c23f1d65e2516dbddee

Observation 1e311480-4017-468b-932f-7df4c5757651 · inbound

Flexible Agent Alignment with Goal Inference from Open-Ended Dialog cites this paper.

Flexible Agent Alignment with Goal Inference from Open-Ended Dialog RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:36:52.365124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T21:35:10.700387Z digest=sha256:e75895a1cfe6d870c70698b0c696357606a0edb40acda11574e91d6f92b7b847

Observation 768289db-0364-4ac0-93cd-24ed24bee1a4 · inbound

PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization cites this paper.

PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T16:30:13.419723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:30:13.419723Z digest=sha256:d25c560bf95dcb9afce904833fadfab2e3b4ae599b7c663c184f9f28e3fee943

Observation 79211895-3090-4cd1-92ae-5e0fbb5630a0 · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:02:39.687603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:a05c700c1df68167aa784a4e800bb45a09b1ff725448e6521b1ed0bf4ac32dee

Observation 875c2ac0-b05f-4cb7-9448-571525ed98bb · inbound

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards cites this paper.

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:48.370736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:15:48.370736Z digest=sha256:b0b0bc9f89295555d985cc5dc863f08cde26d5d619e39b796744068c3fa6582a

Observation 722fc942-83e3-4dc7-bde2-9b6156d27149 · inbound

POPI: Personalizing LLMs via Optimized Natural Language Preference Inference cites this paper.

POPI: Personalizing LLMs via Optimized Natural Language Preference Inference RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:42:24.361109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:41:58.231139Z digest=sha256:c168e1838eaa09ce356b349db6c1ef0feacd4f21c445f527bed40d40ea89794e

Observation 30d48663-d659-4a54-bc46-2fd92d99faf2 · inbound

OOWM: Structuring Embodied Reasoning and Planning via Object-Oriented Programmatic World Modeling cites this paper.

OOWM: Structuring Embodied Reasoning and Planning via Object-Oriented Programmatic World Modeling RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:30:16.442330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T19:29:55.075825Z digest=sha256:8ad9b87aded540263175a795d78fb069b6204bd9b9d372eb38719a7bb54fa0fc

Observation 31089a63-6ed0-4a01-8500-42cf20d59c9d · inbound

GroupDPO: Memory efficient Group-wise Direct Preference Optimization cites this paper.

GroupDPO: Memory efficient Group-wise Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:48.777256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T09:43:18.432084Z digest=sha256:747dce12b8994da5f870003040d492581c2e27d85b5f334cd41fe43deecc4deb

Observation 9cfa3750-ac93-497a-965b-151d7f1f4984 · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:56:05.413247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T16:57:49.396570Z digest=sha256:d928a34074f72bb2a86c0d914c3e32689b6861219d7abd5687ba6d25873dd99a

Observation 54258816-4d2a-4120-b1c3-b6ea9f080ad3 · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:21:26.380112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:26:54.426050Z digest=sha256:e7093fbdd528c77411b5d6a88b716d4aafecb7344f78a1a579d34975c5df9efd

Observation 7bb9b7cd-b001-41b4-9ec4-06c1ae9352fe · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:12:28.792093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:08:39.328446Z digest=sha256:ecbf92bb4f83f43d4d9690dc19cda5fb7b0d3b367ce64055bfce9ee6c8fbfbc2

Observation 985730b5-b7d3-4c31-88ed-4b0322f4ae3e · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T14:53:22.597671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:53:22.597671Z digest=sha256:afb6e93213b8b0510213f920cd8500c01852ca45e1e30ce667d6dc4a15a9a9f9

Observation 9bc0e629-2ba4-4e2f-bfeb-16dbda53b79a · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.706755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:25:52.847432Z digest=sha256:413bb9a12dbfa186f70f0e949d91f3bc60df204c882941f879b23dee59251985

Observation 04d0c94e-1d3a-44d2-8183-29975cecc6ec · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:15:07.887823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T23:14:32.494076Z digest=sha256:b5f75f292dd2cbd3fba2468a65d906efa3ad75817960ae11722dc78235905d18

Observation 96dbc87c-d15d-4ab1-a18a-ff08132f1f16 · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:45:59.770405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:8a85b74c935faccdb883bb8ac832f08993493c64da42911334ff155f31e26991

Observation da909d21-cd01-46ed-9ba2-568e47c4b775 · inbound

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs cites this paper.

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:15.540831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:10:27.595446Z digest=sha256:f027a5717710a524267f0ffbd276b96f6f587973efecce00ad58f13d496a4638

Observation 34a8d2c4-5ccb-461c-abae-f3fc672c673d · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:17.252509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:b0a03b9b47dfe6b6fb5aeca2d67dcc3c489f0cce08fd8ace1f5cfbd85ba012cf

Observation 987a504b-d28b-4e5f-b1fc-d8f7db65f81c · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:06.707485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:a2f8a61ef937b32e9cb8b5252b9405d0ea231acc9bf830c663c750008cf7a772

Observation 70559131-adfd-4148-8834-7d0046e98dde · inbound

Convex Optimization for Alignment and Preference Learning on a Single GPU cites this paper.

Convex Optimization for Alignment and Preference Learning on a Single GPU RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:06:38.375261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T05:01:31.560963Z digest=sha256:a35965f616ac20e8a6223363a3cf1286b572996150aa62f878f52e669e81f3a0

Observation 79fba758-3eae-47f0-aff4-3efa0a651e42 · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:44:40.106556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:5913392aac1ff369274245a75add57fb2402a902514da464221d29f05e52d98a

Observation 70a8db04-0fd0-4fd6-be39-2f830736006c · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 157

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:9801c8dfcf6aa5cfce5e9d26e0bb0f32d543d2f94fb2844569bf495bbae3009a

Observation cc17eec0-cf70-4c10-bc40-2a677e905421 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 158

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:50.066836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:50.066836Z digest=sha256:44834efaa0ddec5940d0f13a3c00225da8da017392341495698d3be67da83538

Observation 98eb984c-4f54-4d6e-aa35-87962f7db8cd · inbound

Every Sample Counts: Supervised Fine-Tuning of Language Models with Pointwise Constraints cites this paper.

Every Sample Counts: Supervised Fine-Tuning of Language Models with Pointwise Constraints RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-13T05:26:06.829382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:26:06.829382Z digest=sha256:6401275e288c9b19b5a6779179f4a47bdeed312c4c5c7dfde72174caa8c98a37

Observation d3585521-9000-4a22-9543-274e88dc66bb · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:27.299214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:27.299214Z digest=sha256:38c1e228b45dbcbea35b2847d19db849c8782df44aee97335b4aba7c36e89d22

Observation da752a54-d624-4528-859d-44f132c7ca38 · inbound

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text cites this paper.

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T13:36:56.283101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:36:56.283101Z digest=sha256:69f3f93794669343e19520cdff4685624f23ceb105cb6ff4ee624dc0c746204d

Observation e36ee06a-daf0-477f-a23d-7923d915f246 · inbound

Quo Vadis, World Modeling? cites this paper.

Quo Vadis, World Modeling? RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 211

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:16.365766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:16.365766Z digest=sha256:0cdaaa647bbcb3819b904066b64236f9718e6086bcec062e671cb1f927309e8f