Pith. sign in

Paper Citation Record · LEDGER

Direct Language Model Alignment from Online AI Feedback

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 56 inbound Pith citation observations for arXiv:2402.04792.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04792 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 56 of 56 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:45:12.246781Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:09:46.394396Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9fc03707-b6ce-471f-84eb-4d75de5d9fed · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Direct Language Model Alignment from Online AI Feedback

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.577360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:2c151d67d5aba11ad66f711e59c2e7f7ab06102290aa278cc7b3337733db554f

Observation 525eec75-07ad-4631-9a9f-da21d0e2f561 · inbound

Aligning LLMs with Domain Invariant Reward Models cites this paper.

Aligning LLMs with Domain Invariant Reward Models Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.246781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.246781Z digest=sha256:243a83ccaae9f1b755fd580ea9b735eac96ffa8c8246aff7bee6612edd69eb03

Observation bb455c0a-59fe-442c-8611-7e6a6b9419db · inbound

Online Preference Alignment for Language Models via Count-based Exploration cites this paper.

Online Preference Alignment for Language Models via Count-based Exploration Direct Language Model Alignment from Online AI Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.426898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.426898Z digest=sha256:4a29f88dd9d8dc289a927a464da70c184912987c8dd6a3b786a6bad685141e40

Observation 202baf4c-ee47-40ec-a812-86f42ce925fb · inbound

Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment cites this paper.

Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Direct Language Model Alignment from Online AI Feedback

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-09T19:56:28.723018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:56:28.723018Z digest=sha256:d578f64f0de853ad2d93ce4ec0b81c5484fed3979c52bebbe75d4717680bf978

Observation c124081d-487d-4d44-9c0f-09880e7e44e4 · inbound

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective cites this paper.

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective Direct Language Model Alignment from Online AI Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T04:08:51.439639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:08:51.439639Z digest=sha256:45349a5c09935a30ba6d539d60f645dc3023b4babfe52321734a8d83324ec9df

Observation 45ccfa6a-012b-4955-a9d8-ff9a8851a250 · inbound

PILAF: Optimal Human Preference Sampling for Reward Modeling cites this paper.

PILAF: Optimal Human Preference Sampling for Reward Modeling Direct Language Model Alignment from Online AI Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.079636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.079636Z digest=sha256:621b936713be2a6e57e34341441814b0371ace9058a5b1f4ac24290f07ae81f7

Observation 35a8d537-a80e-4db4-a8d6-ec6e39ef4f5e · inbound

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator cites this paper.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.442396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.442396Z digest=sha256:4fcd84d9bac8c0c26281cb9227e414c902f19bd2d9c4b077e69c0f91d52eb8a3

Observation 238780b0-f35b-4771-8497-9b696b970edf · inbound

Process Reward Models for LLM Agents: Practical Framework and Directions cites this paper.

Process Reward Models for LLM Agents: Practical Framework and Directions Direct Language Model Alignment from Online AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.864500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.864500Z digest=sha256:e025667b94f4c7d5a59fcd901ea68524ee2a283f68e7c52973172742ada9ea19

Observation 5724ed26-706e-4021-a240-b524f4676b17 · inbound

Preference learning made easy: Everything should be understood through win rate cites this paper.

Preference learning made easy: Everything should be understood through win rate Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.022389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.022389Z digest=sha256:c368d26aacbecf02aba0325937c70e28f6a354464b6048c9949be8bb6d7eb34f

Observation 01d71fb7-6b99-405a-9bc9-81f4537e1e12 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Direct Language Model Alignment from Online AI Feedback

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.394165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:3620b6e6497945dc8b3e6cc4af7e09dc7b502d6ae9ae506ac7e224d70d85935a

Observation 9cb827c2-9926-4f6e-8717-aabe5f2532a0 · inbound

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models cites this paper.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.852749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.852749Z digest=sha256:eccb8ae22e3cc05798984656d23c73cb103f3b62fe8df816a109a711f3f9cc52

Observation 37779663-14c7-4844-8c03-42accb2ae545 · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization Direct Language Model Alignment from Online AI Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:26.468933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:26.468933Z digest=sha256:d7a41aaed730ede7f69d8881670842f63e9e0c68b1eb01c52be081d51d7b1be4

Observation 9ec77987-1d39-4229-8a6e-f57c7ae1920e · inbound

Self-Training Large Language Models with Confident Reasoning cites this paper.

Self-Training Large Language Models with Confident Reasoning Direct Language Model Alignment from Online AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:13.280315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:51:13.280315Z digest=sha256:cfa337513aa07ad23f2fcffac2774431a9346ce219b78109c0aa776682d20667

Observation 49e448cb-6a11-4e17-ae8d-5ecd6eb7f0f3 · inbound

Online Knowledge Distillation with Reward Guidance cites this paper.

Online Knowledge Distillation with Reward Guidance Direct Language Model Alignment from Online AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.270567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.270567Z digest=sha256:15cd9bffaff7bd058046bbe9b313a2d9fd2bc9febfc7a0a2f7b8c79faa7724fb

Observation 70ae6733-b1b2-4dfb-88e2-ee7e74a226c5 · inbound

MOSLIM:Align with diverse preferences in prompts through reward classification cites this paper.

MOSLIM:Align with diverse preferences in prompts through reward classification Direct Language Model Alignment from Online AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:04.853691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:04.853691Z digest=sha256:0796c963a2a537c24deefa3e74895aa6c0e9c845d29c4545811929c67708a72c

Observation 7f1ca33b-face-4559-ab79-aec833b99512 · inbound

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment cites this paper.

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment Direct Language Model Alignment from Online AI Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:41.934550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:41.934550Z digest=sha256:0f9819fd356bdb5eeaedc8f029c0a07eb85704c980bd69a55fdbf6bb782b7aab

Observation 23f9ea8d-a73f-4b8a-a799-0a9968480331 · inbound

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function cites this paper.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Language Model Alignment from Online AI Feedback

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:15.560054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:15.560054Z digest=sha256:c865e28b8ded9a9c8f66180a7a9b19225a65e2fa63ef72a6e499e7ca32f4b2e0

Observation 258e424f-f7ff-40c9-a1f4-26cfa3be0f3d · inbound

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling cites this paper.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.185377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.185377Z digest=sha256:76d96a366b6ec2347a49872eeb98f36b798f08f35a947a2fc8543ba17cf61f0e

Observation 4d66ed61-f61e-490b-ae51-adb4cca2b999 · inbound

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling cites this paper.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.189540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.189540Z digest=sha256:f9ee9ed7d8c9a2d7d7555649513794a163af8ce7e8122d79029152a74297f2f8

Observation 9e272212-0a67-482e-b5a1-9831005a09d1 · inbound

Customizing Speech Recognition Model with Large Language Model Feedback cites this paper.

Customizing Speech Recognition Model with Large Language Model Feedback Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:19.820010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:23:19.820010Z digest=sha256:0efd33415505e755d5e65bbb4456524a4cf2a92d64c847d0020758e8f742dd0b

Observation 84f5acc2-ba3b-43e3-946b-d31459bd35a3 · inbound

Bridging Offline and Online Reinforcement Learning for LLMs cites this paper.

Bridging Offline and Online Reinforcement Learning for LLMs Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.981374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.981374Z digest=sha256:10cc2ba6757e8e6c185878fa46990792486826d0784f08e181dc2dbfaac7cb72

Observation a140d1b5-77dc-4353-9028-07de6ffee18f · inbound

Data Diversification Methods In Alignment Enhance Math Performance In LLMs cites this paper.

Data Diversification Methods In Alignment Enhance Math Performance In LLMs Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:26.226540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:43:26.226540Z digest=sha256:36f432dcf5edb47087cd8c2f10f2c87b0e3cb122050d619d1791ad48a03ebbcf

Observation 1f54acb1-611e-49c7-ba95-7ebfe0a7f823 · inbound

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought cites this paper.

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought Direct Language Model Alignment from Online AI Feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:37.182865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:37.182865Z digest=sha256:8070fdaf13aa048f95a2170f126524793d7f5962aaf919cab13a81a9783d7cd8

Observation d796d31d-9f3b-4109-8e94-f04c9c540e35 · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling Direct Language Model Alignment from Online AI Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.788905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:7a3b9d4a59889394ad5c496c85d1c1aab9f11d7d175d449d7abcf29cf665ca35

Observation 2e45fe2a-0649-4a2e-88a4-cf1875fbf5d2 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Direct Language Model Alignment from Online AI Feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.021993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.021993Z digest=sha256:d84b522bfdbcb1a5b458adff526f4aa3b70fb14eab818f16368740e5d9473309

Observation a9e3b838-9524-4038-8924-5fd067ad16f7 · inbound

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion cites this paper.

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T04:36:13.417768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:36:13.417768Z digest=sha256:c3e28a3d972e2ef5af62629bfdf51ef351b2b03b6a7911d1871cdb0763a4b8f1

Observation 7012f92b-6880-4eea-b3e9-6e20e0838a5f · inbound

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization cites this paper.

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Direct Language Model Alignment from Online AI Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T09:48:05.297616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:48:05.297616Z digest=sha256:8737a67f9622efec137439dd168354cc70b4f872c95feb4ceb8025904ef12790

Observation d1431a99-d1bd-46e6-a6c6-9c58c7d34fcd · inbound

IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation cites this paper.

IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation Direct Language Model Alignment from Online AI Feedback

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:30:35.839277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T01:27:00.824546Z digest=sha256:6073bfb0a8798c3957d54cf0921e90336ebe9e4b87ceec184b7e77c9140c6725

Observation a917cb65-00e6-4af7-8631-da3807faf630 · inbound

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs cites this paper.

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs Direct Language Model Alignment from Online AI Feedback

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:51:17.792019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T21:49:21.785096Z digest=sha256:89bef859ca9651a733e12bcae0a0b7dce87c8f02db9fcbabd2c2ffd2eba2efcd

Observation 50f7b995-cb23-4176-9e28-933b1a7e0e1e · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Direct Language Model Alignment from Online AI Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:21.154235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:21.154235Z digest=sha256:b151c6e5f1821e9313646c6b35e89ba42f274e1f3a0b3b3c15b89c33facbcea0

Observation 38f530cf-96ab-414d-bebf-65dc1657cdfd · inbound

Ministral 3 cites this paper.

Ministral 3 Direct Language Model Alignment from Online AI Feedback

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:12:24.704031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:b44d11ee9e1fbc316ba8fff2b333f55740455ab5f2d70a9e39efc8ef65319635

Observation 2215a8c4-a41f-4730-822e-34c9055982d1 · inbound

Noisy Pairwise-Comparison Random Search for Smooth Nonconvex Optimization cites this paper.

Noisy Pairwise-Comparison Random Search for Smooth Nonconvex Optimization Direct Language Model Alignment from Online AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T07:11:58.083080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:11:58.083080Z digest=sha256:136c673ce77b1bd8c991b354de55b3d3ac0e7b98ccd1299b88d7a81065ca7f3f

Observation 9c503c54-32c4-4024-a4e6-2084b118b98c · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:37:28.599444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:a398819f7a82185db5063bd930934f19110bfeab6fbf9bf5e7aa699d5b5fa208

Observation 33020fae-9d6d-41e6-aa54-059289939857 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:10:10.410196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:fec9fdb56e13052536ab70c884d7fc6fce5902275f0a469fffd580fbdb50b618

Observation c194993b-207c-4c12-8995-6db66a858ee6 · inbound

Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs cites this paper.

Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs Direct Language Model Alignment from Online AI Feedback

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:53:11.873009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T19:48:46.245576Z digest=sha256:0025d9f2ef0bcbf316b149dbbc5b1a1bc578f5b3c42a69c46d606560b28acaba

Observation 07d33688-9cad-4a73-bed4-c96ca2a55bfb · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Direct Language Model Alignment from Online AI Feedback

Reference 164

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.597043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:117610e789296c23c0a85eabbbb6ee036a4be6f279866cc4c6ab40135cd5ddb7

Observation 99e214aa-2fbe-46cd-a9d3-90077f9be17e · inbound

CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization cites this paper.

CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:53:04.770013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T08:48:12.941821Z digest=sha256:dd06ee28d91cd8bb74eeec50d6aba7d4a4333e0d9afeac3b3a9d4db848881187

Observation 5777ec4c-c515-4c3b-8017-c2c5d5c9c181 · inbound

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models cites this paper.

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models Direct Language Model Alignment from Online AI Feedback

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:26.012326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:06:53.959061Z digest=sha256:57abd1985dbb8b532a00a9478544e09142e3db49901165240fe235182925d614

Observation 40b92d08-c052-432c-b463-ba6f952ed9bb · inbound

Mind the Gap: Structure-Aware Consistency in Preference Learning cites this paper.

Mind the Gap: Structure-Aware Consistency in Preference Learning Direct Language Model Alignment from Online AI Feedback

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.582713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-07T06:41:18.311986Z digest=sha256:87ed3b8a4399d1175590689110c715b344a7a0a04ff193126cf1a7e241ee6364

Observation 6efc0a37-4cae-4f1e-8eb2-e890ecd24a08 · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.769327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:6301c7ab5059ab94c3459a75ef64ac69f3f0cd9b55d62eb342a77d9d7be6db95

Observation d527bc9d-9315-4726-ab55-5629f13f5ac3 · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Direct Language Model Alignment from Online AI Feedback

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:09.228349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T17:01:04.571087Z digest=sha256:e1245bc5dd9f11b16ab924443a68e9c79a8566ae73ae23f8df60b1db0ce48844

Observation 1a5303a5-7923-4d03-9a3c-e3cf9e802322 · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Direct Language Model Alignment from Online AI Feedback

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:35:10.418679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:3bdace6dacacb78d3743303a55a2eeac6c32739f1f26304e0bcb8115b675a296

Observation 485bf674-827a-4d27-9811-e11f5725fa08 · inbound

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control cites this paper.

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:14.624333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:50:42.836549Z digest=sha256:6a462f69366cee61cb1929a6fafd1429d13aacbc5e0d340df53612bd1e367baf

Observation 0cfec670-b681-4450-bb79-82a96b6e339a · inbound

dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models cites this paper.

dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models Direct Language Model Alignment from Online AI Feedback

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:51:29.686453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T03:52:05.779559Z digest=sha256:8a6c1a73294e81054a5a299e5fb0ea6466b2b06911fb342f6be99550501236d8

Observation bcc9ea60-ea4a-4e1d-8449-e318756e9fb0 · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Direct Language Model Alignment from Online AI Feedback

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:42.998151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:28e6fd1370a0174c787b88eaf90b6879372fbebb359589700b24df29a1c3a5b4

Observation dec601cd-24b2-4258-86b4-a06000049b14 · inbound

VSPO: Vector-Steered Policy Optimization for Behavioral Control cites this paper.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:44.127323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:3dedb61a0c2a41a953314168e348c88f67f06c5c1c5ba40804ced38d2d89e99b

Observation a3a6c154-8285-4758-8d63-6191ecf29af2 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Direct Language Model Alignment from Online AI Feedback

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.516908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:ca0e085d2c1e6d5645f4595d2b4880a4ccc690149cc0764fada3daf97d8fecb3

Observation d6bf4218-d171-4b0f-8898-8565ca3ab6ad · inbound

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations cites this paper.

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.069649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T21:30:39.993873Z digest=sha256:7a6c7481b1d17ad8c84622658e7d542a618420329a6f63760176ee118c635ec0

Observation 303e8719-2e68-4ce7-a158-36d186402e70 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Direct Language Model Alignment from Online AI Feedback

Reference 203

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.560428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:f80479f9074953f460ff5547ecd590e172eeef7a0aa660177598ee4b09e44a18

Observation 3e34989b-6faf-4738-b3ee-e8305e6abe1c · inbound

BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization cites this paper.

BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization Direct Language Model Alignment from Online AI Feedback

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:36:48.330733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T05:55:00.629666Z digest=sha256:590a049dfe7dde1eef5106e6d776880071c7b984494e388341408ae5c7fcc2fd

Observation 3fbc77f9-7207-494e-bd86-7239b59ed9bd · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Direct Language Model Alignment from Online AI Feedback

Reference 219

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.508306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:a5cc59fb8816f3c247c20c76abc9f2297c557e8a553bdc27e1cca5d09c666047

Observation 5cb7cf2e-02e5-44ff-884a-07d8e2198004 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Direct Language Model Alignment from Online AI Feedback

Reference 209

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.396105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:a74ee1781984018cfd0ac10fdeec9cb5472bff48c155f95c4c4a7737cf385df7

Observation c6fb746f-6fae-42ad-84f8-f9b65d2ac202 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Direct Language Model Alignment from Online AI Feedback

Reference 195

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.465337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.465337Z digest=sha256:210e1022c8b7f5e24d165e8eecb3f51ac7d21a5bc8e56cb9b7570c942911c02e

Observation bd29c9fa-d3eb-4312-9572-46e39fe177cc · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Direct Language Model Alignment from Online AI Feedback

Reference 156

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:14b1462a68c375a4d8a69de23fbe2832573b1ea2bd64da9d2b4f5cfbf4484287

Observation 459a5410-667f-4eaf-9155-22e366d49116 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Direct Language Model Alignment from Online AI Feedback

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:49.906858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:49.906858Z digest=sha256:9ef0dd275c60fe9863f99eec08550a99b8210f11f940d2f11bb32b05b70ac6ca

Observation bd08babf-de01-4076-b7cf-6777e4df8941 · inbound

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding cites this paper.

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding Direct Language Model Alignment from Online AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T21:13:05.334534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:13:05.334534Z digest=sha256:171c27d01e54c2a93dd7976e7b4a0e823d35a3b505bad174bb2fd58bc97fdc82