Pith. sign in

Paper Citation Record · LEDGER

Direct Language Model Alignment from Online AI Feedback

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2402.04792.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04792 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:39:37.864500Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:09:46.394396Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9fc03707-b6ce-471f-84eb-4d75de5d9fed · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Direct Language Model Alignment from Online AI Feedback

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.577360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:148c139fc5f43ca1d946032f20cc9cf88a844774f596bb56728b428a5f62347a

Observation 238780b0-f35b-4771-8497-9b696b970edf · inbound

Process Reward Models for LLM Agents: Practical Framework and Directions cites this paper.

Process Reward Models for LLM Agents: Practical Framework and Directions Direct Language Model Alignment from Online AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:37.864500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:37.864500Z digest=sha256:a456e6df993f6bd3ef2634cfbe2bb3c46f9e62cee98113eadad47f52df21374d

Observation 5724ed26-706e-4021-a240-b524f4676b17 · inbound

Preference learning made easy: Everything should be understood through win rate cites this paper.

Preference learning made easy: Everything should be understood through win rate Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.022389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.022389Z digest=sha256:f950d45ca21b76f2f398230e1951caa8e6ef124f0e6de254a1f9e15540357a6c

Observation 01d71fb7-6b99-405a-9bc9-81f4537e1e12 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Direct Language Model Alignment from Online AI Feedback

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.394165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:7811de92b2b82529db5255526f29bf996139ca5bd72137dfab4a1fb1b48f1f0d

Observation 9cb827c2-9926-4f6e-8717-aabe5f2532a0 · inbound

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models cites this paper.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.852749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.852749Z digest=sha256:bb922667f659bb51d81b4f936e9102e99895c1f652953f393d6f6c9459fbab81

Observation 37779663-14c7-4844-8c03-42accb2ae545 · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization Direct Language Model Alignment from Online AI Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:26.468933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:26.468933Z digest=sha256:48a257be98cdb8aa1a291ef7fb8b086cd4441c7dd89b03884d7e4923d9fe3e70

Observation 9ec77987-1d39-4229-8a6e-f57c7ae1920e · inbound

Self-Training Large Language Models with Confident Reasoning cites this paper.

Self-Training Large Language Models with Confident Reasoning Direct Language Model Alignment from Online AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:13.280315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:51:13.280315Z digest=sha256:0761f436ec0c223fb1aee63a4c4f185a602c571f8fa948dc53290e706e90c5c0

Observation 49e448cb-6a11-4e17-ae8d-5ecd6eb7f0f3 · inbound

Online Knowledge Distillation with Reward Guidance cites this paper.

Online Knowledge Distillation with Reward Guidance Direct Language Model Alignment from Online AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.270567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.270567Z digest=sha256:4606f56a18b6aaf5826feb04e362df489293c8949272127c2748f308091a44ff

Observation 70ae6733-b1b2-4dfb-88e2-ee7e74a226c5 · inbound

MOSLIM:Align with diverse preferences in prompts through reward classification cites this paper.

MOSLIM:Align with diverse preferences in prompts through reward classification Direct Language Model Alignment from Online AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:04.853691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:04.853691Z digest=sha256:3f3fabf99e4c5f9c5239a2855dc44a3245d8b7a785209289e4ab3cced1b62bde

Observation 7f1ca33b-face-4559-ab79-aec833b99512 · inbound

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment cites this paper.

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment Direct Language Model Alignment from Online AI Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:41.934550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:41.934550Z digest=sha256:a6923a242ae7791b5a2e8955258f0420ccd0f188d7f8469ceda430b380c65b16

Observation 23f9ea8d-a73f-4b8a-a799-0a9968480331 · inbound

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function cites this paper.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Language Model Alignment from Online AI Feedback

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:15.560054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:15.560054Z digest=sha256:ad83d664ca9243238f75777a1b61e11e576b5490d63a15c15c77e444cb1eeb29

Observation 258e424f-f7ff-40c9-a1f4-26cfa3be0f3d · inbound

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling cites this paper.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.185377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.185377Z digest=sha256:570b02c585be8290379d5da7e7334911419f9c9e3fd862eaba1a2959db7b7928

Observation 4d66ed61-f61e-490b-ae51-adb4cca2b999 · inbound

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling cites this paper.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.189540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.189540Z digest=sha256:30e946ecfd0368a75315df3b291cd491034601c442f5e2dd9d4668f92f11b844

Observation 9e272212-0a67-482e-b5a1-9831005a09d1 · inbound

Customizing Speech Recognition Model with Large Language Model Feedback cites this paper.

Customizing Speech Recognition Model with Large Language Model Feedback Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:19.820010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:23:19.820010Z digest=sha256:5b0b1c0e0ea42cb71dc383daaed37e29d434d351ee526e6bccec6ae8a6c9c6f1

Observation 84f5acc2-ba3b-43e3-946b-d31459bd35a3 · inbound

Bridging Offline and Online Reinforcement Learning for LLMs cites this paper.

Bridging Offline and Online Reinforcement Learning for LLMs Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.981374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.981374Z digest=sha256:7dbc745bd2e88de2b9358ecc2b031ab452be6a028a0a79cb5b4bf7295daec07b

Observation a140d1b5-77dc-4353-9028-07de6ffee18f · inbound

Data Diversification Methods In Alignment Enhance Math Performance In LLMs cites this paper.

Data Diversification Methods In Alignment Enhance Math Performance In LLMs Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:26.226540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:43:26.226540Z digest=sha256:0560487e532b970e0677477f197eab306fd73cae8ca02e073dcc0dc0124e533f

Observation 1f54acb1-611e-49c7-ba95-7ebfe0a7f823 · inbound

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought cites this paper.

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought Direct Language Model Alignment from Online AI Feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:37.182865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:37.182865Z digest=sha256:f9f6e434ff7e74eea5f57bb0506b11b26d6854a6773cfd9c004443aceaed306e

Observation d796d31d-9f3b-4109-8e94-f04c9c540e35 · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling Direct Language Model Alignment from Online AI Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.788905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:b3e5bc57fff9d98ead229b23599273012d25338afc42062ce09d13cd6dfcf6bb

Observation 2e45fe2a-0649-4a2e-88a4-cf1875fbf5d2 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Direct Language Model Alignment from Online AI Feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.021993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.021993Z digest=sha256:132e7387f7b420c8620f055b09df2c0f0a55b852a6a3bf01adf9dcabab4fc70c

Observation a9e3b838-9524-4038-8924-5fd067ad16f7 · inbound

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion cites this paper.

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T04:36:13.417768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:36:13.417768Z digest=sha256:d96e9f985f07f4a4adb5b6591e732e609cd256abb92e3c1f29873d2faf17f794

Observation 7012f92b-6880-4eea-b3e9-6e20e0838a5f · inbound

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization cites this paper.

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Direct Language Model Alignment from Online AI Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T09:48:05.297616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:48:05.297616Z digest=sha256:2cba8f5b3957c71fc85303a6f88ac50479d10d8c3e46528cc21368cad491668e

Observation d1431a99-d1bd-46e6-a6c6-9c58c7d34fcd · inbound

IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation cites this paper.

IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation Direct Language Model Alignment from Online AI Feedback

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:30:35.839277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T01:27:00.824546Z digest=sha256:4cd1b2c2aaa70f57367b0baac94c0c7e023657fe27d00787681ebdbc50d55920

Observation a917cb65-00e6-4af7-8631-da3807faf630 · inbound

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs cites this paper.

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs Direct Language Model Alignment from Online AI Feedback

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:51:17.792019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T21:49:21.785096Z digest=sha256:4b37950c9a1d01872dd7bbf03ec0e63e3ec98a6e12b131a1f70228d97af48660

Observation 50f7b995-cb23-4176-9e28-933b1a7e0e1e · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Direct Language Model Alignment from Online AI Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:21.154235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:21.154235Z digest=sha256:b27045ffbcb976dac490639b2a09d7c31ab778a981540ab08317304aa4b8d6b3

Observation 38f530cf-96ab-414d-bebf-65dc1657cdfd · inbound

Ministral 3 cites this paper.

Ministral 3 Direct Language Model Alignment from Online AI Feedback

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:12:24.704031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:6e27007e3abfb7d587bdc5c0afb0cd9578af1c274f0e13f6fb2a4b5bfb553741

Observation 2215a8c4-a41f-4730-822e-34c9055982d1 · inbound

Noisy Pairwise-Comparison Random Search for Smooth Nonconvex Optimization cites this paper.

Noisy Pairwise-Comparison Random Search for Smooth Nonconvex Optimization Direct Language Model Alignment from Online AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T07:11:58.083080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:11:58.083080Z digest=sha256:be6b145ebaff8c97693bde63e80143be7d28c03cca8682c299f752a1a8a5cec2

Observation 9c503c54-32c4-4024-a4e6-2084b118b98c · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:37:28.599444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:b4e2df7749f6ef72e2bee7bc6c9099de9c3c3961d1b2e5bde839f6c61e46d93a

Observation 33020fae-9d6d-41e6-aa54-059289939857 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:10:10.410196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:1838db4df076e9fd26509376b9af3ee8dae8a827a450a7037915905bce97c9bf

Observation c194993b-207c-4c12-8995-6db66a858ee6 · inbound

Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs cites this paper.

Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs Direct Language Model Alignment from Online AI Feedback

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:53:11.873009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T19:48:46.245576Z digest=sha256:aae966fdf69d7460a82c8ca6dfaf6ec14337efb34d94785870f33fa42105922a

Observation 07d33688-9cad-4a73-bed4-c96ca2a55bfb · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Direct Language Model Alignment from Online AI Feedback

Reference 164

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.597043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:96d6d5c17176c9f4b1153c8d3b12320c8e1c121217d18ac63fe7ea225843e7dd

Observation 99e214aa-2fbe-46cd-a9d3-90077f9be17e · inbound

CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization cites this paper.

CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:53:04.770013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T08:48:12.941821Z digest=sha256:c2b106c7ca49f517a9b2daa9e4b316aad4e8a4a2c5a815579ec9394b9e4c6d58

Observation 5777ec4c-c515-4c3b-8017-c2c5d5c9c181 · inbound

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models cites this paper.

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models Direct Language Model Alignment from Online AI Feedback

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:26.012326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T06:06:53.959061Z digest=sha256:6bee81a71ab27243f1bd01c254882aa8a86717c7b6e9a52ff2a39db80052c392

Observation 40b92d08-c052-432c-b463-ba6f952ed9bb · inbound

Mind the Gap: Structure-Aware Consistency in Preference Learning cites this paper.

Mind the Gap: Structure-Aware Consistency in Preference Learning Direct Language Model Alignment from Online AI Feedback

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.582713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-07T06:41:18.311986Z digest=sha256:03f6b45afbdde9e2ba531114e20efdc159f1c188592e4362ff5dbd414d5e3cc2

Observation 6efc0a37-4cae-4f1e-8eb2-e890ecd24a08 · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.769327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:0655faf24c0041fe006fb463275aa3b00295c886ad67feaa3564754fbe19b825

Observation d527bc9d-9315-4726-ab55-5629f13f5ac3 · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Direct Language Model Alignment from Online AI Feedback

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:09.228349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T17:01:04.571087Z digest=sha256:6fd12fa8f0c0bdab627cff008a261820c5c4bf1bebf3ff7501a43f8dadba9b36

Observation 1a5303a5-7923-4d03-9a3c-e3cf9e802322 · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Direct Language Model Alignment from Online AI Feedback

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:35:10.418679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:4ab8c852694df5e72434b1b499995cc5128c00d1b5c843b384256de899dbfece

Observation 485bf674-827a-4d27-9811-e11f5725fa08 · inbound

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control cites this paper.

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:14.624333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T00:50:42.836549Z digest=sha256:449684a1015bb284429535d2515fb762b8a9958ee49fe25262b71dc2fc8b7bb7

Observation 0cfec670-b681-4450-bb79-82a96b6e339a · inbound

dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models cites this paper.

dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models Direct Language Model Alignment from Online AI Feedback

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:51:29.686453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:52:05.779559Z digest=sha256:62973e5d426eedbeda45d9ee3882133d9184d582433c36905da1d7dc1626fce2

Observation bcc9ea60-ea4a-4e1d-8449-e318756e9fb0 · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Direct Language Model Alignment from Online AI Feedback

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:42.998151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:6e783144dcebc71873f5d7553fea92f49827a87e36ece5bbfdd1436cd09e40d3

Observation dec601cd-24b2-4258-86b4-a06000049b14 · inbound

VSPO: Vector-Steered Policy Optimization for Behavioral Control cites this paper.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:44.127323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:9799b61bed032f5cd924a916cfdc774c6b4a54e07f77834c953db2abfaddd4db

Observation a3a6c154-8285-4758-8d63-6191ecf29af2 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Direct Language Model Alignment from Online AI Feedback

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.516908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:ce75a3036a0aa9afb24be6b7ae092489b1b6a08040eb91364b5e67a480cfddf6

Observation d6bf4218-d171-4b0f-8898-8565ca3ab6ad · inbound

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations cites this paper.

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.069649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T21:30:39.993873Z digest=sha256:e28e5ac61a85d628ddcad854fb7283f740496720d7e6bdbf053f52ef512eedef

Observation 303e8719-2e68-4ce7-a158-36d186402e70 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Direct Language Model Alignment from Online AI Feedback

Reference 203

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.560428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:88f4f8469071fc661217133c438343736ae15f43d3675e3f43f7ffc1a57ceb8d

Observation 3e34989b-6faf-4738-b3ee-e8305e6abe1c · inbound

BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization cites this paper.

BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization Direct Language Model Alignment from Online AI Feedback

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:36:48.330733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T05:55:00.629666Z digest=sha256:2dbcba9cceed66430e308601b7e4fdd1001435a7dc1501512c80e3dfdc7e88a0

Observation 3fbc77f9-7207-494e-bd86-7239b59ed9bd · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Direct Language Model Alignment from Online AI Feedback

Reference 219

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.508306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:c50acf43b426e5bdb3ff3507310a5122472bede9c0f37f55a1e5d5705e41056b

Observation 5cb7cf2e-02e5-44ff-884a-07d8e2198004 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Direct Language Model Alignment from Online AI Feedback

Reference 209

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.396105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:815aa4e94cd9f08a72782e932127c4cfe04020520cceab744bf0f19c1bb4d2b3

Observation c6fb746f-6fae-42ad-84f8-f9b65d2ac202 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Direct Language Model Alignment from Online AI Feedback

Reference 195

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.465337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.465337Z digest=sha256:babf6704723b1134d298c529a34ee39211835454dd9c1fe10d7301035640f24f

Observation bd29c9fa-d3eb-4312-9572-46e39fe177cc · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Direct Language Model Alignment from Online AI Feedback

Reference 156

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:9b607a12df6f63fcec66603dd954c03324fe382cb6078aa64551315904edc979

Observation 459a5410-667f-4eaf-9155-22e366d49116 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Direct Language Model Alignment from Online AI Feedback

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:49.906858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:49.906858Z digest=sha256:dcd6e2fe15e922b65f0a01a2704b5dc702331e8ea7a56ac3717d47e0cfd8f01c

Observation bd08babf-de01-4076-b7cf-6777e4df8941 · inbound

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding cites this paper.

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding Direct Language Model Alignment from Online AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T21:13:05.334534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:13:05.334534Z digest=sha256:9e9a89f4720dfb53066301dbd55de6d26e27160c673693cd978603483960b67a