Pith. sign in

Paper Citation Record · LEDGER

Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2307.15217.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.15217 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 100 of 133 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:06:54.024699Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

92
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d8a95642-eca0-4579-8bb7-7b6ef46e5aa6 · inbound

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment cites this paper.

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 126

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:46:56.825133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T00:46:56.664582Z digest=sha256:d6d358aa5d9a968c3391a28faec831d57d812c46fed18bede9d9816bde7d6e28

Observation a07e6ed0-a364-4ab9-97be-515c3ad34b6b · inbound

Active teacher selection for reward learning cites this paper.

Active teacher selection for reward learning Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T05:56:01.733753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-24T05:54:43.873174Z digest=sha256:43f90e78cd56d0b224687fc216cc95e6c249d3bce678ccdf754451d9390cf19f

Observation 828da731-537b-4454-aea0-1201131bcf23 · inbound

A Roadmap to Pluralistic Alignment cites this paper.

A Roadmap to Pluralistic Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 154

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:37:53.479800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-16T14:37:53.279275Z digest=sha256:81a6535bedbf233d39a6207e749980814cc95295440091d6a5ff9ac20cc2fe94

Observation 25cf5175-e1be-4c62-b356-a41a5f68d7b9 · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:05ffa77485aca9f6cbbb40dffc371a7ae6a094576711d787ea26c1d9dd78a828

Observation 31597a6d-0474-4ec5-96f5-788952ad2efb · inbound

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models cites this paper.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 264

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T14:43:30.279272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:c40c015991d481fd302f9a1b52ac512abcb7ba79fa4df2724158137218805750

Observation 3c04b55c-8786-430e-bb16-57578f39a27c · inbound

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions cites this paper.

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:55:49.986978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T21:54:26.670284Z digest=sha256:6dec7399bdc13020c235dbc6a1a1a3477f61bf458ad3ecd8c5b345b263d7c636

Observation b586926b-b09e-49cc-83be-fd9e878bacb0 · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 141

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T12:04:10.685287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:926a290e61c083bf9c5587e5d640dbc81062a30044f94ed75998273b96af8e10

Observation d1131658-53a2-4c1b-8da7-02f27aafd9fb · inbound

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs cites this paper.

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T16:18:01.605404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T16:18:01.560780Z digest=sha256:5048a4ab882fe8d0645a1b481464d40bbdb6390ee143231b5eec50102c699c06

Observation 00bf9566-314e-43d9-b38e-03e081c3ab0b · inbound

Towards Data Governance of Frontier AI Models cites this paper.

Towards Data Governance of Frontier AI Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:54.024699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:54.024699Z digest=sha256:e08c1871c984c7771720baebed9ac4024d6595d9a8cf14ee79c1f4b4db17608d

Observation a718bade-719b-4609-8a77-e6cf523f3bdb · inbound

ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks cites this paper.

ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:21.698370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:19:21.698370Z digest=sha256:948921078069b43ffd7d55abab8c58566f4d32f1676861a1277b00fafbbb9233

Observation 98458dd8-32be-44da-89a8-380a9eab4476 · inbound

Active Inference and Human--Computer Interaction cites this paper.

Active Inference and Human--Computer Interaction Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:59:02.638813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:59:02.638813Z digest=sha256:10f2d3a6bdf98444bbd40e929ab453ba091a48af3200c2e672b983298df8d93e

Observation 97938775-c86a-4530-a403-76c05dd628f5 · inbound

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples cites this paper.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.363100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.363100Z digest=sha256:df4da9a954ce79d550f96e088756f8cdf926951c7e86cebc028e0164d57fb1ed

Observation 453f6f5d-cbc5-4d96-a03a-fd10d120fdb0 · inbound

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe cites this paper.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:35.901723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:45:35.901723Z digest=sha256:7784399d16819f9f39923e799120c5517e45eeac3a9e44be182c83d558592350

Observation 71e97d38-5174-4b94-ba1f-49e4796217cf · inbound

AI Agent for Education: von Neumann Multi-Agent System Framework cites this paper.

AI Agent for Education: von Neumann Multi-Agent System Framework Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:06:05.525086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:06:05.525086Z digest=sha256:6b0b25d51edb4a2cbb2366db013e3c4edbc79880771aa8ce4b7779a329ace8b9

Observation e75027bd-2ea8-4057-830a-9d5dc3d72fa9 · inbound

Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information cites this paper.

Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:56.354957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:56.354957Z digest=sha256:24225b6d700ecf1f50827f3b01689c74346a15fcbbab3a8f83f6f8b64c3baeb6

Observation c945e325-1d29-4ab8-b70e-07f7936b83cd · inbound

AlphaPO: Reward Shape Matters for LLM Alignment cites this paper.

AlphaPO: Reward Shape Matters for LLM Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:51:07.514828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:51:07.514828Z digest=sha256:3f943658eb51d074f3a992ddcfc85d106320562c195991797efe6e39a4c95181

Observation e054cc71-f857-4a29-84a1-efc37b332d2f · inbound

Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction cites this paper.

Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:48.341490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:17:48.341490Z digest=sha256:1f5ecf49c0e5388a0a3ec13cf06b0387b881da395c5e0e2af0907690887e3bbf

Observation 769e38ec-1831-4bfe-a1e9-89e1706b2c4f · inbound

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment cites this paper.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.506817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.506817Z digest=sha256:68ec68cf884e0d8bd7a46fcd59f14545779c02a054d0ba6cbe7e9582db4dd6a9

Observation a2973a96-1e78-4e54-8240-458dcbcedce9 · inbound

Debate Helps Weak-to-Strong Generalization cites this paper.

Debate Helps Weak-to-Strong Generalization Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:50:56.305623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:50:56.305623Z digest=sha256:40933d9de6e0e802384b9adf94bb111b7ae6441aaa219b3eb020055825b399e2

Observation d8f884b5-31a6-40bb-89b3-0c29f5c8b440 · inbound

Trustworthiness in Stochastic Systems: Towards Opening the Black Box cites this paper.

Trustworthiness in Stochastic Systems: Towards Opening the Black Box Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T13:10:07.857096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:10:07.857096Z digest=sha256:0b1e4edf431bd766bdbfeb791838da7c5c1c6062cc1cbc1e93c620ec2039cc4e

Observation db8e2882-4d5e-4719-b39d-617a0ab8bac5 · inbound

Offline Learning for Combinatorial Multi-armed Bandits cites this paper.

Offline Learning for Combinatorial Multi-armed Bandits Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T20:45:51.240862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:45:51.240862Z digest=sha256:846c93fbc5880635c55f6bc32532f3e7eb2aa6b05b46d2bc794449ca972070f6

Observation e1bf8590-3090-406b-a708-657b7bd30e25 · inbound

The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking cites this paper.

The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.238611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.238611Z digest=sha256:963aad2cc516ba63cdfd266be19c48b47271ec6f54a417599834b03998219f28

Observation 81b336dd-1a86-4be9-b198-ad2b53e51433 · inbound

Process-Supervised Reinforcement Learning for Code Generation cites this paper.

Process-Supervised Reinforcement Learning for Code Generation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T15:10:38.772031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:10:38.772031Z digest=sha256:8703977e7c02fbfa9915ad1bbfa50ff611f02b1e8eb4ea992cc60f5d2988c61b

Observation 19ba90b6-35f8-4d14-996f-9b0f03fd7e62 · inbound

Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning cites this paper.

Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T01:04:04.115081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:04:04.115081Z digest=sha256:8fc16f300454fe1143187cdd0c345ab9f0e44eaf26cb32a4111afc02f7e809d9

Observation 8f12cd1d-928d-4dad-8eb7-ce13dc188a92 · inbound

TruthFlow: Truthful LLM Generation via Representation Flow Correction cites this paper.

TruthFlow: Truthful LLM Generation via Representation Flow Correction Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T22:23:27.507483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:23:27.507483Z digest=sha256:1209d63d2c20bc74693273c49d0505c1c781475d53acd007157661933f6ae906

Observation 0a55315f-54a4-494a-955d-74c839955e2f · inbound

Use of Winsome Robots for Understanding Human Feedback (UWU) cites this paper.

Use of Winsome Robots for Understanding Human Feedback (UWU) Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T20:13:24.242292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:13:24.242292Z digest=sha256:f2eb0bd579ddf16eeefb702660a5dba74fb0784d825d5384a2067a5f06bfd714

Observation ceb3ae33-a214-48ea-a9b7-9747da5c0bc8 · inbound

Probabilistic Artificial Intelligence cites this paper.

Probabilistic Artificial Intelligence Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-08T20:51:10.056960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:51:10.056960Z digest=sha256:06fb83d21fcab09fe0200711365d4f61d565bbfd9b33ee7fecf947105aefdb4c

Observation 1163456b-9200-41ef-9695-08a2fd6e6d56 · inbound

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models cites this paper.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.803268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.803268Z digest=sha256:cf3281d72392bb8a7a90d7ee20663ceb8a29f0b16d3a7e59d0e7d7f662ea74cb

Observation 7dc7c782-b221-45b8-b989-3fd527b53881 · inbound

Thinking beyond the anthropomorphic paradigm benefits LLM research cites this paper.

Thinking beyond the anthropomorphic paradigm benefits LLM research Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T22:21:52.483818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:21:52.483818Z digest=sha256:c1d139fdd624c46933b6a1880916f0bd2308a484dd96b1fc9a5685d6d25c1757

Observation 6c96f6ae-cab5-45d5-bba6-5fbf314e9818 · inbound

Kaleidoscope Gallery: Exploring Ethics and Generative AI Through Art cites this paper.

Kaleidoscope Gallery: Exploring Ethics and Generative AI Through Art Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.805406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.805406Z digest=sha256:bc49d7d227d6a863b8b86bff96a6ab8a60985ce8bb998cc99dc13badbb277185

Observation 9bf3d758-8c50-40ea-88e2-3edf4d488e02 · inbound

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO cites this paper.

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:08.255598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:08.255598Z digest=sha256:ff01bd260f9f214931de3a40990bedb35510cbb44f82777d29cd6fb010eac1e3

Observation c234d425-3000-4e85-8a67-7c03b38df923 · inbound

Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects cites this paper.

Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:34.921482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:34.921482Z digest=sha256:9f52d7c548df40b987cf57f145dedd2960b6f7db058a4e1fbfc44cf9effde978

Observation a629e6ad-7be2-4c03-81fd-d1473e5d3b6d · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.826681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.826681Z digest=sha256:56e50223f32f4888415612313681038079eeacd6ea02438be9a78cc9c2e4c79d

Observation b351382a-f6ee-491e-9b37-d8518700b8c6 · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:31.042608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:31.042608Z digest=sha256:19dfbf46ab9015814b508ff6434f65de387ef1aff2c9fb010ae8ba2a47a2cd2d

Observation 1c0b73b2-aa5c-49b1-990e-9906ac6bb7ed · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.614974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.614974Z digest=sha256:01b4a7b7b22d4a9fe7046517ad3ddec89f98034b478e5718c31b038a21dd62a9

Observation aaebfc88-f326-44ee-abbd-07aab08942f9 · inbound

Risks of AI-driven product development and strategies for their mitigation cites this paper.

Risks of AI-driven product development and strategies for their mitigation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:11.081752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:11.081752Z digest=sha256:acf8da00d602c8547b386a42029e2f75cb407bb21bcb64e5d7d3db766b335c87

Observation 1b0746f2-8130-4e13-81be-37aa6ee289b5 · inbound

Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising cites this paper.

Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:57:42.237392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:57:42.237392Z digest=sha256:38e5b32013ed1ba0bcb909eca84e2b0894e2773dd8bb075b18ce607aec8190f2

Observation ef7b58ca-d551-4fcd-bb97-aa7a6aee3c69 · inbound

Crowd-SFT: Crowdsourcing for LLM Alignment cites this paper.

Crowd-SFT: Crowdsourcing for LLM Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:09.014935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:54:09.014935Z digest=sha256:24f8b79d8a967edae1d870001b7994cb5d6234942306427eb0cb37140b19bdb2

Observation 22e083f8-5051-4f3b-aa9e-db05d3809bee · inbound

AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models cites this paper.

AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:18.687477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:18.687477Z digest=sha256:ca870695d0e8602a1134a175b74d61c6611d5fe84d7a572024a1c5939b793691

Observation a262c9d3-2641-4e1a-af74-b2c73504f47d · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:42:14.023050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:e5c75f1804c751e4f5aefb9c071b029b4febd6d4602b083a2b7101694700fdb9

Observation cab35d18-a447-4352-ac03-fd37e5153ac7 · inbound

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values cites this paper.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.123823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.123823Z digest=sha256:950d97bac0dadd549e557d551d14502657cd3189fbe16a5265314d0e163e85f0

Observation cddfd517-9ac4-4910-a2d7-25882bebf5e6 · inbound

Collaborative Editable Model cites this paper.

Collaborative Editable Model Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:12.912770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:24:12.912770Z digest=sha256:db2626d43f8e55fdefab0e668230d1dd30d9572b6dd027a1c33ea00d5e7bf76c

Observation df5fe3e3-889a-428e-a270-c6de61da79d2 · inbound

Affective-CARA: A Knowledge Graph Driven Framework for Culturally Adaptive Emotional Intelligence in HCI cites this paper.

Affective-CARA: A Knowledge Graph Driven Framework for Culturally Adaptive Emotional Intelligence in HCI Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:28.071431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:28.071431Z digest=sha256:acc89e281be9e78acf4839675a23729b860750541cbe9960488baf4148117506

Observation 1dd5e96f-9931-4ee4-84e8-729017ed3d0f · inbound

PrefPaint: Enhancing Medical Image Inpainting through Expert Human Feedback cites this paper.

PrefPaint: Enhancing Medical Image Inpainting through Expert Human Feedback Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:32:11.600714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T08:31:24.549747Z digest=sha256:33f2a9bc7791fc3e3cc75c4809aa90171f430ad1476c8996889956e9d68c58e8

Observation 77d01e76-e017-48ab-9720-a3eca423d311 · inbound

Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language cites this paper.

Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:18:15.468808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:18:15.468808Z digest=sha256:c7be919f7da8a65a109a2f6ed5244b98d867b57ef593c9eb45f888943bb87966

Observation 1cd622ca-9fb5-4544-a414-a41cc0fa6ecd · inbound

Exploring a Gamified Personality Assessment Method through Interaction with LLM Agents Embodying Different Personalities cites this paper.

Exploring a Gamified Personality Assessment Method through Interaction with LLM Agents Embodying Different Personalities Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:37:07.656554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T06:35:06.890058Z digest=sha256:3c7ba4b10657085b95a977e158be1833cc0f89c0df3d3637203a2001d8ccf3b6

Observation ac2747d2-67f7-42a0-9154-a4e87dbab96e · inbound

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis cites this paper.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:53.661417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:53.661417Z digest=sha256:38ca7aaa21d9af4f3b00b356e0314272bd8c002a0d9bf54c5769290690fb4027

Observation 79a2f52d-fd27-4fbb-9f10-74f12a55cf3c · inbound

Granular feedback merits sophisticated aggregation cites this paper.

Granular feedback merits sophisticated aggregation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T17:01:18.225322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:01:18.225322Z digest=sha256:31eedb50bd241dce819a775a9e1be8d38efc55f8342f27ab212bd48f188c6644

Observation 81e681ba-1912-40da-ba7b-1d2f75ac1ce6 · inbound

PrefPalette: Personalized Preference Modeling with Latent Attributes cites this paper.

PrefPalette: Personalized Preference Modeling with Latent Attributes Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-06T16:27:57.204698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:27:57.204698Z digest=sha256:aad2f3c4d2eccd5859d9ce84e27d98a255607d4256432d15f9b7073e9879cc0f

Observation 8a5a0256-0ed2-4a11-a324-0a6ae5dfe732 · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.002781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.002781Z digest=sha256:d95386986a1c480f96c92bed50599319149f95d571b71617b94365b1bb898687

Observation 5b05867a-b2bb-41e9-bdae-b1cfd9d5b789 · inbound

Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor cites this paper.

Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:37:50.967677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:37:50.967677Z digest=sha256:6ae1c60f4fb74ab473d216ed34cb90e6a86111fd8526819860c1cadffdd936da

Observation e8b2e8f2-8d63-476c-91d6-e9393a28425a · inbound

RecGPT Technical Report cites this paper.

RecGPT Technical Report Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:04.349891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:04.349891Z digest=sha256:452f93049cf19e824d5d73f6a76cc4e4d41079846425f0603ab4c39fd02414bd

Observation 5d235c00-dbf4-4475-9b75-70eb77925cd5 · inbound

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap cites this paper.

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-21T23:50:47.581965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T23:46:24.208438Z digest=sha256:f8c4ad89a3bbf77c38fabc88d02dbbed8ed0507536e4a5ba2aa7c049b4b2bf35

Observation 05948fb1-cc9e-4c62-833a-44cbbf013ac9 · inbound

Beyond Prediction: Reinforcement Learning as the Defining Leap in Healthcare AI cites this paper.

Beyond Prediction: Reinforcement Learning as the Defining Leap in Healthcare AI Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:08:47.910667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:08:47.910667Z digest=sha256:6217aa93f6f808a09fd486c32967f701c901290409aa258cf90e55caf3d024b5

Observation dd4f5822-fdd6-4f1e-b279-7c9a34e7a1ff · inbound

FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus cites this paper.

FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:53:38.242707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:53:38.242707Z digest=sha256:56af0c55d3134fd8ba9d8ad9fcfb382861fb9135b2597594719b57b2df2c003d

Observation 97bb6238-1d2c-4042-b3e8-2ea3e4dddebe · inbound

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences cites this paper.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:10.974544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:10.974544Z digest=sha256:c17aa655bb3b8770ab83ca690363d031008440015dcd1324d6fdd83ee028b0a0

Observation 204f11e3-f7e5-475c-9fe1-17de2c24ec84 · inbound

The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback cites this paper.

The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:04:04.797646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:04:04.797646Z digest=sha256:644830e300a5c99a7a8ade35f56713f59c3096e7286acb45dd2b19fdc7b0fd9e

Observation 95956f12-4d1f-464e-a2b7-b46706ebc560 · inbound

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey cites this paper.

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T19:19:22.387405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:19:22.387405Z digest=sha256:40c882701d1980d55b3d65e959bac3136d7aca5a74896fa7ac361518be1b18eb

Observation a2f5499a-556a-4cb9-b502-68a2b81a64ec · inbound

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance cites this paper.

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T14:44:13.604742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:44:13.604742Z digest=sha256:e8d59f358965f73d367bfb77fb99ecad67f24920f53db7875b1babd2badb6c82

Observation b2174ce8-2c66-4898-a84a-1f6dcd1579e8 · inbound

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives cites this paper.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.700516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.700516Z digest=sha256:db57b356a6c29f051566404a96f76b16613032330cb8c215952b66e0b164b3d1

Observation 4bec7b9b-952a-4327-a2fe-21ec7b80fcf2 · inbound

A Multifaceted Analysis of Social Biases in Large Language Models cites this paper.

A Multifaceted Analysis of Social Biases in Large Language Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T16:19:31.037475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:19:31.037475Z digest=sha256:2704fce9473169bec8a03d9f4793b0a2f25117a5aa908eaa0c39779a0838ae4b

Observation 1e93959a-afe3-4f8a-977d-a8e32867b405 · inbound

Beyond Context: Large Language Models' Failure to Grasp Users' Intent cites this paper.

Beyond Context: Large Language Models' Failure to Grasp Users' Intent Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:11:13.642122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T20:09:25.827452Z digest=sha256:8e56cb11f54925feda4c485c3dd90350aa76f456a423de83fccc9d6db6b8e17d

Observation ee760c3f-a2e5-4896-b32f-a326426bfb41 · inbound

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment cites this paper.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:40:49.108927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T09:37:57.120779Z digest=sha256:aacbd7f6aca1390ff4f9a60df05cff1a0f8225ac485d13913cd7d8b5bf23a850

Observation 114285a7-d4c2-4d75-b1a5-181c265bfa08 · inbound

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment cites this paper.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T14:10:13.449970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:82bdc8ce61bbbbc9d71ba3d883ee11d05aa34df0ed538efa8514ea51a3e2be03

Observation 0e13fa8f-ac8b-46b0-a53d-0f724e22b70f · inbound

Acquiring Human-Like Data-Efficient Mechanics Prediction from Deep Reinforcement Learning cites this paper.

Acquiring Human-Like Data-Efficient Mechanics Prediction from Deep Reinforcement Learning Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:49:03.255485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:49:03.255485Z digest=sha256:efcf7eadbc817dc52898a51a1273fc2b58f58d2b92c7200144be387bbafee8b3

Observation 14962ea5-6d68-42a0-b81d-f9a56741eaab · inbound

After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions cites this paper.

After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:54:12.936609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:54:12.936609Z digest=sha256:a00dd4a7376f928d707810d50830b536b390b5c2e89b94acd38a7d1e155da9d9

Observation 2c3f6a3f-121e-4ebc-bd54-8368f0cd73ba · inbound

Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent cites this paper.

Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:46:35.697989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:46:15.275441Z digest=sha256:73416e015ff06aa4c2943a944ef746b7f415085188019a51297abe9bd686d235

Observation 42c1d9f3-25d4-4821-a6a5-414e6fc24da4 · inbound

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models cites this paper.

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T20:29:31.354743Z digest=sha256:6f1809d1d0f5dc94533f2113e2808138a20a48a5c7aca6147cb9d57aa24ff490

Observation 8b647c86-fe95-4a56-842f-3bf8c19fe435 · inbound

The Theorems of Dr. David Blackwell and Their Contributions to Artificial Intelligence cites this paper.

The Theorems of Dr. David Blackwell and Their Contributions to Artificial Intelligence Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T18:33:50.296933Z digest=sha256:a00a8fdbdb6521d906423a9ee4d2af08be13874b944667049ef641d9e71a7a5b

Observation 2e66137a-7514-48f5-b823-6564a678bd7a · inbound

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training cites this paper.

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:18:56.476698Z digest=sha256:ca2887a46a9e50f12eea1e4de5034de948a7ef405b7c1da5f36bd4ab45b06be8

Observation f2b2cbde-4eca-4447-9353-aeacf297aa06 · inbound

PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification cites this paper.

PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:19:09.341399Z digest=sha256:0461357c337f32608505401a1fcf5835096fb1c585082fa89ff2a9b722427dbb

Observation 13c48fb7-dd71-4edc-8adc-4923ed9218ac · inbound

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation cites this paper.

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T13:09:35.407790Z digest=sha256:3fc767e943ee49f6bea9e3a93e22d42ce02c12772ea4965a468a4aefc92a8508

Observation 36751cd0-2cc4-4a3e-8169-3fa3be83a811 · inbound

AIRA: AI-Induced Risk Audit: A Structured Inspection Framework for AI-Generated Code cites this paper.

AIRA: AI-Induced Risk Audit: A Structured Inspection Framework for AI-Generated Code Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T05:14:19.185184Z digest=sha256:c1c242e4f42c968bd23d245eb9ceaf489481246a8f9d20c8833b54792386910b

Observation f7d9e837-e024-45cd-b794-36bf595eac1b · inbound

Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback cites this paper.

Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T02:23:20.208976Z digest=sha256:02b1f660ca4be783657a6b378a55dd2f1fc927c64fb95ab81c214cc05764d0a4

Observation 21963a03-975b-4c12-8210-0430445ba567 · inbound

Post-AGI Economies: Autonomy and the First Fundamental Theorem of Welfare Economics cites this paper.

Post-AGI Economies: Autonomy and the First Fundamental Theorem of Welfare Economics Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T13:13:01.920636Z digest=sha256:28ba958f33f2ae32bf41e9ecddf98ba094edb9da8b9f470a6b58626372750c9c

Observation d85e55c0-fb36-4e7e-8db8-167d471bb613 · inbound

Three Models of RLHF Annotation: Extension, Evidence, and Authority cites this paper.

Three Models of RLHF Annotation: Extension, Evidence, and Authority Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T14:28:31.451466Z digest=sha256:de959cf217f38f23b5a9876191f6c5c8ad213b96717bb87b5a1d973232262081

Observation f4fbdf6f-ea2f-4df3-9781-2c01b6941981 · inbound

Learning from Disagreement: Clinician Overrides as Implicit Preference Signals for Clinical AI in Value-Based Care cites this paper.

Learning from Disagreement: Clinician Overrides as Implicit Preference Signals for Clinical AI in Value-Based Care Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T06:26:57.714846Z digest=sha256:2a36bc24c25e82c74ce0bef499dfe8fefb9f3af4f0e70c4c195528a5cd68b9f6

Observation 594bfc7a-b5ee-4499-9129-3b49da21f9ec · inbound

Learning from Disagreement: Clinician Overrides as Implicit Preference Signals for Clinical AI in Value-Based Care cites this paper.

Learning from Disagreement: Clinician Overrides as Implicit Preference Signals for Clinical AI in Value-Based Care Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:50.990306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T23:43:37.688604Z digest=sha256:c0f347426ad0997969826476bb2c06ce548ff64b9fd370dee0132d7beb1136ab

Observation 92b09ebb-7d35-401e-935b-ef0397d63fbc · inbound

Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading cites this paper.

Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T17:08:46.405278Z digest=sha256:df12d89aae322800a390db0ed16ec01a18e8404f2cb1aa0b7597648638a21f34

Observation 10cbd289-f5b3-4176-b54d-eef634840de4 · inbound

Efficient Preference Poisoning Attack on Offline RLHF cites this paper.

Efficient Preference Poisoning Attack on Offline RLHF Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-08T19:29:25.000361Z digest=sha256:bfa535a2df280e877659fd1e204ad5c58d74a85dd7cae7a3e2729898f460989f

Observation 7bd7c218-ff88-40ea-9006-d14104726a17 · inbound

Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training cites this paper.

Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-11T01:57:40.347786Z digest=sha256:2514c80c2bd479353945d9775a840afab2a71b6015d805533744da245a06cbe7

Observation a7882be5-e14b-4d10-9a28-b824a2c66a79 · inbound

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences cites this paper.

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-11T02:30:14.693348Z digest=sha256:c5479f3a9231cd8797d5c3055a45f22a1c8d635e9eaeb5e7ebc3bfcd4b0d9b00

Observation a228f44e-35ad-4a0f-88c3-af473d5152d3 · inbound

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems cites this paper.

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:47:40.772146Z digest=sha256:dfce130288700fa04272cc40c6a65d1fd2681b38dba6a539e9e3993a1e9c60fd

Observation 97b9a3c1-143e-47c4-be16-64e2753cb0cd · inbound

Can Revealed Preferences Clarify LLM Alignment and Steering? cites this paper.

Can Revealed Preferences Clarify LLM Alignment and Steering? Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T02:10:50.298626Z digest=sha256:61e66be8b504adca791648c2c62c47b5fe0d98f96a1db25b47955314e9721987

Observation 28c1bddf-928e-457b-96f3-c92bf7e3840d · inbound

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping cites this paper.

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:33:40.994346Z digest=sha256:4005de7048efed51938419cd8cc280b3bd2744d75dd3d897e2d7a322dcf59f1c

Observation 9d69cc25-5418-41e5-97c0-3b07659707fb · inbound

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training cites this paper.

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T06:30:51.812541Z digest=sha256:3f2b6815c9ba31d4965d67f7c3b5aa06e4231531e577953dbb9bbb4a979a2b7b

Observation 91045b5c-6b56-470a-80bc-2d0ff7c0f63f · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 215

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:44:22.602157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:a28eab94326ee9ab3707ed52575f5c2c8f7205b267e81096926fc53e50de1354

Observation 57819ffd-b01a-45c3-895e-8b7debcb0fda · inbound

Common-agency Games for Multi-Objective Test-Time Alignment cites this paper.

Common-agency Games for Multi-Objective Test-Time Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 184

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:15:06.439800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T06:14:53.685486Z digest=sha256:90ad1cc08ae6c80f8a7c7af248bec4ac498cf49f7beca3925d233aa7aecf9b77

Observation ba3bf406-2a29-4790-b9ba-99a65925cc69 · inbound

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination cites this paper.

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T18:02:42.314028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-19T18:01:06.649723Z digest=sha256:31480b42dc09a5bb0bdc99edbea719df57a3477520bf53e85273fdfe6d805c2b

Observation 9e08c13b-3780-4ef4-b8d5-341dd7ec9d55 · inbound

Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems cites this paper.

Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-20T17:28:48.024488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T17:26:02.346222Z digest=sha256:6359991d0b85cb88474fe0126c6dcf1d27d7493857c047b8c38f29b1f7a14da2

Observation d47a6bbf-8f01-48ce-ab89-8f3b3b24f7a1 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.204764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:0b49d9e2d4586a1b082da7554f3cc17c5987d0a08e830d65744e5c20f3b8255a

Observation 542eb250-0ca5-4b23-b1cc-734e8f95d18e · inbound

Some[Body] Must Receive That Pain for Agent Accountability cites this paper.

Some[Body] Must Receive That Pain for Agent Accountability Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-19T19:47:44.480439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-19T19:46:03.728266Z digest=sha256:f541102d2a36345b2bc3887bf0c56cf353a06e8d0a144443cbc0ee23c988ccba

Observation 134094df-a128-4114-b62e-dc477c9eadb9 · inbound

ClaHF: A Human Feedback-inspired Reinforcement Learning Framework for Improving Classification Tasks cites this paper.

ClaHF: A Human Feedback-inspired Reinforcement Learning Framework for Improving Classification Tasks Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T15:13:24.966375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T15:11:27.420642Z digest=sha256:3b455668385f3c43102dd8fcfb3eed05cc1c2a85a7231c5e66c46616d17b4b3e

Observation da158d90-1513-4cff-b6ca-bf2fc83d9b8f · inbound

Base Models Look Human To AI Detectors cites this paper.

Base Models Look Human To AI Detectors Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:13:22.514511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T06:13:10.815204Z digest=sha256:216b20d325d0c1533a4235d56fec1036717b900e3035e9e3ad8751be1ec7b7d0

Observation ae5723ab-12d5-4757-9601-5cb5fb29229b · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T06:19:42.015118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:1d4ae51f40e3daf4155bda4c52ae574ae36b41a26e383009f8c1e3590e403005

Observation 85dbb226-d4be-4a22-8545-5ed867158fa0 · inbound

Echo: Learning from Experience Data via User-Driven Refinement cites this paper.

Echo: Learning from Experience Data via User-Driven Refinement Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:41:10.792145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T06:37:34.840129Z digest=sha256:6c46c589d11fd05e8f8161407fec549fbdc40b433d7bfbfcfd626d3de81ff727

Observation b4c8e3d3-9f7a-4a62-8a41-4f51d03fa8c7 · inbound

Emotional intelligence in large language models is fragmented across perception, cognition, and interaction cites this paper.

Emotional intelligence in large language models is fragmented across perception, cognition, and interaction Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:24:40.448632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T13:16:38.848193Z digest=sha256:d09af13e12c99233973f283c0e435fb5325d19d62abb7e2e676a527c6c88d508

Observation 17887a60-a748-4acf-a1e5-5c02444d73f5 · inbound

Cultivating Machine Intelligence: The OMEGA Shift from Top-Down Optimization to Autopoietic Cognitive Ecologies cites this paper.

Cultivating Machine Intelligence: The OMEGA Shift from Top-Down Optimization to Autopoietic Cognitive Ecologies Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-30T00:04:06.706440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T23:49:06.077148Z digest=sha256:9c7b073501d31d296c8815d887371437a312ba077ff8d4f7a947e7f2ed213644

Observation 988ba260-04d3-4ee6-a0e9-134ab4555582 · inbound

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm cites this paper.

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.799213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T12:54:36.818698Z digest=sha256:e5785059d7aba858d89af747583210678c1db5189562c82a4524f67d22fe72fe

Observation 4b911027-c79e-418a-a753-723082778d06 · inbound

In-Context Reward Adaptation for Robust Preference Modeling cites this paper.

In-Context Reward Adaptation for Robust Preference Modeling Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:23:15.204055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T08:19:55.177440Z digest=sha256:845fe517d8e454af67b86dde16106f58b7ea25a638a6d7852256f3161e5810ae