Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:40:09.533305Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 3 inbound Pith citation observations for arXiv:2504.20157.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:40:09.533305Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:29:18.049711Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-13T06:07:56.781634Z
78 of 78 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 84014ee3-0a43-46f5-b326-d6970b87b155 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Ahmadian, C
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4306b0ca-f5f8-4992-8f57-d901d427a5b4 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Amodei, C
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a224da52-802f-42ba-b81f-db6ed05552aa · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Concrete Problems in AI Safety
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 556e3ce8-83e2-4569-969f-8804d3cc9109 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10bd2885-16aa-4cf9-9b8a-fa048e6f08a9 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Buckley, T
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 91c60746-0dce-4ca3-a41b-b26db3800892 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models ODIN: Disentangled Reward Mitigates Hacking in RLHF
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0345da-19bc-4bd4-81e0-f040ef5bffa9 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 50b31a4f-c1f1-43c9-a793-b8ba75fb26f4 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32ef45af-468c-4eac-8426-7e477df43f2b · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Coste, U
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 836b619c-f099-4806-bf18-aa63d08c53d9 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ae088c77-19a0-410f-835a-817d8ce11952 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66984877-6b63-4282-ba84-e1fa7d93d77a · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 192bcb4f-0014-4344-86d5-96a3f0016df2 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f2d9a440-6136-4c47-9383-4967c4f8086f · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Efklides
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e1af0d2c-812f-41fd-9a01-dd322fcab31e · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Eisenstein, C
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f4431f81-cd12-44f6-bb02-e79f62bbd885 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models KTO: Model Alignment as Prospect Theoretic Optimization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d57dbda2-dc2e-4c8b-8c38-bf3675930d1a · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Everitt, M
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 69bd380b-f029-45cd-b87c-8e663f6923cc · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1088a2dc-d7d6-4b3d-aafb-dbb2db24837c · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 622f78c8-129d-48ff-a38f-179609261787 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c964a05-3372-4c03-81cd-062246a087ae · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 972e2a40-d167-4af7-9275-05ca3f9ee066 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Hamner, J
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9a360b64-3cbf-43bd-8b13-02fa13c14d9c · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Hendrycks, C
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4b777d64-80d6-4460-95d2-25b7b5664f22 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4d5521-3aff-4dd4-b14c-0c29b73a4b73 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ad9198-c9d1-4078-ab47-40aee5bc693b · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Kornilova and V
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f38cfb96-0bc9-4eea-a2aa-15bd1d42b31b · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Krakovna, J
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 433df192-df45-420f-9bb0-346297af1b6e · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5f8948f6-c1b9-43c5-a354-fc31c7dbf8ca · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 65d092b0-0a66-4b03-80bc-38d1630052be · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb6f3bd-3dd9-4bce-8e61-4dd24a6dfc98 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2b6b871e-def0-4485-8e08-94be64156b19 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2a05bfa3-4df1-45af-a08b-56cb3225c0e6 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models RRM: Robust Reward Model Training Mitigates Reward Hacking
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35f4bdc6-319e-43c5-9303-46afe503f3c0 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8facfa69-c929-44de-be2e-537ca11185c1 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Lourie, R
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a83fbe96-8807-4a2b-83df-d972bad9ed09 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a9e4a206-746d-4a89-bc50-5f3bbb6f1d3e · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7b1aaf42-37b4-471c-ae6c-3b964cf90611 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 527e7dc8-f5f6-4d90-8a32-520471788536 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Metcalfe and N
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a58260a0-9af0-45cf-8590-43ca5d4e1164 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2294f5e1-732b-40e7-910a-3e74daaa4d19 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68bc6e34-2ae4-4707-a5da-c381600b1957 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Ouyang, J
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1671728b-5791-4991-8af3-1404c6cb7c89 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Ouyang, J
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b11e6543-9da9-4bbd-a492-8597831bab04 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Ouyang, J
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bcb71774-95aa-431c-bfb5-00ae317dee4e · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b014d8c6-a3ac-4f19-95ae-579bfc4cf962 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd4b6a38-51b6-4861-8ce8-b0683ac52398 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cdf3ab89-eaf0-420f-a33d-d487e4804a56 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11df0c3d-e5ee-40f6-9677-7e0438431df9 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b76ba7ae-bc15-47c3-ad44-1e5664272e99 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Perez, S
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0d625eb-b103-4a56-b950-355b7800b92d · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Qwen2.5 Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b2ac55e-fa62-4471-a7a6-955281384036 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ce64f43f-f56f-4ceb-a41b-946b213c5fc1 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ac2ff3d-b2bc-488d-a9f0-6c7e6e41f189 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Verbosity Bias in Preference Labeling by Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ffa1d05-eec1-4e60-afd8-78932d7724dc · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 945a2281-cd88-478f-a281-cabf37e7366b · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Sharma, M
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0af90142-28ef-42f0-af13-5cf3c4d47bfc · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Singhal, T
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e811439f-75ea-4476-9b60-b3b631382a62 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 22eb0913-acdf-4a0e-b0da-3ae9645d77bd · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Stiennon, L
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d0391c6e-1b24-4948-9b00-a8c2744aa454 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ede0c600-1e5e-4546-8b8e-15db8f5c538f · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 78c37b0b-b751-4d8c-812b-8f2c0eb3be84 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 88e587cd-4c3b-4ad9-93cb-8ef92bdc6926 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models von Werra, Y
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceb0dfb0-e5f0-41c0-b2a2-4499233940d3 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 81672bab-2540-4b7e-a93a-9d2bdb0a2356 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8087916d-e8cb-4716-a863-e32a161a2792 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9725a20f-ecbb-4e9d-b395-b5b8eff154be · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d1fea079-aa42-499f-9b4d-ffc4df35466f · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e727195b-d96b-45c0-b113-21d25c98ee16 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30273e3a-f987-4b7d-9c49-dd0e691a5f71 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Qwen2 Technical Report
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a049343-5f77-4b12-ba85-9963fe750d34 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Self-Rewarding Language Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f871c059-471c-480a-a2df-30632b8255e1 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Zhang, C
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16fbac64-21f1-4c7b-be01-69dcc5635eaa · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91dd55db-cff6-4a6f-b309-b58a9a88d1c6 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models SGLang: Efficient Execution of Structured Language Model Programs
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28dd2998-59d5-438a-8394-5222b09e945c · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Fine-Tuning Language Models from Human Preferences
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aa3b570-ab53-40e2-bba8-cbe2d4dd9940 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models @esa (Ref
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 725ee1b4-4440-437b-a44b-19c7a9fe4673 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e22da6-d4ce-4a22-b621-5a20094e4d05 · outbound
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a11841a-1d3f-494b-b92f-976e3a3625a3 · inbound
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 514714e3-6952-4251-a15e-d7d09cae203f · inbound
Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models
Reference 222
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7112acee-2383-4d63-9ac7-fbb30715048d · inbound
Improving Generalization Robustness of Multimodal RLVR Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.