Pith. sign in

Paper Citation Record · LEDGER

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models

As of 16 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 3 inbound Pith citation observations for arXiv:2504.20157.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.20157 v2

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:40:09.533305Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:29:18.049711Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T06:07:56.781634Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84014ee3-0a43-46f5-b326-d6970b87b155 · outbound

This paper cites Ahmadian, C.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Ahmadian, C

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.128421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.128421Z digest=sha256:ab48a17b999e93f5713855cebec9209649ba9b6bb610ead6079c5ba31cc4ea10

Observation 4306b0ca-f5f8-4992-8f57-d901d427a5b4 · outbound

This paper cites Amodei, C.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Amodei, C

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.884168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.135897Z digest=sha256:20fa743748ed01cd0f25a31363817f3e804d9bf82880c1b3a401b3b07514844a

Observation a224da52-802f-42ba-b81f-db6ed05552aa · outbound

This paper cites Concrete Problems in AI Safety.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Concrete Problems in AI Safety

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.141124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.141124Z digest=sha256:b66922ee133754424f0171447339e198577f3a2fc61f2adebdc5c5f17dfd34d9

Observation 556e3ce8-83e2-4569-969f-8804d3cc9109 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.146832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.146832Z digest=sha256:f027ab8088dd9793b07131ad96ed0874d95aba5c4d7e1d77a15b9afb72d26159

Observation 10bd2885-16aa-4cf9-9b8a-fa048e6f08a9 · outbound

This paper cites Buckley, T.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Buckley, T

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.868225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.152614Z digest=sha256:73f9f90a594c44b4d7b012f5437c739fb879c47e4f1c3d437ff8a78b113bd0a4

Observation 91c60746-0dce-4ca3-a41b-b26db3800892 · outbound

This paper cites ODIN: Disentangled Reward Mitigates Hacking in RLHF.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.158251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.158251Z digest=sha256:493eb6e14c642791c6538705c3e252efe85b56344293c91d44ab95896d2bd715

Observation ae0345da-19bc-4bd4-81e0-f040ef5bffa9 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.852486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.164194Z digest=sha256:b7d16100a41379dd8766119660993b132da78e30b259d2539a603b9e9145161e

Observation 50b31a4f-c1f1-43c9-a793-b8ba75fb26f4 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.169269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.169269Z digest=sha256:d2cb1efc8e258e9773d0fbda7dd0d8c34566abeef58daf866ebd26f65dccd8c9

Observation 32ef45af-468c-4eac-8426-7e477df43f2b · outbound

This paper cites Coste, U.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Coste, U

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.836698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.174620Z digest=sha256:dcb01c121feca00db3cf36f5c023646718b5f9b882ac6a8a583399900556df6a

Observation 836b619c-f099-4806-bf18-aa63d08c53d9 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.820404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.179775Z digest=sha256:482ddbf099c694dfb86e11a565d551ea4b865cafe442c9d1a3da751607033f61

Observation ae088c77-19a0-410f-835a-817d8ce11952 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.184730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.184730Z digest=sha256:c507be565512af108f737ba1f044a29e3a09a52c33133f08a293d7f37b5fe106

Observation 66984877-6b63-4282-ba84-e1fa7d93d77a · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.189989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.189989Z digest=sha256:35ba80d3a971f0fa105789dad12a2305a4e8ac9975ef216fefe4a09190a67323

Observation 192bcb4f-0014-4344-86d5-96a3f0016df2 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.804793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.195069Z digest=sha256:afa1a33c2066dfc7e73b8c325b1173e3927d91081c14aa596f5fc2eed6d4fb8d

Observation f2d9a440-6136-4c47-9383-4967c4f8086f · outbound

This paper cites Efklides.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Efklides

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.788673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.199747Z digest=sha256:90164d7f30adf0702a4f35deb89644c4928a577c7d77cd9788d8f6eaf25fad58

Observation e1af0d2c-812f-41fd-9a01-dd322fcab31e · outbound

This paper cites Eisenstein, C.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Eisenstein, C

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.772380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.204248Z digest=sha256:71c99ff1fc446b9ea073425db5aad6bcf58a8614a89a243ddc473d15e8162a79

Observation f4431f81-cd12-44f6-bb02-e79f62bbd885 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models KTO: Model Alignment as Prospect Theoretic Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.208936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.208936Z digest=sha256:e012ffed94d505e0111902a0cbc6e50030c6052b197040944c86af701bee24c1

Observation d57dbda2-dc2e-4c8b-8c38-bf3675930d1a · outbound

This paper cites Everitt, M.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Everitt, M

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.754664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.213859Z digest=sha256:318513c8a9e61228cef9acce3aacb82bb3342656ae27383148b3335272c61bc0

Observation 69bd380b-f029-45cd-b87c-8e663f6923cc · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.738502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.219104Z digest=sha256:dbff7784a468ea48ac46ffb2eee2d0f9636a4e54c45ad5fda014dcd52135a005

Observation 1088a2dc-d7d6-4b3d-aafb-dbb2db24837c · outbound

This paper cites The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.223889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.223889Z digest=sha256:b05a72990d20652c631e88fdaf1d5757ff2b66c28076905ce07045e30d78c79c

Observation 622f78c8-129d-48ff-a38f-179609261787 · outbound

This paper cites Reward Shaping to Mitigate Reward Hacking in RLHF.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.228950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.228950Z digest=sha256:dbc296ecf6f0b566f4811900e644be22ad836f9d8824fd371bbcffb939f31458

Observation 8c964a05-3372-4c03-81cd-062246a087ae · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.722228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.233843Z digest=sha256:a79c0502130854db4501bbd9a312a2c1e4e78667cf0524a6476bcb2943c68f1f

Observation 972e2a40-d167-4af7-9275-05ca3f9ee066 · outbound

This paper cites Hamner, J.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Hamner, J

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.706274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.238853Z digest=sha256:f29e5415a87bff4f30c28acadc7558472b1a3a1da4925551d17d20fed13af4eb

Observation 9a360b64-3cbf-43bd-8b13-02fa13c14d9c · outbound

This paper cites Hendrycks, C.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Hendrycks, C

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.689981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.243702Z digest=sha256:a42b5b08fc473eb6e43ed3b96eaf5bd28ee3e968d66066536d8543303a8c780b

Observation 4b777d64-80d6-4460-95d2-25b7b5664f22 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.249327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.249327Z digest=sha256:57fa340c15d3fb3d077ce9bf8263f79b0ac8bf429e5ea46b8760bcc4d9d3d39c

Observation 4d4d5521-3aff-4dd4-b14c-0c29b73a4b73 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.254238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.254238Z digest=sha256:88b33a25060c5f1b74e5d1c24d05a9f409f7d3a7ccc080c25f81767b27eefd9b

Observation a1ad9198-c9d1-4078-ab47-40aee5bc693b · outbound

This paper cites Kornilova and V.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Kornilova and V

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.259920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.259920Z digest=sha256:3e532b84076a514cb0dad5f6c42d4c39e4b3ceef6e03c2fe55576290818e6976

Observation f38cfb96-0bc9-4eea-a2aa-15bd1d42b31b · outbound

This paper cites Krakovna, J.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Krakovna, J

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.265135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.265135Z digest=sha256:de4656f0c7fa4ba52444bd28d4362d1e7ade1b9bab3e2eb5f909ef025c7e94d9

Observation 433df192-df45-420f-9bb0-346297af1b6e · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.663066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.270126Z digest=sha256:baa0bc492be9945eec0c11bfc997839dbaf0842a96917bfb55e6a0a380bcf34b

Observation 5f8948f6-c1b9-43c5-a354-fc31c7dbf8ca · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.647596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.274936Z digest=sha256:5149f118009f2dfcabb27bea093d715bdc3d3a713f1386121450d041823aefbf

Observation 65d092b0-0a66-4b03-80bc-38d1630052be · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.284787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.284787Z digest=sha256:c4f0607823b02a0a2de462272c0fd5a790cb16b177aa7a09056defc9c247369f

Observation 5fb6f3bd-3dd9-4bce-8e61-4dd24a6dfc98 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.619821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.289547Z digest=sha256:27ef2f131854c5deb28529bef2c65ab9f15cf600b5cfd5cd374044fdb5cc2450

Observation 2b6b871e-def0-4485-8e08-94be64156b19 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.604313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.294328Z digest=sha256:1f8651fb1c64cd1c6d2770f1ff08d08cb5678cd8ac830e2292a27a7cda006553

Observation 2a05bfa3-4df1-45af-a08b-56cb3225c0e6 · outbound

This paper cites RRM: Robust Reward Model Training Mitigates Reward Hacking.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.300373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.300373Z digest=sha256:514b886c2c8af3549a65e6b94fffa475654bcfae4f08953d80c6828e918a43bf

Observation 35f4bdc6-319e-43c5-9303-46afe503f3c0 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.588317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.305542Z digest=sha256:d8c74c7ccf2ff21dfea7c38c81c6a56a5299fbcec8edfe9a29e8ae3402e9c19f

Observation 8facfa69-c929-44de-be2e-537ca11185c1 · outbound

This paper cites Lourie, R.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Lourie, R

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.572366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.310275Z digest=sha256:e55e25b5be0103b38e4b3737b26232ec74e7c1fdad4ec107fb77634df4c43b39

Observation a83fbe96-8807-4a2b-83df-d972bad9ed09 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.556775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.315294Z digest=sha256:cf4ea6d24bf6f70ae2f48918e54b5c310dcb0057540d7428bd9beceba213673b

Observation a9e4a206-746d-4a89-bc50-5f3bbb6f1d3e · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.540436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.319942Z digest=sha256:542d116834b35d14e18a474e806efa8daa3fda60ff83a821c453e6f9cc672306

Observation 7b1aaf42-37b4-471c-ae6c-3b964cf90611 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.524150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.324685Z digest=sha256:14d28ccace10bba0b18a0f6d33e7f0c6d3215bcbda8878a859f246189ebd239f

Observation 527e7dc8-f5f6-4d90-8a32-520471788536 · outbound

This paper cites Metcalfe and N.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Metcalfe and N

Reference 40

Resolution
verified exact
doi, observed 2026-08-16T05:40:09.598519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.329272Z digest=sha256:f1a80cad7a357921e120c101136fc43d83d24b57e92d280edb24f254579e34b5

Observation a58260a0-9af0-45cf-8590-43ca5d4e1164 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.506102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.334058Z digest=sha256:e9b173a659b5c0f1aca03a056fa8f9a21a030ece8f60bfcab7fde6c8089debdb

Observation 2294f5e1-732b-40e7-910a-3e74daaa4d19 · outbound

This paper cites The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.338975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.338975Z digest=sha256:1c9ca553c5157e99eae0288a9eb8a4a9517e0be322c6834e9c21e007b3f70b5a

Observation 68bc6e34-2ae4-4707-a5da-c381600b1957 · outbound

This paper cites Ouyang, J.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Ouyang, J

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.488096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.344202Z digest=sha256:d78c2016e634aee398f83ac39fedfa2d935c4c2602ba11db5272aa6a52b5b6d1

Observation 1671728b-5791-4991-8af3-1404c6cb7c89 · outbound

This paper cites Ouyang, J.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Ouyang, J

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.471645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.349041Z digest=sha256:8dd06c09fbdbf582f90319ec2c7d9ffa931d7966750b74f9a7b90e3df0ba5372

Observation b11e6543-9da9-4bbd-a492-8597831bab04 · outbound

This paper cites Ouyang, J.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Ouyang, J

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.454776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.353973Z digest=sha256:550c4241b39adc42489d328b4dc3c42e217555a6f9c631254ac39994d9dff4b9

Observation bcb71774-95aa-431c-bfb5-00ae317dee4e · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.437918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.358880Z digest=sha256:4390457d4bf53a89fd2f87c94a119c352ff89ff33cbc5a492518946be3796dc3

Observation b014d8c6-a3ac-4f19-95ae-579bfc4cf962 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.363769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.363769Z digest=sha256:5218213bc26270c3f1e6157dd61d98bb572a87b3c0014449641ca6da9275ab2c

Observation cd4b6a38-51b6-4861-8ce8-b0683ac52398 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.422198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.369179Z digest=sha256:846d429fa5463bd12381108a476b627af1e5786c52ef0a62aca4cc8c6623444a

Observation cdf3ab89-eaf0-420f-a33d-d487e4804a56 · outbound

This paper cites RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.374182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.374182Z digest=sha256:83189488d542d37ef4c60151d0a90e963f5ce14b5299175e8c9bcd6f1aa58b9a

Observation 11df0c3d-e5ee-40f6-9677-7e0438431df9 · outbound

This paper cites MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.379491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.379491Z digest=sha256:97ab4a60c1c068d7006ebe464c5a504c0a3e4a8f1409e20f3b3f687cae76b911

Observation b76ba7ae-bc15-47c3-ad44-1e5664272e99 · outbound

This paper cites Perez, S.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Perez, S

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.384702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.384702Z digest=sha256:660d11d0edc64cfaf91a0063ed0579b464f18044c1431e6d5a61279f5bee991c

Observation b0d625eb-b103-4a56-b950-355b7800b92d · outbound

This paper cites Qwen2.5 Technical Report.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Qwen2.5 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.389446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.389446Z digest=sha256:b953f7d7a4131406232f5b3b37d383a871461b60b5221840617abbf2eb6108a6

Observation 8b2ac55e-fa62-4471-a7a6-955281384036 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.394787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.394121Z digest=sha256:5dc191a2ef76d51dc8804ca303349fac0232d60db4ac90debca7d4d78c9e3f0f

Observation ce64f43f-f56f-4ceb-a41b-946b213c5fc1 · outbound

This paper cites Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.399353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.399353Z digest=sha256:b6ec154be151631fdeb4f5e2e763ae2c0d667b4accf7bc70a11acf25e9a7507f

Observation 9ac2ff3d-b2bc-488d-a9f0-6c7e6e41f189 · outbound

This paper cites Verbosity Bias in Preference Labeling by Large Language Models.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Verbosity Bias in Preference Labeling by Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.404427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.404427Z digest=sha256:faf081f9123305930a66671b86978fc5da8be2dad07d56e75bce4beccedc98df

Observation 1ffa1d05-eec1-4e60-afd8-78932d7724dc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.414460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.414460Z digest=sha256:b179a2b5520c98b132e3c9509f63a6840afcdf7b1671939d5643cbcda83d3cc4

Observation 945a2281-cd88-478f-a281-cabf37e7366b · outbound

This paper cites Sharma, M.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Sharma, M

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.378237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.419402Z digest=sha256:1c2f9a1bd31e43d60656d4040e0b7accfc722db8933f47a289127fca00203623

Observation 0af90142-28ef-42f0-af13-5cf3c4d47bfc · outbound

This paper cites Singhal, T.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Singhal, T

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.361213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.424140Z digest=sha256:676d1475dceeb06f2e4e9aa5b0267e2ca29fd43c3b83b8866dd512403f2a367c

Observation e811439f-75ea-4476-9b60-b3b631382a62 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.342362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.429286Z digest=sha256:17634a536c9d3043444323b0a51824a515d0515a59b8d306ae99e65840fdf714

Observation 22eb0913-acdf-4a0e-b0da-3ae9645d77bd · outbound

This paper cites Stiennon, L.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Stiennon, L

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:40:10.324517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.434243Z digest=sha256:223beace08f22e9c54d096f2cd4ac64fd3194dd3363b40e8166c038d55f8ca8b

Observation d0391c6e-1b24-4948-9b00-a8c2744aa454 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.306564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.439183Z digest=sha256:b80d7c8bfad602f15651bb0ad3aaf41f567e31a238b4c1387d165b81e03feb7c

Observation ede0c600-1e5e-4546-8b8e-15db8f5c538f · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.290637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.443899Z digest=sha256:df2fc272fd15697f044ccb6897c060d87140685d40ac8f2b44bfc6f1e8fe46c2

Observation 78c37b0b-b751-4d8c-812b-8f2c0eb3be84 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.274336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.448407Z digest=sha256:e6853cdde60d7b816c56a759cafda7825a29fda9323d2980535d03d7c4bec8de

Observation 88e587cd-4c3b-4ad9-93cb-8ef92bdc6926 · outbound

This paper cites von Werra, Y.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models von Werra, Y

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.453207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.453207Z digest=sha256:39b3ebf7dfb05c8d25937ecddfac7c7a42fde64081dff954fe0aed131b9bec52

Observation ceb0dfb0-e5f0-41c0-b2a2-4499233940d3 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.248070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.458205Z digest=sha256:f107d5b337707d0ac9f00e68cdc8e724a1fdb771e45752d74d986beaf050334a

Observation 81672bab-2540-4b7e-a93a-9d2bdb0a2356 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.231946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.462785Z digest=sha256:3ebb77c00609d9ff195a8518067342ad9de2e8bee7e06e7961e0b15b3aeee032

Observation 8087916d-e8cb-4716-a863-e32a161a2792 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.467576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.467576Z digest=sha256:03af75700891c205e41ff9087bcdee39d609b3164bcf9ff2b2601c159690b103

Observation 9725a20f-ecbb-4e9d-b395-b5b8eff154be · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:40:10.215955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T05:40:09.472576Z digest=sha256:a4698c678921c7ca59d726321762d6f7005f5e31ca0b8accc980fcb1bb85084c

Observation d1fea079-aa42-499f-9b4d-ffc4df35466f · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.477323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.477323Z digest=sha256:6062c37427854f9cd8c978482d6f298eeac6af8dded5dbb2113de399b31452d6

Observation e727195b-d96b-45c0-b113-21d25c98ee16 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.483189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.483189Z digest=sha256:5a7a568de2ff78cb4253fb343dcdd75609dfbf72d0115d0968bffc02d166a466

Observation 30273e3a-f987-4b7d-9c49-dd0e691a5f71 · outbound

This paper cites Qwen2 Technical Report.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Qwen2 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.488443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.488443Z digest=sha256:a7153dd58f62fedd70e2ca946ab3d58b4b6fa99c9c916e9b98dc0fc8b9e51e77

Observation 3a049343-5f77-4b12-ba85-9963fe750d34 · outbound

This paper cites Self-Rewarding Language Models.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Self-Rewarding Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.493614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.493614Z digest=sha256:5269d7d1ac94123282296edbc2fc62ae39306944226a15072fff7c777a8a95c1

Observation f871c059-471c-480a-a2df-30632b8255e1 · outbound

This paper cites Zhang, C.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Zhang, C

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.499053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.499053Z digest=sha256:2745319bcd6cbfbc4a11eb0e2aa79a4904f20a29d4e9caf8dce5c7b94543a790

Observation 16fbac64-21f1-4c7b-be01-69dcc5635eaa · outbound

This paper cites Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.503684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.503684Z digest=sha256:19a4e53a00212c085e257b30066ac4182fb1a9862e5cc4f61cd3f509ce83c5c2

Observation 91dd55db-cff6-4a6f-b309-b58a9a88d1c6 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.508737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.508737Z digest=sha256:0582e4545f7bf9ed299ce4bdb4c187d69bdc86906aed4891460a06fbc3c522fa

Observation 28dd2998-59d5-438a-8394-5222b09e945c · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Fine-Tuning Language Models from Human Preferences

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.518689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.518689Z digest=sha256:91b532837d4c6010ed6568c850b949affe5115a201ff208d0355d9d5d2202cf5

Observation 6aa3b570-ab53-40e2-bba8-cbe2d4dd9940 · outbound

This paper cites @esa (Ref.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models @esa (Ref

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.523386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.523386Z digest=sha256:ad695712ae8de3ab2e4ab2febfe28b5a365b2e1047ec89ed9b807030fc7bb5e8

Observation 725ee1b4-4440-437b-a44b-19c7a9fe4673 · outbound

This paper cites an unresolved cited work.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.528530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.528530Z digest=sha256:86118bc6dc1814b935b1f40223fb4bf16b8ff9e2d279baa16a4ce88e2714fc08

Observation 87e22da6-d4ce-4a22-b621-5a20094e4d05 · outbound

This paper cites A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.533305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.533305Z digest=sha256:188c4b453ea97bb94ca1ac18ed230e5d9186d6c250d549b0323b8c408b0e07c7

Pith citing papers

Observation 4a11841a-1d3f-494b-b92f-976e3a3625a3 · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.785000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:dbe6c2d045ec94cf60d5d74b915f29e078ecf3adba9bc37d16dbf54f0d863067

Observation 514714e3-6952-4251-a15e-d7d09cae203f · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models

Reference 222

Resolution
unresolved
no resolver link, observed 2026-07-31T23:52:09.180660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:52:09.180660Z digest=sha256:000d7728d8f9f183f9a8fac4c29fd1f4c9bf14fa07ad2b0cdd76655dfb373a2a

Observation 7112acee-2383-4d63-9ac7-fbb30715048d · inbound

Improving Generalization Robustness of Multimodal RLVR cites this paper.

Improving Generalization Robustness of Multimodal RLVR Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.049711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.049711Z digest=sha256:d31aa66947a08e4e673bf24ed529ba3092187029d80c4a5c772cf799d0968011