Pith. sign in

Paper Citation Record · LEDGER

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization

As of 23 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2505.12763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12763 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:34:33.006327Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:06.813430Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:04:08.237530Z

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved36
  • parse uncertain1
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3dd74b66-7ef7-4216-9965-3c3ee9b83011 · outbound

This paper cites This implies that240k should be in the form ofm3 for some integerm.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization This implies that240k should be in the form ofm3 for some integerm

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.566895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.855508Z digest=sha256:7ed60f39eafdbf1b680580be432b68edddfefc8ff3219274a6ceda53d76b44ed

Observation 1e2f23df-1cb7-4085-bcd7-841fd5c2f693 · outbound

This paper cites Rejected(Unalinged GPT-4) 1.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Rejected(Unalinged GPT-4) 1

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.994342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.725498Z digest=sha256:691cc3683dac7a8625c798e8b4978a9c6a612226824dc82df4534e407bfd58a0

Observation 34f32bd1-3029-449a-bce9-36d47b9bf3dd · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:32.693379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:32.693379Z digest=sha256:6af6aea63f17bb35ab7e3f718e688d0c76b0e31c6ed1148184d3716931e38bdb

Observation df890ed8-4339-4fc9-bcc3-c743a8b679ab · outbound

This paper cites RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:32.698392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:32.698392Z digest=sha256:a4f9361a886d6c7e4495c63525edf0d5bd57173296dd704a9701b3448aad3f63

Observation 154d232b-4591-4a66-825d-5bc2c356a39e · outbound

This paper cites Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:32.703118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:32.703118Z digest=sha256:ed580db1ca248b1e4619b261e4992e6d7089a69ddf24ac676b2f69d2bfd1c015

Observation 551757cb-7d42-4d87-9df2-a53eae1d4dcc · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:32.707744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:32.707744Z digest=sha256:bcff8aff8bee414d833e8e84427282f7c5a1a11a3b0ffa29c3779021a1498821

Observation f376e46e-84f6-40f2-9f86-c4455c1e8f43 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Secrets of RLHF in Large Language Models Part I: PPO

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:32.712424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:32.712424Z digest=sha256:5eaca56a0823c50d15a486afdf857e89270e6095a5b7348532268ce4d3e13243

Observation 1b4755f8-f836-45ba-a6ce-69d439f43737 · outbound

This paper cites Table 6: An example of chosen and rejected solution from RewardBench.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Table 6: An example of chosen and rejected solution from RewardBench

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.904089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.751706Z digest=sha256:c3df2ea1cc11a6e2a8f362423d90ebdb04d402de80ceaa75042a5484d1d0ad19

Observation 505c06a6-df15-4fdf-85e6-c1edca76788f · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 9

Resolution
malformed identifier
no resolver link, observed 2026-08-15T20:34:32.721081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:32.721081Z digest=sha256:95067a79aa410147d60f6c1ac7d91001c568fc7326d431c1f79b019b6db8bc59

Observation 4e2e51d5-d15b-49c5-8c03-27c148e7872a · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.982224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.729099Z digest=sha256:860858c243e99e3b3372b114c1161ab6bb34fd31c55b039c252a8995c6fbfbef

Observation 5a24994f-0bb7-4761-9b51-7bb3c5fe96d7 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.970124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.732985Z digest=sha256:882792f50e1c419e95700fab0f263bf466556f267cbaf74e93cefd0140c40b42

Observation 2629dcc6-529e-4d6e-9572-eef3a6809f1e · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.957557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.736568Z digest=sha256:0633b2dbc1c1e3b52a776f0c2529ba1ed352fb2370b206b59f94004b92536d6f

Observation d94c65db-2420-4d5d-87de-18cb3936d8c0 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.943575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.740347Z digest=sha256:332cf26478ba56650c7dbbaec24354f0aca4955f5a3108ff891cf263d978a383

Observation b907e826-497f-4ef8-bd13-2b8d8327d812 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.932442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.744056Z digest=sha256:a2093bae6d8374b981c8283305aee998f4fb457512d54ae5cbb8e35c0de4fc3e

Observation 24d136fd-2521-4b43-a4c0-1143a379696d · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.920457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.748005Z digest=sha256:54fcf4d761cc7f8bc4e9b97c3956a0de59199fb5665618eb7b73af9a4f12b1c9

Observation f735ec23-2f34-4a8d-aa21-9c0be9615ba4 · outbound

This paper cites So, 240 = 24 ×31 ×51.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization So, 240 = 24 ×31 ×51

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.888723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.755474Z digest=sha256:dceabb9d4672c32c1d4d3d2059ff1f7d92d502a81ed5d1384806283578026d7e

Observation 855849ce-812e-4cf3-8aa0-8e723ffbd518 · outbound

This paper cites For31, add 2 (1 + 2 = 3).

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization For31, add 2 (1 + 2 = 3)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.874841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.759018Z digest=sha256:54168e84b957c165e549618d42a27864c2e88145a87f9feaf3e79f98af77f7be

Observation fc9b169c-68ef-4c55-bfed-6b2e908e2f60 · outbound

This paper cites Chosen Detailed 1.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Chosen Detailed 1

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.861146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.762536Z digest=sha256:67fe7e4213a6ac9c963903e8906236a6c05977de52a916c0ce7538dab5032439

Observation de2ccbe1-f7b1-4e80-b991-2c04a01643fe · outbound

This paper cites We do this by progressively dividing by the smallest prime numbers.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization We do this by progressively dividing by the smallest prime numbers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.849579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.766805Z digest=sha256:1197b138b65abba3b8803f9373de3b9db85c7e4b4988de3cb1d1b2c8cef895d3

Observation 6f970ad8-fe7d-407a-a238-b761a00a16ee · outbound

This paper cites 120 is even, so divide by 2:120÷2 = 60.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization 120 is even, so divide by 2:120÷2 = 60

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.838370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.770722Z digest=sha256:9183f7eca53c0612297a4d398aa9ae784e6019da83bc3c116d4eac6afd902d20

Observation ea9fa6f1-ffa3-4d88-96f5-7b4f0bb4a3f8 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.825971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.774746Z digest=sha256:862d65718901595424dc60a68958ca449b63428345fbf980f06c5843853ae562

Observation 3b64f641-fcf1-4b97-92e5-d5a1aaede84a · outbound

This paper cites Here, a = 2 (since 4 + 2 = 6).

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Here, a = 2 (since 4 + 2 = 6)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.814134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.778491Z digest=sha256:088d24a3dd8307c7146520a8c65af7d8f1298f7e10b73ce4d8ee4d7efd6d6663

Observation c4d75061-7ad2-48cb-8054-739c4c7d93fc · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.801610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.781956Z digest=sha256:6c3f4749ac35fbd194dbc7a29f590b60fc3d5ee8e07215c75eccd714d08aab21

Observation 2b795c59-fa15-4a91-ab2f-01c7c1277ef9 · outbound

This paper cites Thus,k = 4×9×25.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Thus,k = 4×9×25

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.789886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.785906Z digest=sha256:020f958b309f7ba3acb2e5532a9d2f9624f87645fab01a14055bc8df8c1fd504

Observation 4912dc05-fe63-4f88-8f0e-31229895b0af · outbound

This paper cites Then,36×25 = 900.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Then,36×25 = 900

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.777421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.790811Z digest=sha256:564bead481508db2c64b2eec2f374ad4457635334623066283c67a0a758a93e7

Observation 7847273f-5578-4459-bca7-284d422d287d · outbound

This paper cites To solve the problem, we need to find the smallest positive integerk such that the product 240k is a perfect cube.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization To solve the problem, we need to find the smallest positive integerk such that the product 240k is a perfect cube

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.765399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.794952Z digest=sha256:0cb2c459671dc473b361c0c189260bc4b52f0f46eae5bc8b6b9cacd0beaa4194

Observation 75a3dfad-5f39-4cfd-a7b0-a643be3aa5e6 · outbound

This paper cites We do this by progressively dividing by the smallest prime numbers.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization We do this by progressively dividing by the smallest prime numbers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.753614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.798793Z digest=sha256:ce367a1de8f95a40970bf7a99d2c18a3cb0162393a7bd181795720f46501748a

Observation 31d78d8f-55c3-4979-ac13-cba078e7aec4 · outbound

This paper cites - For24, we need4 +a ≡0 (mod 3).

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization - For24, we need4 +a ≡0 (mod 3)

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.741544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.802559Z digest=sha256:a13cdbee6cab1015205b769bb8d998d83fd8f3fbf38fcab9a76a2e2a34aeb8ec

Observation 1e25aef0-1c1f-46c0-b7ec-38956c57f036 · outbound

This paper cites Problem The product of integers 240 andk is a perfect cube.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Problem The product of integers 240 andk is a perfect cube

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.729112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.806224Z digest=sha256:94dd7558d33f4e86894c9df3d817bbfad54f8e538f81231acd29e3d05aa42287

Observation 5f52b9ae-297a-47ff-b7cf-03f622e81c81 · outbound

This paper cites -51 needs two more factors of5(52) to reach53.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization -51 needs two more factors of5(52) to reach53

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.714769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.810182Z digest=sha256:f0785c67bd986089033d2f15be509744521d6ab56c5502757ac4b839ab0ff830

Observation ae7be370-cd0d-4c2b-96f0-6945423a9e95 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.701140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.813850Z digest=sha256:55fadb455b6bfc4186f3469ec4ae8393fa49995c26cb0c3a5b052cd973a04684

Observation 99191700-7604-4364-9a83-ab726bd6a51f · outbound

This paper cites Rejected Detailed 1.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Rejected Detailed 1

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.688399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.817564Z digest=sha256:6b63d3ac76c5f8b3a2aff8412629b39437e05a54261cee2bc5a32e9e5b1069ee

Observation 80fca837-f18f-4228-a09c-3bdd673aa0cb · outbound

This paper cites Each prime power in the factorization of a number that forms a perfect cube should be a multiple of 3.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Each prime power in the factorization of a number that forms a perfect cube should be a multiple of 3

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.675697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.821763Z digest=sha256:8652ccec773150a2b0249d76420d15359252190e55e0b1462ca9171f4128f124

Observation 0e7054af-0219-4777-9c33-38c5c94ec6b7 · outbound

This paper cites For24, we need at least25 to have a power that is a multiple of 3.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization For24, we need at least25 to have a power that is a multiple of 3

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.663641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.825611Z digest=sha256:dd62610b6100aa26671f0a6611b2a7803933966a399be0b0c95b0ab09164d19b

Observation 84abcaca-dc95-4aaf-86f7-2a44634b76ec · outbound

This paper cites Calculate each component:21 = 2,32 = 9,52 = 25.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Calculate each component:21 = 2,32 = 9,52 = 25

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.651757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.829139Z digest=sha256:40f24570b90794086d670e65ccc260edd71f0f504a636fae1e50dc65c1b98de6

Observation 9d3876c3-af4f-4d1f-9446-2ad716c05573 · outbound

This paper cites Markdown Format 1.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Markdown Format 1

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.638195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.833521Z digest=sha256:e4cd689752ff0ecc2d654f8b6ccbc9172524bb39aef9c109021eda82090697ad

Observation 6e09b614-765c-4a2e-a9bf-b124d98b34ec · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.625716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.837324Z digest=sha256:0b96254847738c2f85c1cc2b682e980ae2a5f24536b77ccd56b4b702c2652055

Observation 10994be8-6693-4517-b3cb-41195dced9ab · outbound

This paper cites - For24, we need at least25 to have a power that is a multiple of 3.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization - For24, we need at least25 to have a power that is a multiple of 3

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.612811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.840890Z digest=sha256:85475960e7aa03cfd312a114cbe6cc3ee03a138a63a0c46cc1381335903de082

Observation dbbac7d5-29c2-4176-8260-1f103a2e06d7 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.601064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.844717Z digest=sha256:233f44caaa7123ad5515a4c2b132a5d3502941897b1c6ad3a219af87010af9fb

Observation 23d50643-8f09-4019-a96a-aadec19cf07f · outbound

This paper cites Problem The product of integers 240 andk is a perfect cube.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Problem The product of integers 240 andk is a perfect cube

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.590047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.848223Z digest=sha256:42d99701718be20c6da8f6141125eb81004fa3e4045318e4e4da1193912d8773

Observation 5ff82e89-91af-42e8-b0d7-2bc5c789dbea · outbound

This paper cites GPT-3.5-turbo- 0125.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization GPT-3.5-turbo- 0125

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.578491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.851609Z digest=sha256:9c7d82cc63d98d0b92d4a97161e46deef78f4f07d2bc97b461914fa024cd6568

Observation 5f1b8d29-e7ac-44d2-9284-fa65c093df67 · outbound

This paper cites To make240k a perfect cube, we need to add multiples of 2, 3, and 5 to make the powers of all prime factors multiples of 3.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization To make240k a perfect cube, we need to add multiples of 2, 3, and 5 to make the powers of all prime factors multiples of 3

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.554566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.859783Z digest=sha256:a976e652eddd4707d6252965e2a7779e5475750680c5e97b95501fa2e25cb4c3

Observation 4d997e61-e49d-4072-a64c-52b7933b1309 · outbound

This paper cites Chosen(Random) Llama3-70B- Instruct.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Chosen(Random) Llama3-70B- Instruct

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.542122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.863384Z digest=sha256:d395519a2462522e31861fc3c2b5058fddda0a793d04492f4aa6bdccbb7b0c2a

Observation 4bd195e9-77b5-4eef-9332-8263de86b118 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.529873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.867343Z digest=sha256:a0a0e40dbd8a4ce5d1e3f55492aac4b1a2d7db2abf3c3763d78e4300965244cb

Observation 8b2e8ed3-a837-4499-ae81-b63a37a81711 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.518831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.871074Z digest=sha256:0422dc28a0a4342a2053b48e4cdc08a338d191e8e910e2eb846bac0da3d3aec0

Observation 391bed6f-1f80-40c5-af4b-60fac25cd7db · outbound

This paper cites 4.Therefore, the smallest possible positive value ofk is 900.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization 4.Therefore, the smallest possible positive value ofk is 900

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.506949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.874743Z digest=sha256:6a519b5566fd81210a512d439f565efe652f1fb52511540eebf8bcf72b8fe956

Observation de89d9d5-268f-4704-9dee-9f070e3468b1 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.494883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.879094Z digest=sha256:c51f526bffa603bab670479aedbfc5350de9f99de52e8b3b96e3dffae8a65613

Observation cddb7c31-da16-497a-b6cb-49aeec47f5e7 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.482096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.883544Z digest=sha256:064a40d6d41ea6565e096a8a6966181736e4a8bf6e5ff3113441fbd3c7000055

Observation fb113322-4031-4806-bb5e-baca09bc4945 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.470101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.887816Z digest=sha256:43c64a61ac2f44e2316ad152ef515906d47aecf4c033bb3b268b51f1bc196ccd

Observation caaab7bb-ad4a-42f6-a51b-de73af03b9ff · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 53

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T20:34:33.458329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.891782Z digest=sha256:768881d0707d9dfe7e0e3ff1ca9aebd3fe2d6c6ed2758acf0696b1ded9a26139

Observation f41d9989-5ce4-4bd4-ad1e-a2f337ad80c9 · outbound

This paper cites GPT-3.5-turbo- 0125.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization GPT-3.5-turbo- 0125

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.446079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.895909Z digest=sha256:1a0eceac30e2a78e80e822b926cc34458d56685895fde103d47542a7ef6af7f7

Observation 6c455fbb-dbc3-4411-a169-3910a9677535 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.435644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.900746Z digest=sha256:9a0ea33538f0c8d3cee69d9a07dd949211926715d6f26c7de0394cd3b72cdc55

Observation 8cbff164-8eb0-4d92-9627-080090220899 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.423898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.905147Z digest=sha256:98c60423c6bff309cae80853986dd724fb678a312a722126ebd86aa2d917b3bd

Observation 31faad9f-7f9b-49dc-890c-a33b4411652a · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.411527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.909464Z digest=sha256:29faae50e74dbaec4c003f39cd724bd5a1c9b1c10cbafca748e8dbbd18449938

Observation 126a26e4-9f3a-4046-90a6-d31edcf8682b · outbound

This paper cites Rejected(Random) Gemma2-27B-it 1.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Rejected(Random) Gemma2-27B-it 1

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.399390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.913567Z digest=sha256:2111d35ef3d09475808e27ef85bb444d0869aa6753be2c640944e1c29fe1f00d

Observation b5f1df76-756c-43fd-93c6-1cc808406dbc · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.386522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.917624Z digest=sha256:962401d912423addbc27d8305ce9ba45faa3059c27db4bc02efb6d3ccacfc91b

Observation 6fb6afe6-4da2-4e71-bc0d-88a3dded69b1 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.373945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.921670Z digest=sha256:a4435c6bf325f34834f6e9cf2d51a2e6cdd13a326623e0a80988346e6936a0a5

Observation 69560c16-f8b8-49da-8fe5-45c43ac47b66 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.362211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.925415Z digest=sha256:4f80cef1324957501c4b6aecc9ed77607449cadf2aef15ad69c7ea872b396c30

Observation 206ecfbb-ed79-4635-ba48-0ff05e28bca1 · outbound

This paper cites WizardMath-7B- v1.1.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization WizardMath-7B- v1.1

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.350361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.929376Z digest=sha256:d26ce11ecefd7c0bff8c6c488925c8a494639610281288cbc3ad78157a9705d3

Observation 4720a234-aa17-4da8-a114-9f83e9af44e9 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.338614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.933230Z digest=sha256:443595f82934ad717422d949dc6c676c84cc54618870455191205bc038ff162b

Observation ee2c1def-7bbd-437a-a671-07bb4c30ce04 · outbound

This paper cites First, we need to find the prime factorization of 240.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization First, we need to find the prime factorization of 240

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.327564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.936934Z digest=sha256:91ef609685862b59baeb1869e0a93a6107d6c6a1c40ec900f16084ed563370c5

Observation da0ca5b2-0dcf-4156-a171-5713dae68018 · outbound

This paper cites A perfect cube is a number that can be expressed as the cube of an integer.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization A perfect cube is a number that can be expressed as the cube of an integer

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.316094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.940984Z digest=sha256:a0dc7e954545aa4584e826d34cc03d48ce36319ce3919837ea31d242f2fee187

Observation f3860637-8c7e-45bd-8ac2-d152c343f50a · outbound

This paper cites We need to find the smallestk such that the product of 240 andk is a perfect cube.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization We need to find the smallestk such that the product of 240 andk is a perfect cube

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.304030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.944569Z digest=sha256:58da44f0b65c742b54cc1307150c1b804a414cfa77b1b5ac4fe87c35108e59c7

Observation 63417d51-e455-4bb3-8782-69dc2ad548a0 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.291832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.948023Z digest=sha256:57a74d787045cb07b0736c8007e8f50dfc7bcc25114baebb4d140bb46f111c96

Observation df20eb16-ad6d-471f-835d-82f63c05d870 · outbound

This paper cites The smallest possible values fora,b, andc that make3a +b +c divisible by 3 area = 0,b = 1, andc = 0.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization The smallest possible values fora,b, andc that make3a +b +c divisible by 3 area = 0,b = 1, andc = 0

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.281117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.951837Z digest=sha256:4c9db5b948ce08ea41ed9c2a77ade4ffd18f4a5c4fb01bb3f03c0f9307cb6995

Observation 5604e2a9-70d6-40e1-9206-98680906ae84 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.269971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.955599Z digest=sha256:73c1b47ee99ec10750df60ef4fbeea98af5df1e50bbc06c6cee2ed306567c05a

Observation f7dfcf50-6a95-4b0c-8b57-b7b74a19b52d · outbound

This paper cites Table 9: An example of chosen (GPT-4* and Random, 3) and rejected (Random, 3) solution.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Table 9: An example of chosen (GPT-4* and Random, 3) and rejected (Random, 3) solution

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T20:34:33.257901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.958937Z digest=sha256:576a93cdb0228dc9693104cbcacdb6a58e9da89946a28560cedf474806d88e00

Observation 79d6aaae-668a-460b-9d19-bf2b57dff193 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.246855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.962617Z digest=sha256:3ac596cb3b00b6dd9104bd282623080a16a2c547da9c77ba110efe1e8411db71

Observation 94b05ac8-9864-4a51-b11c-54eb7f0780c2 · outbound

This paper cites overrefusal.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization overrefusal

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.236642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.966043Z digest=sha256:64845b5baa3126f018cbc0443da62711fa72c40f9f999672c075a565d1482a00

Observation 447548d9-cb33-464a-9219-90a04e58592d · outbound

This paper cites - There are no unnecessary disclaimers or limitations.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization - There are no unnecessary disclaimers or limitations

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.226022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.969947Z digest=sha256:ed6c217ad4d17b570d19ea21f36ecc1bd0c90acf5cc764c71d8baf81b63e9a9f

Observation 3c0d095a-885f-4def-9ac8-0ce67e76b543 · outbound

This paper cites - Some answer is provided, yet it is needlessly restricted or accompanied by refusal-like statements.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization - Some answer is provided, yet it is needlessly restricted or accompanied by refusal-like statements

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.214589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.973909Z digest=sha256:cb7cd0cca98c1cd632530bf6bf729b97d7941dae6b9762fd7510f17cc2d364bd

Observation 6fdf8bca-12bd-42ec-9c13-e9e06e893c50 · outbound

This paper cites - The refusal is not aligned with policy since the prompt is actually allowed.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization - The refusal is not aligned with policy since the prompt is actually allowed

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.201653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.977989Z digest=sha256:85de528e43c8ba4ec143198c369d014dd24502534509e313a39390090b193426

Observation 0369e169-7881-4b11-b0ff-f1f1d779f04f · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.188623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.982075Z digest=sha256:cadde42afb56d5907b927e82b41338c12370c5716b32b105cc22f8eb5353c8ec

Observation 8f1f476a-48c4-462d-a63c-1d600bbffd93 · outbound

This paper cites an unresolved cited work.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:33.176364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.986126Z digest=sha256:72d269bdb97ed8aa6c480cc414fc44dd57489ad95786188d6dd833b8b3ef939e

Observation 49ac354b-0334-407b-bc20-488365de4c3f · outbound

This paper cites Your task is to determine how the response handles the harmful prompt.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Your task is to determine how the response handles the harmful prompt

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.163252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.990717Z digest=sha256:33e3522a9e11a8b06c8863293c45718d0d287146bbececd876244438f1918a20

Observation 99194f9a-9080-4553-982a-6f49fadbc8b4 · outbound

This paper cites - There is no language indicating refusal.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization - There is no language indicating refusal

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.151992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.994458Z digest=sha256:65e0177a508063c0f5404235ed91befbab404d82737ccca837d5598769b90b49

Observation 378be4bb-433c-407e-89c2-74b42b6e527b · outbound

This paper cites - However, it may be incomplete or include disclaimers, indicating it is not fully complying.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization - However, it may be incomplete or include disclaimers, indicating it is not fully complying

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.141121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:32.998114Z digest=sha256:2d01cd8af844efa5a9d2813d3f7c70cd09ecfce889401530158e6c42254750ab

Observation 996f3d35-73a0-4055-a4d3-865200753fa0 · outbound

This paper cites - It does not fully comply by giving the requested harmful information.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization - It does not fully comply by giving the requested harmful information

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.129658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:33.002430Z digest=sha256:a46d780b4ec5cf5070832d12f3b481b7c602f8471cba6bd923bd72c99a74c7b4

Observation 68a8f1e9-f7f8-4720-bb5d-ea0ebf970520 · outbound

This paper cites - It does not provide information or instructions that fulfill the harmful request.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization - It does not provide information or instructions that fulfill the harmful request

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:33.117867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:34:33.006327Z digest=sha256:1b4094253de0025afa857e623cd77d331a8e25a703ec62a7761b3c0d7a143f6d

Observation 07251166-3685-4dca-8b10-33a14f639ed7 · outbound

This paper cites In gen- eral, the reward model is obtained by replacing the final output layer of causal language model with a linear head to predict a scalar.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization In gen- eral, the reward model is obtained by replacing the final output layer of causal language model with a linear head to predict a scalar

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:32.716697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:32.716697Z digest=sha256:bddbd3030940ed65c2523898541abbb8d6da882063b76e93ac46a5de0e72fa2b

Observation 604f7e18-35cc-4fd8-92d9-7c37da7c66a2 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:32.688492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:32.688492Z digest=sha256:0deefcdfa04ce362d017f3e9efeda7587f1e7690b455b6376f406eb228bc9e49

Observation f6a8fccf-9412-48af-956e-32eb4e0390cb · outbound

This paper cites Critique-out-Loud Reward Models.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Critique-out-Loud Reward Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:32.682880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:32.682880Z digest=sha256:49d8bba14fd8bd7afd1e17d38377d52a77d1c09c4938e8536d74ebdea481823a

Pith citing papers

Observation 4de5ecbc-5476-40f2-ae93-403b085cc640 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:04:08.243595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:04:06.813430Z digest=sha256:628934b56358ffb424d7d20a80f39d2b1e893d321f4d8a19910510f7d5d8e1dd