Pith. sign in

Paper Citation Record · LEDGER

RewardAnything: Generalizable Principle-Following Reward Models

As of 15 August 2026, this Paper Citation Record lists 100 of 112 outbound references and 8 inbound Pith citation observations for arXiv:2506.03637.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03637 v2

Coverage vector

measured 100 of 112 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:07.168303Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:46.427272Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 112 outbound references displayed

  • verified exact5
  • verified fuzzy2
  • unresolved93
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9b279bdf-f850-4297-90fd-6741b5310583 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

RewardAnything: Generalizable Principle-Following Reward Models Fine-Tuning Language Models from Human Preferences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.668341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.668341Z digest=sha256:fefd741cad805f65445fe0da7192ba49d80f8c05a5461a43ffb36b068a298552

Observation 56102ffb-be24-4cca-8208-538f0502a5a6 · outbound

This paper cites Training language models to follow instructions with human feedback,.

RewardAnything: Generalizable Principle-Following Reward Models Training language models to follow instructions with human feedback,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.674152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.674152Z digest=sha256:a025f498ccd6c2f5a4595db6fc9be4dd415b4aa030e4ed07bebbc4e35b090e19

Observation f72af75e-4930-41bf-851a-696ec6e5388e · outbound

This paper cites Deep reinforcement learning from human preferences,.

RewardAnything: Generalizable Principle-Following Reward Models Deep reinforcement learning from human preferences,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.679780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.679780Z digest=sha256:409241ff74c584045853d03a01ca059b7411a771ab49bdc9ede5f5506ba56d79

Observation 85a9511f-355b-47bb-aee9-2d38ad65ee30 · outbound

This paper cites Learning to summarize with human feedback,.

RewardAnything: Generalizable Principle-Following Reward Models Learning to summarize with human feedback,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.684775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.684775Z digest=sha256:ff8d0dc797b86995939865c964c93fbc01090c3fc59409320f472138c2a1f82e

Observation 9312a7b4-5134-49da-aaa8-dd7b93a8885d · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

RewardAnything: Generalizable Principle-Following Reward Models A General Language Assistant as a Laboratory for Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.689596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.689596Z digest=sha256:46999d0ee862677dac46603b933e97aac3880f3f3c74bd11ba5a54c772fba48a

Observation 882b93d0-697f-4cf3-b839-d141d3bdef4e · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

RewardAnything: Generalizable Principle-Following Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.694852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.694852Z digest=sha256:8c65642b4f379eecb36f29f471666aa332872a6038dc7c163dba941f280c9874

Observation 6f079aa2-70ae-4aeb-8524-49e57e6fd4b7 · outbound

This paper cites RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style.

RewardAnything: Generalizable Principle-Following Reward Models RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.700804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.700804Z digest=sha256:b649284f67811c91678e3e652e588e208be695e6ab5bcf02d1317a1e1bbd76bb

Observation 4eab97cf-aac4-410b-a980-cb5daf48fca0 · outbound

This paper cites A survey of reinforcement learning from human feedback,.

RewardAnything: Generalizable Principle-Following Reward Models A survey of reinforcement learning from human feedback,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.705995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.705995Z digest=sha256:f7c456388daf66de30ad79bcb37a3ecdb32a2188e0c16a768922cab0849a9d89

Observation ba520f4b-043f-456e-a640-7cc474ca65e9 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

RewardAnything: Generalizable Principle-Following Reward Models Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.710698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.710698Z digest=sha256:7b4c7bc86893f09570e248daec2c8e6789db1e31a0f06aec3894acdbbc257a0e

Observation ddcae839-bad5-4a40-90aa-3b24b2270704 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

RewardAnything: Generalizable Principle-Following Reward Models Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.715361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.715361Z digest=sha256:a34e58f89d3c7aa7f0beff602c1ebced1fb95ecf92a0cbbe88d51b3ea65fef0e

Observation 9e02376c-6f72-43ad-90fd-b9cf08e846b3 · outbound

This paper cites Evaluating Large Language Models at Evaluating Instruction Following.

RewardAnything: Generalizable Principle-Following Reward Models Evaluating Large Language Models at Evaluating Instruction Following

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.720772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.720772Z digest=sha256:8af8aa28a4983d9c788fe65a12c5e0f52a0894c4a3b4db4251ba36194c5e0f7b

Observation 647a9c0a-a442-4d10-ba1b-77b7e8c222b5 · outbound

This paper cites KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models.

RewardAnything: Generalizable Principle-Following Reward Models KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.725726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.725726Z digest=sha256:c875f27190f7c27b284d28b92354138496e4355c09f5fe36e39e8c2aa8976671

Observation 5e54f08c-bccc-4bfa-94f7-a66e374e2086 · outbound

This paper cites Helpsteer2: Open-source dataset for training top-performing reward models,.

RewardAnything: Generalizable Principle-Following Reward Models Helpsteer2: Open-source dataset for training top-performing reward models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.731714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.731714Z digest=sha256:9b79e14c7869ac7593b66984df2e9f506cfeda0872d2f306b0bc2f3c393ce608

Observation 31155be0-5e2c-4c8f-baf2-db78c69e4cb1 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback,.

RewardAnything: Generalizable Principle-Following Reward Models Alpacafarm: A simulation framework for methods that learn from human feedback,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.736332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.736332Z digest=sha256:7557b7270b876515bf8c11795fe459864c735558c4bef813320e779bfec82c06

Observation 95a650e7-7f8c-4641-ae0c-28e9bd3de960 · outbound

This paper cites Let’s verify step by step,.

RewardAnything: Generalizable Principle-Following Reward Models Let’s verify step by step,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.741111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.741111Z digest=sha256:e9a1116631a12387089dcea00c11aed21e39d60954abe295046e24b1ef6e4e00

Observation a61d097b-10a7-4888-b9bb-c5df544d0818 · outbound

This paper cites Rank analysis of incomplete block designs: I. the method of paired comparisons,.

RewardAnything: Generalizable Principle-Following Reward Models Rank analysis of incomplete block designs: I. the method of paired comparisons,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.745739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.745739Z digest=sha256:b0bffebd8c845794e5a63ae58a9ee9dbc06d811c348911efecacd41ad6f7f8e2

Observation 897d129d-c17a-4482-b52e-b6f713c33977 · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

RewardAnything: Generalizable Principle-Following Reward Models PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.750627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.750627Z digest=sha256:ce61d4aca5281c489d8901a792852e096dd4eda5db1dae1cb294f125e5eca744

Observation 7b388b80-12ab-40f5-bf29-edf188dc48eb · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models,.

RewardAnything: Generalizable Principle-Following Reward Models Prometheus: Inducing fine-grained evaluation capability in language models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.755519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.755519Z digest=sha256:15562cf42731b7d6d8fb47c6888d57b407e16358881385239ef280c700d7ce4c

Observation 35da3c3b-5ae6-42f6-ae7e-6063175a4487 · outbound

This paper cites Understanding dataset difficulty with v-usable information,.

RewardAnything: Generalizable Principle-Following Reward Models Understanding dataset difficulty with v-usable information,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.760307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.760307Z digest=sha256:7c92b0c5ec9d02e998db269b454c85aa9167375c799a872f8eb2746ebf2e6007

Observation f3b495a0-ac32-4338-a9eb-9404eed09896 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

RewardAnything: Generalizable Principle-Following Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.764837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.764837Z digest=sha256:53de3e597e25227e0727bc8215f01e3000942ce3d2f5f54fac5de3b7a91f136d

Observation 284e87a8-4613-42db-9312-a2f323338b5d · outbound

This paper cites How to Evaluate Reward Models for RLHF.

RewardAnything: Generalizable Principle-Following Reward Models How to Evaluate Reward Models for RLHF

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.769759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.769759Z digest=sha256:5bfce806b436c0be6cf84732e7e1060d2afd2b7d720b8e82803ae889e49b3a00

Observation e30463e9-60e3-492e-8b33-8d61471e0f3a · outbound

This paper cites MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment.

RewardAnything: Generalizable Principle-Following Reward Models MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.774840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.774840Z digest=sha256:fe77c5d4a1bab66fe0391d208c4c2390bd020596538f4a007a6a9e2b373857b0

Observation 1cbe3cfc-5f4c-4f0c-b08c-fceb2f81098c · outbound

This paper cites SALMON: Self-Alignment with Instructable Reward Models.

RewardAnything: Generalizable Principle-Following Reward Models SALMON: Self-Alignment with Instructable Reward Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.780144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.780144Z digest=sha256:ba8999fe7f15ff651ed16fed7704812d8ba455a1af834e01e19b7eb6a567f0c1

Observation 06bf0455-4fe3-404f-9c02-5543eb4f7fd5 · outbound

This paper cites Inference-time scaling for generalist reward modeling,.

RewardAnything: Generalizable Principle-Following Reward Models Inference-time scaling for generalist reward modeling,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.785130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.785130Z digest=sha256:3e4426d3c101a6e132fb08f152924f8b9116ae8d94c3616040e96fe10573d52d

Observation 6846a318-8e26-4a35-a5f3-c3e54524ba47 · outbound

This paper cites Rm-r1: Reward modeling as reasoning,.

RewardAnything: Generalizable Principle-Following Reward Models Rm-r1: Reward modeling as reasoning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.789757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.789757Z digest=sha256:408f9dbb158f3b5d47752bda73e5266d15a5a746ef709ec34580f80c2cc387d5

Observation 385a7958-a8a7-4e24-8094-3a29b2cc48a5 · outbound

This paper cites Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems.

RewardAnything: Generalizable Principle-Following Reward Models Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.794352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.794352Z digest=sha256:76eb7b4759d1cd4a59fe9fda8c7ee169f99c6b475ec7d17047cf50986b4b0a42

Observation 7c913abf-9885-4020-8c26-9ea29f1c011b · outbound

This paper cites Proximal Policy Optimization Algorithms.

RewardAnything: Generalizable Principle-Following Reward Models Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.799281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.799281Z digest=sha256:ed12e105b7cc11f8e0f59ac1c321f2cd64b09a160a7a13c7ae79eab7bb7209a5

Observation 585bd097-d7a4-4279-a548-f10e14c4d904 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RewardAnything: Generalizable Principle-Following Reward Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.804029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.804029Z digest=sha256:3df589408fc14080c7d403a0e7a8ee5449ce640bf60aef0ae35111919cce3edb

Observation 5e9c6673-1a60-41f9-af83-2bb8bfbdd2bc · outbound

This paper cites What makes a reward model a good teacher? an optimization perspective,.

RewardAnything: Generalizable Principle-Following Reward Models What makes a reward model a good teacher? an optimization perspective,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.808458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.808458Z digest=sha256:b78aabed59126b9cd6203d133d1551218f39f45a226e21d22fb524599e1e90a2

Observation 4de5ecbc-5476-40f2-ae93-403b085cc640 · outbound

This paper cites Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization.

RewardAnything: Generalizable Principle-Following Reward Models Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:04:08.243595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:04:06.813430Z digest=sha256:ffd32d7c0cf982cfcca678d3a2075bdd0859dea9aef46b373c1ce23949f2725e

Observation a27bbd21-8242-4092-b6a8-dd026c14e7e7 · outbound

This paper cites OffsetBias: Leveraging Debiased Data for Tuning Evaluators.

RewardAnything: Generalizable Principle-Following Reward Models OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.819184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.819184Z digest=sha256:dd994ac6e030d626e34268ce2b593dc3da189b3aa13b29651e3599084dbac292

Observation eed00d0e-6f1d-4820-a85f-fc1adb3eada8 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

RewardAnything: Generalizable Principle-Following Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.823967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.823967Z digest=sha256:63bb62db5afed999f71cdf389d624eb37e0b3ddac058b1366ec0b6ee6865bf19

Observation ff7853db-ca4b-4728-a698-24b306bfa837 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

RewardAnything: Generalizable Principle-Following Reward Models XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.829011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.829011Z digest=sha256:377896e704be5b7f1668cb0b2983e9a850af6196fa3834c5e3c434693621d3e9

Observation 3caeb390-76d5-4292-abbc-7d1102c841bb · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

RewardAnything: Generalizable Principle-Following Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.833906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.833906Z digest=sha256:6dd3249190f7481b554f2039ad06ae5235aea2bf8a8d10e7268fa70eb321388d

Observation c7c2a000-22cd-47a4-8c5e-76e767e818af · outbound

This paper cites Transforming and Combining Rewards for Aligning Large Language Models.

RewardAnything: Generalizable Principle-Following Reward Models Transforming and Combining Rewards for Aligning Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.839256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.839256Z digest=sha256:e4eb4539c465d5933463f7629c255cfd4aaad66c5747d94ac257cc07d73146db

Observation 1f4e3b2f-1766-4840-8825-57e7aacc2330 · outbound

This paper cites Heimdall: test-time scaling on the generative verification.

RewardAnything: Generalizable Principle-Following Reward Models Heimdall: test-time scaling on the generative verification

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.844166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.844166Z digest=sha256:385201e80007cb13086aa7e3a9b5acff4840497c9f8471408aee6164475bdf6e

Observation a9c81cf0-383d-47b2-8810-b36116b58808 · outbound

This paper cites When to solve, when to verify: Compute-optimal problem solving and generative verification for llm reasoning,.

RewardAnything: Generalizable Principle-Following Reward Models When to solve, when to verify: Compute-optimal problem solving and generative verification for llm reasoning,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.849638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.849638Z digest=sha256:233cdbd0f2199623bb615ebfbcab6418912a3c060830e8d9e2635304efc7ecef

Observation d4655333-ac14-4745-9507-ef5406937add · outbound

This paper cites Large Language Models are Better Reasoners with Self-Verification.

RewardAnything: Generalizable Principle-Following Reward Models Large Language Models are Better Reasoners with Self-Verification

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.854412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.854412Z digest=sha256:d090d06218efc053fc989fddea65b0c9dd38aa58ceec2aad29c28d62546c2f3a

Observation 20a05919-d3b7-4492-afd8-954af94a293b · outbound

This paper cites GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning.

RewardAnything: Generalizable Principle-Following Reward Models GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.860706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.860706Z digest=sha256:ab43528ca25deccfa21bc8253cc7024fc47d5f4384745d59c8872ab4ff224877

Observation 67a8321e-6415-4aa3-bdc3-7eefe6eaae70 · outbound

This paper cites Dynamic Multi-Reward Weighting for Multi-Style Controllable Generation.

RewardAnything: Generalizable Principle-Following Reward Models Dynamic Multi-Reward Weighting for Multi-Style Controllable Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.866772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.866772Z digest=sha256:48ecf138fdf44059f839242812dda509e950c5f4cb873f3f2435b1775e4edd46

Observation 036900c6-a90b-497d-a489-085704e09552 · outbound

This paper cites Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment.

RewardAnything: Generalizable Principle-Following Reward Models Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.872561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.872561Z digest=sha256:0d2982d2a5cd61ef014acabf4d290b3400fd16997ce323c28544842cdd8182f2

Observation 3838075c-3529-4139-a0c1-84cdfd69ff06 · outbound

This paper cites Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning,.

RewardAnything: Generalizable Principle-Following Reward Models Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.878071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.878071Z digest=sha256:6dead68c644882e444d790da4e9859ff76c9990376192bc4c105570a52aabb7b

Observation d1464911-8e0c-4a5c-9e45-9069c5ea5cba · outbound

This paper cites Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation.

RewardAnything: Generalizable Principle-Following Reward Models Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.882676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.882676Z digest=sha256:c0e724d4bb6702152ffbc2bb649c16fef157ecdd5d748082566b70ffaeb95eb0

Observation 85489b8f-af8a-4d90-a8ff-eac910f95769 · outbound

This paper cites An Empirical Analysis of Uncertainty in Large Language Model Evaluations.

RewardAnything: Generalizable Principle-Following Reward Models An Empirical Analysis of Uncertainty in Large Language Model Evaluations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.887470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.887470Z digest=sha256:525e1005cd98ded3995b77669a6a49eeb1d2399f6682ba8a8b35ef449bd9429a

Observation 9a17a9eb-7e34-45b2-871d-73a3db0b36cf · outbound

This paper cites Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models.

RewardAnything: Generalizable Principle-Following Reward Models Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.892322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.892322Z digest=sha256:fe48bb1bf728c055738a724208a76860f46b730cc40455ca7dad710231350d47

Observation 5c3df9b4-277a-4135-b4c4-75ffb9785030 · outbound

This paper cites Reward Shaping to Mitigate Reward Hacking in RLHF.

RewardAnything: Generalizable Principle-Following Reward Models Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.897214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.897214Z digest=sha256:a601878aeef613f1f8cae2e8324e56d00077730faad9a707ae39f2606727db0a

Observation 726b7ead-80f1-4df5-a9f2-480b78a979cd · outbound

This paper cites Reward Model Ensembles Help Mitigate Overoptimization.

RewardAnything: Generalizable Principle-Following Reward Models Reward Model Ensembles Help Mitigate Overoptimization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.902238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.902238Z digest=sha256:91f24dc6adb602770d0da2172180188b98bd5262f61e7225924ba594a851f15e

Observation 8b1f6274-a680-4173-82d8-64c2de1361b1 · outbound

This paper cites Language models are few-shot learners,.

RewardAnything: Generalizable Principle-Following Reward Models Language models are few-shot learners,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.906934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.906934Z digest=sha256:11a94650e4966a923c043c5e84d4986bd73a4c820577a2dfe46119ce6d3a1e2b

Observation fda1a968-0a31-4e89-a696-40536e1fd5b8 · outbound

This paper cites PaLM 2 Technical Report.

RewardAnything: Generalizable Principle-Following Reward Models PaLM 2 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.911443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.911443Z digest=sha256:d94aaaaffbc227fce9234bffb9b66268ce0645b322ded6e7cd512608f2a7ce63

Observation 789a1e01-ae63-467d-9284-3fd6be070583 · outbound

This paper cites Gpt-4 technical report,.

RewardAnything: Generalizable Principle-Following Reward Models Gpt-4 technical report,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.916486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.916486Z digest=sha256:dc61f247b23231033b414f90798a6bcdbfc493fd479874648a174cbef16d5ffb

Observation 4d096eb2-2766-4bb7-8d6f-d29ce7e81c16 · outbound

This paper cites A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity.

RewardAnything: Generalizable Principle-Following Reward Models A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.921103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.921103Z digest=sha256:0a632260034904b3d022dc75de4a1c450c89f8f7953722777576c7e99072e77f

Observation 85a8a507-cc31-4994-bf67-281478cd8ce0 · outbound

This paper cites A fast learning algorithm for deep belief nets,.

RewardAnything: Generalizable Principle-Following Reward Models A fast learning algorithm for deep belief nets,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.925730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.925730Z digest=sha256:b9a46821763fa2021cd5be21b86479e65a545fba5a4fee3123951b6b2e51fb53

Observation 29c94fe2-a672-4bff-8d4d-b1458d710815 · outbound

This paper cites Goodfellow, Y.

RewardAnything: Generalizable Principle-Following Reward Models Goodfellow, Y

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.930247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.930247Z digest=sha256:2935bfe73552d03005525dec009822902cf4377122a4fc35c1da0a3ca3ecb710

Observation 979e9e92-d7d7-4316-88ec-3ec571473c55 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

RewardAnything: Generalizable Principle-Following Reward Models Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.934924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.934924Z digest=sha256:7f4160f18fdd8033a5b4edfc514ed175bf63b6577ec75078199209472deed609

Observation 79f2caed-bde8-4a7a-88e4-8b78accb4f8e · outbound

This paper cites Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping.

RewardAnything: Generalizable Principle-Following Reward Models Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.940032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.940032Z digest=sha256:d84319c8a9f5bf4a2e8a792831b7e7546bc3844113de97072858b105d3951d83

Observation 1337053c-cbc1-4d73-99c8-b887f3c611c2 · outbound

This paper cites How to fine-tune bert for text classification?.

RewardAnything: Generalizable Principle-Following Reward Models How to fine-tune bert for text classification?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.944747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.944747Z digest=sha256:bbc16aca003484d286507bbcaaad0e50145e13877eb30e1bc817289f4a8b6b6b

Observation f68efd9f-1b6f-4349-9070-c4fa21b6fbbc · outbound

This paper cites Natural language question answering: the view from here,.

RewardAnything: Generalizable Principle-Following Reward Models Natural language question answering: the view from here,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.949212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.949212Z digest=sha256:3cd00cc9a544ca2f88bdcede161a11adaba6529cba8ee2526af73c617e9e2c54

Observation 16119788-8a5d-4f1a-a803-8ea0c2bdd1f5 · outbound

This paper cites Natural questions: a benchmark for question answering research,.

RewardAnything: Generalizable Principle-Following Reward Models Natural questions: a benchmark for question answering research,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.953766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.953766Z digest=sha256:04e68021871407c445e08da0a03e2079ac04cb438f9ade831f5d470e6075f1ec

Observation a8d7d9a4-f9c8-494e-afb9-e36450842e50 · outbound

This paper cites Attention is all you need,.

RewardAnything: Generalizable Principle-Following Reward Models Attention is all you need,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.959425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.959425Z digest=sha256:9bfb0d1f970e04a2c604aa495d3f0343e7a50e4c7fcf015053c50b5aacacd2d5

Observation 23240761-a0d3-47f9-8aaa-2ffa09540afd · outbound

This paper cites A Survey on Evaluation of Large Language Models.

RewardAnything: Generalizable Principle-Following Reward Models A Survey on Evaluation of Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.964744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.964744Z digest=sha256:163382e3a6aded7aeb5abea6b16059edadbcfec8e7a76077b3a99825e95d8d60

Observation e48dfae5-2c14-485d-910a-44cfe1603345 · outbound

This paper cites O’Reilly Media, Inc.

RewardAnything: Generalizable Principle-Following Reward Models O’Reilly Media, Inc

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.969461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.969461Z digest=sha256:7cce5fdf1317e0e0eefe27243fb173b15cf56b3d571ecf5563548d840d260433

Observation 4bb07342-09f8-4b42-b70f-ebacad01aa41 · outbound

This paper cites Deep learning tuning playbook,.

RewardAnything: Generalizable Principle-Following Reward Models Deep learning tuning playbook,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.974250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.974250Z digest=sha256:b575a4f903d1c5883828f9c1657349b512c7449a4597945dac51745188ffeffd

Observation 34600fe8-cbf0-4e47-bebe-81c353a602b5 · outbound

This paper cites Glue: A multi-task bench- mark and analysis platform for natural language understanding,.

RewardAnything: Generalizable Principle-Following Reward Models Glue: A multi-task bench- mark and analysis platform for natural language understanding,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.979209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.979209Z digest=sha256:e94083151111515c6f8a40150a3e10ea98e6f063a4d7ae56f55c3f9e24b53a56

Observation 10379657-7d58-4459-81fd-de8febfd5de2 · outbound

This paper cites GLUE-X: Evaluating Natural Language Understanding Models from an Out-of-distribution Generalization Perspective.

RewardAnything: Generalizable Principle-Following Reward Models GLUE-X: Evaluating Natural Language Understanding Models from an Out-of-distribution Generalization Perspective

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.983862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.983862Z digest=sha256:ba4b0b0d3275b4d17ed0925be023e9b2cbe403ffe1758c14f7e8fb2e2384249e

Observation 8d338659-03da-4d26-88a0-ba50fec7d0f2 · outbound

This paper cites The Llama 3 Herd of Models.

RewardAnything: Generalizable Principle-Following Reward Models The Llama 3 Herd of Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.988674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.988674Z digest=sha256:a3a867ac9c0a0e2e9486183d69f853656ba3bc7675b183a10e7858a61eb448c8

Observation 2afbe2cd-240a-4140-a316-0e57d8408983 · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

RewardAnything: Generalizable Principle-Following Reward Models Lora: Low-rank adaptation of large language models,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.993399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.993399Z digest=sha256:134b013798b2e6f360a0c5cf1fd02b086362dfe2837acc476316c88132c88ab8

Observation 91c101b8-0f39-427b-a6c4-5a9c707b8781 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

RewardAnything: Generalizable Principle-Following Reward Models Training Verifiers to Solve Math Word Problems

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.997891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.997891Z digest=sha256:35f81f1c3dc78e6f503c79386f4dc7b376eee2206bcfba8ae7c78e197a60c6da

Observation 9171a74f-991e-4ff4-b18d-8ed9ca0db66c · outbound

This paper cites NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark.

RewardAnything: Generalizable Principle-Following Reward Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.003448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.003448Z digest=sha256:c262c8ebf0a5e481b0325eca709ab1fcc03b596f31c2c50a9aaa891e7c122feb

Observation b62b8a95-8eb5-48b6-a927-22fa4dd3188b · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,.

RewardAnything: Generalizable Principle-Following Reward Models Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.008038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.008038Z digest=sha256:120e465dbe36b42f6ec352cc0c1d2038db68d44125fd3e4ed4f8807155ababb4

Observation a201190d-f322-4fac-8904-48bca019f5e7 · outbound

This paper cites Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning,.

RewardAnything: Generalizable Principle-Following Reward Models Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.012476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.012476Z digest=sha256:b24bf613d8e3dd02eb5b2d05e4d8bf12a12a3dec9d2c5f70a02a779952773056

Observation 18585e90-09fc-4b26-8905-2e866c416456 · outbound

This paper cites Supervised Knowledge Makes Large Language Models Better In-context Learners.

RewardAnything: Generalizable Principle-Following Reward Models Supervised Knowledge Makes Large Language Models Better In-context Learners

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:04:07.758392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:04:07.017382Z digest=sha256:a2b8220f9a088315b30d4110d563576010723c483000815c0a7b6a8a3e43067a

Observation 8225683a-db5b-4222-a75b-5c71558b5fc6 · outbound

This paper cites Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity.

RewardAnything: Generalizable Principle-Following Reward Models Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.022136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.022136Z digest=sha256:0864ef3943e2b4d1e46187479a8a10d392f1731c4d248dfe28a16e10f94f1af9

Observation 404cbede-a3c2-48d9-a35f-4fe3734b43e8 · outbound

This paper cites Large language models are zero- shot reasoners,.

RewardAnything: Generalizable Principle-Following Reward Models Large language models are zero- shot reasoners,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.027152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.027152Z digest=sha256:967bfc57187bd0a4acee74e70391547802de198ea19abe6a963eaae46d51d5aa

Observation 83ab8c1a-22ac-4a8f-9eac-3cb3de3b8b80 · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

RewardAnything: Generalizable Principle-Following Reward Models LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.031741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.031741Z digest=sha256:ce9b7e1bafb8b600bad211b25c830dae9c059807cc64f5f896e9c469deece341

Observation 9bcc8e31-95e2-4215-9e50-b123a60c7ee0 · outbound

This paper cites The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models.

RewardAnything: Generalizable Principle-Following Reward Models The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.036636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.036636Z digest=sha256:1cd1a37f425e5f0a2f970cceb7abb1ed4c5fb1a110bfc174416555f9abbae09d

Observation f446cdd1-8045-4f10-97d8-1fdee4fe94fa · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

RewardAnything: Generalizable Principle-Following Reward Models UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.042216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.042216Z digest=sha256:1fec8b5b8860cadb7fab289c479cacd5819948ab465bbc7bafb0b13a098e8fc0

Observation 1afc5850-2586-4b53-af6a-4fe07349a318 · outbound

This paper cites Qwen2 Technical Report.

RewardAnything: Generalizable Principle-Following Reward Models Qwen2 Technical Report

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.051915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.051915Z digest=sha256:13f2d78cee1dd4f3e36cfd4d2f9ad490ba9f4c1c33709142e3c7739cf8537bfb

Observation 32811580-272d-4435-afb6-65154165c245 · outbound

This paper cites Team, “Qwen3,” April 2025.

RewardAnything: Generalizable Principle-Following Reward Models Team, “Qwen3,” April 2025

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.056870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.056870Z digest=sha256:92328b313b987d1f0a842a7cb167bdb7fa2c19b1fbf78b7841fef35c83a5ec17

Observation 7f7d287e-2326-4723-9c4c-36296bc5f83c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

RewardAnything: Generalizable Principle-Following Reward Models HybridFlow: A Flexible and Efficient RLHF Framework

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.061627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.061627Z digest=sha256:9fc95e8f952154ba2b37ae5c965fddc87cb4bda4f0679b437e05573119f8f025

Observation 2ed5251c-e61f-408c-89ff-b365ffdb9234 · outbound

This paper cites A framework for training large language models for code generation via proximal policy optimization,.

RewardAnything: Generalizable Principle-Following Reward Models A framework for training large language models for code generation via proximal policy optimization,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.066384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.066384Z digest=sha256:287fcb87768a165e8b6901a379001711af8670deefcf25c65511670fe76b4c40

Observation 08592c85-be79-4986-bd43-1d4168896b0e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

RewardAnything: Generalizable Principle-Following Reward Models Gemini: A Family of Highly Capable Multimodal Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.075157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.075157Z digest=sha256:2f35d0a78baffb0361a9982899ab378d4cba182e2cde2a03d34966d3fe3a93e6

Observation 6c5e91f6-313d-488d-af17-339cb0313ca2 · outbound

This paper cites Generative Reward Models.

RewardAnything: Generalizable Principle-Following Reward Models Generative Reward Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.079919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.079919Z digest=sha256:193549358151df328a86e50b82ce237aca715edef336973fc8c053e93600dd25

Observation db7ccc68-7dbe-4d98-b54a-14efed3aedd2 · outbound

This paper cites Improving context-aware preference modeling for language models,.

RewardAnything: Generalizable Principle-Following Reward Models Improving context-aware preference modeling for language models,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.084937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.084937Z digest=sha256:bd5277a8e43a2e0585fc77e5f61b528769e349ed3363a38a61530902f4c0bfa0

Observation 8a617fe4-8e3b-498f-b90c-1866807daa7d · outbound

This paper cites Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation.

RewardAnything: Generalizable Principle-Following Reward Models Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:04:07.578996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:04:07.089789Z digest=sha256:b33a8f23865ff38197fb3df490d9202040da23d594d67aa676ea07b07973d7b8

Observation 87488f2b-c604-425b-8a52-626fdfa5666f · outbound

This paper cites Best practices for the human evaluation of automatically generated text,.

RewardAnything: Generalizable Principle-Following Reward Models Best practices for the human evaluation of automatically generated text,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.094686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.094686Z digest=sha256:ea8f2a832820adafeda13acc3ea71d391c8efa6426d01864216a14e183ff08e1

Observation fa1a059a-bdf8-4893-9288-213356b36821 · outbound

This paper cites DeepSeek-V3 Technical Report.

RewardAnything: Generalizable Principle-Following Reward Models DeepSeek-V3 Technical Report

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.099307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.099307Z digest=sha256:203f99e53cbbbd5423bc7a3c382405d10689fe706ef769fdfa55902c9c000a33

Observation 4be1c353-9fea-41a9-b0d1-faf73251904c · outbound

This paper cites Leveraging Large Language Models for NLG Evaluation: Advances and Challenges.

RewardAnything: Generalizable Principle-Following Reward Models Leveraging Large Language Models for NLG Evaluation: Advances and Challenges

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:04:07.539350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:04:07.104919Z digest=sha256:f010788d9992e650211cfbc7f87e1e137843ab4b4c6b75146b13b266d579dd0c

Observation aac12416-9c89-4f07-a085-38a7e4404567 · outbound

This paper cites FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models.

RewardAnything: Generalizable Principle-Following Reward Models FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.109652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.109652Z digest=sha256:5ae5bea91d993e3057538141184d5d3c31deb2e1460f8949883d8b0ae28e463d

Observation 27f6e241-7a3b-4db7-86b9-4d656f555e4c · outbound

This paper cites From generation to judgment: Opportunities and challenges of llm-as-a-judge,.

RewardAnything: Generalizable Principle-Following Reward Models From generation to judgment: Opportunities and challenges of llm-as-a-judge,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.114445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.114445Z digest=sha256:a81c39de88021965f6e39d86421481dc0e93296fd6d6d93dc6f2f7f1330b9535

Observation 77f5f3d2-ef74-4f68-a451-c39f90ca04b3 · outbound

This paper cites Automating Dataset Updates Towards Reliable and Timely Evaluation of Large Language Models.

RewardAnything: Generalizable Principle-Following Reward Models Automating Dataset Updates Towards Reliable and Timely Evaluation of Large Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.119152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.119152Z digest=sha256:c38a424a529a7854494d7977f94acb88e2af5a7b86c97f6c8ff506ed3319156d

Observation 3787a9e0-2fa5-4f7c-a854-d2cd1140ca1d · outbound

This paper cites A Comprehensive Survey of Contamination Detection Methods in Large Language Models.

RewardAnything: Generalizable Principle-Following Reward Models A Comprehensive Survey of Contamination Detection Methods in Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.124211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.124211Z digest=sha256:d7418ccd2e3ff61901e8a18ba76f42ec00505e3a1a2f32992dcb4246b881ee99

Observation 1a4e773d-0972-41b0-b0c5-f77d719df2f4 · outbound

This paper cites Prompt-to-Leaderboard.

RewardAnything: Generalizable Principle-Following Reward Models Prompt-to-Leaderboard

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.129328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.129328Z digest=sha256:aca28c4ff9784ae15fe00115454d93675147f1cbccd8fad0c31c2a18524413f5

Observation 16fb42d0-8d71-4213-994e-65da12dda5f6 · outbound

This paper cites Scaling laws for reward model overoptimization,.

RewardAnything: Generalizable Principle-Following Reward Models Scaling laws for reward model overoptimization,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.135219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.135219Z digest=sha256:9580a8747cee1108f3ee3762d0b908ae871593bba71ded1398daf43dff10a7f0

Observation a978274c-3e04-4b41-9561-e696d770c2d0 · outbound

This paper cites Self-Generated Critiques Boost Reward Modeling for Language Models.

RewardAnything: Generalizable Principle-Following Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.139918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.139918Z digest=sha256:afd7dac9692ddc3a6cebc99fc9f34f064b32443a5b2c04bf71fa9caca5994165

Observation e2782c99-9211-4367-ad11-4a63a5f59a4f · outbound

This paper cites Critique-out-Loud Reward Models.

RewardAnything: Generalizable Principle-Following Reward Models Critique-out-Loud Reward Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.144776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.144776Z digest=sha256:3f9964ea30de8efcc41a2a4f2820fb06c4fb32444c35006b9a8b9b9a3e0401a5

Observation 7f3f9ad9-0e3f-4bcf-b043-4f51188e8a69 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

RewardAnything: Generalizable Principle-Following Reward Models Constitutional AI: Harmlessness from AI Feedback

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.149524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.149524Z digest=sha256:e9d8ac51d65c9b8a475ff315cdc5c59d1d7ea0bcc85701c1b50426550f01b516

Observation bedcc321-eb0c-4115-b9dd-2143ccb00f25 · outbound

This paper cites Is Elo Rating Reliable? A Study Under Model Misspecification.

RewardAnything: Generalizable Principle-Following Reward Models Is Elo Rating Reliable? A Study Under Model Misspecification

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:04:07.278235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:04:07.154095Z digest=sha256:79e04a8eabd22caed2988c7c1f7b987d215d25afa70b64662745143710413389

Observation d3eb81df-c2b1-4ff0-bf88-6ac5df644d5b · outbound

This paper cites On the biology of a large language model,.

RewardAnything: Generalizable Principle-Following Reward Models On the biology of a large language model,

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.159228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.159228Z digest=sha256:af3e58748a9717e37d3b43f9f523dd2bc3d7c9ff9d4ddc5cd6bf862c5f1e52b3

Observation 2ea09ea2-b1c6-4bdd-a03b-49a91e102ab8 · outbound

This paper cites Gpt-4.1 and gpt-4.1 nano overview,.

RewardAnything: Generalizable Principle-Following Reward Models Gpt-4.1 and gpt-4.1 nano overview,

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:08.978013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:04:07.163892Z digest=sha256:ea4cd01035500258ba7613a1f57248586bd8a9a71642058eb529d5a2c78a833a

Observation 3ec7d72a-c526-44ee-a919-72b4c4096480 · outbound

This paper cites Gemini 2.5: Our most intelligent ai model,.

RewardAnything: Generalizable Principle-Following Reward Models Gemini 2.5: Our most intelligent ai model,

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:04:08.962777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:04:07.168303Z digest=sha256:4b95a4472acff7b10eb76e7fd9d482cf4a1955015e80c2eb42736ea240ab758a

Pith citing papers

Observation b69cfabc-835e-44b2-8e20-58f323790bed · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle RewardAnything: Generalizable Principle-Following Reward Models

Reference 231

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:46.427272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:46.427272Z digest=sha256:a8c9eb02421e0f0b5290c96fba92811d5fefb95798586afaaff74aa7a88d8a92

Observation f72b0cf3-5632-41cc-82c9-10b554b62a62 · inbound

RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards cites this paper.

RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards RewardAnything: Generalizable Principle-Following Reward Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:10:42.036446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-21T22:09:47.649346Z digest=sha256:eb5ed54d02580cd210ebabb4675f38c65e34b954068327eb65c26e8645fc6d34

Observation 4dbe7bcd-ddd2-4f60-b0a1-8deb8beb1af6 · inbound

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents cites this paper.

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents RewardAnything: Generalizable Principle-Following Reward Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:25:35.741067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T12:22:13.551709Z digest=sha256:62736b89e818d4d4eaf6f94e3c5aae5406ab57ebd3b99c7125eeec62890af0df

Observation 2db5892d-8c19-45ac-9fec-1ef606dfc1dc · inbound

Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization cites this paper.

Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization RewardAnything: Generalizable Principle-Following Reward Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:51:45.655359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T06:47:34.623340Z digest=sha256:d92f77c6c1fe320794d990dd330d37c3c3b4061edb1e5afb6653d129f683e85e

Observation 3c8adfd8-f24e-475f-b765-777ab48bdc7f · inbound

DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation cites this paper.

DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation RewardAnything: Generalizable Principle-Following Reward Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:31:07.401146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T17:29:22.089174Z digest=sha256:91cb63c97d54c7eb753da28488ead5b096ca4e4e25433becd793a98ec633c750

Observation 6aac035e-19a1-41be-b03e-d3bcaed68a42 · inbound

DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation cites this paper.

DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation RewardAnything: Generalizable Principle-Following Reward Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:15:09.597128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T00:05:28.004125Z digest=sha256:cf1c466784d30d805f7866313c118f5bc11b37bce93a543eb5f4f4111c7f7fe9

Observation 47d74668-400a-4c4f-97b1-af72d8570a03 · inbound

Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill cites this paper.

Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill RewardAnything: Generalizable Principle-Following Reward Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T11:02:01.384983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T10:53:22.633115Z digest=sha256:22f3e31f00c9d967792475df5cb0a58f75a76c61654977821bf589a5c21dd7e1

Observation 4d9bc8bb-3426-4345-92d4-2e42081b3aeb · inbound

Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics cites this paper.

Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics RewardAnything: Generalizable Principle-Following Reward Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:47:23.335587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T20:06:12.523067Z digest=sha256:7e7f89259b34e20b0496752df351e74d30cb29a42fe7913930e7234d2def6f18