Pith. sign in

Paper Citation Record · LEDGER

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment

As of 20 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 1 inbound Pith citation observation for arXiv:2505.10597.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10597 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:19:41.527363Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T08:32:38.883019Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T08:32:51.695159Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3030dd8e-2d36-4d2c-8741-5ed4fca3c232 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.263687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.263687Z digest=sha256:2a30fc43731456a011f2357f632372cd06c63b69441518348534685b47653635

Observation 7ca5174f-0ea4-406a-a676-b463e0ca9104 · outbound

This paper cites Training language models to follow instructions with human feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Training language models to follow instructions with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.268318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.268318Z digest=sha256:6f102a9a2477181e08f0e729e03feff890483f835f804ceed3ff1989f11156be

Observation fa577b07-d0a0-418e-97dd-183650ecf221 · outbound

This paper cites A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.271849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.271849Z digest=sha256:bec3359d159fa2d787a0adb9fd4cd9033f2d27ff1b42062e760b397dd593ba65

Observation 00e2721e-cb01-40d5-ae25-630bdfb76a8c · outbound

This paper cites A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.275259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.275259Z digest=sha256:fc68113883ea7fc757cc1c34785771296e9e19e66440a2678e502d384410800a

Observation 92402796-48dc-41fc-9aea-a29af5247491 · outbound

This paper cites Better Process Supervision with Bi-directional Rewarding Signals.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Better Process Supervision with Bi-directional Rewarding Signals

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.278716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.278716Z digest=sha256:09a7aef2dde7fbc493ee8181af9cb4b4448a7960deda85d5add6864a0c95eade

Observation 4d239447-3e69-4824-bafa-4fad15c5596c · outbound

This paper cites Reward Function Design in Reinforcement Learning, pages 25–33.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward Function Design in Reinforcement Learning, pages 25–33

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.248402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.282169Z digest=sha256:e6471c829cce7d03b01bac051e51d735192013bd9d04ec0d2a3ee85b26861508

Observation be8753fc-b9cb-4731-9515-b7a1cb6c1dab · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.285668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.285668Z digest=sha256:d641077efbd062293080a3002f8e9ff1ec4eff9fa46b91107b7c1bdcb665e4a6

Observation 3289ef25-83b8-40ca-baa2-b91abff2fa82 · outbound

This paper cites Gpt-4 technical report, 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Gpt-4 technical report, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.288969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.288969Z digest=sha256:6f180b87d1d9cd860fcff0ec4d63d6c267a86ce0b2245528166fa60e25d34100

Observation 21ad8af1-130a-4f9f-bb6f-4d01402cdb91 · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.292618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.292618Z digest=sha256:7f5ac03febdcc9a160d781a56e83dabab14ae85cfbb9fa867ce05949bb1dcebd

Observation 1c8d58a6-9201-4524-b469-2666a30a531d · outbound

This paper cites Skywork-reward: Bag of tricks for reward modeling in llms, October 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Skywork-reward: Bag of tricks for reward modeling in llms, October 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.235122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.296080Z digest=sha256:f1c0c534ae790fa986bf55baecd237cdb52e84ea21647f668c8a75d38f4e4f84

Observation ba77e19c-df9a-46d0-81a4-eeb13a0ea201 · outbound

This paper cites RMB: Comprehensively benchmarking reward models in LLM alignment.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment RMB: Comprehensively benchmarking reward models in LLM alignment

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.227695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.300521Z digest=sha256:f34e9e61eb9fc69b10e641678fcae2ae6c73ec8b468e1244ea05dcebce3e9bf1

Observation 082da1a0-2646-43a7-a0a0-e67ca2f656d6 · outbound

This paper cites Helpsteer 2: Open-source dataset for training top-performing reward models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Helpsteer 2: Open-source dataset for training top-performing reward models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.219817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.304537Z digest=sha256:7ef412bcb3efe4a42638eec57b4741cf7399b978722638b7d5017a308bb90ea7

Observation 60ff893e-87d5-4441-bd2e-aab7d1ffb549 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.211777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.307570Z digest=sha256:878ae11385c5f6cac8edcdaac619ac5bef6b8dc604600ceb3daa3a805ff2bdce

Observation 38fba78c-52a5-4611-8f9d-0db0dae359b5 · outbound

This paper cites Impact of preference noise on the alignment performance of generative language models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Impact of preference noise on the alignment performance of generative language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.203441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.310840Z digest=sha256:f08a2f52ee6f34d33115aed51d48628b0f10b32033384a21093e4250d4551e5b

Observation 53221f72-75e5-41ce-849e-7ecbc4f10ae5 · outbound

This paper cites Improving reinforcement learning from human feedback using contrastive rewards, March 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Improving reinforcement learning from human feedback using contrastive rewards, March 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.195356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.313689Z digest=sha256:afa419c182943cca7a9147ae5881bd5c6ed4e518ce982fa9a7f826de47e387ab

Observation 988419a2-f71a-4481-8f66-eeacda53a897 · outbound

This paper cites Goal misgeneralization in deep reinforcement learning.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Goal misgeneralization in deep reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.186985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.316987Z digest=sha256:7ad1c023f8aca8e732d87452da51af98ad192fb57d10ad734327214511ad2cab

Observation fde9f6c0-a8b0-4a72-8da0-9a1a444cf8ee · outbound

This paper cites Scaling laws for reward model overoptimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Scaling laws for reward model overoptimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.178575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.320021Z digest=sha256:650257a89b2a94c72ff2dd813c6199834d456b6bf5c461b1db4986d92fddcdbf

Observation f48b4cb1-5d4a-40ed-9ea3-02446c1a41dc · outbound

This paper cites Improving discriminative capability of reward models in rlhf using contrastive learning.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Improving discriminative capability of reward models in rlhf using contrastive learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.170632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.323132Z digest=sha256:a3f3339e855386fc9aa0b714fb4bdd50c70f1d259c8dacb9762cc9ea064bbd23

Observation d58fa8a7-78b0-417d-8794-d3fe9ad0c387 · outbound

This paper cites Reward Generalization in RLHF: A Topological Perspective.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward Generalization in RLHF: A Topological Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.326063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.326063Z digest=sha256:b8a6ade5866646c2a9315cc2b43f6cf250662fb1e25c0cb72f8f8bf0d06e7293

Observation 2a2abe0b-2f93-4fdb-9f83-54a3fb57ff3b · outbound

This paper cites A note on dpo with noisy preferences & relationship to ipo, 2023.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A note on dpo with noisy preferences & relationship to ipo, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.329460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.329460Z digest=sha256:2927df48ecd461f5e2874360f7a0f558e93bca78db787f5704341efcbee0a6d3

Observation 16d62cf5-fd55-4a94-83e5-bc4e82b705ba · outbound

This paper cites Provably robust dpo: aligning language models with noisy feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Provably robust dpo: aligning language models with noisy feedback

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.156044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.332351Z digest=sha256:1018ad751a6941812e9eb3371ed2bb1b1542588f63d6af670fd8f993ae60f8fc

Observation 18218c88-4406-4f6b-8e3e-56db5d47e382 · outbound

This paper cites Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.335604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.335604Z digest=sha256:2b15fc1994e55088dccc9e0ff9871fca97609046ffd65f5deab7237d28a6eff2

Observation 4157e457-ac76-4f1e-8d66-a8849691436d · outbound

This paper cites ROPO: Robust Preference Optimization for Large Language Models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment ROPO: Robust Preference Optimization for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.338848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.338848Z digest=sha256:5d7b27d91fe63d71f50d80d20049afa6de263452408a518b0b7cc79180e2c1f1

Observation 96fc20ea-3b02-4aca-b22f-12ad442ef022 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Rank analysis of incomplete block designs: I

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.342349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.342349Z digest=sha256:73b482fa3de602af2b81a379cd140436c87ce88414f3357b129d8314c6a48248

Observation 9ff9cb68-246c-472b-a123-41e13de42a5b · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.345277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.345277Z digest=sha256:4563e74233bdcc9cba2181baa4ee1fe4dfdccb01669ba9732e5b09e7b4714289

Observation b8ccced3-09c5-4b76-b69f-9e5cb79926c2 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Direct preference optimization: Your language model is secretly a reward model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.137207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.348463Z digest=sha256:916e6bd6955c532a95d1fcf6e9ffa3f2a886bd1a2d5c9a756ed688b7dbf2905b

Observation 8ef22d1c-5541-41d8-a330-f0e54ccc63c1 · outbound

This paper cites Confirma- tion bias in human reinforcement learning: Evidence from counterfactual feedback processing.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Confirma- tion bias in human reinforcement learning: Evidence from counterfactual feedback processing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.128912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.351616Z digest=sha256:bfe2773f03c40c143855e00daa15ccef7ca4c257a2b228d91b95d0b40d744364

Observation ba666142-e16b-42e1-9f92-c3b045e2efdb · outbound

This paper cites Pseudo- labeling and confirmation bias in deep semi-supervised learning.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Pseudo- labeling and confirmation bias in deep semi-supervised learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.119261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.354598Z digest=sha256:f00bc78d89adb87b4b01b35361370b9fe820ecc3c814c19af91a5e05bf5dd46a

Observation cabc74fc-328b-4ccb-9ac2-c5c7083b72f4 · outbound

This paper cites Zephyr: Direct distillation of lm alignment.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Zephyr: Direct distillation of lm alignment

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.109684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.357680Z digest=sha256:aaf234f62fbbe087c7d184ebcddff9561cf3e02b5193d72393b77c8e79103292

Observation cf523574-1ae8-40ac-9f1a-f61880f0780e · outbound

This paper cites Smith, and Hannaneh Hajishirzi.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Smith, and Hannaneh Hajishirzi

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.361790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.361790Z digest=sha256:78eaae667a146782002a2b40d360222296e7fc655cab3ee1b43a6fe7ac03be3d

Observation b15471eb-e204-4329-b44c-4f008bcf0341 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf, 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Rlhf workflow: From reward modeling to online rlhf, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.365456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.365456Z digest=sha256:700eab3312879d1ef250f3d357e2d493a6c1924365e08aef0dc7b4627f9b1e7d

Observation f2e51a55-b2c5-4d1c-9d37-1647e826b773 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.368380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.368380Z digest=sha256:c4c75e793b9cffd85886adaa61dc58de0ba5852be8e39ef157a1b48ed2e9e9a9

Observation d782e6c3-2cd6-47a9-abcb-73ed3b3fe353 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.371363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.371363Z digest=sha256:a2bd5d1a0c990d7b5f87f2a6d56fd51fe6db5acea1a26e9587bef32c17d3f928

Observation 54fb0898-f685-49e7-ab98-c4c3a1911a67 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Secrets of RLHF in Large Language Models Part I: PPO

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.374445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.374445Z digest=sha256:69e9417dfe0182c5232c09e5df9591b07d28b6a51af0ebe8401cafe2d673862a

Observation b550525c-d12b-4d96-b34b-92ca68151aa5 · outbound

This paper cites RM-bench: Benchmarking reward models of language models with subtlety and style.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment RM-bench: Benchmarking reward models of language models with subtlety and style

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.378271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.378271Z digest=sha256:01ab652f5c8b6fe5c007b78a4982b67869d82c568d840e7e1ceaf680c16cae7e

Observation 8ff587ba-c6d5-4423-8122-eff6a4e55f94 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.381195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.381195Z digest=sha256:3a2185b9ebfd7cb92f82bfc2e6fcdbd62d71a1570b18f9e376e24a7517b5eb81

Observation 5003dbc6-318f-4334-a41e-5cfc9bc965f0 · outbound

This paper cites Weak-to-strong generalization: Eliciting strong capabilities with weak supervision.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Weak-to-strong generalization: Eliciting strong capabilities with weak supervision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.080708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.384743Z digest=sha256:72391677083e30d2811b0be86879a554e401f4334e46569b559d1fac6e499c53

Observation 8a0e4840-f774-4ab0-9b3e-0d7ffa901a37 · outbound

This paper cites Weak-to-strong prefer- ence optimization: Stealing reward from weak aligned model.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Weak-to-strong prefer- ence optimization: Stealing reward from weak aligned model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.073280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.387657Z digest=sha256:b0a2faa703dade1c33771a1784b2964d1a08a564f67b593caca57d44752abc49

Observation eca98603-0169-41fa-8fa7-63551354f620 · outbound

This paper cites Adversarial Training of Reward Models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Adversarial Training of Reward Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.390489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.390489Z digest=sha256:ba09cbd45c9b17457bf87431ddb51d2e0b658ebaf8c565d446e10c7892a357f0

Observation c3ad53a6-8e84-4009-8174-79ef207ada62 · outbound

This paper cites Defining and charac- terizing reward gaming.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Defining and charac- terizing reward gaming

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.065101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.393496Z digest=sha256:77f93ab5204103281fda51380005cd54b0885bd3cdc616c2cb67cd5e1b6e8e18

Observation b4d11414-6125-41d7-a0b6-b2b415bb0752 · outbound

This paper cites Reward Shaping to Mitigate Reward Hacking in RLHF.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.396471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.396471Z digest=sha256:5a3de6970af58db24596e473449566199059f5387bec2f6805365dd69325c76a

Observation bff2a27b-8fc5-41b3-b048-998c8d611d2c · outbound

This paper cites Odin: disentangled reward mitigates hacking in rlhf.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Odin: disentangled reward mitigates hacking in rlhf

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.057629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.399094Z digest=sha256:93422b52803d658ab7f97223a0e19f815b0979a7183cb075701934fd5ee2beb4

Observation a467c576-3136-40ee-849e-12726229087b · outbound

This paper cites Taming Overconfidence in LLMs: Reward Calibration in RLHF.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Taming Overconfidence in LLMs: Reward Calibration in RLHF

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.402572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.402572Z digest=sha256:4ed8fcab3917fc5370eb916b4dbd285d92da1873361a5b880beaca66ec35247b

Observation 6c73bc40-df62-4342-a705-47d6cbc6fc6c · outbound

This paper cites Overcom- ing reward overoptimization via adversarial policy optimization with lightweight uncertainty estimation, July 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Overcom- ing reward overoptimization via adversarial policy optimization with lightweight uncertainty estimation, July 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.049800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.406480Z digest=sha256:8f2e59fc3b90488dea8f20df22c9427c6277688bb6f558a68c9dc6d235db3c79

Observation eb3802f2-ce7b-42ab-a53c-98f3a8f4f09c · outbound

This paper cites The energy loss phenomenon in rlhf: A new perspective on mitigating reward hacking, February 2025.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment The energy loss phenomenon in rlhf: A new perspective on mitigating reward hacking, February 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.041660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.409426Z digest=sha256:0fff2a9b6c7689b11275ae428b653da174850ee6ca4a5faad955d2e4f5f9f23f

Observation 1c655f76-79e9-47ad-b6d7-7eb3abed0f85 · outbound

This paper cites Reward model ensembles help mitigate overoptimizatio.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward model ensembles help mitigate overoptimizatio

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.034242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.412385Z digest=sha256:dd42b1a7b11ecb7f017a7453f8f7507e18cc0f639a15019c9436d8704ed5e55c

Observation 52e9f12e-fc13-45a0-885e-f9273de244d7 · outbound

This paper cites Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.026701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.416160Z digest=sha256:325a5a85bd8250d82ca354bfbe28aecf03c6acce19c43cc472274984c5321ba2

Observation 100c5a23-f993-43dc-8f28-a38a16053b51 · outbound

This paper cites Reward-robust rlhf in llms, October 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward-robust rlhf in llms, October 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.019019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.419845Z digest=sha256:4c30df23bbfcb279414b55d56ac30120db1e4a9b49e363a1d169df325b23e237

Observation 6b50f217-957a-487a-96e1-3f8ebff88faf · outbound

This paper cites Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles, December 2023.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles, December 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.011300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.422374Z digest=sha256:f34566cf8f01baa0b01dd6ebc3517b34badd813e72926abc81f1f36cb316342f

Observation 6b62b461-8298-4123-b4e8-66c177a4693e · outbound

This paper cites Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.425267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.425267Z digest=sha256:3bc0ed8b7431164397446b89596947c42883702821c67ad9367523d341d3b2e0

Observation 57244e02-3310-442d-959f-f0d6e96fe115 · outbound

This paper cites RRM: Robust Reward Model Training Mitigates Reward Hacking.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.428176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.428176Z digest=sha256:e683fb297e327fcc064d2cae0acdeae01a104322875f15d51370deba2cb8ebc2

Observation f8052a73-4469-45a1-af82-b118f48354af · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A general theoretical paradigm to understand learning from human preferences

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.430957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.430957Z digest=sha256:bae77d1bd44e0abf6e22547a8ea254db6479a693c0997e11e08b3c4019d9c118

Observation 94c0d798-619d-468d-a22b-5d7c26826fb4 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment KTO: Model Alignment as Prospect Theoretic Optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.434016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.434016Z digest=sha256:7f6627b8f4736724fb4642d78339ebece720f45a9c921574146a00e9d058f0a2

Observation 951e9b08-4d2f-40db-9183-2f049dbd036b · outbound

This paper cites Orpo: Monolithic preference optimization without reference model.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Orpo: Monolithic preference optimization without reference model

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.993237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.437862Z digest=sha256:0dcc2f6afae19e36012c3085962cddf26f9a4aa5b2f60cd2723b7702e6edecc2

Observation 5be45d8c-f341-4446-a2aa-549e0898254b · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Simpo: Simple preference optimization with a reference-free reward

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.440975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.440975Z digest=sha256:5de7b2c5249b85825ccef31225bbd110a15df092b45c53db53e336729b52a8a7

Observation 0020e4f9-2fa6-4555-a8a1-c7af9ab23f5b · outbound

This paper cites Is dpo superior to ppo for llm alignment? a comprehensive study.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Is dpo superior to ppo for llm alignment? a comprehensive study

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.978016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.443647Z digest=sha256:f7f33bd9e6bb6050029cdc32809a28c03194efe3087ce0d4c8533ae5da2d38b0

Observation 3b7cb61e-1722-40ea-97b9-284057dc51b6 · outbound

This paper cites Smith, Yejin Choi, and Hannaneh Hajishirzi.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Smith, Yejin Choi, and Hannaneh Hajishirzi

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.968885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.448068Z digest=sha256:9ec944dff7c8c8853e0ede3873da3b2ed9010935f9365cd14cf065b5ad416097

Observation c9c7c98f-214b-4e3c-9397-7cc677826137 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Direct Language Model Alignment from Online AI Feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.451174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.451174Z digest=sha256:747909cac20d729b14a9f696ebd0af9b94965c0554198e6fe3fefd33c6a5c6c0

Observation 5d220a40-9759-4e58-a761-1d70cbc63392 · outbound

This paper cites DPO-Shift: Shifting the Distribution of Direct Preference Optimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment DPO-Shift: Shifting the Distribution of Direct Preference Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.454892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.454892Z digest=sha256:a5a403266efe85cedbb183918bdab3aeae7921389e42419b4a856d506fa086c5

Observation 9557c46c-3769-4070-9a87-bd0f00de5dd9 · outbound

This paper cites Understanding generalization of preference optimization under noisy feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Understanding generalization of preference optimization under noisy feedback

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.959084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.458038Z digest=sha256:2bae2eac8118f80adafdb59c3fdde8db2c821e0682b5c93bf8ef07780cec2311

Observation 90ca6071-d427-4f0c-a23d-4cdaa5e807cb · outbound

This paper cites Robust reinforcement learning from corrupted human feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Robust reinforcement learning from corrupted human feedback

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.949483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.461332Z digest=sha256:05bec9d41d8b79c8860f62b9261b90ca52f264d884bf8914d98b110055ecb03b

Observation 19fd781e-d323-47f1-9ff3-051373fba9cc · outbound

This paper cites Theory of games and economic behavior, 60th- anniversary, 2007.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Theory of games and economic behavior, 60th- anniversary, 2007

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.941040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.465070Z digest=sha256:d18e07ca2745d654f2f79aaab51992e2d442abd61200af825305990e7c228ac5

Observation 24cdd193-71e1-4a6d-aec3-67c2dbc7e614 · outbound

This paper cites Multiagent systems: Algorithmic, game-theoretic, and logical foundations.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Multiagent systems: Algorithmic, game-theoretic, and logical foundations

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.468046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.468046Z digest=sha256:606bc724a9b3686900b90d1f5aad477723c8f1beba058ff2cbc77c87e3c16bad

Observation f30cc94a-75ce-452a-a767-917d0e9a53e8 · outbound

This paper cites [Yes] " is generally preferable to.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment [Yes] " is generally preferable to

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.926924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.472095Z digest=sha256:0f5f0d099e5bcb062d66aed1720dc73d874ed4debb01b656aec5238f61566c96

Observation 587f0f17-064a-44f2-adff-90fc31e59d3d · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.918496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.475861Z digest=sha256:8f8bd167e7e6e22d6bc3c1a933d236b2a677d7ec96541495cc53d6f37a5bbfab

Observation 2587aa68-18fb-45d6-b3f9-7186901ca14d · outbound

This paper cites Limitations.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Limitations

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.910990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.479196Z digest=sha256:36a270357fb85edf2c8262e4dad53cab115d547315b18fb32cb5664d59566c2e

Observation 0ae68a1d-93c7-4ca9-bf0f-efb0910d40e3 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.902892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.483442Z digest=sha256:3cd2792458670227f6d92b53424fc59e3e8e8feee18b28b946c44c14b19b0f44

Observation 45575863-75fc-491c-b97b-d3828cd3544c · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.894349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.486539Z digest=sha256:5ecc116b01a0494c4b25406dcc318410554d00db4c8652f89394a2eb77530a82

Observation 28b3ac97-92c3-44fd-808d-72b200a62c21 · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.884914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.490836Z digest=sha256:1af995b3da703910b965764179285b34a543b8459b25805a93230826477bdca4

Observation ca42c146-6f32-41b1-a962-51a0703b86ee · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.876158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.494510Z digest=sha256:9bc9b4e2673ec3ea69287400d2e5292d294a939fdf5df6c453a60085ef8bc945

Observation 6e4a06e4-3921-4646-909d-6cfebcadef07 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.866146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.497449Z digest=sha256:ffad27aa06327ae56b522582ed7e0b2be8e2c84fd26c8f26e1f7fc90c69dd09e

Observation e0aa576d-be3d-41bc-a356-eb2cdc6abde9 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.855662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.500484Z digest=sha256:9e700e3bbde47cb6e1797fbf0212984a4429c9560e02abfc21db594035d6b88a

Observation c43dba8f-bd3f-4f05-ad03-59ed9980f531 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.843927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.505139Z digest=sha256:cb5d6664b35cbed6703a7afc4b1f775af1f0397a5d6b889e8532c39178f0467b

Observation 76be5127-3bb9-4fb1-bdfe-168239b3e32c · outbound

This paper cites an unresolved cited work.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:19:41.828840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.508480Z digest=sha256:7b1aa67884952d928ffcc90a0c041a32a175b584f132527acec2fd0cab8ca2ad

Observation 187b16d1-8534-48d7-ba67-1575806c6f2b · outbound

This paper cites an unresolved cited work.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:19:41.811120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.512195Z digest=sha256:ec76fb994cafece0678821865b79d5513b72e587344c262b92f44271a0a05be6

Observation a29dce32-996e-4b09-9b85-e560dff6dfb2 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not use existing assets

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.792568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.515056Z digest=sha256:ee560de005c2ee348a7b4e6379640d72a67f51e856c992545a5cfd199b134a4b

Observation 27395a9e-b895-4c14-88eb-bfb3742071e2 · outbound

This paper cites • Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment • Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.772923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.517980Z digest=sha256:1a7e1490038b883942754bad4e5a76c1c3bb3bf4a73518dc0a39c4e99a8a038f

Observation 4ec17451-33bb-4202-ad24-41121e2d824b · outbound

This paper cites 32 Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment 32 Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.754738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.520830Z digest=sha256:d0ca4b1c07e34f29921905c0a4a93dc1a0b15e9553041bfc03d3381bd746c4a3

Observation 009638ff-9737-402e-ab8d-ae28df7ecc55 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.737498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.524216Z digest=sha256:c270a647cc28830addee6daffebb78f03a8076f2b13fa16cee4f8637071b71e7

Observation 767017fa-c7a9-451f-8aa4-13ab29e25803 · outbound

This paper cites Answer: [NA] Justification: LLM is used only for editing, or formatting and does not impact the core methodology, scientific rigorousness.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Answer: [NA] Justification: LLM is used only for editing, or formatting and does not impact the core methodology, scientific rigorousness

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.722866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:19:41.527363Z digest=sha256:db18f32f54cfb5c79d96cbc4328394e2808942c8c0ce493d60d546dde0e7baec

Pith citing papers

Observation b431773f-cf7b-4d93-9727-f0dfce4a409b · inbound

AgentV-RL: Scaling Reward Modeling with Agentic Verifier cites this paper.

AgentV-RL: Scaling Reward Modeling with Agentic Verifier Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:32:51.696658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T08:32:38.883019Z digest=sha256:e1fa851c6c0d1c3ded69c8e8bdb2b2226f794c9d71448ee407a203db9ae24078