Pith. sign in

Paper Citation Record · LEDGER

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment

As of 16 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 1 inbound Pith citation observation for arXiv:2505.10597.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10597 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:19:41.527363Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T08:32:38.883019Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T08:32:51.695159Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3030dd8e-2d36-4d2c-8741-5ed4fca3c232 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.263687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.263687Z digest=sha256:49d9ca356b149ab692662d63a8b90bc4fe5505695666f393fb8a09ff16bee793

Observation 7ca5174f-0ea4-406a-a676-b463e0ca9104 · outbound

This paper cites Training language models to follow instructions with human feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Training language models to follow instructions with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.268318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.268318Z digest=sha256:c82c48dadabac81edef05ebd2bd42a7cdc336d4231c44acde4698630b2491070

Observation fa577b07-d0a0-418e-97dd-183650ecf221 · outbound

This paper cites A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.271849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.271849Z digest=sha256:46c44c62b1d34b27fc93fde4fcc43feefe1788c04a9d5857bba0fe8d9dbb28b6

Observation 00e2721e-cb01-40d5-ae25-630bdfb76a8c · outbound

This paper cites A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.275259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.275259Z digest=sha256:593cb95c02e98c8b1c54f2eb7b8ece1e54dbbdbdb606c32c75fe35b4febbc440

Observation 92402796-48dc-41fc-9aea-a29af5247491 · outbound

This paper cites Better Process Supervision with Bi-directional Rewarding Signals.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Better Process Supervision with Bi-directional Rewarding Signals

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.278716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.278716Z digest=sha256:442b75ad0673986166923b8ab6292aa36c981587f1085946511eb67a2372f557

Observation 4d239447-3e69-4824-bafa-4fad15c5596c · outbound

This paper cites Reward Function Design in Reinforcement Learning, pages 25–33.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward Function Design in Reinforcement Learning, pages 25–33

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.248402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.282169Z digest=sha256:93aaae220c7d8a822f85038f75bbba30ce4c87be5b3c04cd279e9d34aea6e883

Observation be8753fc-b9cb-4731-9515-b7a1cb6c1dab · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.285668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.285668Z digest=sha256:f6f07c49f5f8df48a2a829a014976b9c64eeb996b1e66ee5dd7ed4a976a83de5

Observation 3289ef25-83b8-40ca-baa2-b91abff2fa82 · outbound

This paper cites Gpt-4 technical report, 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Gpt-4 technical report, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.288969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.288969Z digest=sha256:8e6af15953fae5f0953f28f1cbccdf7a21829cef0b30918b9e06f55820f98e74

Observation 21ad8af1-130a-4f9f-bb6f-4d01402cdb91 · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.292618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.292618Z digest=sha256:4bf3ae1f49af05555100293a95caee299e0368e40959050e408327277fee06e3

Observation 1c8d58a6-9201-4524-b469-2666a30a531d · outbound

This paper cites Skywork-reward: Bag of tricks for reward modeling in llms, October 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Skywork-reward: Bag of tricks for reward modeling in llms, October 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.235122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.296080Z digest=sha256:9ca2a3d1b694672576e1cab027e00a2241b15562d36fdb78da36f581c232bcbd

Observation ba77e19c-df9a-46d0-81a4-eeb13a0ea201 · outbound

This paper cites RMB: Comprehensively benchmarking reward models in LLM alignment.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment RMB: Comprehensively benchmarking reward models in LLM alignment

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.227695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.300521Z digest=sha256:707380d83fb3f898a084dfcaf4ad9fbc7135cba94a9bae2b61f29c29cabc781b

Observation 082da1a0-2646-43a7-a0a0-e67ca2f656d6 · outbound

This paper cites Helpsteer 2: Open-source dataset for training top-performing reward models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Helpsteer 2: Open-source dataset for training top-performing reward models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.219817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.304537Z digest=sha256:32d067ca8bafe61884d1f506328d6e24f989492490d4f6e88952999fc8eca636

Observation 60ff893e-87d5-4441-bd2e-aab7d1ffb549 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.211777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.307570Z digest=sha256:92cdad3cbde3f06cddb56ac3ff0271b454a7a07c4b24241b5fb99fbbed5a809b

Observation 38fba78c-52a5-4611-8f9d-0db0dae359b5 · outbound

This paper cites Impact of preference noise on the alignment performance of generative language models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Impact of preference noise on the alignment performance of generative language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.203441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.310840Z digest=sha256:8c43f36f14d689927fbc13cd127dc7c097de3586d50ff267081e4969b68ef4ee

Observation 53221f72-75e5-41ce-849e-7ecbc4f10ae5 · outbound

This paper cites Improving reinforcement learning from human feedback using contrastive rewards, March 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Improving reinforcement learning from human feedback using contrastive rewards, March 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.195356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.313689Z digest=sha256:f7ec1464684518750280085909f7192ee61dae2fff9767014d82f469559e949e

Observation 988419a2-f71a-4481-8f66-eeacda53a897 · outbound

This paper cites Goal misgeneralization in deep reinforcement learning.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Goal misgeneralization in deep reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.186985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.316987Z digest=sha256:093ba8132d19dd5f7c6f565d2a9220280313a0cde55895a82884acc9c3b2a26f

Observation fde9f6c0-a8b0-4a72-8da0-9a1a444cf8ee · outbound

This paper cites Scaling laws for reward model overoptimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Scaling laws for reward model overoptimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.178575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.320021Z digest=sha256:7ee4c00c750e58c9aa9360eb3a00445bf1270361cbf064bb8de03a260b839079

Observation f48b4cb1-5d4a-40ed-9ea3-02446c1a41dc · outbound

This paper cites Improving discriminative capability of reward models in rlhf using contrastive learning.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Improving discriminative capability of reward models in rlhf using contrastive learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.170632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.323132Z digest=sha256:24794d5a150eb6877589914b5345182b87f7fe12db5378b45e7182bd6d5e97e5

Observation d58fa8a7-78b0-417d-8794-d3fe9ad0c387 · outbound

This paper cites Reward Generalization in RLHF: A Topological Perspective.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward Generalization in RLHF: A Topological Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.326063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.326063Z digest=sha256:fcfdc35199ae949ed09ac5d2e9f37e053a4734b2354d6babf48c0f8814f737ae

Observation 2a2abe0b-2f93-4fdb-9f83-54a3fb57ff3b · outbound

This paper cites A note on dpo with noisy preferences & relationship to ipo, 2023.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A note on dpo with noisy preferences & relationship to ipo, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.329460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.329460Z digest=sha256:4bff43e2b439dda25b8196c23c47e9759dd171c6a690beb727489c6bc556a678

Observation 16d62cf5-fd55-4a94-83e5-bc4e82b705ba · outbound

This paper cites Provably robust dpo: aligning language models with noisy feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Provably robust dpo: aligning language models with noisy feedback

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.156044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.332351Z digest=sha256:b9861d286a9a322f237a97f6640c32a32295f8a0e31d7fe0621a660e0b4b2d42

Observation 18218c88-4406-4f6b-8e3e-56db5d47e382 · outbound

This paper cites Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.335604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.335604Z digest=sha256:19123c4b528079b59942dbad646b947052b4e5a6e8c3ab5ee5d863b5368a1908

Observation 4157e457-ac76-4f1e-8d66-a8849691436d · outbound

This paper cites ROPO: Robust Preference Optimization for Large Language Models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment ROPO: Robust Preference Optimization for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.338848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.338848Z digest=sha256:cb85b50ea0b012a21fa5a363525542deef8e0ff12a11eaef4c04d95fea7fcfb7

Observation 96fc20ea-3b02-4aca-b22f-12ad442ef022 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Rank analysis of incomplete block designs: I

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.342349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.342349Z digest=sha256:9062ef9be1b4bddbb42a89bf03348b94e35ca59fe440c96e9b31d5f6accbfa66

Observation 9ff9cb68-246c-472b-a123-41e13de42a5b · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.345277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.345277Z digest=sha256:76252a332c41deb94886da5592c91d26b9e767082a98101c07da6b63ab702c2e

Observation b8ccced3-09c5-4b76-b69f-9e5cb79926c2 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Direct preference optimization: Your language model is secretly a reward model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.137207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.348463Z digest=sha256:123796ff0fc873e01674b7187bb12f078ad7f4f00995acb6836fc88711bf421e

Observation 8ef22d1c-5541-41d8-a330-f0e54ccc63c1 · outbound

This paper cites Confirma- tion bias in human reinforcement learning: Evidence from counterfactual feedback processing.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Confirma- tion bias in human reinforcement learning: Evidence from counterfactual feedback processing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.128912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.351616Z digest=sha256:ed30ad9d24dcfe1a144fda372b12cd08cc8222994835d0a128a3b379d75c4d1c

Observation ba666142-e16b-42e1-9f92-c3b045e2efdb · outbound

This paper cites Pseudo- labeling and confirmation bias in deep semi-supervised learning.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Pseudo- labeling and confirmation bias in deep semi-supervised learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.119261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.354598Z digest=sha256:28079d75cdc96ceb3d5ade2c49fab354e683d3d4bf7e090c41bdeb209fccd281

Observation cabc74fc-328b-4ccb-9ac2-c5c7083b72f4 · outbound

This paper cites Zephyr: Direct distillation of lm alignment.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Zephyr: Direct distillation of lm alignment

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.109684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.357680Z digest=sha256:c606450528a1ae92e5eb3b053beb2310d07c77dbcb6a5baad57f324c1f70fe97

Observation cf523574-1ae8-40ac-9f1a-f61880f0780e · outbound

This paper cites Smith, and Hannaneh Hajishirzi.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Smith, and Hannaneh Hajishirzi

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.361790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.361790Z digest=sha256:509e879955319b18d92eb687cc604cb810f921554e3c791ed1ec97778471f433

Observation b15471eb-e204-4329-b44c-4f008bcf0341 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf, 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Rlhf workflow: From reward modeling to online rlhf, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.365456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.365456Z digest=sha256:75a98325865ddcc791ad0f6fc2d110ed3d3937b3191be3262b7023e92e9fd941

Observation f2e51a55-b2c5-4d1c-9d37-1647e826b773 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.368380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.368380Z digest=sha256:4bce5e9eeb017ff354f0902ab33772f450aabfe98d60252e095b8d2b0f1e56bb

Observation d782e6c3-2cd6-47a9-abcb-73ed3b3fe353 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.371363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.371363Z digest=sha256:4b08ac9a8fa8351640fca4d1d4cd46d4fe493138f2233fdabdcc351a6d208ca9

Observation 54fb0898-f685-49e7-ab98-c4c3a1911a67 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Secrets of RLHF in Large Language Models Part I: PPO

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.374445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.374445Z digest=sha256:253285232f72886919add47a9402e9a91aa9415a09111ca25a135e8326af4991

Observation b550525c-d12b-4d96-b34b-92ca68151aa5 · outbound

This paper cites RM-bench: Benchmarking reward models of language models with subtlety and style.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment RM-bench: Benchmarking reward models of language models with subtlety and style

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.378271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.378271Z digest=sha256:1389dbac6143e370aefac66dc8ba7eed1645af23194b19022440f9f590939ef5

Observation 8ff587ba-c6d5-4423-8122-eff6a4e55f94 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.381195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.381195Z digest=sha256:763894a95ff3091451cbc20ed24c8f417a5300dd64fe60e339b371a6e9c04aa3

Observation 5003dbc6-318f-4334-a41e-5cfc9bc965f0 · outbound

This paper cites Weak-to-strong generalization: Eliciting strong capabilities with weak supervision.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Weak-to-strong generalization: Eliciting strong capabilities with weak supervision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.080708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.384743Z digest=sha256:f324747c9467c33bc496e851906d79af43063e6f4345e8923aadbc241ed80284

Observation 8a0e4840-f774-4ab0-9b3e-0d7ffa901a37 · outbound

This paper cites Weak-to-strong prefer- ence optimization: Stealing reward from weak aligned model.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Weak-to-strong prefer- ence optimization: Stealing reward from weak aligned model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.073280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.387657Z digest=sha256:a0652b9afb551e0fea1b0b5f57ef6277334144d3f9ca2dfc46aaa4c38a79c50d

Observation eca98603-0169-41fa-8fa7-63551354f620 · outbound

This paper cites Adversarial Training of Reward Models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Adversarial Training of Reward Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.390489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.390489Z digest=sha256:48aff87b5b59225e96fc17e796ae5b6932aa36ed17c9223e971eca84b94f97e8

Observation c3ad53a6-8e84-4009-8174-79ef207ada62 · outbound

This paper cites Defining and charac- terizing reward gaming.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Defining and charac- terizing reward gaming

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.065101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.393496Z digest=sha256:06417eb4d2079b49fefe37ca7b71f4c10aa2b29a6a420fbd5d7882e7038a246b

Observation b4d11414-6125-41d7-a0b6-b2b415bb0752 · outbound

This paper cites Reward Shaping to Mitigate Reward Hacking in RLHF.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.396471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.396471Z digest=sha256:bd7faf66d9ec208f977ff9a5059df28fca88fa91b255e623bd00dcc7ed501360

Observation bff2a27b-8fc5-41b3-b048-998c8d611d2c · outbound

This paper cites Odin: disentangled reward mitigates hacking in rlhf.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Odin: disentangled reward mitigates hacking in rlhf

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.057629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.399094Z digest=sha256:3199acc7c1148503ec4d356459afd0b7390ebf1fdd0222ca21e7506bfca27260

Observation a467c576-3136-40ee-849e-12726229087b · outbound

This paper cites Taming Overconfidence in LLMs: Reward Calibration in RLHF.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Taming Overconfidence in LLMs: Reward Calibration in RLHF

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.402572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.402572Z digest=sha256:4515059300fe502562ec23286a73d76c27f760db5a5f840c45faf95faac6b044

Observation 6c73bc40-df62-4342-a705-47d6cbc6fc6c · outbound

This paper cites Overcom- ing reward overoptimization via adversarial policy optimization with lightweight uncertainty estimation, July 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Overcom- ing reward overoptimization via adversarial policy optimization with lightweight uncertainty estimation, July 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.049800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.406480Z digest=sha256:f3dfa345ddb6e66930981a321d58f8006e0c24b69f8fd319fd2f85fee8f88bff

Observation eb3802f2-ce7b-42ab-a53c-98f3a8f4f09c · outbound

This paper cites The energy loss phenomenon in rlhf: A new perspective on mitigating reward hacking, February 2025.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment The energy loss phenomenon in rlhf: A new perspective on mitigating reward hacking, February 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.041660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.409426Z digest=sha256:ccaca464134bc27dda4ea064589f1b0d13824467eb4d0fb67e520e11c71e17e3

Observation 1c655f76-79e9-47ad-b6d7-7eb3abed0f85 · outbound

This paper cites Reward model ensembles help mitigate overoptimizatio.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward model ensembles help mitigate overoptimizatio

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.034242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.412385Z digest=sha256:8339e69388b1c951728342adc9663b3f015c785545a07180a230f31e70351a9f

Observation 52e9f12e-fc13-45a0-885e-f9273de244d7 · outbound

This paper cites Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.026701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.416160Z digest=sha256:808624476036ca1ab1e912bdd62ecc35259fdd2b5a9fb3b0f033d5a10c21b8b2

Observation 100c5a23-f993-43dc-8f28-a38a16053b51 · outbound

This paper cites Reward-robust rlhf in llms, October 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward-robust rlhf in llms, October 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.019019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.419845Z digest=sha256:f8c9c7c417aac59347aacf901b645814c05d01d92a29c63c779ac9d7154de440

Observation 6b50f217-957a-487a-96e1-3f8ebff88faf · outbound

This paper cites Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles, December 2023.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles, December 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.011300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.422374Z digest=sha256:d09e1dad57e59391c552a8822125893c8942e3b0faf963379619041d305cc9e7

Observation 6b62b461-8298-4123-b4e8-66c177a4693e · outbound

This paper cites Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.425267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.425267Z digest=sha256:0048387be3d025959dfab6e03e5bee9b80b709867c1fcfbba0ddd965f12fa705

Observation 57244e02-3310-442d-959f-f0d6e96fe115 · outbound

This paper cites RRM: Robust Reward Model Training Mitigates Reward Hacking.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.428176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.428176Z digest=sha256:ea4723e5ca9692594413e9a8e5631669decc4edb6ff9711054027affa6dc0b84

Observation f8052a73-4469-45a1-af82-b118f48354af · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A general theoretical paradigm to understand learning from human preferences

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.430957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.430957Z digest=sha256:27bd599390a7e38591a629a7ea82fdf225cf674c121baed6ea14d6740d71fc5a

Observation 94c0d798-619d-468d-a22b-5d7c26826fb4 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment KTO: Model Alignment as Prospect Theoretic Optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.434016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.434016Z digest=sha256:892fa98c58dd7fda4207c2a54bea91fa7ec29f684edce30f5d3b948d9b45cbf7

Observation 951e9b08-4d2f-40db-9183-2f049dbd036b · outbound

This paper cites Orpo: Monolithic preference optimization without reference model.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Orpo: Monolithic preference optimization without reference model

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.993237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.437862Z digest=sha256:277e1c4b6c98a75e258e3d9c19bfd98ae3f92254a82d2e95dfd06258a03b37b4

Observation 5be45d8c-f341-4446-a2aa-549e0898254b · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Simpo: Simple preference optimization with a reference-free reward

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.440975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.440975Z digest=sha256:7f2e0aca15d6ec7eb3682b36ac7371b709f60bbd0ce179baef078cf26169c366

Observation 0020e4f9-2fa6-4555-a8a1-c7af9ab23f5b · outbound

This paper cites Is dpo superior to ppo for llm alignment? a comprehensive study.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Is dpo superior to ppo for llm alignment? a comprehensive study

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.978016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.443647Z digest=sha256:89d48f010f3c27e8c29b870473790fdd604d785db2d311e7c5fb104017fc8994

Observation 3b7cb61e-1722-40ea-97b9-284057dc51b6 · outbound

This paper cites Smith, Yejin Choi, and Hannaneh Hajishirzi.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Smith, Yejin Choi, and Hannaneh Hajishirzi

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.968885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.448068Z digest=sha256:5c334fce978c694b38ac5cfe9bbb2fc4381ddbf1ec8d39aad740dccedf125937

Observation c9c7c98f-214b-4e3c-9397-7cc677826137 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Direct Language Model Alignment from Online AI Feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.451174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.451174Z digest=sha256:cf3f3cc41e0169e1f0f1f9820d292d35b3d1c22cfa59af2f3ce9cfc394e5e1a7

Observation 5d220a40-9759-4e58-a761-1d70cbc63392 · outbound

This paper cites DPO-Shift: Shifting the Distribution of Direct Preference Optimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment DPO-Shift: Shifting the Distribution of Direct Preference Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.454892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.454892Z digest=sha256:5e83c48edfe39b381a2bb6717459eceb009b6b3279a36557cd558e390ba620c7

Observation 9557c46c-3769-4070-9a87-bd0f00de5dd9 · outbound

This paper cites Understanding generalization of preference optimization under noisy feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Understanding generalization of preference optimization under noisy feedback

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.959084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.458038Z digest=sha256:f00d3bf842c3810031bf8a2ece4d778113286bb4a9ef14c69273fe3a64c953ce

Observation 90ca6071-d427-4f0c-a23d-4cdaa5e807cb · outbound

This paper cites Robust reinforcement learning from corrupted human feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Robust reinforcement learning from corrupted human feedback

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.949483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.461332Z digest=sha256:4d2474cb97f18e215b5cef61473f953af43fd17f885f2dae93476d75b9f8e6b5

Observation 19fd781e-d323-47f1-9ff3-051373fba9cc · outbound

This paper cites Theory of games and economic behavior, 60th- anniversary, 2007.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Theory of games and economic behavior, 60th- anniversary, 2007

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.941040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.465070Z digest=sha256:ade441be8556083d5035bda65cae125706b48ff49213b60d90e66df32b7688a0

Observation 24cdd193-71e1-4a6d-aec3-67c2dbc7e614 · outbound

This paper cites Multiagent systems: Algorithmic, game-theoretic, and logical foundations.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Multiagent systems: Algorithmic, game-theoretic, and logical foundations

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.468046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.468046Z digest=sha256:fadaf916dbbf3e3b12826b199e8c686488e48d1ea2cc80824ba72520c15a33aa

Observation f30cc94a-75ce-452a-a767-917d0e9a53e8 · outbound

This paper cites [Yes] " is generally preferable to.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment [Yes] " is generally preferable to

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.926924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.472095Z digest=sha256:ba06aa7d1cfa5f05eae3b09dbdf0fde6824a91afcf6830bde1ba5744bd4057a5

Observation 587f0f17-064a-44f2-adff-90fc31e59d3d · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.918496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.475861Z digest=sha256:07feabeef05c7f5b8e949bb9ac96940d5a3badcea82a13223c1b06038549e4c4

Observation 2587aa68-18fb-45d6-b3f9-7186901ca14d · outbound

This paper cites Limitations.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Limitations

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.910990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.479196Z digest=sha256:92ce4a96c7d433e4ef19b4da4a09c65c6a4bb644e554fe582aa33cb8a09b864f

Observation 0ae68a1d-93c7-4ca9-bf0f-efb0910d40e3 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.902892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.483442Z digest=sha256:9c1b15fbf2a8b27feb84b4ac2990c0764f71b8576befbf915e7c9c9c07123749

Observation 45575863-75fc-491c-b97b-d3828cd3544c · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.894349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.486539Z digest=sha256:b909febcb6fc79427b40e8f870f685590e8556f5d4feb08f8e531e3577921120

Observation 28b3ac97-92c3-44fd-808d-72b200a62c21 · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.884914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.490836Z digest=sha256:45fcf42872952612745fb4066db229b205908beea94d35f8468f4f0508646c51

Observation ca42c146-6f32-41b1-a962-51a0703b86ee · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.876158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.494510Z digest=sha256:d969251037dd92f495a7f80b8825b72d85caa3fd9895f594955ff38ad1f82586

Observation 6e4a06e4-3921-4646-909d-6cfebcadef07 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.866146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.497449Z digest=sha256:16abbf271a324de26cc51cb258bb2d177d8f5c98db31889d5edb06ce2f68242e

Observation e0aa576d-be3d-41bc-a356-eb2cdc6abde9 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.855662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.500484Z digest=sha256:ebf1a1d3ac169e642ea06a0e25aad607cb89193f7450d4d8994fad5f7dc2808d

Observation c43dba8f-bd3f-4f05-ad03-59ed9980f531 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.843927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.505139Z digest=sha256:e7de2f9ce5eeebf7917774759da6964f0c27e10ea007afd04986bde2a814a46b

Observation 76be5127-3bb9-4fb1-bdfe-168239b3e32c · outbound

This paper cites an unresolved cited work.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:19:41.828840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.508480Z digest=sha256:4b5fa967f11b6c80fd4596ecd7f945f5065422d7d9a86d3a6fbce9a1c45872c8

Observation 187b16d1-8534-48d7-ba67-1575806c6f2b · outbound

This paper cites an unresolved cited work.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:19:41.811120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.512195Z digest=sha256:e0ee6e06fe442dee21da1a8a291f3cb56bad7c6f8789d6c6d5536232706e3252

Observation a29dce32-996e-4b09-9b85-e560dff6dfb2 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not use existing assets

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.792568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.515056Z digest=sha256:0a6504763255a56b27a6755a8955f8c5ebe13a63306a3fc7edfbf41a26d501cc

Observation 27395a9e-b895-4c14-88eb-bfb3742071e2 · outbound

This paper cites • Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment • Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.772923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.517980Z digest=sha256:3b597c7e6222396626696b17206a1b579a1e49fb382476739bee4db92e10ec6a

Observation 4ec17451-33bb-4202-ad24-41121e2d824b · outbound

This paper cites 32 Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment 32 Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.754738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.520830Z digest=sha256:9cd600728ffefe4fc32540453c2a1f8743d52f77cc8715bcbe148562e9956283

Observation 009638ff-9737-402e-ab8d-ae28df7ecc55 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.737498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.524216Z digest=sha256:d1df9f832ccd94ed42da1e406f4944d7aa86f27d1c921d98352839de8262090e

Observation 767017fa-c7a9-451f-8aa4-13ab29e25803 · outbound

This paper cites Answer: [NA] Justification: LLM is used only for editing, or formatting and does not impact the core methodology, scientific rigorousness.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Answer: [NA] Justification: LLM is used only for editing, or formatting and does not impact the core methodology, scientific rigorousness

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.722866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:19:41.527363Z digest=sha256:84a6b168d870022ace1864d180188a432595789be4ba1e956817703806b0eff5

Pith citing papers

Observation b431773f-cf7b-4d93-9727-f0dfce4a409b · inbound

AgentV-RL: Scaling Reward Modeling with Agentic Verifier cites this paper.

AgentV-RL: Scaling Reward Modeling with Agentic Verifier Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:32:51.696658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T08:32:38.883019Z digest=sha256:c85ed3de315e921d958b25cc8d314e9d147844ed25053199a00e0a6cea3c204c