Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:19:41.527363Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 1 inbound Pith citation observation for arXiv:2505.10597.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:19:41.527363Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T08:32:38.883019Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T08:32:51.695159Z
80 of 80 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3030dd8e-2d36-4d2c-8741-5ed4fca3c232 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ca5174f-0ea4-406a-a676-b463e0ca9104 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Training language models to follow instructions with human feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa577b07-d0a0-418e-97dd-183650ecf221 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00e2721e-cb01-40d5-ae25-630bdfb76a8c · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92402796-48dc-41fc-9aea-a29af5247491 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Better Process Supervision with Bi-directional Rewarding Signals
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d239447-3e69-4824-bafa-4fad15c5596c · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward Function Design in Reinforcement Learning, pages 25–33
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation be8753fc-b9cb-4731-9515-b7a1cb6c1dab · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3289ef25-83b8-40ca-baa2-b91abff2fa82 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Gpt-4 technical report, 2024
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21ad8af1-130a-4f9f-bb6f-4d01402cdb91 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8d58a6-9201-4524-b469-2666a30a531d · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Skywork-reward: Bag of tricks for reward modeling in llms, October 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ba77e19c-df9a-46d0-81a4-eeb13a0ea201 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment RMB: Comprehensively benchmarking reward models in LLM alignment
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 082da1a0-2646-43a7-a0a0-e67ca2f656d6 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Helpsteer 2: Open-source dataset for training top-performing reward models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 60ff893e-87d5-4441-bd2e-aab7d1ffb549 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 38fba78c-52a5-4611-8f9d-0db0dae359b5 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Impact of preference noise on the alignment performance of generative language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 53221f72-75e5-41ce-849e-7ecbc4f10ae5 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Improving reinforcement learning from human feedback using contrastive rewards, March 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 988419a2-f71a-4481-8f66-eeacda53a897 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Goal misgeneralization in deep reinforcement learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fde9f6c0-a8b0-4a72-8da0-9a1a444cf8ee · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Scaling laws for reward model overoptimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f48b4cb1-5d4a-40ed-9ea3-02446c1a41dc · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Improving discriminative capability of reward models in rlhf using contrastive learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d58fa8a7-78b0-417d-8794-d3fe9ad0c387 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward Generalization in RLHF: A Topological Perspective
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a2abe0b-2f93-4fdb-9f83-54a3fb57ff3b · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A note on dpo with noisy preferences & relationship to ipo, 2023
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16d62cf5-fd55-4a94-83e5-bc4e82b705ba · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Provably robust dpo: aligning language models with noisy feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 18218c88-4406-4f6b-8e3e-56db5d47e382 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4157e457-ac76-4f1e-8d66-a8849691436d · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment ROPO: Robust Preference Optimization for Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96fc20ea-3b02-4aca-b22f-12ad442ef022 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Rank analysis of incomplete block designs: I
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ff9cb68-246c-472b-a123-41e13de42a5b · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8ccced3-09c5-4b76-b69f-9e5cb79926c2 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Direct preference optimization: Your language model is secretly a reward model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8ef22d1c-5541-41d8-a330-f0e54ccc63c1 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Confirma- tion bias in human reinforcement learning: Evidence from counterfactual feedback processing
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ba666142-e16b-42e1-9f92-c3b045e2efdb · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Pseudo- labeling and confirmation bias in deep semi-supervised learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cabc74fc-328b-4ccb-9ac2-c5c7083b72f4 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Zephyr: Direct distillation of lm alignment
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cf523574-1ae8-40ac-9f1a-f61880f0780e · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Smith, and Hannaneh Hajishirzi
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b15471eb-e204-4329-b44c-4f008bcf0341 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Rlhf workflow: From reward modeling to online rlhf, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2e51a55-b2c5-4d1c-9d37-1647e826b773 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d782e6c3-2cd6-47a9-abcb-73ed3b3fe353 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54fb0898-f685-49e7-ab98-c4c3a1911a67 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Secrets of RLHF in Large Language Models Part I: PPO
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b550525c-d12b-4d96-b34b-92ca68151aa5 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment RM-bench: Benchmarking reward models of language models with subtlety and style
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ff587ba-c6d5-4423-8122-eff6a4e55f94 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5003dbc6-318f-4334-a41e-5cfc9bc965f0 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8a0e4840-f774-4ab0-9b3e-0d7ffa901a37 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Weak-to-strong prefer- ence optimization: Stealing reward from weak aligned model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eca98603-0169-41fa-8fa7-63551354f620 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Adversarial Training of Reward Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3ad53a6-8e84-4009-8174-79ef207ada62 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Defining and charac- terizing reward gaming
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b4d11414-6125-41d7-a0b6-b2b415bb0752 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bff2a27b-8fc5-41b3-b048-998c8d611d2c · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Odin: disentangled reward mitigates hacking in rlhf
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a467c576-3136-40ee-849e-12726229087b · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Taming Overconfidence in LLMs: Reward Calibration in RLHF
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c73bc40-df62-4342-a705-47d6cbc6fc6c · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Overcom- ing reward overoptimization via adversarial policy optimization with lightweight uncertainty estimation, July 2024
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eb3802f2-ce7b-42ab-a53c-98f3a8f4f09c · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment The energy loss phenomenon in rlhf: A new perspective on mitigating reward hacking, February 2025
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1c655f76-79e9-47ad-b6d7-7eb3abed0f85 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward model ensembles help mitigate overoptimizatio
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 52e9f12e-fc13-45a0-885e-f9273de244d7 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 100c5a23-f993-43dc-8f28-a38a16053b51 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward-robust rlhf in llms, October 2024
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6b50f217-957a-487a-96e1-3f8ebff88faf · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles, December 2023
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6b62b461-8298-4123-b4e8-66c177a4693e · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57244e02-3310-442d-959f-f0d6e96fe115 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment RRM: Robust Reward Model Training Mitigates Reward Hacking
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8052a73-4469-45a1-af82-b118f48354af · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A general theoretical paradigm to understand learning from human preferences
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c0d798-619d-468d-a22b-5d7c26826fb4 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment KTO: Model Alignment as Prospect Theoretic Optimization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 951e9b08-4d2f-40db-9183-2f049dbd036b · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Orpo: Monolithic preference optimization without reference model
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5be45d8c-f341-4446-a2aa-549e0898254b · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Simpo: Simple preference optimization with a reference-free reward
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0020e4f9-2fa6-4555-a8a1-c7af9ab23f5b · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Is dpo superior to ppo for llm alignment? a comprehensive study
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3b7cb61e-1722-40ea-97b9-284057dc51b6 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Smith, Yejin Choi, and Hannaneh Hajishirzi
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c9c7c98f-214b-4e3c-9397-7cc677826137 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Direct Language Model Alignment from Online AI Feedback
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d220a40-9759-4e58-a761-1d70cbc63392 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment DPO-Shift: Shifting the Distribution of Direct Preference Optimization
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9557c46c-3769-4070-9a87-bd0f00de5dd9 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Understanding generalization of preference optimization under noisy feedback
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 90ca6071-d427-4f0c-a23d-4cdaa5e807cb · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Robust reinforcement learning from corrupted human feedback
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 19fd781e-d323-47f1-9ff3-051373fba9cc · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Theory of games and economic behavior, 60th- anniversary, 2007
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 24cdd193-71e1-4a6d-aec3-67c2dbc7e614 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Multiagent systems: Algorithmic, game-theoretic, and logical foundations
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f30cc94a-75ce-452a-a767-917d0e9a53e8 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment [Yes] " is generally preferable to
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 587f0f17-064a-44f2-adff-90fc31e59d3d · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2587aa68-18fb-45d6-b3f9-7186901ca14d · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Limitations
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0ae68a1d-93c7-4ca9-bf0f-efb0910d40e3 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include theoretical results
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 45575863-75fc-491c-b97b-d3828cd3544c · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 28b3ac97-92c3-44fd-808d-72b200a62c21 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that paper does not include experiments requiring code
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ca42c146-6f32-41b1-a962-51a0703b86ee · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6e4a06e4-3921-4646-909d-6cfebcadef07 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e0aa576d-be3d-41bc-a356-eb2cdc6abde9 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c43dba8f-bd3f-4f05-ad03-59ed9980f531 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 76be5127-3bb9-4fb1-bdfe-168239b3e32c · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 187b16d1-8534-48d7-ba67-1575806c6f2b · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a29dce32-996e-4b09-9b85-e560dff6dfb2 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not use existing assets
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 27395a9e-b895-4c14-88eb-bfb3742071e2 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment • Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4ec17451-33bb-4202-ad24-41121e2d824b · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment 32 Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 009638ff-9737-402e-ab8d-ae28df7ecc55 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 767017fa-c7a9-451f-8aa4-13ab29e25803 · outbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Answer: [NA] Justification: LLM is used only for editing, or formatting and does not impact the core methodology, scientific rigorousness
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b431773f-cf7b-4d93-9727-f0dfce4a409b · inbound
AgentV-RL: Scaling Reward Modeling with Agentic Verifier Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.