Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:07.168303Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 100 of 112 outbound references and 8 inbound Pith citation observations for arXiv:2506.03637.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:07.168303Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:46.427272Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
100 of 112 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 9b279bdf-f850-4297-90fd-6741b5310583 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Fine-Tuning Language Models from Human Preferences
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56102ffb-be24-4cca-8208-538f0502a5a6 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Training language models to follow instructions with human feedback,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f72af75e-4930-41bf-851a-696ec6e5388e · outbound
RewardAnything: Generalizable Principle-Following Reward Models Deep reinforcement learning from human preferences,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85a9511f-355b-47bb-aee9-2d38ad65ee30 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Learning to summarize with human feedback,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9312a7b4-5134-49da-aaa8-dd7b93a8885d · outbound
RewardAnything: Generalizable Principle-Following Reward Models A General Language Assistant as a Laboratory for Alignment
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 882b93d0-697f-4cf3-b839-d141d3bdef4e · outbound
RewardAnything: Generalizable Principle-Following Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f079aa2-70ae-4aeb-8524-49e57e6fd4b7 · outbound
RewardAnything: Generalizable Principle-Following Reward Models RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eab97cf-aac4-410b-a980-cb5daf48fca0 · outbound
RewardAnything: Generalizable Principle-Following Reward Models A survey of reinforcement learning from human feedback,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba520f4b-043f-456e-a640-7cc474ca65e9 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddcae839-bad5-4a40-90aa-3b24b2270704 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Judging llm-as-a-judge with mt-bench and chatbot arena,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e02376c-6f72-43ad-90fd-b9cf08e846b3 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Evaluating Large Language Models at Evaluating Instruction Following
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 647a9c0a-a442-4d10-ba1b-77b7e8c222b5 · outbound
RewardAnything: Generalizable Principle-Following Reward Models KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e54f08c-bccc-4bfa-94f7-a66e374e2086 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Helpsteer2: Open-source dataset for training top-performing reward models,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31155be0-5e2c-4c8f-baf2-db78c69e4cb1 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Alpacafarm: A simulation framework for methods that learn from human feedback,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a650e7-7f8c-4641-ae0c-28e9bd3de960 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Let’s verify step by step,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a61d097b-10a7-4888-b9bb-c5df544d0818 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Rank analysis of incomplete block designs: I. the method of paired comparisons,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 897d129d-c17a-4482-b52e-b6f713c33977 · outbound
RewardAnything: Generalizable Principle-Following Reward Models PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b388b80-12ab-40f5-bf29-edf188dc48eb · outbound
RewardAnything: Generalizable Principle-Following Reward Models Prometheus: Inducing fine-grained evaluation capability in language models,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35da3c3b-5ae6-42f6-ae7e-6063175a4487 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Understanding dataset difficulty with v-usable information,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3b495a0-ac32-4338-a9eb-9404eed09896 · outbound
RewardAnything: Generalizable Principle-Following Reward Models RewardBench: Evaluating Reward Models for Language Modeling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 284e87a8-4613-42db-9312-a2f323338b5d · outbound
RewardAnything: Generalizable Principle-Following Reward Models How to Evaluate Reward Models for RLHF
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e30463e9-60e3-492e-8b33-8d61471e0f3a · outbound
RewardAnything: Generalizable Principle-Following Reward Models MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cbe3cfc-5f4c-4f0c-b08c-fceb2f81098c · outbound
RewardAnything: Generalizable Principle-Following Reward Models SALMON: Self-Alignment with Instructable Reward Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06bf0455-4fe3-404f-9c02-5543eb4f7fd5 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Inference-time scaling for generalist reward modeling,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6846a318-8e26-4a35-a5f3-c3e54524ba47 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Rm-r1: Reward modeling as reasoning,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 385a7958-a8a7-4e24-8094-3a29b2cc48a5 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c913abf-9885-4020-8c26-9ea29f1c011b · outbound
RewardAnything: Generalizable Principle-Following Reward Models Proximal Policy Optimization Algorithms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 585bd097-d7a4-4279-a548-f10e14c4d904 · outbound
RewardAnything: Generalizable Principle-Following Reward Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e9c6673-1a60-41f9-af83-2bb8bfbdd2bc · outbound
RewardAnything: Generalizable Principle-Following Reward Models What makes a reward model a good teacher? an optimization perspective,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4de5ecbc-5476-40f2-ae93-403b085cc640 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a27bbd21-8242-4092-b6a8-dd026c14e7e7 · outbound
RewardAnything: Generalizable Principle-Following Reward Models OffsetBias: Leveraging Debiased Data for Tuning Evaluators
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed00d0e-6f1d-4820-a85f-fc1adb3eada8 · outbound
RewardAnything: Generalizable Principle-Following Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff7853db-ca4b-4728-a698-24b306bfa837 · outbound
RewardAnything: Generalizable Principle-Following Reward Models XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3caeb390-76d5-4292-abbc-7d1102c841bb · outbound
RewardAnything: Generalizable Principle-Following Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7c2a000-22cd-47a4-8c5e-76e767e818af · outbound
RewardAnything: Generalizable Principle-Following Reward Models Transforming and Combining Rewards for Aligning Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f4e3b2f-1766-4840-8825-57e7aacc2330 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Heimdall: test-time scaling on the generative verification
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c81cf0-383d-47b2-8810-b36116b58808 · outbound
RewardAnything: Generalizable Principle-Following Reward Models When to solve, when to verify: Compute-optimal problem solving and generative verification for llm reasoning,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4655333-ac14-4745-9507-ef5406937add · outbound
RewardAnything: Generalizable Principle-Following Reward Models Large Language Models are Better Reasoners with Self-Verification
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a05919-d3b7-4492-afd8-954af94a293b · outbound
RewardAnything: Generalizable Principle-Following Reward Models GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67a8321e-6415-4aa3-bdc3-7eefe6eaae70 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Dynamic Multi-Reward Weighting for Multi-Style Controllable Generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 036900c6-a90b-497d-a489-085704e09552 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3838075c-3529-4139-a0c1-84cdfd69ff06 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1464911-8e0c-4a5c-9e45-9069c5ea5cba · outbound
RewardAnything: Generalizable Principle-Following Reward Models Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85489b8f-af8a-4d90-a8ff-eac910f95769 · outbound
RewardAnything: Generalizable Principle-Following Reward Models An Empirical Analysis of Uncertainty in Large Language Model Evaluations
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a17a9eb-7e34-45b2-871d-73a3db0b36cf · outbound
RewardAnything: Generalizable Principle-Following Reward Models Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c3df9b4-277a-4135-b4c4-75ffb9785030 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 726b7ead-80f1-4df5-a9f2-480b78a979cd · outbound
RewardAnything: Generalizable Principle-Following Reward Models Reward Model Ensembles Help Mitigate Overoptimization
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b1f6274-a680-4173-82d8-64c2de1361b1 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Language models are few-shot learners,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fda1a968-0a31-4e89-a696-40536e1fd5b8 · outbound
RewardAnything: Generalizable Principle-Following Reward Models PaLM 2 Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 789a1e01-ae63-467d-9284-3fd6be070583 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Gpt-4 technical report,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d096eb2-2766-4bb7-8d6f-d29ce7e81c16 · outbound
RewardAnything: Generalizable Principle-Following Reward Models A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85a8a507-cc31-4994-bf67-281478cd8ce0 · outbound
RewardAnything: Generalizable Principle-Following Reward Models A fast learning algorithm for deep belief nets,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29c94fe2-a672-4bff-8d4d-b1458d710815 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Goodfellow, Y
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 979e9e92-d7d7-4316-88ec-3ec571473c55 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79f2caed-bde8-4a7a-88e4-8b78accb4f8e · outbound
RewardAnything: Generalizable Principle-Following Reward Models Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1337053c-cbc1-4d73-99c8-b887f3c611c2 · outbound
RewardAnything: Generalizable Principle-Following Reward Models How to fine-tune bert for text classification?
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f68efd9f-1b6f-4349-9070-c4fa21b6fbbc · outbound
RewardAnything: Generalizable Principle-Following Reward Models Natural language question answering: the view from here,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16119788-8a5d-4f1a-a803-8ea0c2bdd1f5 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Natural questions: a benchmark for question answering research,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d7d9a4-f9c8-494e-afb9-e36450842e50 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Attention is all you need,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23240761-a0d3-47f9-8aaa-2ffa09540afd · outbound
RewardAnything: Generalizable Principle-Following Reward Models A Survey on Evaluation of Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e48dfae5-2c14-485d-910a-44cfe1603345 · outbound
RewardAnything: Generalizable Principle-Following Reward Models O’Reilly Media, Inc
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bb07342-09f8-4b42-b70f-ebacad01aa41 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Deep learning tuning playbook,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34600fe8-cbf0-4e47-bebe-81c353a602b5 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Glue: A multi-task bench- mark and analysis platform for natural language understanding,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10379657-7d58-4459-81fd-de8febfd5de2 · outbound
RewardAnything: Generalizable Principle-Following Reward Models GLUE-X: Evaluating Natural Language Understanding Models from an Out-of-distribution Generalization Perspective
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d338659-03da-4d26-88a0-ba50fec7d0f2 · outbound
RewardAnything: Generalizable Principle-Following Reward Models The Llama 3 Herd of Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2afbe2cd-240a-4140-a316-0e57d8408983 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Lora: Low-rank adaptation of large language models,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91c101b8-0f39-427b-a6c4-5a9c707b8781 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Training Verifiers to Solve Math Word Problems
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9171a74f-991e-4ff4-b18d-8ed9ca0db66c · outbound
RewardAnything: Generalizable Principle-Following Reward Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b62b8a95-8eb5-48b6-a927-22fa4dd3188b · outbound
RewardAnything: Generalizable Principle-Following Reward Models Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a201190d-f322-4fac-8904-48bca019f5e7 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning,
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18585e90-09fc-4b26-8905-2e866c416456 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Supervised Knowledge Makes Large Language Models Better In-context Learners
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8225683a-db5b-4222-a75b-5c71558b5fc6 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 404cbede-a3c2-48d9-a35f-4fe3734b43e8 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Large language models are zero- shot reasoners,
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83ab8c1a-22ac-4a8f-9eac-3cb3de3b8b80 · outbound
RewardAnything: Generalizable Principle-Following Reward Models LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bcc8e31-95e2-4215-9e50-b123a60c7ee0 · outbound
RewardAnything: Generalizable Principle-Following Reward Models The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f446cdd1-8045-4f10-97d8-1fdee4fe94fa · outbound
RewardAnything: Generalizable Principle-Following Reward Models UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1afc5850-2586-4b53-af6a-4fe07349a318 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Qwen2 Technical Report
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32811580-272d-4435-afb6-65154165c245 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Team, “Qwen3,” April 2025
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f7d287e-2326-4723-9c4c-36296bc5f83c · outbound
RewardAnything: Generalizable Principle-Following Reward Models HybridFlow: A Flexible and Efficient RLHF Framework
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed5251c-e61f-408c-89ff-b365ffdb9234 · outbound
RewardAnything: Generalizable Principle-Following Reward Models A framework for training large language models for code generation via proximal policy optimization,
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08592c85-be79-4986-bd43-1d4168896b0e · outbound
RewardAnything: Generalizable Principle-Following Reward Models Gemini: A Family of Highly Capable Multimodal Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c5e91f6-313d-488d-af17-339cb0313ca2 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Generative Reward Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db7ccc68-7dbe-4d98-b54a-14efed3aedd2 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Improving context-aware preference modeling for language models,
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a617fe4-8e3b-498f-b90c-1866807daa7d · outbound
RewardAnything: Generalizable Principle-Following Reward Models Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 87488f2b-c604-425b-8a52-626fdfa5666f · outbound
RewardAnything: Generalizable Principle-Following Reward Models Best practices for the human evaluation of automatically generated text,
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa1a059a-bdf8-4893-9288-213356b36821 · outbound
RewardAnything: Generalizable Principle-Following Reward Models DeepSeek-V3 Technical Report
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4be1c353-9fea-41a9-b0d1-faf73251904c · outbound
RewardAnything: Generalizable Principle-Following Reward Models Leveraging Large Language Models for NLG Evaluation: Advances and Challenges
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aac12416-9c89-4f07-a085-38a7e4404567 · outbound
RewardAnything: Generalizable Principle-Following Reward Models FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27f6e241-7a3b-4db7-86b9-4d656f555e4c · outbound
RewardAnything: Generalizable Principle-Following Reward Models From generation to judgment: Opportunities and challenges of llm-as-a-judge,
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f5f3d2-ef74-4f68-a451-c39f90ca04b3 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Automating Dataset Updates Towards Reliable and Timely Evaluation of Large Language Models
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3787a9e0-2fa5-4f7c-a854-d2cd1140ca1d · outbound
RewardAnything: Generalizable Principle-Following Reward Models A Comprehensive Survey of Contamination Detection Methods in Large Language Models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a4e773d-0972-41b0-b0c5-f77d719df2f4 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Prompt-to-Leaderboard
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16fb42d0-8d71-4213-994e-65da12dda5f6 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Scaling laws for reward model overoptimization,
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a978274c-3e04-4b41-9561-e696d770c2d0 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2782c99-9211-4367-ad11-4a63a5f59a4f · outbound
RewardAnything: Generalizable Principle-Following Reward Models Critique-out-Loud Reward Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f3f9ad9-0e3f-4bcf-b043-4f51188e8a69 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Constitutional AI: Harmlessness from AI Feedback
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bedcc321-eb0c-4115-b9dd-2143ccb00f25 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Is Elo Rating Reliable? A Study Under Model Misspecification
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d3eb81df-c2b1-4ff0-bf88-6ac5df644d5b · outbound
RewardAnything: Generalizable Principle-Following Reward Models On the biology of a large language model,
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea09ea2-b1c6-4bdd-a03b-49a91e102ab8 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Gpt-4.1 and gpt-4.1 nano overview,
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3ec7d72a-c526-44ee-a919-72b4c4096480 · outbound
RewardAnything: Generalizable Principle-Following Reward Models Gemini 2.5: Our most intelligent ai model,
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b69cfabc-835e-44b2-8e20-58f323790bed · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle RewardAnything: Generalizable Principle-Following Reward Models
Reference 231
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f72b0cf3-5632-41cc-82c9-10b554b62a62 · inbound
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards RewardAnything: Generalizable Principle-Following Reward Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4dbe7bcd-ddd2-4f60-b0a1-8deb8beb1af6 · inbound
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents RewardAnything: Generalizable Principle-Following Reward Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2db5892d-8c19-45ac-9fec-1ef606dfc1dc · inbound
Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization RewardAnything: Generalizable Principle-Following Reward Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3c8adfd8-f24e-475f-b765-777ab48bdc7f · inbound
DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation RewardAnything: Generalizable Principle-Following Reward Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6aac035e-19a1-41be-b03e-d3bcaed68a42 · inbound
DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation RewardAnything: Generalizable Principle-Following Reward Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 47d74668-400a-4c4f-97b1-af72d8570a03 · inbound
Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill RewardAnything: Generalizable Principle-Following Reward Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4d9bc8bb-3426-4345-92d4-2e42081b3aeb · inbound
Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics RewardAnything: Generalizable Principle-Following Reward Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.