Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2403.13787.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T19:52:50.233278Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
5
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 4837305e-d645-41f6-abde-f621aa5cc087 · inbound
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing RewardBench: Evaluating Reward Models for Language Modeling
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f60d7cd4-f05a-43a8-bbe0-7b35578a78e3 · inbound
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs RewardBench: Evaluating Reward Models for Language Modeling
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation be6c4a43-db3f-4f82-a220-d8163efb1b52 · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods RewardBench: Evaluating Reward Models for Language Modeling
Reference 120
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 936ed788-f2ff-4077-8299-61906f93abfb · inbound
Qwen2.5 Technical Report RewardBench: Evaluating Reward Models for Language Modeling
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c045ae24-dc61-44d1-badd-76157314aa2a · inbound
Improving Video Generation with Human Feedback RewardBench: Evaluating Reward Models for Language Modeling
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7d46cc0-4d97-4312-9f6c-32857585d627 · inbound
RewardBench 2: Advancing Reward Model Evaluation RewardBench: Evaluating Reward Models for Language Modeling
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7a841bc7-63da-41a7-ba87-bb019948e034 · inbound
Exploring the Secondary Risks of Large Language Models RewardBench: Evaluating Reward Models for Language Modeling
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ecefc4e-1d42-41d6-8b40-73afba4c166b · inbound
TabArena: A Living Benchmark for Machine Learning on Tabular Data RewardBench: Evaluating Reward Models for Language Modeling
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 02e09a74-85eb-419c-b489-28852c0dc2f3 · inbound
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap RewardBench: Evaluating Reward Models for Language Modeling
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3ec0fdf-f59d-440d-9450-2477156f23d6 · inbound
Controlling Multimodal LLMs via Reward-guided Decoding RewardBench: Evaluating Reward Models for Language Modeling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe8d6d21-146a-42de-9908-d1c58720b217 · inbound
Hermes 4 Technical Report RewardBench: Evaluating Reward Models for Language Modeling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b4df1d4-e9d8-435f-ae1b-2cbf8d1d07bf · inbound
Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning RewardBench: Evaluating Reward Models for Language Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf22cafc-b38e-42bd-a6d0-9e4780abe1ed · inbound
HEAL: A Hypothesis-Based Preference-Aware Analysis Framework RewardBench: Evaluating Reward Models for Language Modeling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3727717-6844-4be7-9d66-1c15b74f0283 · inbound
Evalet: Evaluating Large Language Models through Functional Fragmentation RewardBench: Evaluating Reward Models for Language Modeling
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1e925bf3-af47-4214-bfe5-5b02f232054a · inbound
Adaptive Margin RLHF via Preference over Preferences RewardBench: Evaluating Reward Models for Language Modeling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ff3ac9-20ca-4966-b64f-d54c27368c1f · inbound
Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images RewardBench: Evaluating Reward Models for Language Modeling
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2ec61eb6-808a-4173-9cb2-004f4a719cde · inbound
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling RewardBench: Evaluating Reward Models for Language Modeling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bb48b874-e55b-4f2c-8572-c3f6a9ed6cb4 · inbound
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning RewardBench: Evaluating Reward Models for Language Modeling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7edbe698-b7a3-458e-9e3d-ec620a6efa60 · inbound
SAM 3D: 3Dfy Anything in Images RewardBench: Evaluating Reward Models for Language Modeling
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb3c57c4-7e89-44c6-8440-31393d531545 · inbound
A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data RewardBench: Evaluating Reward Models for Language Modeling
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad7b4d95-86db-494a-8959-c9d62fcc80f7 · inbound
AI Can Learn Scientific Taste RewardBench: Evaluating Reward Models for Language Modeling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e046eb5-bcad-4b7c-a6db-61232767e668 · inbound
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems RewardBench: Evaluating Reward Models for Language Modeling
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 79d4d7a1-b5a7-4336-ad5c-715649dda6db · inbound
Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents RewardBench: Evaluating Reward Models for Language Modeling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed37f525-9e2d-4ef8-9279-cf37a138580b · inbound
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model RewardBench: Evaluating Reward Models for Language Modeling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb242eb0-85d6-4c4d-ab0b-1890bbbbcbea · inbound
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines RewardBench: Evaluating Reward Models for Language Modeling
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2bcb6ca-2584-4584-beb6-d2a9e5679d51 · inbound
Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning RewardBench: Evaluating Reward Models for Language Modeling
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15ae18c7-113b-450f-b49e-deaa22192fbb · inbound
Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning RewardBench: Evaluating Reward Models for Language Modeling
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e458652d-1cae-4cf4-ac8c-a05762a10c21 · inbound
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone RewardBench: Evaluating Reward Models for Language Modeling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3255c4e-82e1-4cae-b4c5-c13fa899701c · inbound
Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders RewardBench: Evaluating Reward Models for Language Modeling
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc8c028b-4999-46e6-98cb-c0ec40f0b283 · inbound
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR RewardBench: Evaluating Reward Models for Language Modeling
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 54af2848-6ec6-4303-abc2-74d4261a815c · inbound
Pairwise Reference Alignment as a Model-Level Ordinal Observable RewardBench: Evaluating Reward Models for Language Modeling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f55d6e4-a7a1-4529-8c44-ed14ca1e86b8 · inbound
PReMISE: Policy Rubrics as Measurement Specifications for LLM Judges RewardBench: Evaluating Reward Models for Language Modeling
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 363a698c-d10d-4f83-a758-774e45d93061 · inbound
EST-PRM: Stress-Testing Process Reward Models Before They Become Load-Bearing RewardBench: Evaluating Reward Models for Language Modeling
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation add9a932-09bc-4dfd-9171-9ca2dde3b2a5 · inbound
A Finite-Calibration Regime Map for LLM Judge Panels RewardBench: Evaluating Reward Models for Language Modeling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 288c43ac-8ee4-4bce-89b5-c9943aeebd86 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization RewardBench: Evaluating Reward Models for Language Modeling
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 074aa224-c373-46aa-ae16-23f30a1aab78 · inbound
Discriminatory Compliance: How LLMs Answer Queries from Protected Groups RewardBench: Evaluating Reward Models for Language Modeling
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29dd4dd1-dfcd-457f-a273-5deadf5ffba2 · inbound
Addressing Over-Refusal in LLMs with Competing Rewards RewardBench: Evaluating Reward Models for Language Modeling
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 157aae37-3740-48b6-9780-a27d40860516 · inbound
QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents RewardBench: Evaluating Reward Models for Language Modeling
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50d8a596-f9d8-4d93-b405-c6eefd3e9c9d · inbound
Multi-Turn On-Policy Distillation with Prefix Replay RewardBench: Evaluating Reward Models for Language Modeling
Reference 143
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d5d7f8e-46e4-4991-8960-a65454de3acb · inbound
Multi-Turn On-Policy Distillation with Prefix Replay RewardBench: Evaluating Reward Models for Language Modeling
Reference 144
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46da7b85-3aed-42ef-a3a3-0aba1766dafe · inbound
When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability RewardBench: Evaluating Reward Models for Language Modeling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96b1c4e1-4c5b-43a2-a34c-d67a402cb7ea · inbound
SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning RewardBench: Evaluating Reward Models for Language Modeling
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32b72cbd-74b6-441e-9202-a0963abf23a3 · inbound
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias RewardBench: Evaluating Reward Models for Language Modeling
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ee9dde9-c887-4ede-b44e-8cfc9300ca8b · inbound
Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary RewardBench: Evaluating Reward Models for Language Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41b013a8-943d-46c9-b0db-4fe04a4b6371 · inbound
Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis RewardBench: Evaluating Reward Models for Language Modeling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1d217c-8280-4bb1-8336-802bfeaecb64 · inbound
Test-Time Scaling via Error Localization RewardBench: Evaluating Reward Models for Language Modeling
Reference 119
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d10d2cd2-a4c2-4cb7-88a7-3f85aa373f0c · inbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback RewardBench: Evaluating Reward Models for Language Modeling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d6e2847-5870-4ffc-acac-8822098b0191 · inbound
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges RewardBench: Evaluating Reward Models for Language Modeling
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.