Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2406.08673.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:53:03.339487Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation d655d4cf-9b93-415e-84c2-e0a144d85aaf · inbound
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs HelpSteer2: Open-source dataset for training top-performing reward models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6a029ccd-4f98-48ef-9ea0-a947deed1b80 · inbound
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach HelpSteer2: Open-source dataset for training top-performing reward models
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1109fac1-d0c1-485c-ba78-d9ec08f5d82d · inbound
Reinforcement Learning from Human Feedback HelpSteer2: Open-source dataset for training top-performing reward models
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93f23228-ef50-42f8-802c-ad68b2f37369 · inbound
Discriminative Policy Optimization for Token-Level Reward Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e57198dc-bc3b-4e25-9d4e-e0450c9c8766 · inbound
TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence HelpSteer2: Open-source dataset for training top-performing reward models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad03ac79-0afa-486b-b655-aa3c4bee344f · inbound
RewardBench 2: Advancing Reward Model Evaluation HelpSteer2: Open-source dataset for training top-performing reward models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3b23bd3-640a-4e84-89ff-091dd43c11d9 · inbound
PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization HelpSteer2: Open-source dataset for training top-performing reward models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb732b7-d282-465f-b157-e4328d374676 · inbound
SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e077090a-7ba5-4c32-8c37-46c2b97c1ef3 · inbound
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs HelpSteer2: Open-source dataset for training top-performing reward models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a415e80b-7a3e-428c-b68d-885bbdbf7a28 · inbound
OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique HelpSteer2: Open-source dataset for training top-performing reward models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be4873b-40e3-45a8-896b-91a648f1633d · inbound
Tiny Reward Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab21d6ac-d056-44f4-9ca7-01b06844b234 · inbound
Second-Order Bounds for [0,1]-Valued Regression via Betting Loss HelpSteer2: Open-source dataset for training top-performing reward models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abf26adc-4c7d-4de0-b06e-55ef50970bdd · inbound
Improving Large Vision and Language Models by Learning from a Panel of Peers HelpSteer2: Open-source dataset for training top-performing reward models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc0d8b94-5e21-4b50-841c-5a4a769fafd5 · inbound
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases HelpSteer2: Open-source dataset for training top-performing reward models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation db8f3726-5db8-4eb9-b9e8-983974304b8c · inbound
PolyAlign: Conditional Human-Distribution Alignment HelpSteer2: Open-source dataset for training top-performing reward models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation efbb4249-af8a-485e-903b-8c4416d37995 · inbound
PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration HelpSteer2: Open-source dataset for training top-performing reward models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 26726097-ee01-4fec-bbb3-dc053c65c315 · inbound
Step-Level Preference Learning for Generative Agents in Social Simulations HelpSteer2: Open-source dataset for training top-performing reward models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff47f6f4-94a1-4420-ae85-74801ece0ce2 · inbound
Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization HelpSteer2: Open-source dataset for training top-performing reward models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.