Pith. sign in

Paper Citation Record · LEDGER

HelpSteer2: Open-source dataset for training top-performing reward models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2406.08673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08673 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:53:03.339487Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d655d4cf-9b93-415e-84c2-e0a144d85aaf · inbound

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs cites this paper.

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs HelpSteer2: Open-source dataset for training top-performing reward models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:18:01.675521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T16:18:01.560780Z digest=sha256:e864fe19bbc093483c7d908632c1f89ccdee1bd3ee5f0b37dfc93d4b2e95cd9a

Observation 6a029ccd-4f98-48ef-9ea0-a947deed1b80 · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach HelpSteer2: Open-source dataset for training top-performing reward models

Reference 163

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:39:41.285505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:3294f1b6ca3222ac22d1db6db8cdd31e64a7904c8b2ca4427d67743e7e01080d

Observation 1109fac1-d0c1-485c-ba78-d9ec08f5d82d · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback HelpSteer2: Open-source dataset for training top-performing reward models

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.189460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:1ceeec13d2a9b819e03c3713b8323aecc503c9964646c5f0a80623ce74ddb0b4

Observation 93f23228-ef50-42f8-802c-ad68b2f37369 · inbound

Discriminative Policy Optimization for Token-Level Reward Models cites this paper.

Discriminative Policy Optimization for Token-Level Reward Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.339487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.339487Z digest=sha256:b78ca3b93492868a0de007618c284ec75e815f517530fd8db753a000242e1dcf

Observation e57198dc-bc3b-4e25-9d4e-e0450c9c8766 · inbound

TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence cites this paper.

TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence HelpSteer2: Open-source dataset for training top-performing reward models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:50.908285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:50.908285Z digest=sha256:8bb589c51e4e3ac7366e94a6ffcde0684ab79b4b89c67a23a13616633ca7fef4

Observation ad03ac79-0afa-486b-b655-aa3c4bee344f · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation HelpSteer2: Open-source dataset for training top-performing reward models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:22:16.733814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:a6b31589d5d45cd0831ac28198eb7e23ff0f7dff81564cd66b0572a5632b89a4

Observation b3b23bd3-640a-4e84-89ff-091dd43c11d9 · inbound

PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization cites this paper.

PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization HelpSteer2: Open-source dataset for training top-performing reward models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:47.880878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:42:47.880878Z digest=sha256:bfec5a76d9f95b55f68fe443d3860a4ef63347c9fd91f8e1d062944f781aa9c5

Observation 0eb732b7-d282-465f-b157-e4328d374676 · inbound

SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models cites this paper.

SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:33.902602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:33.902602Z digest=sha256:5faf7071ad331739220b0e8158157eb6b195b85c5cbf74a1cc31e6354b9daf16

Observation e077090a-7ba5-4c32-8c37-46c2b97c1ef3 · inbound

When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs cites this paper.

When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs HelpSteer2: Open-source dataset for training top-performing reward models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:51:33.947555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:51:33.947555Z digest=sha256:894ba439114c55a65b51916da2ea1902abb1a69446f76ee16434fe05e8a9e88d

Observation a415e80b-7a3e-428c-b68d-885bbdbf7a28 · inbound

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique cites this paper.

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique HelpSteer2: Open-source dataset for training top-performing reward models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:19.396424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:19.396424Z digest=sha256:c8911aae9ef41b8c43b87f60d9830872dd2754a1265cc701c638da620312b80d

Observation 5be4873b-40e3-45a8-896b-91a648f1633d · inbound

Tiny Reward Models cites this paper.

Tiny Reward Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.201291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.201291Z digest=sha256:243db25fed372f368b36d2816d12a1014ae5cfcc019f159c331baffaeca5accc

Observation ab21d6ac-d056-44f4-9ca7-01b06844b234 · inbound

Second-Order Bounds for [0,1]-Valued Regression via Betting Loss cites this paper.

Second-Order Bounds for [0,1]-Valued Regression via Betting Loss HelpSteer2: Open-source dataset for training top-performing reward models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:58:14.283812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:58:14.283812Z digest=sha256:a5d0ed7b8326a666d72828e210089bfb8642dc4a24303e6bdcb3f28c577b466a

Observation abf26adc-4c7d-4de0-b06e-55ef50970bdd · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers HelpSteer2: Open-source dataset for training top-performing reward models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.632845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.632845Z digest=sha256:f1b7e96e1e48150cfce7beb021d59a2d4ddfe59368bee1f2175993a072213632

Observation bc0d8b94-5e21-4b50-841c-5a4a769fafd5 · inbound

Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases cites this paper.

Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases HelpSteer2: Open-source dataset for training top-performing reward models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:21:07.946949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T17:18:10.045283Z digest=sha256:26d9c4c3b91923cc53bbf62ce53447857ced613c9a09bc5c2d3f3c29b3450073

Observation db8f3726-5db8-4eb9-b9e8-983974304b8c · inbound

PolyAlign: Conditional Human-Distribution Alignment cites this paper.

PolyAlign: Conditional Human-Distribution Alignment HelpSteer2: Open-source dataset for training top-performing reward models

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T15:08:33.613033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T06:38:55.992652Z digest=sha256:0d13f423783706e357cf7dd3533d2c86cba5df1cba40e05e86467457df1b73a4

Observation efbb4249-af8a-485e-903b-8c4416d37995 · inbound

PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration cites this paper.

PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration HelpSteer2: Open-source dataset for training top-performing reward models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:58.586206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T01:26:00.228782Z digest=sha256:2e5c42e357527810ec2e40406b43b30edad2d375fd2b70a37d710ee91138bf84

Observation 26726097-ee01-4fec-bbb3-dc053c65c315 · inbound

Step-Level Preference Learning for Generative Agents in Social Simulations cites this paper.

Step-Level Preference Learning for Generative Agents in Social Simulations HelpSteer2: Open-source dataset for training top-performing reward models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T02:00:27.655165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:00:27.655165Z digest=sha256:7c41e09c16b175a0403f66c9ec76d3a5f2d57374f0481260fe27876fa610617a

Observation ff47f6f4-94a1-4420-ae85-74801ece0ce2 · inbound

Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization cites this paper.

Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization HelpSteer2: Open-source dataset for training top-performing reward models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T00:30:21.547440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:30:21.547440Z digest=sha256:077aeeafc7179a9c1b5bbe7a9caab28db2c3f5228b512f8547417e33d541c522