Pith. sign in

Paper Citation Record · LEDGER

HelpSteer2: Open-source dataset for training top-performing reward models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2406.08673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08673 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:53:03.339487Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d655d4cf-9b93-415e-84c2-e0a144d85aaf · inbound

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs cites this paper.

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs HelpSteer2: Open-source dataset for training top-performing reward models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:18:01.675521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T16:18:01.560780Z digest=sha256:480f4c64d6ab7eabdc38dfa3ecc49b6541bccbc022f70652f0caa73988b1dd9c

Observation 6a029ccd-4f98-48ef-9ea0-a947deed1b80 · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach HelpSteer2: Open-source dataset for training top-performing reward models

Reference 163

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:39:41.285505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:5669ab0b4deab4aeb9c0e3418993d7630b744313da0c78bad93317174b232a96

Observation 1109fac1-d0c1-485c-ba78-d9ec08f5d82d · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback HelpSteer2: Open-source dataset for training top-performing reward models

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.189460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:7dc4fd9356ee60db62c0b3c7bcaf68136153a465e6e8b2c4664bbb6ab9620c88

Observation 93f23228-ef50-42f8-802c-ad68b2f37369 · inbound

Discriminative Policy Optimization for Token-Level Reward Models cites this paper.

Discriminative Policy Optimization for Token-Level Reward Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.339487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.339487Z digest=sha256:079c3867fb048c1c47c227fcc4d8927f8e285c5f38be97c72f7b9a6387055d48

Observation e57198dc-bc3b-4e25-9d4e-e0450c9c8766 · inbound

TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence cites this paper.

TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence HelpSteer2: Open-source dataset for training top-performing reward models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:50.908285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:50.908285Z digest=sha256:ddadd0cdfb9c3ae80c9b76c7b480c251a2a6fd9e0a6e994b82b61f0693c27538

Observation ad03ac79-0afa-486b-b655-aa3c4bee344f · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation HelpSteer2: Open-source dataset for training top-performing reward models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:22:16.733814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:d3cddc17bd99ba7217f68ff58a29bb9567eacd26cfcf1394745838947dc8c43b

Observation b3b23bd3-640a-4e84-89ff-091dd43c11d9 · inbound

PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization cites this paper.

PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization HelpSteer2: Open-source dataset for training top-performing reward models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:47.880878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:42:47.880878Z digest=sha256:8473078c8d5ec6b243b654d5236615ffaf2283292dbcba0563a9585dda454d93

Observation 0eb732b7-d282-465f-b157-e4328d374676 · inbound

SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models cites this paper.

SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:33.902602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:33.902602Z digest=sha256:e17fd4d19bf22eb09a75d94c390defb28ba83f19c66e6a7e2b37c5671f722476

Observation e077090a-7ba5-4c32-8c37-46c2b97c1ef3 · inbound

When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs cites this paper.

When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs HelpSteer2: Open-source dataset for training top-performing reward models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:51:33.947555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:51:33.947555Z digest=sha256:a05ee2c48ef12345bacac6da317050c45cfcb475c11806426fd2792f26ad2304

Observation a415e80b-7a3e-428c-b68d-885bbdbf7a28 · inbound

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique cites this paper.

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique HelpSteer2: Open-source dataset for training top-performing reward models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:19.396424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:19.396424Z digest=sha256:314df3f805a0c4e353545c4e9331a3b65783d75832f8a90b301d610685ea7335

Observation 5be4873b-40e3-45a8-896b-91a648f1633d · inbound

Tiny Reward Models cites this paper.

Tiny Reward Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.201291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.201291Z digest=sha256:6ebaccb26af16a544152cede119f83a0aab1bcab715c9605b4b27d48f6c7a19e

Observation ab21d6ac-d056-44f4-9ca7-01b06844b234 · inbound

Second-Order Bounds for [0,1]-Valued Regression via Betting Loss cites this paper.

Second-Order Bounds for [0,1]-Valued Regression via Betting Loss HelpSteer2: Open-source dataset for training top-performing reward models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:58:14.283812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:58:14.283812Z digest=sha256:30891403d04c7e6e9ac3e329d86adc65db54d1b12df97b8df7abb2734af7f3f0

Observation abf26adc-4c7d-4de0-b06e-55ef50970bdd · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers HelpSteer2: Open-source dataset for training top-performing reward models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.632845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.632845Z digest=sha256:8eb533a3523805074cf67ca61c50a5cbb837f7eb048356bfa4aa5db07d63fad8

Observation bc0d8b94-5e21-4b50-841c-5a4a769fafd5 · inbound

Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases cites this paper.

Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases HelpSteer2: Open-source dataset for training top-performing reward models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:21:07.946949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T17:18:10.045283Z digest=sha256:86f3ae8ecd96aad7c9c6e84b94981af32fe91319caef8d0a6382039d8516a558

Observation db8f3726-5db8-4eb9-b9e8-983974304b8c · inbound

PolyAlign: Conditional Human-Distribution Alignment cites this paper.

PolyAlign: Conditional Human-Distribution Alignment HelpSteer2: Open-source dataset for training top-performing reward models

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T15:08:33.613033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T06:38:55.992652Z digest=sha256:b24086cbf74a962b4597378657d47ef173269d0f38fc823c9f2149c1cc821faf

Observation efbb4249-af8a-485e-903b-8c4416d37995 · inbound

PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration cites this paper.

PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration HelpSteer2: Open-source dataset for training top-performing reward models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:58.586206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T01:26:00.228782Z digest=sha256:def88b224c28ecd50ffb9384e9dd400d840e8ba9d820ec82977195d3534a6754

Observation 26726097-ee01-4fec-bbb3-dc053c65c315 · inbound

Step-Level Preference Learning for Generative Agents in Social Simulations cites this paper.

Step-Level Preference Learning for Generative Agents in Social Simulations HelpSteer2: Open-source dataset for training top-performing reward models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T02:00:27.655165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:00:27.655165Z digest=sha256:9a53f8b98be31c1c3ee63aa3bbe9404a26924a0bdd79b82253588c085783d58d

Observation ff47f6f4-94a1-4420-ae85-74801ece0ce2 · inbound

Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization cites this paper.

Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization HelpSteer2: Open-source dataset for training top-performing reward models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T00:30:21.547440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:30:21.547440Z digest=sha256:37446fd5c5a0b92e459d6d54c6da4220af91b0c4bb57ea2744e89770cffe4885