Pith. sign in

Paper Citation Record · LEDGER

HelpSteer2-Preference: Complementing Ratings with Preferences

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2410.01257.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.01257 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:18:39.220697Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:37:30.542430Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0f408b60-18da-45ec-990d-743993458d17 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 246

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.666129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:09c509dd9bde7a3e7a585ff9de96fdc8ee853efb109bc7e272895d7759b87c68

Observation 42666668-3a53-487f-97a0-02efe5af086b · inbound

From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning cites this paper.

From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:39.220697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:39.220697Z digest=sha256:c3cfe46161d37df7655ef4e1606ec7306dc924322e6322fe9ec2bc7346dc6f95

Observation 16781b5f-b038-4ecf-9483-0c3d591e4a18 · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:22:16.659478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:7f7e5d5cf754ac23a8ff7778f8c8b42520964375eb7a47242e89aaf5b66d201c

Observation 38ed5144-7e39-40df-b130-03924ff07c98 · inbound

Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding cites this paper.

Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:59.874401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:59.874401Z digest=sha256:44cce8e3ad0b93774329f42ab739e98ff8620484a57f4f4edd7006e6ea22b04a

Observation a39d22f1-4ff5-4198-b742-58d1984e87cb · inbound

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models cites this paper.

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:58.189073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:58.189073Z digest=sha256:f906d6b76fd710bb250d17795af2f2b1f1cace44f28dc367048471ab0af882eb

Observation 755f426c-4980-4b25-9170-c6be89aa72b4 · inbound

Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems cites this paper.

Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:22:15.929822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T09:20:12.827871Z digest=sha256:49ebec4f7bedb200c37e639981c89f31a96dda38a059b045222f94bc6e10d3fb

Observation e9bba9b6-ce48-4fcb-83d2-d29a073e870d · inbound

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique cites this paper.

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:19.287789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:19.287789Z digest=sha256:67fc935d7b29e538e0a75a481b41e501cd38dc57e65091031583772a9fd656e1

Observation 0941c192-6f1f-491f-82f9-a7d3820bddc5 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.787636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.787636Z digest=sha256:cb921d83a45538f853de24e187ed6b2f48657769c4d35812bcdc689940b40292

Observation 0d92521d-7af7-4991-839f-b0e355584ce3 · inbound

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation cites this paper.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.619519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.619519Z digest=sha256:39906f1c047ad4acaeeb6edc8a23d3305d45e68d75a6ab42db802fe21b5473cc

Observation e57d647d-c45f-4c8d-8daf-a9bf780f2275 · inbound

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs cites this paper.

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:14.482448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:14.482448Z digest=sha256:b72b754e9d7d287f3828bb7991ba95511c0814681a68c76961501ba01e7e47ce

Observation 2a379593-88ca-4d08-bb53-4ebb4da394de · inbound

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework cites this paper.

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:35.388592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:35.388592Z digest=sha256:f66be2d00d748fcb4ff96143cb449c75c27587ffc962915cf0d1e7aee018cb8d

Observation b4039719-d289-4f68-abe4-e1b8e027db86 · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:23.416543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:23.416543Z digest=sha256:c962d0ef8f15374d96bfa3927b24e9ee9c844bbe97ff228685a92770b7d9b935

Observation 49e3ca38-f659-4cf0-8018-470ee97cc75a · inbound

Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents cites this paper.

Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:30.732674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:57:20.186356Z digest=sha256:f71db14b7cfe18cd324e4a2eacad2373d20a7701430e6c05d9f32710dcb2ea84

Observation 0d6e39b8-e452-49dc-b58d-4e4be105394e · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.543769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:0194bcb10474f9e97b0afbba44337aae24eb3343f77eb724e3ed7791ae35a272

Observation ce7f2b4a-8299-4273-a58c-2adf792cb670 · inbound

Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis cites this paper.

Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T10:30:46.777170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:30:46.777170Z digest=sha256:7ee8022f67aaab398237d8fd99bf31e4fb7f2e8a50c425b39212627369d03774