Pith. sign in

Paper Citation Record · LEDGER

GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2410.08193.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.08193 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:22:59.901678Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T01:20:52.147219Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 513ae582-aa7c-4aba-a346-c1826eabce63 · inbound

Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs cites this paper.

Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:52.150446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T01:16:51.288077Z digest=sha256:96f0e00b5c09db2a2500d44f1f7c98093be58fc4a525c7091bf5ae0ba901d7e4

Observation e8b48618-9ec6-41d2-b9c0-7b2731102926 · inbound

UCD: Unlearning in LLMs via Contrastive Decoding cites this paper.

UCD: Unlearning in LLMs via Contrastive Decoding GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:59.901678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:59.901678Z digest=sha256:2624bf4e9e5a8e81071ee1b585352c5ff9678f74e9f8182dcc6dd3837ea31c75

Observation affda21e-2a61-4a2e-a931-ae8f8241e175 · inbound

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment cites this paper.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.770058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.770058Z digest=sha256:2ea44e314e6a61907da67fc47247b21e4d081e1a579b5ba78699cd6ac5c548d4

Observation 2bf46a25-5c75-4d75-ae39-4efc3705e560 · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.942423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:5f2b8be90b934387c8e09d39cec67e3dd57c1c8ce17a75053456612a2d8ecffd

Observation 7e914c68-0a1e-45aa-85aa-b74a332cbd70 · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:12:52.077537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:12:52.077537Z digest=sha256:3906ddb0c288d891f4020042e36ef45ec8a69ca82e6234c95623b5578ee35bbc

Observation e491323c-cac4-4d1f-9eb0-fb71d021fbfb · inbound

A Survey on Training-free Alignment of Large Language Models cites this paper.

A Survey on Training-free Alignment of Large Language Models GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:49.339379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:18:49.339379Z digest=sha256:feae507e38bbe9ee0ab89f2b28b75bf1ab4404415125df873cbe47db9170d3d7

Observation 9d2e5f6e-7f7c-4fb8-a578-6658cb4e8d47 · inbound

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards cites this paper.

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:48.047495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:15:48.047495Z digest=sha256:f882e873f2a3e4e7a894b09f8fa55e3922ac66fcf9e510f6d43c949021bb9063

Observation 1272e9f4-5e14-47ea-9842-b35fa6d61fee · inbound

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective cites this paper.

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:24.052889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:24.052889Z digest=sha256:0b2402bbc36f2519fedabdecf338b07e6041c78c3fe23b0de47953c3e44797a8

Observation 31c699c7-3224-44c8-b4ca-1fc5a30c9f91 · inbound

Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models cites this paper.

Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:55:03.840041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T10:54:57.732141Z digest=sha256:f1b98d93ae59168bba38d90f955623aff7ac3863379ea30f89c72380e906e1ce

Observation 8a88186b-a2c7-42e4-8f82-912266d671c2 · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:53.855338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:aca7c7319223c02955122b0daeeda24797e74cede3711ca40ec34b5929fc98f9

Observation c5b03e13-64cd-4850-af1b-d7723e7ba66f · inbound

Common-agency Games for Multi-Objective Test-Time Alignment cites this paper.

Common-agency Games for Multi-Objective Test-Time Alignment GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 237

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:15:06.493487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:14:53.685486Z digest=sha256:c22a19d930bc0e7c6e7d4181b228b526a07179848d417ee019196339f6cdff07

Observation 536ccdae-99c9-4ac8-a60e-d67c75ff6468 · inbound

Inference-Time Policy Alignment for Fair Reinforcement Learning cites this paper.

Inference-Time Policy Alignment for Fair Reinforcement Learning GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-04T01:10:45.032296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:10:45.032296Z digest=sha256:30bcb81612693ff25f1a82f1db468faa51547e0b13894ca4f8c5db555064ea88

Observation 4fefa55d-22c4-41d3-947e-649ac10207e5 · inbound

Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach cites this paper.

Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T23:40:44.240475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:40:44.240475Z digest=sha256:1a543f92ad441532741d32b6d025c7186e6c34f95ac39ec80f96e69b87caef10