Pith. sign in

Paper Citation Record · LEDGER

ARGS: Alignment as Reward-Guided Search

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2402.01694.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.01694 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:31:57.817590Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:56:44.378731Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 898702c6-e737-4651-8b91-939f6f63ebeb · inbound

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution cites this paper.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution ARGS: Alignment as Reward-Guided Search

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.653892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.653892Z digest=sha256:39805d3a382a3b975d56a9a6ebeeb9d970c6e8ac73669f53beb9e8d3e8a9cd38

Observation 965e2f08-0769-4155-9b02-14dad57a1b4e · inbound

Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models cites this paper.

Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models ARGS: Alignment as Reward-Guided Search

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T21:28:20.798133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:28:20.798133Z digest=sha256:219103fd0244a53d21656ef2993a31b073bee4c2bbafc5bb0503d7a979a5c49d

Observation 69102108-2525-434e-b73e-5e9fd28869b8 · inbound

Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review cites this paper.

Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review ARGS: Alignment as Reward-Guided Search

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:50:33.654941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:50:33.654941Z digest=sha256:5a3b687d6f59b4400f48562e62002c52a95d0f241a21ed1edff391a23df626a4

Observation 124f1ebe-35b4-4a99-8e30-73cfad6e3a45 · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time ARGS: Alignment as Reward-Guided Search

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.679002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.679002Z digest=sha256:84e85fadb08a172d908ad384d2966d9b757c5a104e282b49ce5634d5c91cde04

Observation 572fb1eb-4d61-48a5-b9af-2e6be92ca496 · inbound

Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment cites this paper.

Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment ARGS: Alignment as Reward-Guided Search

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:31:57.817590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:31:57.817590Z digest=sha256:73dce824b61e1b222be8df308dc22c36ec2c8c835b5c843baf0b74ec7971bceb

Observation 64c21c09-81fb-4ae8-ba5e-d5da690a9317 · inbound

CEC-Zero: Chinese Error Correction Solution Based on LLM cites this paper.

CEC-Zero: Chinese Error Correction Solution Based on LLM ARGS: Alignment as Reward-Guided Search

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:44:05.484491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:44:05.484491Z digest=sha256:f5aec7b5e82c3b197883703adf07e8c7eed0e5e4a95d05e90db19223a083206e

Observation e3f9c18b-c540-4974-8201-5feb17b9aa2d · inbound

Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment cites this paper.

Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment ARGS: Alignment as Reward-Guided Search

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:08.687134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:08.687134Z digest=sha256:1087e93436ce688ddedaeb07ff85b7b7dea8cf13637705ba30f9c35ce00c98f0

Observation 8f7ffe76-90bb-44b7-ad54-10d5e720510a · inbound

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO cites this paper.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ARGS: Alignment as Reward-Guided Search

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.192292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.192292Z digest=sha256:1a9f4299a7674112ce7651128174931438054339dffeb034ad1f4a187951fb26

Observation c3258aed-6092-456c-8e05-fe4a1ca085de · inbound

Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time cites this paper.

Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time ARGS: Alignment as Reward-Guided Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:20.211826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:20.211826Z digest=sha256:311d9ce952aa37f916e54b4f97d6507a9e0c2dff8d7fb5c41b99a9df2aa27ea8

Observation 9eabd63f-8d6d-4fa7-a945-faa52e58f4af · inbound

BiasFilter: An Inference-Time Debiasing Framework for Large Language Models cites this paper.

BiasFilter: An Inference-Time Debiasing Framework for Large Language Models ARGS: Alignment as Reward-Guided Search

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:28.390947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:21:28.390947Z digest=sha256:5c9e8478e441946840b2662607f9ea02c2a3cb6898e302c81def71c826d45039

Observation 8ae2857e-1b3d-4a3a-bc91-16c3ace39994 · inbound

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment cites this paper.

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment ARGS: Alignment as Reward-Guided Search

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:53:03.597106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T11:52:36.688263Z digest=sha256:9b07bec0ae287a57d779d8108ca41789089c14d7148e03c45288d5ab8865f56b

Observation 55bc5fd5-0f8d-40b5-8ddd-d4c87946db01 · inbound

LLM-ML Teaming: Integrated Symbolic Decoding and Gradient Search for Valid and Stable Generative Feature Transformation cites this paper.

LLM-ML Teaming: Integrated Symbolic Decoding and Gradient Search for Valid and Stable Generative Feature Transformation ARGS: Alignment as Reward-Guided Search

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:57.257974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:57.257974Z digest=sha256:b76ea6c461cb2c7c1b232daffc79288539d2a789c6977e3501f83aea848632bc

Observation ba331d66-217c-4371-a38e-96c144afed5f · inbound

Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples cites this paper.

Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples ARGS: Alignment as Reward-Guided Search

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:11.808816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:32:11.808816Z digest=sha256:fd93667bbe282b463e6c1815cfd16ec895b7b64002378d3a4219bb04fee7b09c

Observation 0db04d10-d604-4cc7-a8b4-31080f448bea · inbound

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach cites this paper.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach ARGS: Alignment as Reward-Guided Search

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.361484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.361484Z digest=sha256:19ce4dbd472b04a3a80d3f533b6455ac044cb5e1121af71f53b1b7aac303b176

Observation a3223f1b-673c-43ab-ae62-693610b0990b · inbound

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis cites this paper.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis ARGS: Alignment as Reward-Guided Search

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:55.771105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:55.771105Z digest=sha256:b53859d8080a83ce2e37d1917e099e3200afe15e9b36ae12ea06c2672f0333c3

Observation 041852b0-1b03-4c76-b9c6-1bf402807f70 · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling ARGS: Alignment as Reward-Guided Search

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.872167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:a658653d46cdc0a662138ebe3aa4339e8ffcc93b9df5898c0b573ce42cc62bc5

Observation e6a67e93-2552-4e32-97b0-f01c49869cd0 · inbound

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary cites this paper.

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary ARGS: Alignment as Reward-Guided Search

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:32.704726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:32.704726Z digest=sha256:a2c905877e5f6fe32382a420209f7deb36183b9f213680b1347aa7d303a474d3

Observation 50e73708-3b7d-4fc0-90bc-a99fa3c9ceb4 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities ARGS: Alignment as Reward-Guided Search

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.063586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.063586Z digest=sha256:58cf1a7232a72ebe9f4d1f3e4c54f7f3ef937ac82177dee69415d88b55844de6

Observation 0b960119-1126-497f-8a76-8dfbd9da11f7 · inbound

A Survey on Training-free Alignment of Large Language Models cites this paper.

A Survey on Training-free Alignment of Large Language Models ARGS: Alignment as Reward-Guided Search

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:42.488047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:18:42.488047Z digest=sha256:59a93b6a4fe776d5a5e36e2908d66eb174e63cd995ce124ad1bf83321f1c7adf

Observation 58f9ff3c-6405-491a-bbab-d6af7d01fcbe · inbound

Virtual Agent Economies cites this paper.

Virtual Agent Economies ARGS: Alignment as Reward-Guided Search

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T18:09:37.443627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:09:37.443627Z digest=sha256:a7e562ff0a0725357c6c34241d21391cc50571a13b80c0ce64583e175e0542b6

Observation 94dfb535-c47f-463f-9aa0-cb9d3c44c5c4 · inbound

T-POP: Test-Time Personalization with Online Preference Feedback cites this paper.

T-POP: Test-Time Personalization with Online Preference Feedback ARGS: Alignment as Reward-Guided Search

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T13:52:11.128628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:52:11.128628Z digest=sha256:32ed1976a5ae619249b5fd0881d56c0db70c500e69126e676c8271bd3fa3201b

Observation b8888d26-c4bb-474f-8d6e-a3042504f5fd · inbound

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards cites this paper.

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards ARGS: Alignment as Reward-Guided Search

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:44.167888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:15:44.167888Z digest=sha256:5ebcc1716d795e859be17a9a0fb67d75d851d261e87f6044b5d650250873fb0e

Observation 80061e44-3d5b-4f76-975c-ae0241157149 · inbound

Representation-Based Exploration for Language Models: From Test-Time to Post-Training cites this paper.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training ARGS: Alignment as Reward-Guided Search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.535591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.535591Z digest=sha256:9fe27e2a5bde3bafa84d328bf0236c0e9cc4c011ac7a9762ef8795961f31ab30

Observation 0ce5fda3-b4e0-4e8f-b7b2-94d90910bdd8 · inbound

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective cites this paper.

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective ARGS: Alignment as Reward-Guided Search

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:23.362311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:23.362311Z digest=sha256:d6e823fddbcd8b0a8749e4cc672ada8fa6b147eabc81a964fa84fea9f48389e4

Observation 63eda618-2b8d-4180-ae03-67eaa70e4fea · inbound

Meet Dynamic Individual Preferences: Resolving Conflicting Human Value with Paired Fine-Tuning cites this paper.

Meet Dynamic Individual Preferences: Resolving Conflicting Human Value with Paired Fine-Tuning ARGS: Alignment as Reward-Guided Search

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:05.309215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T15:06:43.463303Z digest=sha256:f45725d94d9b7f3d0c095352a5a85adc64ae7b0395cf0f85853146140162b64f

Observation 8b261bb4-3bc0-405d-8574-37ed6cabbf66 · inbound

Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control cites this paper.

Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control ARGS: Alignment as Reward-Guided Search

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:01:03.967708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T02:31:07.932802Z digest=sha256:7e5038889b6eefba5eb311a4e54a2c3a1cb93d27963ecbe6a6d8e235ccc30373

Observation 68bba25e-0882-401e-95c8-7161af716b3f · inbound

Pref-CTRL: Preference Driven LLM Alignment using Representation Editing cites this paper.

Pref-CTRL: Preference Driven LLM Alignment using Representation Editing ARGS: Alignment as Reward-Guided Search

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:19.384158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T06:28:14.378602Z digest=sha256:6746caf785253e61b82507a9ca31bd9023d50ee7aab643c52bb980edb35a9db5

Observation 0fdb043b-ea26-481c-b498-e96acf9505b8 · inbound

Training-Free Cultural Alignment of Large Language Models via Persona Disagreement cites this paper.

Training-Free Cultural Alignment of Large Language Models via Persona Disagreement ARGS: Alignment as Reward-Guided Search

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:23:48.490580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:19:33.582024Z digest=sha256:f6491f5aa8adc4ddf3731ba69c60ae73d72eb9b6b2864c2bb9712a58b171090d

Observation 6e3afd58-137b-463b-ad8d-35702c7b44c6 · inbound

Spectral Souping: A Unified Framework for Online Preference Alignment cites this paper.

Spectral Souping: A Unified Framework for Online Preference Alignment ARGS: Alignment as Reward-Guided Search

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.948373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T07:54:56.356555Z digest=sha256:38c2536352d8141558ae1f458ac858be61b0cfb7a8ab048cb595792a546914aa

Observation 7b6cdbd5-0343-4c1f-ae14-6ef2bd6f0982 · inbound

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning cites this paper.

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning ARGS: Alignment as Reward-Guided Search

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:11:17.008018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T08:10:55.720464Z digest=sha256:0c1e3267aa115c86971b0a557a481875a2069e8020c5529ccda049410c136ee3

Observation 3bbb5d27-7c5d-49d2-b8c2-2b40751759b5 · inbound

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning cites this paper.

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning ARGS: Alignment as Reward-Guided Search

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:50:24.337367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T05:47:33.925413Z digest=sha256:55a23ef651ba8016430809156bcc6fa9682d64e9f903eb8575114e2a2a3f0350

Observation 2daf7f82-38aa-4714-815e-050469ebd7b7 · inbound

Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control cites this paper.

Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control ARGS: Alignment as Reward-Guided Search

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:56:44.380450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T07:18:51.066384Z digest=sha256:37d523db53005a88ab7949739618163854a401b280d711b635a829afe4777535

Observation 2aec05c3-0fba-4c6c-9285-d088541337c5 · inbound

Safe Inference-Time Alignment via Lagrangian Reward Augmentation cites this paper.

Safe Inference-Time Alignment via Lagrangian Reward Augmentation ARGS: Alignment as Reward-Guided Search

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T07:05:47.150308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T07:05:47.150308Z digest=sha256:8e43db541380a6aff45846ca3c89b8c0b4d6634ba0df603c104e9df1ff357345

Observation 71560dde-6950-480f-8a10-0553bd785f97 · inbound

IFHierBench: Hierarchical Instruction Following for Large Language Models cites this paper.

IFHierBench: Hierarchical Instruction Following for Large Language Models ARGS: Alignment as Reward-Guided Search

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T23:14:22.758776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:14:22.758776Z digest=sha256:84b72408300d50dd8b6a2d96ebd19e810e08a354d6c7bb336f75c35b8a5b142e