Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T06:06:24.769271Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2602.02572.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T06:06:24.769271Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T16:26:34.918099Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T01:37:30.345559Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a9b43f2f-0b80-497a-8fff-94c0d560866c · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Persistent anti-muslim bias in large language models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ff10d8f-afe3-464b-885a-22ea0178fcc6 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective In contrast, when B is large, expected user utility initially increases with α but eventually degrades
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9622f0a8-fff5-4af3-bc82-b2a5c317209b · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90cc93f8-760a-4fcb-ade0-71c938838865 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Stackelberg game preference optimization for data-efficient alignment of language models.arXiv preprint arXiv:2502.18099,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a78b3777-1720-4aac-8f48-19dd7acb5fa7 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ef0cd4-3f38-4d63-9c94-f90d62084d57 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8a2cea7-0a82-4ab3-b5de-826e458c5c16 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Principal-Agent Reinforcement Learning: Orchestrating AI Agents with Contracts
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a845c286-8195-4ab7-bcbb-6aaeb6d0cce3 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82323386-dc3b-43fb-a29d-cb314b8dfd28 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Regularized best-of-n sampling to mitigate reward hacking for language model alignment
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ce5fda3-b4e0-4e8f-b7b2-94d90910bdd8 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective ARGS: Alignment as Reward-Guided Search
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377f2fb5-94e8-49db-a955-713eaa5d4c48 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Aligning Large Language Models with Representation Editing: A Control Perspective
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6568c070-7823-4fa9-961a-a6cbb678375e · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22810edb-4870-4918-8036-4b85ba90efa3 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Red: Unleashing token-level rewards from holistic feedback via reward redistribution
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e97ca9-2072-4292-bbb5-ae7ba90057e8 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc54a390-813a-40cb-8573-e4fe20d25397 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Controlled Decoding from Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e8ef3f-021f-44f6-93f5-41c99cdb020c · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Of Models and Tin Men: A Behavioural Economics Study of Principal-Agent Problems in AI Alignment using Large-Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f59d078f-612b-4b29-96c8-2b6e0079f15a · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d21238ad-6117-4d70-b6d8-711f799a14eb · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Verbosity Bias in Preference Labeling by Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf15c008-06bc-4e12-9c32-6ed24ac2e395 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b6b150-4c10-4454-83f9-c1108497a144 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Robust multi-objective controlled decoding of large language models.arXiv preprint arXiv:2503.08796,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee96b49f-3cc8-47e0-9d31-f77649e93a17 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1378a620-b05f-49a2-b875-d8cf129227b0 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Transforming and Combining Rewards for Aligning Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1272e9f4-5e14-47ea-9842-b35fa6d61fee · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fbfd065-7f1f-43a4-81ba-b22b67e1d733 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Qwen3 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3367dc2-ea39-4b96-bdba-d8f844872624 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f350215-e90e-4d8c-8481-074525852ad5 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Fair-PP: A Synthetic Dataset for Aligning LLM with Personalized Preferences of Social Equity
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 376c0293-8980-46a2-b094-0c5def82af85 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d396c6c-ed68-4576-a0a5-8507d036bfcf · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective [2024], Ivanov et al
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a068fe7e-abfb-4518-b45f-1a76a3203c66 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09309d26-586a-432a-87ee-ab2c36c32fc2 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective The model generates responses for 1,000 test prompts, of which 300 are randomly selected for GPT-4 evaluation against base-policy answers generated without alignment
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beff5586-f56f-4b10-a48f-b6c89cb1a75f · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective This behavior indicates that the learned value function ˆQ(st,y ) fails to provide a meaningful estimate of the reward for expected model completions
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f475aac-43d3-4cff-be7b-a8c450338468 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Recall we also have ˆQr ˆm∗,α(st, yt)≤r max ≤B
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93d8e8a8-237e-479a-8b2b-f3ce778c67c8 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective ""USER_PROMPT =
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2040902-8b98-4aa3-af64-9a23b8da6c4c · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Incomplete contracting and ai alignment
Reference 1992
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a86dd72-d410-4993-82f2-8181493fb18b · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f658e658-2ded-4b04-b9a1-2270e55554e3 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Reward shaping to mitigate reward hacking in rlhf.arXiv preprint arXiv:2502.18770,
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b613c101-70ea-4240-afab-6f13686e4075 · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Simple versus optimal contracts
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98511649-daee-476a-8d0a-5445a8933dcd · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Strategyproof reinforcement learning from human feedback.arXiv preprint arXiv:2503.09561,
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 259c1238-4e92-4686-b67a-8bcbf4f5ffce · outbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bbed268-4dad-4fd9-ad83-4b80970051af · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d57426d3-0ff6-4fe3-8245-e818bf522e44 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective
Reference 211
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.