Pith. sign in

Paper Citation Record · LEDGER

Scalable agent alignment via reward modeling: a research direction

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 87 inbound Pith citation observations for arXiv:1811.07871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1811.07871 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 87 of 87 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:28:47.669686Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

125
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b94fdaa7-3843-4324-8855-114572c29861 · inbound

Risks from Learned Optimization in Advanced Machine Learning Systems cites this paper.

Risks from Learned Optimization in Advanced Machine Learning Systems Scalable agent alignment via reward modeling: a research direction

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:21:53.104437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T12:21:53.058760Z digest=sha256:59673e407dce6ba6fc0c221e605b71fc086b6c582c456f243b750b3fe13366e3

Observation d1e9dfd4-2010-4583-89d1-ed849309bbce · inbound

Modeling AGI Safety Frameworks with Causal Influence Diagrams cites this paper.

Modeling AGI Safety Frameworks with Causal Influence Diagrams Scalable agent alignment via reward modeling: a research direction

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-25T19:51:10.658921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-25T19:47:10.770836Z digest=sha256:adb2b80437ffe995628530a3a7eb7f055dad964e67582084dbc8f1c28752057b

Observation d203dae5-4607-4bfa-86b9-7b75b60e10cb · inbound

Towards Empathic Deep Q-Learning cites this paper.

Towards Empathic Deep Q-Learning Scalable agent alignment via reward modeling: a research direction

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-25T15:35:59.002481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T15:33:05.781171Z digest=sha256:1022dcb7c0650a782bc2368f1a1711ea0e8e7614632beb8ff6cafc8ca08d3b30

Observation f26a7609-fa77-4a48-9253-ea8d790485c0 · inbound

Requisite Variety in Ethical Utility Functions for AI Value Alignment cites this paper.

Requisite Variety in Ethical Utility Functions for AI Value Alignment Scalable agent alignment via reward modeling: a research direction

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-25T12:15:47.069147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T12:15:10.057411Z digest=sha256:e82577a64b8bce46d8cb11cddf64302611246c109eec5ccf583bee2cc877d3a1

Observation 52c0de86-d26c-4b11-9896-78bc6a622517 · inbound

Fine-Tuning Language Models from Human Preferences cites this paper.

Fine-Tuning Language Models from Human Preferences Scalable agent alignment via reward modeling: a research direction

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:59:58.773063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T20:59:58.294355Z digest=sha256:ec5012a1139dc12c5e6da59169339bbf1d419589827d40824c8cae93ff0ea96b

Observation d87632bb-493b-4517-bc0f-7ec226ed6527 · inbound

Learning to summarize from human feedback cites this paper.

Learning to summarize from human feedback Scalable agent alignment via reward modeling: a research direction

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:46:18.597410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T01:46:18.486086Z digest=sha256:9d8e5cdbbcbba8c8210003d111c4e29bcbe6b5aa34a395d8b31c5ed5cdc68919

Observation bb487044-62be-4736-a970-e6ac7f7712f4 · inbound

Scaling Laws for Reward Model Overoptimization cites this paper.

Scaling Laws for Reward Model Overoptimization Scalable agent alignment via reward modeling: a research direction

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:04:53.315297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T09:04:53.129737Z digest=sha256:e382b4550a0f4a510f7dec78b83a91a027be414e9743889c8298d6beff1be671

Observation bc09f87c-8d8a-4b89-85ed-780447feea9e · inbound

Measuring Progress on Scalable Oversight for Large Language Models cites this paper.

Measuring Progress on Scalable Oversight for Large Language Models Scalable agent alignment via reward modeling: a research direction

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:01:41.314000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-17T15:01:41.161487Z digest=sha256:3fcadfb43cea70530d93cf6e13e22447ed07208483f868d4a634ce8ef70514b7

Observation 838a7ba4-b93f-4f1f-9f8e-aa2f772f759b · inbound

Discovering Latent Knowledge in Language Models Without Supervision cites this paper.

Discovering Latent Knowledge in Language Models Without Supervision Scalable agent alignment via reward modeling: a research direction

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:34:08.324344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T20:34:08.207848Z digest=sha256:5984f49fc66aeeadbdf309ae60d20d86766bbc01ce5e15dd88c4231c6b674135

Observation 376e6a77-db9e-484a-a804-6d0eca1b892b · inbound

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment cites this paper.

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Scalable agent alignment via reward modeling: a research direction

Reference 114

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:46:56.776915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T00:46:56.664582Z digest=sha256:abd851a966021c8d31f5e439f7ae35e3f6b4c1af8e475a9adb8577494bf7bdf9

Observation 321cb681-049f-486e-96dc-72cd871218ab · inbound

Universal and Transferable Adversarial Attacks on Aligned Language Models cites this paper.

Universal and Transferable Adversarial Attacks on Aligned Language Models Scalable agent alignment via reward modeling: a research direction

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.548111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:970d5db80bf21dc154882fcb809b512056835eff9adc508f3d0eca8e1e42dc61

Observation 086e11e0-497c-40dd-850b-e057234545be · inbound

A Roadmap to Pluralistic Alignment cites this paper.

A Roadmap to Pluralistic Alignment Scalable agent alignment via reward modeling: a research direction

Reference 271

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T14:37:53.465693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-16T14:37:53.279275Z digest=sha256:572d98876679ef6910ff35a624c325be5ad0960e7d9edfaf785ba25ea3442871

Observation cb019c6b-8076-4151-a9e9-5fb79ce59bcc · inbound

LLM Evaluators Recognize and Favor Their Own Generations cites this paper.

LLM Evaluators Recognize and Favor Their Own Generations Scalable agent alignment via reward modeling: a research direction

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:44:28.828850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:237c72d2ea60603514bfc88d2a7bae0c653b761cfb196218f419582626ed5187

Observation 63a74004-5198-45c2-a7ab-c65270c68565 · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Scalable agent alignment via reward modeling: a research direction

Reference 278

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T06:38:37.136877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:47708c020a54bc0d487beb1c5c89a5d4656e1c291d327c5130f8425fdc4a5d12

Observation eb7444a7-eb29-4c93-9394-6d45ee2f02de · inbound

Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering cites this paper.

Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering Scalable agent alignment via reward modeling: a research direction

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:47.669686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:28:47.669686Z digest=sha256:b35f13fc4f88b0936d6dce3cd0dca7ca2fde8023f8d93743faf80eabd2853cbb

Observation 69fad9f2-d5b7-4d57-b5ce-393613f57313 · inbound

DLBacktrace: A Model Agnostic Explainability for any Deep Learning Models cites this paper.

DLBacktrace: A Model Agnostic Explainability for any Deep Learning Models Scalable agent alignment via reward modeling: a research direction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:23:13.690025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:23:13.690025Z digest=sha256:0a27a4a9171026d6a6ec7578f4607fe05644531c56701badc26885bfa53ca170

Observation b76db881-6c47-4909-abd7-0cb3edd0fab4 · inbound

Effective Reward Specification in Deep Reinforcement Learning cites this paper.

Effective Reward Specification in Deep Reinforcement Learning Scalable agent alignment via reward modeling: a research direction

Reference 188

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:54.741369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:54.741369Z digest=sha256:f6976e5b5e88ce4ae24c275fb642fedbbaa5a5f2a1ffa0da9f0bc6905b9990c8

Observation a28e9c60-0409-4a0e-ade6-96397527604a · inbound

The Superalignment of Superhuman Intelligence with Large Language Models cites this paper.

The Superalignment of Superhuman Intelligence with Large Language Models Scalable agent alignment via reward modeling: a research direction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:18:17.792623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:18:17.792623Z digest=sha256:96fe5212dca2c0cc4274e46c507d17b5ab738fb88a1cad7694b4a63f4c2cd519

Observation 146ee1b6-60df-4fb6-a233-fbaab0d3e5fd · inbound

The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment cites this paper.

The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment Scalable agent alignment via reward modeling: a research direction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T10:36:17.396795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:36:17.396795Z digest=sha256:abe0a78b96799be6bdf1ae35f8ce8dc5073090f92bde0abe8ef533b14b09c6b8

Observation 2855cfac-44af-460d-9951-9d95e4373c8c · inbound

Aligning LLMs with Domain Invariant Reward Models cites this paper.

Aligning LLMs with Domain Invariant Reward Models Scalable agent alignment via reward modeling: a research direction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.304282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.304282Z digest=sha256:ebe38d94e749a08f5a855c08e2d4d29672d70bcc77515b64d2b265e2ea937b55

Observation 18412e5e-0a41-4f6d-a3dc-2769650d39ec · inbound

Debate Helps Weak-to-Strong Generalization cites this paper.

Debate Helps Weak-to-Strong Generalization Scalable agent alignment via reward modeling: a research direction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T17:50:56.420130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:50:56.420130Z digest=sha256:f2260218e29a2fc27e6fa69020b534a3426a1ebabd8b7584ea47c66e539e4718

Observation 067cdd38-c458-4409-b732-4577e4929598 · inbound

Leveraging Sparsity for Sample-Efficient Preference Learning: A Theoretical Perspective cites this paper.

Leveraging Sparsity for Sample-Efficient Preference Learning: A Theoretical Perspective Scalable agent alignment via reward modeling: a research direction

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T00:11:59.191696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:11:59.191696Z digest=sha256:c62ba21e14fad3b7b761580459e44451513a3f07d3dd8aff677f75117cc192d9

Observation e076956e-6c58-435c-8b9f-f2c327833128 · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards Scalable agent alignment via reward modeling: a research direction

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:23:31.059761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:dd116fd9e513473afede909fea2779bdefe30d8bdbe3e3ba44fe6db3bd10dba6

Observation ab5c7713-fe20-4ef9-a9f2-ad7da4b506d9 · inbound

Learning from Active Human Involvement through Proxy Value Propagation cites this paper.

Learning from Active Human Involvement through Proxy Value Propagation Scalable agent alignment via reward modeling: a research direction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:13.217444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:02:13.217444Z digest=sha256:97803329c3cea77c182e80c5b5e22d74bafa5539193b73e3b6d8bccc23cf9a7b

Observation 9f76cd35-3b8d-44d2-9919-f1446876d4a9 · inbound

Active Inference through Incentive Design in Markov Decision Processes cites this paper.

Active Inference through Incentive Design in Markov Decision Processes Scalable agent alignment via reward modeling: a research direction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:08.912648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:58:08.912648Z digest=sha256:3debd593805af33f92a1a799c0b523501d17cfe895ae179b448bb6f933d7e2ac

Observation f8dccded-8c51-48f6-b699-731479c7493b · inbound

DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization cites this paper.

DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Scalable agent alignment via reward modeling: a research direction

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T13:29:10.640998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:29:10.640998Z digest=sha256:24f762d916d8ff872b93e49c9090ee3df634dd6a7572ce622aaa69f866a4af2c

Observation bbb23fa1-8626-40e0-8e4b-0195ccf6c77d · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Scalable agent alignment via reward modeling: a research direction

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:32:01.324270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:3ece9e83bd676e207f7823350f9b513f5fd5a4f8e323af85e3a49e5400b27257

Observation 2667f632-86dc-4ae1-b8b1-24eca8b177bb · inbound

Exploring Societal Concerns and Perceptions of AI: A Thematic Analysis through the Lens of Problem-Seeking cites this paper.

Exploring Societal Concerns and Perceptions of AI: A Thematic Analysis through the Lens of Problem-Seeking Scalable agent alignment via reward modeling: a research direction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:20.853268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:20.853268Z digest=sha256:c977d7fc71c9de2802bc5fa4b0eed0141effe77c34bd11f4a27e725676b3ac52

Observation 45a0b585-35d0-4e79-b213-a2c76e17bc74 · inbound

Adversarial Attacks on Robotic Vision Language Action Models cites this paper.

Adversarial Attacks on Robotic Vision Language Action Models Scalable agent alignment via reward modeling: a research direction

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:40.565294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:40.565294Z digest=sha256:a999cb1c39f09dbe6c662ccbfaa08be503eaaf5be3baef99a6f1ab92638132b4

Observation 5d2bb129-4ecb-4225-92d3-85b487ae0178 · inbound

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving cites this paper.

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving Scalable agent alignment via reward modeling: a research direction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:47.141529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:47.141529Z digest=sha256:1d56d7c6bf656816676551139e5d0eac927b234c1999319d4df8bc0c49966d4e

Observation 853a6945-caa1-4574-bb72-90775fe66610 · inbound

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models cites this paper.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Scalable agent alignment via reward modeling: a research direction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.800735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.800735Z digest=sha256:57adb1ed85aaa13ad12da7de4b063ed1d6d0edbd44e0b6ffeb64546f5ae75f3b

Observation 4fc3d327-5031-45a9-bc38-86ab82cea280 · inbound

Towards Efficient and Effective Alignment of Large Language Models cites this paper.

Towards Efficient and Effective Alignment of Large Language Models Scalable agent alignment via reward modeling: a research direction

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:42.722920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:42.722920Z digest=sha256:c71adbb03aee7fc1d318f5ba1db2d958c3e8d9176a394cfcc8200dd8b882c145

Observation a75905cd-e4ab-40db-969b-80a1713f02c9 · inbound

Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack) cites this paper.

Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack) Scalable agent alignment via reward modeling: a research direction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:57.362934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:57.362934Z digest=sha256:696ab6018b89ddca3741fe96851cd873ab7d6e82ba3c97dc5714f79bbe43229e

Observation f0891b52-ab44-40d8-b5bb-2d89b1cb5696 · inbound

Optimising Language Models for Downstream Tasks: A Post-Training Perspective cites this paper.

Optimising Language Models for Downstream Tasks: A Post-Training Perspective Scalable agent alignment via reward modeling: a research direction

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-06T22:44:44.175109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:44:44.175109Z digest=sha256:e32bc86dee29efafe9ad0ce6b4475eed4f72fbe6cab3845361c812e1d69fe4ff

Observation ca30404d-f858-4923-bbef-4c3b8c65c32b · inbound

Data Diversification Methods In Alignment Enhance Math Performance In LLMs cites this paper.

Data Diversification Methods In Alignment Enhance Math Performance In LLMs Scalable agent alignment via reward modeling: a research direction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:27.141756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:43:27.141756Z digest=sha256:b83c706236840c6288d15ee2cd242772a55ccb1de686ad5c138e90e86ef09e53

Observation de72642e-059b-4375-af75-3174116c1109 · inbound

5C Prompt Contracts: A Minimalist, Creative-Friendly, Token-Efficient Design Framework for Individual and SME LLM Usage cites this paper.

5C Prompt Contracts: A Minimalist, Creative-Friendly, Token-Efficient Design Framework for Individual and SME LLM Usage Scalable agent alignment via reward modeling: a research direction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:13.642658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:13.642658Z digest=sha256:5a926dc2c591f3689fbf8035176f4317db910864a6a6087a45092ccdc8510d45

Observation 00f4f874-3bec-43ef-a11b-88b37a366202 · inbound

One Token to Fool LLM-as-a-Judge cites this paper.

One Token to Fool LLM-as-a-Judge Scalable agent alignment via reward modeling: a research direction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:37.722075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:15:37.722075Z digest=sha256:2604beecd12baed8163793deb9a9d0d0312a12f8992bdbb3b49ac3e28221cead

Observation f13d4424-1aed-4394-9070-21d3e2c96fd9 · inbound

The AI Ethical Resonance Hypothesis: The Possibility of Discovering Moral Meta-Patterns in AI Systems cites this paper.

The AI Ethical Resonance Hypothesis: The Possibility of Discovering Moral Meta-Patterns in AI Systems Scalable agent alignment via reward modeling: a research direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:04.754632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:04.754632Z digest=sha256:eed9eebcc001cc5588d2ce4cfd887ed3c439d3e4ed1a9cb84dd8d2cc0239eaf2

Observation 6893840f-8913-456c-8bfe-03992331d013 · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable agent alignment via reward modeling: a research direction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.260173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.260173Z digest=sha256:091f4faaebab5271953209a9289ce32b9ae9b2962d789b0fca4f707ef0cef36d

Observation 968bd471-84e3-41b6-900d-994e907fdce0 · inbound

Learning the Value Systems of Societies from Preferences cites this paper.

Learning the Value Systems of Societies from Preferences Scalable agent alignment via reward modeling: a research direction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:26:21.098024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:26:21.098024Z digest=sha256:10bbc0fba3a5880a6d77f4f9d266e9c1f0a72315279e1ce1dcd2835120196a43

Observation 7ea747f1-0089-4256-a17b-1ce8ffe9fec6 · inbound

Sample-efficient LLM Optimization with Reset Replay cites this paper.

Sample-efficient LLM Optimization with Reset Replay Scalable agent alignment via reward modeling: a research direction

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:41:54.702361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T23:38:08.044450Z digest=sha256:ebcc4115220605991c7dc4966651379e1b60e0f73a590c6eb10705bb646d50a3

Observation 71bf551d-680e-43a1-a83b-17835cde46b4 · inbound

SSRL: Self-Search Reinforcement Learning cites this paper.

SSRL: Self-Search Reinforcement Learning Scalable agent alignment via reward modeling: a research direction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.710424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.710424Z digest=sha256:48ad26c6585cc7cb9c35d90148b47d432ebbd4546e7aba7e6149b97d463d2b40

Observation 5251e13c-1174-463d-a131-8903d4b20d22 · inbound

Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement cites this paper.

Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement Scalable agent alignment via reward modeling: a research direction

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:54.908954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:17:54.908954Z digest=sha256:55fdd64a9263a7cae7cad1b6b72a83ba918c1bf35d6afc5a3c9f65592743473d

Observation 2a40b2d6-0391-4062-8c0f-1bc4c973d466 · inbound

Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities cites this paper.

Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities Scalable agent alignment via reward modeling: a research direction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:36.893471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:36.893471Z digest=sha256:35838a00ea76d2933588c017bfb9223f81168633d92300318d3aa0701c7c94d8

Observation d1f490a5-886d-44a7-ae8a-6e568a64fc81 · inbound

Safety Alignment Should Be Made More Than Just A Few Attention Heads cites this paper.

Safety Alignment Should Be Made More Than Just A Few Attention Heads Scalable agent alignment via reward modeling: a research direction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:37:43.446319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:37:43.446319Z digest=sha256:5164add1b30d8ea0550aa2288b6ed76855dafedb231952049edd9cde1c228baf

Observation a6500af1-ac20-43ba-9e3a-d7c9f5d72e2d · inbound

What Fundamental Structure in Reward Functions Enables Efficient Sparse-Reward Learning? cites this paper.

What Fundamental Structure in Reward Functions Enables Efficient Sparse-Reward Learning? Scalable agent alignment via reward modeling: a research direction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T10:44:53.792642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:44:53.792642Z digest=sha256:35b3fed5921c174790148c6f1902cfd8f6ec8975ecbbfa2bf332bd7992b30218

Observation 57594bd1-0c0c-4f0a-9891-1937d44e3e1c · inbound

Contrastive Weak-to-strong Generalization cites this paper.

Contrastive Weak-to-strong Generalization Scalable agent alignment via reward modeling: a research direction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T10:56:37.116353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:56:37.116353Z digest=sha256:00a771ebf138a759cf451a2bdf90ee6dfb9e2a5a9f69b6ab111d4b84cb70a804

Observation 0a71921f-3217-4d3b-ba42-5238a305e796 · inbound

Human-AI Complementarity: A Goal for Amplified Oversight cites this paper.

Human-AI Complementarity: A Goal for Amplified Oversight Scalable agent alignment via reward modeling: a research direction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:22.501890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:22.501890Z digest=sha256:1761dfd48e8c93f799ac74f4f3bc86f848a7a5ec5ec50b88861c6eeb923f8e8e

Observation 6814fecb-f832-43ca-a1e4-42b35894e8e7 · inbound

An Onto-Relational-Sophic Framework for Governing Synthetic Minds cites this paper.

An Onto-Relational-Sophic Framework for Governing Synthetic Minds Scalable agent alignment via reward modeling: a research direction

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:59:53.040763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T08:57:58.369793Z digest=sha256:888e009c931490ffa47d73ce9c263f5c9e20c18166f280902028315f31bf04d8

Observation 6f694e5f-4080-4fd4-a56e-e8d684ac58a4 · inbound

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems cites this paper.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems Scalable agent alignment via reward modeling: a research direction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:3c1252c066dbe516fd0dbb51512b41fc52958cdc7ed51d7f53eae9b8b648ce66

Observation 346dd22f-f42d-4ff7-83ba-9109cb697879 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Scalable agent alignment via reward modeling: a research direction

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.208095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:cc9f8652fc0087e2e66016953fb7c22dd900b9294588d393cbd774806017a768

Observation 290b8545-bbe7-4356-8202-21e39b3928a5 · inbound

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers cites this paper.

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Scalable agent alignment via reward modeling: a research direction

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:17.604475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T17:58:26.932519Z digest=sha256:6087478c3256900228f4bf259a348b00230bbcbe125758869ca24bb56f61f4cf

Observation 7eb3076e-cc94-4147-bf1f-713ed36275b1 · inbound

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers cites this paper.

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Scalable agent alignment via reward modeling: a research direction

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:22:28.808425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T07:20:52.825158Z digest=sha256:2ecaa541b4e1904a87d346b8fd9a2eeedd74bcbd32ab0d0e581fb41660388942

Observation 8f5f280a-5801-49d8-b7b9-e24a601ce69b · inbound

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers cites this paper.

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Scalable agent alignment via reward modeling: a research direction

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:25:39.883682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T09:24:17.904893Z digest=sha256:8753e4aa70451b34d53681d6d9c904ea272d805804653bb12bc4a5c66f5879c1

Observation ecec322c-54ad-4663-8969-26e9f2fa5233 · inbound

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers cites this paper.

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Scalable agent alignment via reward modeling: a research direction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T15:29:47.901857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:29:47.901857Z digest=sha256:93c5cf227d6867564f249ccfa457b326bfda5a1f00cd679f05af44aefe807e17

Observation ba094ca4-f60c-4c44-9d2b-3237657e92f1 · inbound

Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking cites this paper.

Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking Scalable agent alignment via reward modeling: a research direction

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.976967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T13:44:23.302523Z digest=sha256:96425ba10fea3297ef35b23e2521dcedeb567d3a1d5ceb39e5b8ef4c957c9660

Observation 15c88cbe-4c01-4dec-86a6-9f0b373e65f4 · inbound

Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking cites this paper.

Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking Scalable agent alignment via reward modeling: a research direction

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:35:33.800170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T08:26:03.578737Z digest=sha256:e48d3226e2fd36fe4d585561093da8b2f762a8112ab1acbd8a4854906eea03bc

Observation bd65d651-7eeb-464e-a2e7-6eecf7809936 · inbound

AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries cites this paper.

AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries Scalable agent alignment via reward modeling: a research direction

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:09.210282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T14:23:25.658703Z digest=sha256:cf7b870e3e6041d72d7f7c606a4477042a1b5e59eb1e4a8ec9d0aaaff3bea3f0

Observation 91f3e312-3992-454c-9b1d-5a4be522007a · inbound

AI Alignment via Incentives and Correction cites this paper.

AI Alignment via Incentives and Correction Scalable agent alignment via reward modeling: a research direction

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:01:10.353808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T14:07:44.717260Z digest=sha256:9de8a9e2fe3e82798f55095e567a10a611c4572ffc1d37a7d7c627409247293a

Observation dd840523-7196-4847-812e-c7d2c448744b · inbound

AI Alignment via Incentives and Correction cites this paper.

AI Alignment via Incentives and Correction Scalable agent alignment via reward modeling: a research direction

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:26.656504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T04:51:31.357544Z digest=sha256:25159bd2cfd6ad1d9e02d9659ec95543bf19277e3b4e3127ff26d538615e9548

Observation 11287020-73ed-4169-9207-669247e1193f · inbound

Brainrot: Deskilling and Addiction are Overlooked AI Risks cites this paper.

Brainrot: Deskilling and Addiction are Overlooked AI Risks Scalable agent alignment via reward modeling: a research direction

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:30.840167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T13:30:02.923661Z digest=sha256:d53bc0eb9ce36631d2e17f5f0956f883dca02a94c6b5fe3ba0722b52644b78f6

Observation ee9e7ea6-dfa6-4125-bf1c-c3544f0324dc · inbound

Automated alignment is harder than you think cites this paper.

Automated alignment is harder than you think Scalable agent alignment via reward modeling: a research direction

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:16:10.148716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T09:47:05.526632Z digest=sha256:8a5da1c808e18b78862180cd251337810dbc4cd3fe89afc8abba218e3ec58c50

Observation 0a830ea8-b730-42b0-b3fd-fec8c5087a4d · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Scalable agent alignment via reward modeling: a research direction

Reference 222

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:17:55.512128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:734e4794935460fda315e08685f0972f429a3afa0ba5a2d129b65403b3911f1a

Observation 302cc156-bda3-4183-ae8a-9a7e927eb4c5 · inbound

Silent Collapse in Recursive Learning Systems cites this paper.

Silent Collapse in Recursive Learning Systems Scalable agent alignment via reward modeling: a research direction

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:08:59.149378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-20T20:07:57.071410Z digest=sha256:882037c61b176b8666882dbe3e37c0588d262a588402a24aa9bbc452369cbac7

Observation 1ef68a19-2b98-47dc-ab8a-36b11d113ca7 · inbound

Deep Pre-Alignment for VLMs cites this paper.

Deep Pre-Alignment for VLMs Scalable agent alignment via reward modeling: a research direction

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T16:27:39.609156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-19T16:26:41.094936Z digest=sha256:b9032b747b9ce58307a2b2e860087d4bf3c3908b23e432af6110488e660a1485

Observation b979e1aa-d578-4595-8afc-95e9cbceae06 · inbound

When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift cites this paper.

When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift Scalable agent alignment via reward modeling: a research direction

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T21:23:58.707769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T21:23:35.843524Z digest=sha256:e3d1efd0cadd91f30535a9799640c656383eb9098da6873d28f55e2ba8e17397

Observation 2858b5f3-6215-4f10-9aa0-61270355617d · inbound

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible cites this paper.

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible Scalable agent alignment via reward modeling: a research direction

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:01.895887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T22:28:02.493124Z digest=sha256:0e7fba26abe27bf548be913a0938844d79df000d3a911f8a3e28ed4e656ac1a0

Observation 2096c4f9-dca9-4785-8631-9ae77de9396e · inbound

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible cites this paper.

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible Scalable agent alignment via reward modeling: a research direction

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-02T13:17:57.009987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:17:57.009987Z digest=sha256:8964b17ea5d03c3e480d2f10f52737ac421eeb1d4b17e5dc25f040fa109000dd

Observation c2e71a42-012c-47ec-892c-040bd312a826 · inbound

When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning cites this paper.

When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning Scalable agent alignment via reward modeling: a research direction

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:36:23.867402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T14:05:49.696342Z digest=sha256:bfc611d7ef004f5c80af2387810a6e4e5d76a922ed3366984544c61709c14a9b

Observation 6d0bac49-ca62-46cc-a8b3-0edc1f23e5ce · inbound

Self-Commitment Latency: A Reward-Free Probe for Prompted Implicit Hacking cites this paper.

Self-Commitment Latency: A Reward-Free Probe for Prompted Implicit Hacking Scalable agent alignment via reward modeling: a research direction

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:36:57.243461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T01:57:14.160273Z digest=sha256:4e6a4a84f654441a8a4017d110a683dc86fcb5a6d1f28452bb186723aae3d494

Observation 0a850ee1-20d6-4505-89d1-7c224499ee90 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Scalable agent alignment via reward modeling: a research direction

Reference 243

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:37:30.372342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:5600313c6b63c7f206b5853b70980860ddb9199ad22588318d0b0bacf214b047

Observation 604eccae-230a-4312-8723-db246ae5877e · inbound

AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks cites this paper.

AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks Scalable agent alignment via reward modeling: a research direction

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T13:08:08.763950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T08:26:57.379418Z digest=sha256:7738a4a92de4faf846f57197c73681357e78983a35a17fd57931d872ecf94d26

Observation 34fd7f76-8c20-48b0-a40f-eac988dbbaff · inbound

The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self cites this paper.

The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self Scalable agent alignment via reward modeling: a research direction

Reference 112

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:49:30.871078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T17:34:46.741962Z digest=sha256:e674fdcac3b8808a2ad7d24119a6ed33e89b1899caa757592998d958ac481cf5

Observation 45b653c3-268d-494b-967d-96ed487c5700 · inbound

Objective-Behavior Alignment: Diagnostics for MORL Policy Selection cites this paper.

Objective-Behavior Alignment: Diagnostics for MORL Policy Selection Scalable agent alignment via reward modeling: a research direction

Reference 91

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:19:38.545439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T14:34:58.398278Z digest=sha256:450103cc0f49c102bf313204725b0cf87c9c110985ea1609af2e9e4b2580f2b5

Observation 8bd9b28e-0363-4cd5-b607-d46f485e04d0 · inbound

AI Alignment From Social Choice Perspectives cites this paper.

AI Alignment From Social Choice Perspectives Scalable agent alignment via reward modeling: a research direction

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:49:37.682452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T14:12:36.892697Z digest=sha256:2d32820cdb00644ed840c16a6b4801b33479cef4a386edf89186d14b61f5e3db

Observation 2deda88b-774f-4af2-985e-a446b8904f59 · inbound

The Unverifiability of Artificial General Intelligence (AGI) Alignment, Static and Dynamic: From Trakhtenbrot's Wall to the Safety-Generality Tension cites this paper.

The Unverifiability of Artificial General Intelligence (AGI) Alignment, Static and Dynamic: From Trakhtenbrot's Wall to the Safety-Generality Tension Scalable agent alignment via reward modeling: a research direction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T11:23:07.646178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:23:07.646178Z digest=sha256:0fc5ae919e82c4f77ff073e1dc4d24a58eda20c8247d791a3d9f9c19126dcd2b

Observation f76ab2fe-e0e2-4b9f-873b-193d816330ba · inbound

Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction cites this paper.

Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction Scalable agent alignment via reward modeling: a research direction

Reference 138

Resolution
unresolved
no resolver link, observed 2026-07-13T14:35:55.709732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:35:55.709732Z digest=sha256:9e37ce87596cdfe4a2afcdd875ffc0c6e491636ca2ef8298671dfd1a42ef8ea8

Observation d133d6e5-7d41-4d4a-a333-c5e7d559ee95 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Scalable agent alignment via reward modeling: a research direction

Reference 170

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:7d86fdd810c37c3842a75111047b359c94d50bef5d6ba75cbc59a73a113e1199

Observation fcaf0638-9baf-4b6f-ba54-6d5a5ed255e1 · inbound

How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs cites this paper.

How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs Scalable agent alignment via reward modeling: a research direction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T01:34:45.323277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:34:45.323277Z digest=sha256:490d938d2baa0602e2eebef40473505c8f077a178bbe108b5b90867ae688e5d7

Observation 514601cb-6a7d-4996-94de-f388951ae611 · inbound

Attention Limited Reward Learning cites this paper.

Attention Limited Reward Learning Scalable agent alignment via reward modeling: a research direction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T16:47:52.768236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:47:52.768236Z digest=sha256:c648a3cf3e15b48dbf0067f6c76a44da5987dcdb8a468ceb12a92ccbf114c9ea

Observation a0ef56ba-a906-45f4-8759-0fcb74e602fd · inbound

Weak-to-Strong Generalization via Direct On-Policy Distillation cites this paper.

Weak-to-Strong Generalization via Direct On-Policy Distillation Scalable agent alignment via reward modeling: a research direction

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-07-07T12:33:45.005293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T12:31:42.224094Z digest=sha256:c2bb536ccb609ca2e188d6906fe1538fa99fb3df6523406068947a8428c503dc

Observation 0fbe11a8-925b-4d49-8365-804d4389dcb3 · inbound

Weak-to-Strong Generalization via Direct On-Policy Distillation cites this paper.

Weak-to-Strong Generalization via Direct On-Policy Distillation Scalable agent alignment via reward modeling: a research direction

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:a1fd8d59b1eaef14ae1158873e2c11d340d2df6b743cedd5503c069d46c293c2

Observation ee0177bf-6ff1-4f8b-9caa-d4173551a443 · inbound

Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models cites this paper.

Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models Scalable agent alignment via reward modeling: a research direction

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T06:06:14.400874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:06:14.400874Z digest=sha256:0d749a08bf355f40d535023fa07e948e00b14ece6557d1fc474eaa5ce03d4a68

Observation 5349a75b-0367-4be7-99c1-a8f420d5ca75 · inbound

Relative Value Learning cites this paper.

Relative Value Learning Scalable agent alignment via reward modeling: a research direction

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T08:32:00.178072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:32:00.178072Z digest=sha256:62b13a25f93ac941c75163878a8d46700e056f2c3ac9e3e99340343a2909d38f

Observation 8aba963f-f187-4d10-82b5-07710c08c0d8 · inbound

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops cites this paper.

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops Scalable agent alignment via reward modeling: a research direction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T00:06:00.562062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:06:00.562062Z digest=sha256:bc849f6a7e2ef8108a443692c8cce99715add8a2a87aa1b752c325a688faefd1

Observation e4a00365-adb6-467d-b5f8-9001742877bb · inbound

Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning cites this paper.

Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning Scalable agent alignment via reward modeling: a research direction

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:33:38.747570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:33:38.747570Z digest=sha256:298f0dd70ef7e615a55c337a33119eb2b8b2f19e04b3864594f3b0283377242b

Observation 3cc8b0a5-815d-48cf-9357-d58ee0e680b0 · inbound

Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction cites this paper.

Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction Scalable agent alignment via reward modeling: a research direction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T11:13:43.697532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T11:13:43.697532Z digest=sha256:89fa4ac855ce265a262805014821769a93f0b454faa4fe9a2d1de9dfb2478acb