Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:02:42.537188Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 3 inbound Pith citation observations for arXiv:2507.01368.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:02:42.537188Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T13:46:04.570352Z
75 of 75 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a309e831-ad2e-4be8-83ef-60ed9188627e · outbound
Activation Reward Models for Few-Shot Model Alignment Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3392bfb3-0a1b-4fa3-b2c9-0efd1d22b87f · outbound
Activation Reward Models for Few-Shot Model Alignment Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e39bdd43-c782-4440-85ce-e37b73592ede · outbound
Activation Reward Models for Few-Shot Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca3027ac-b74b-4c0e-b83f-5a15eface521 · outbound
Activation Reward Models for Few-Shot Model Alignment Constitutional AI: Harmlessness from AI Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71de3fff-e5fe-4a2f-8a1e-0c330a8d5a01 · outbound
Activation Reward Models for Few-Shot Model Alignment Capturing individual human preferences with reward features
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e72bd98-4df8-4b46-ab47-2f018686a0d8 · outbound
Activation Reward Models for Few-Shot Model Alignment Network dissection: Quantify- ing interpretability of deep visual representations
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 61cdfbcf-9caf-459b-81d2-7543ce242d59 · outbound
Activation Reward Models for Few-Shot Model Alignment Understanding the role of individual units in a deep neural network
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 53587fdb-b7bb-48fe-99fd-bbac1ee0e9d9 · outbound
Activation Reward Models for Few-Shot Model Alignment Language Models are Few-Shot Learners
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73ebb749-b745-4825-9bd6-38f9b129794b · outbound
Activation Reward Models for Few-Shot Model Alignment RRHF-V: Ranking responses to mitigate hallucinations in multimodal large language models with human feedback
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a7db32e6-3734-400e-ae30-4dc63aedd773 · outbound
Activation Reward Models for Few-Shot Model Alignment Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vlad Mikulik, Sam Bowman, Jan Leike, Jared Kaplan, and Ethan Perez
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bcab6f08-6cb1-457a-992b-c32bdb4344f6 · outbound
Activation Reward Models for Few-Shot Model Alignment Christiano, Jan Leike, Tom B
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 333f3104-8309-43ff-a71a-fa32a444418a · outbound
Activation Reward Models for Few-Shot Model Alignment Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8325f5-f3fe-49db-925f-63231a3714a4 · outbound
Activation Reward Models for Few-Shot Model Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd0d2183-911c-4a10-84e4-5364a54d0916 · outbound
Activation Reward Models for Few-Shot Model Alignment Paint by Word
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 565ec317-f552-4704-8525-3be5fb04fa51 · outbound
Activation Reward Models for Few-Shot Model Alignment A Survey on LLM-as-a-Judge
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b27f7cc-25e1-485c-9580-29ab4868c157 · outbound
Activation Reward Models for Few-Shot Model Alignment M-RewardBench: Evaluating Reward Models in Multilingual Settings
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 701e53e9-6d25-4896-916e-b67d26311382 · outbound
Activation Reward Models for Few-Shot Model Alignment In-context learning creates task vectors
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2b2b4cff-a11b-420f-ad15-f45d39767669 · outbound
Activation Reward Models for Few-Shot Model Alignment In-Context Learning Creates Task Vectors
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8ad2182-a253-43d3-9764-a0c2b1c06f15 · outbound
Activation Reward Models for Few-Shot Model Alignment Inspecting and Editing Knowledge Representations in Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bff11584-e388-4db1-b95c-74ee3e411e88 · outbound
Activation Reward Models for Few-Shot Model Alignment Finding visual task vectors
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 26c76699-badc-4c7e-861a-fd81a78130c5 · outbound
Activation Reward Models for Few-Shot Model Alignment SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8fe77af-5b37-4b9d-ac1d-607631c37add · outbound
Activation Reward Models for Few-Shot Model Alignment Multimodal task vectors enable many-shot multimodal in-context learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 739cb215-0073-4d93-91f7-2a70c9221e3d · outbound
Activation Reward Models for Few-Shot Model Alignment Multimodal task vectors enable many-shot multimodal in-context learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ca3a6e68-0545-4c7a-aadb-9fb082c0ee22 · outbound
Activation Reward Models for Few-Shot Model Alignment RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8acdb000-0686-4932-9d12-8726a2b6d43c · outbound
Activation Reward Models for Few-Shot Model Alignment Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5529c393-9ff4-4b69-8aad-fd628fabf51a · outbound
Activation Reward Models for Few-Shot Model Alignment RewardBench: Evaluating Reward Models for Language Modeling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a064690-906b-4d33-b571-b637800c65ef · outbound
Activation Reward Models for Few-Shot Model Alignment RLAIF vs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1bd11883-f8cd-44a9-b664-ddd46ee6c5ff · outbound
Activation Reward Models for Few-Shot Model Alignment RLAIF vs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 301eebe0-a8d2-4d91-9e22-d502c04e8352 · outbound
Activation Reward Models for Few-Shot Model Alignment The power of scale for parameter-efficient prompt tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d5696a21-e77d-48dd-87c1-d6158f0f5657 · outbound
Activation Reward Models for Few-Shot Model Alignment LLaVA-OneVision: Easy Visual Task Transfer
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ac81866-ea23-4e7a-89f0-c1a62e6a1509 · outbound
Activation Reward Models for Few-Shot Model Alignment BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0495770b-eb30-4685-bb38-1008d01ebe57 · outbound
Activation Reward Models for Few-Shot Model Alignment Evaluating text-to-visual generation with image-to-text generation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 82779433-6465-4df6-9e46-967589565a75 · outbound
Activation Reward Models for Few-Shot Model Alignment Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e87e8d74-2e2e-4461-b28c-1e519c258d34 · outbound
Activation Reward Models for Few-Shot Model Alignment Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2295d038-75a9-4a7b-92da-5174726ff034 · outbound
Activation Reward Models for Few-Shot Model Alignment Rule Based Rewards for Language Model Safety
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f932738d-68f2-4a43-9acd-b74640b2bcbf · outbound
Activation Reward Models for Few-Shot Model Alignment In-context Learning and Induction Heads
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9de138df-d363-4716-a9cf-d5e70f13e740 · outbound
Activation Reward Models for Few-Shot Model Alignment GPT-4 Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e53c4cd-a86c-401b-a243-ea10fb95f939 · outbound
Activation Reward Models for Few-Shot Model Alignment Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ed740f6a-768d-440b-ae61-43cb915c2319 · outbound
Activation Reward Models for Few-Shot Model Alignment Training language models to follow instructions with human feedback
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 09a7c555-8945-4616-ac78-ffc005f697fa · outbound
Activation Reward Models for Few-Shot Model Alignment Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 18c8f537-ed32-4557-a495-86e6bfb803a1 · outbound
Activation Reward Models for Few-Shot Model Alignment Steering Llama 2 via Contrastive Activation Addition
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ef9ecfa-3019-442a-90d9-88d37689425d · outbound
Activation Reward Models for Few-Shot Model Alignment Pytorch: An imperative style, high-performance deep learning library
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdc8313f-0b58-48f9-b918-8a89e203f6ba · outbound
Activation Reward Models for Few-Shot Model Alignment Red Teaming Language Models with Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89618b88-420f-45bd-ae5d-8471c85695b0 · outbound
Activation Reward Models for Few-Shot Model Alignment Improving language understanding by generative pre-training
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0234c10-4552-46d9-a723-ca951d5ca19e · outbound
Activation Reward Models for Few-Shot Model Alignment Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b59a616-ab3d-40e1-a961-41b0ddcf8921 · outbound
Activation Reward Models for Few-Shot Model Alignment GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08a76f31-3f9e-4e7a-923b-1454c5381435 · outbound
Activation Reward Models for Few-Shot Model Alignment Proximal Policy Optimization Algorithms
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68bbd057-2773-4fa5-a82f-ded588030d21 · outbound
Activation Reward Models for Few-Shot Model Alignment Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6b977d1-c958-4b0e-8354-29a257170a05 · outbound
Activation Reward Models for Few-Shot Model Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 416fede6-7ea0-435b-b28f-6c9561e43e39 · outbound
Activation Reward Models for Few-Shot Model Alignment FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01895fb7-6d99-4260-ad20-7e6d22c9fcd4 · outbound
Activation Reward Models for Few-Shot Model Alignment Alpaca: A strong, replicable instruction- following model
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6a231fd1-79dc-4ccd-9ef3-66cb51d24458 · outbound
Activation Reward Models for Few-Shot Model Alignment Learning to summarize with human feedback
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01348201-5f79-4333-a21a-8e69d5d86b58 · outbound
Activation Reward Models for Few-Shot Model Alignment Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, Paul Christiano, Jan Leike, and Others
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7597d84e-84c7-4f0c-ae61-22ea44d0e9ff · outbound
Activation Reward Models for Few-Shot Model Alignment Extracting latent steering vectors from pretrained language models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6c308a2a-af65-4032-94a0-27c9dc16d7ee · outbound
Activation Reward Models for Few-Shot Model Alignment Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87fff3e-fc59-4c55-8b05-953b8182f488 · outbound
Activation Reward Models for Few-Shot Model Alignment Li, Arnab Sen Sharma, Aaron Mueller, Byron C
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5b0d810c-d934-4b46-a0ab-e1a0b4d32077 · outbound
Activation Reward Models for Few-Shot Model Alignment LLaMA: Open and Efficient Foundation Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d292f8cf-a3bb-4424-bba5-1c9b73ac1628 · outbound
Activation Reward Models for Few-Shot Model Alignment Steering Language Models With Activation Engineering
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0166f73-c608-4948-9c7d-e922a1945039 · outbound
Activation Reward Models for Few-Shot Model Alignment Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 885d840c-9f1f-4e78-bc23-e9c322835eef · outbound
Activation Reward Models for Few-Shot Model Alignment Large language models are not fair evaluators
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e8db6c18-6cf8-4e95-9d2c-fd2e7aa05d7e · outbound
Activation Reward Models for Few-Shot Model Alignment Williams
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5c653327-9475-4a00-8979-ce6e5ce22713 · outbound
Activation Reward Models for Few-Shot Model Alignment rewordbench: Benchmarking and improving the robustness of reward models with transformed inputs
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2908560-f838-4916-a340-625aab08f338 · outbound
Activation Reward Models for Few-Shot Model Alignment Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ac0d03-a65d-4f0b-ad1a-586be8f418bb · outbound
Activation Reward Models for Few-Shot Model Alignment Zettlemoyer, and Marjan Ghazvininejad
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a898471a-c5b5-4a9c-bb5c-6c89f20c46d1 · outbound
Activation Reward Models for Few-Shot Model Alignment Which attention heads matter for in-context learning? In Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fdfc61a0-60bb-41b1-8192-1c5893abc6df · outbound
Activation Reward Models for Few-Shot Model Alignment ICPL: Few-shot In-context Preference Learning via LLMs
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d7b7262-f799-465f-a7e2-0578d1c3d2a1 · outbound
Activation Reward Models for Few-Shot Model Alignment Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8f5b9bb-35ec-4f6f-84c6-d5d8950363da · outbound
Activation Reward Models for Few-Shot Model Alignment RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaef4565-ec43-4aa0-8dd9-7c359a83de49 · outbound
Activation Reward Models for Few-Shot Model Alignment Rag-reward: Optimizing rag with reward modeling and rlhf
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 206120a9-5041-4a56-a862-e9e6169bdc5a · outbound
Activation Reward Models for Few-Shot Model Alignment Generative verifiers: Reward modeling as next-token prediction
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0ac44e47-5c44-4ac3-9f80-4bbcdfea8319 · outbound
Activation Reward Models for Few-Shot Model Alignment Generative verifiers: Reward modeling as next-token prediction
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7508d138-a168-4880-8c5b-b1f4a844fd42 · outbound
Activation Reward Models for Few-Shot Model Alignment MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39407f16-bc69-4dcc-afc7-8c670e2c90e6 · outbound
Activation Reward Models for Few-Shot Model Alignment Interpreting deep visual representations via network dissection
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a0b6a405-a8af-46ec-8d6b-3882c7fdc397 · outbound
Activation Reward Models for Few-Shot Model Alignment Fine-Tuning Language Models from Human Preferences
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97ed4056-3f3b-4b9c-8b63-918eb0850aa2 · outbound
Activation Reward Models for Few-Shot Model Alignment Unresolved cited work
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation df15135a-6ad7-4bf5-9d50-2dfb086d01af · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Activation Reward Models for Few-Shot Model Alignment
Reference 222
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c697a69a-b83f-45aa-8222-9158dc93e9e7 · inbound
Building a Precise Video Language with Human-AI Oversight Activation Reward Models for Few-Shot Model Alignment
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e10fc225-4be8-474f-9947-e9332a7ffc41 · inbound
Multimodal Reward Hacking in Reinforcement Learning Activation Reward Models for Few-Shot Model Alignment
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.