Pith. sign in

Paper Citation Record · LEDGER

Activation Reward Models for Few-Shot Model Alignment

As of 14 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 3 inbound Pith citation observations for arXiv:2507.01368.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01368 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:02:42.537188Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T13:46:04.570352Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a309e831-ad2e-4be8-83ef-60ed9188627e · outbound

This paper cites an unresolved cited work.

Activation Reward Models for Few-Shot Model Alignment Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.321263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.321263Z digest=sha256:f8526006f9c815e475b55c31d014ec8ee70cf43639b07196f135be2b704a86a6

Observation 3392bfb3-0a1b-4fa3-b2c9-0efd1d22b87f · outbound

This paper cites Qwen2.5-VL Technical Report.

Activation Reward Models for Few-Shot Model Alignment Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.325145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.325145Z digest=sha256:6142e4f57ae4289c47bde7a89061a241472d1aa51a26baed5ee185a45fbf5278

Observation e39bdd43-c782-4440-85ce-e37b73592ede · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Activation Reward Models for Few-Shot Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.331317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.331317Z digest=sha256:f3884f984f1ebcd3e134a3ed9c7fa1e8669433ff0ac61c9cbae0f6b774b34a76

Observation ca3027ac-b74b-4c0e-b83f-5a15eface521 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Activation Reward Models for Few-Shot Model Alignment Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.337121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.337121Z digest=sha256:c1253e62a15f55853c8be0511e0de7439fc537a572858a5f4b43b613d44fe875

Observation 71de3fff-e5fe-4a2f-8a1e-0c330a8d5a01 · outbound

This paper cites Capturing individual human preferences with reward features.

Activation Reward Models for Few-Shot Model Alignment Capturing individual human preferences with reward features

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.339704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.339704Z digest=sha256:e65389b0c17b8ec9e9fed7c0a2d131e609048d00e8c4f50532285151a85fe92c

Observation 0e72bd98-4df8-4b46-ab47-2f018686a0d8 · outbound

This paper cites Network dissection: Quantify- ing interpretability of deep visual representations.

Activation Reward Models for Few-Shot Model Alignment Network dissection: Quantify- ing interpretability of deep visual representations

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.378656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.342712Z digest=sha256:cb48c765c8145e011a5729ef9582539a49e0a615ed6ec8ed98df3531237b6e9a

Observation 61cdfbcf-9caf-459b-81d2-7543ce242d59 · outbound

This paper cites Understanding the role of individual units in a deep neural network.

Activation Reward Models for Few-Shot Model Alignment Understanding the role of individual units in a deep neural network

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.370334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.345424Z digest=sha256:c0f33a97b49301b43eb0496e8438f41e110c051195f6ac9063afe91376081d0e

Observation 53587fdb-b7bb-48fe-99fd-bbac1ee0e9d9 · outbound

This paper cites Language Models are Few-Shot Learners.

Activation Reward Models for Few-Shot Model Alignment Language Models are Few-Shot Learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.348325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.348325Z digest=sha256:35ad780a34d986b40d36e69481f0c49f19a9653cd409f1b0196ab820c678929a

Observation 73ebb749-b745-4825-9bd6-38f9b129794b · outbound

This paper cites RRHF-V: Ranking responses to mitigate hallucinations in multimodal large language models with human feedback.

Activation Reward Models for Few-Shot Model Alignment RRHF-V: Ranking responses to mitigate hallucinations in multimodal large language models with human feedback

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.361377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.351096Z digest=sha256:af4c1902147231058b1427353c4183d3d5a122f6513951c5c3175fa5dfd5fbb4

Observation a7db32e6-3734-400e-ae30-4dc63aedd773 · outbound

This paper cites Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vlad Mikulik, Sam Bowman, Jan Leike, Jared Kaplan, and Ethan Perez.

Activation Reward Models for Few-Shot Model Alignment Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vlad Mikulik, Sam Bowman, Jan Leike, Jared Kaplan, and Ethan Perez

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.344698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.357004Z digest=sha256:2bcdce13cde03108d82ceb95a5f561d6dc4a61b6924e276c2e8424203f649747

Observation bcab6f08-6cb1-457a-992b-c32bdb4344f6 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Activation Reward Models for Few-Shot Model Alignment Christiano, Jan Leike, Tom B

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.336451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.359782Z digest=sha256:1d8bf8e747ffd2156c4beec1ffdfd38dd81d5dea2cd0e1ff23799c44c6f6d0e5

Observation 333f3104-8309-43ff-a71a-fa32a444418a · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Activation Reward Models for Few-Shot Model Alignment Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.362383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.362383Z digest=sha256:b168a98297610b07cd2b8079fa593aec8ef21fd821235c7b47fad8a74cc36ade

Observation 0b8325f5-f3fe-49db-925f-63231a3714a4 · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Activation Reward Models for Few-Shot Model Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.365297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.365297Z digest=sha256:ba5712f3d44c9af6beb484ab292980f069f93835d3c52beee3de03346fd7d952

Observation dd0d2183-911c-4a10-84e4-5364a54d0916 · outbound

This paper cites Paint by Word.

Activation Reward Models for Few-Shot Model Alignment Paint by Word

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.368595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.368595Z digest=sha256:13953aab7a27bc042b170a82240edac26789e7cf45b6e1963f984dc40e1552bb

Observation 565ec317-f552-4704-8525-3be5fb04fa51 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Activation Reward Models for Few-Shot Model Alignment A Survey on LLM-as-a-Judge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.371326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.371326Z digest=sha256:97e091d4cd5191597daa09e856302aba94fd4f226c3bff1584e9c972b031eca7

Observation 7b27f7cc-25e1-485c-9580-29ab4868c157 · outbound

This paper cites M-RewardBench: Evaluating Reward Models in Multilingual Settings.

Activation Reward Models for Few-Shot Model Alignment M-RewardBench: Evaluating Reward Models in Multilingual Settings

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.374334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.374334Z digest=sha256:47c336201439db3b1840a95ddc3f188d0530bdfcad141d72827986191ef9ef64

Observation 701e53e9-6d25-4896-916e-b67d26311382 · outbound

This paper cites In-context learning creates task vectors.

Activation Reward Models for Few-Shot Model Alignment In-context learning creates task vectors

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.327878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.377311Z digest=sha256:29234225c276d9f0fa859e92b49793e7d0e3220cb0c265ef78651d8baba40468

Observation 2b2b4cff-a11b-420f-ad15-f45d39767669 · outbound

This paper cites In-Context Learning Creates Task Vectors.

Activation Reward Models for Few-Shot Model Alignment In-Context Learning Creates Task Vectors

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.379822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.379822Z digest=sha256:d3fdfc978554c7ea29e7386879b5fceee9716fce535fadc13c3dfd9210ef4d43

Observation c8ad2182-a253-43d3-9764-a0c2b1c06f15 · outbound

This paper cites Inspecting and Editing Knowledge Representations in Language Models.

Activation Reward Models for Few-Shot Model Alignment Inspecting and Editing Knowledge Representations in Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.382687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.382687Z digest=sha256:3d36b5c5cfa9be44a78ae0ed9a9faef8cd0de40e70da80abb6e7be2aab61e1ac

Observation bff11584-e388-4db1-b95c-74ee3e411e88 · outbound

This paper cites Finding visual task vectors.

Activation Reward Models for Few-Shot Model Alignment Finding visual task vectors

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.318883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.385383Z digest=sha256:3024027d1a46057c0f1673461040880b0b88e157dc8838bd7182853144bb9417

Observation 26c76699-badc-4c7e-861a-fd81a78130c5 · outbound

This paper cites SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality.

Activation Reward Models for Few-Shot Model Alignment SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.387975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.387975Z digest=sha256:4b7e3a23165a2d9ba65d14751d1891b742cee56d0197a2d75517665fc608d1c1

Observation d8fe77af-5b37-4b9d-ac1d-607631c37add · outbound

This paper cites Multimodal task vectors enable many-shot multimodal in-context learning.

Activation Reward Models for Few-Shot Model Alignment Multimodal task vectors enable many-shot multimodal in-context learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.309709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.390963Z digest=sha256:2cb73cd6a2a30aa6fe6c5242c2ee644186a3507710bed85c43f04db2c4b40c66

Observation 739cb215-0073-4d93-91f7-2a70c9221e3d · outbound

This paper cites Multimodal task vectors enable many-shot multimodal in-context learning.

Activation Reward Models for Few-Shot Model Alignment Multimodal task vectors enable many-shot multimodal in-context learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.300705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.393812Z digest=sha256:790e5dfaf6f52acf911c05cef020702d301328ea949ec8c376dd3cefc16962b5

Observation ca3a6e68-0545-4c7a-aadb-9fb082c0ee22 · outbound

This paper cites RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment.

Activation Reward Models for Few-Shot Model Alignment RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.396463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.396463Z digest=sha256:17c53be2a1807d0222d6878e52101023926318671b03b5715016343d22265ec5

Observation 8acdb000-0686-4932-9d12-8726a2b6d43c · outbound

This paper cites Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes.

Activation Reward Models for Few-Shot Model Alignment Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.399277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.399277Z digest=sha256:98d87d431e1a970b556832c233edaa5209466355cbd21d928a0580eb09f4ad81

Observation 5529c393-9ff4-4b69-8aad-fd628fabf51a · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Activation Reward Models for Few-Shot Model Alignment RewardBench: Evaluating Reward Models for Language Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.402165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.402165Z digest=sha256:ed576c18ded7666a72cdb17224ed85aebd40a6d7a92f941402555cb3d06a9f00

Observation 1a064690-906b-4d33-b571-b637800c65ef · outbound

This paper cites RLAIF vs.

Activation Reward Models for Few-Shot Model Alignment RLAIF vs

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.291393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.405063Z digest=sha256:2df51f6fad3460346e2d787c5770262283daf45356bf7381e01bea25390ac59b

Observation 1bd11883-f8cd-44a9-b664-ddd46ee6c5ff · outbound

This paper cites RLAIF vs.

Activation Reward Models for Few-Shot Model Alignment RLAIF vs

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.281451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.407759Z digest=sha256:58e77c9f9d1e0937e80a169a97debc17f5cd733ca392afe91d17e0e48a927ba2

Observation 301eebe0-a8d2-4d91-9e22-d502c04e8352 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Activation Reward Models for Few-Shot Model Alignment The power of scale for parameter-efficient prompt tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.272260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.410251Z digest=sha256:6a5df65609060641c5f4c57a499fb09539c3eae07ef2bdabfcbb7800042000aa

Observation d5696a21-e77d-48dd-87c1-d6158f0f5657 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Activation Reward Models for Few-Shot Model Alignment LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.413245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.413245Z digest=sha256:0575a1f36c9bfa880912f2c142ee82bcab919d0cdbe867c3efe14da4a973e28a

Observation 0ac81866-ea23-4e7a-89f0-c1a62e6a1509 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

Activation Reward Models for Few-Shot Model Alignment BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.416301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.416301Z digest=sha256:26802d67fd0eb396796e43a9fa74cbe1af665aced182a9f285c1b5a97f0b8542

Observation 0495770b-eb30-4685-bb38-1008d01ebe57 · outbound

This paper cites Evaluating text-to-visual generation with image-to-text generation.

Activation Reward Models for Few-Shot Model Alignment Evaluating text-to-visual generation with image-to-text generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.256118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.418758Z digest=sha256:c3951d00ac67937a2b1272d333fdf5f9f528d557ec226a6a1f8d9f18170ec70e

Observation 82779433-6465-4df6-9e46-967589565a75 · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.

Activation Reward Models for Few-Shot Model Alignment Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.247216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.421311Z digest=sha256:d6645c73152faec7d21e4a6c0a2e2551bb491fa27502e277b186ccb75af35db3

Observation e87e8d74-2e2e-4461-b28c-1e519c258d34 · outbound

This paper cites Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features.

Activation Reward Models for Few-Shot Model Alignment Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.424201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.424201Z digest=sha256:c5587e2f1bed04adda500b9798c4d877e9c84838dc4175187884d8f82a59dbc0

Observation 2295d038-75a9-4a7b-92da-5174726ff034 · outbound

This paper cites Rule Based Rewards for Language Model Safety.

Activation Reward Models for Few-Shot Model Alignment Rule Based Rewards for Language Model Safety

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.427164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.427164Z digest=sha256:92bdb27b5b68769a9f350b07e720fe878e35c6d174d28f65d748c5afcc537592

Observation f932738d-68f2-4a43-9acd-b74640b2bcbf · outbound

This paper cites In-context Learning and Induction Heads.

Activation Reward Models for Few-Shot Model Alignment In-context Learning and Induction Heads

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.429971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.429971Z digest=sha256:44f72e00af7a1d4aa174be10530edbb1ce93f1d848666a6216ac367f312326cf

Observation 9de138df-d363-4716-a9cf-d5e70f13e740 · outbound

This paper cites GPT-4 Technical Report.

Activation Reward Models for Few-Shot Model Alignment GPT-4 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.432668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.432668Z digest=sha256:3151363eb71a6fe8e571f81c9aa7aeb66d1b5618b54d47d72dcbddf17b5e3832

Observation 2e53c4cd-a86c-401b-a243-ea10fb95f939 · outbound

This paper cites an unresolved cited work.

Activation Reward Models for Few-Shot Model Alignment Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:43.238094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.435269Z digest=sha256:09db25dc08e6d615403dbb50dea76077e8180a9444103b4ab9f78865f45c3a95

Observation ed740f6a-768d-440b-ae61-43cb915c2319 · outbound

This paper cites Training language models to follow instructions with human feedback.

Activation Reward Models for Few-Shot Model Alignment Training language models to follow instructions with human feedback

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.229102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.438395Z digest=sha256:66bc19de6801efa726712f8ee667d66f97521d5db3f5c78034f15436647c6332

Observation 09a7c555-8945-4616-ac78-ffc005f697fa · outbound

This paper cites an unresolved cited work.

Activation Reward Models for Few-Shot Model Alignment Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:43.219895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.440917Z digest=sha256:991f685f8609cc915629c0b6317526a5d03665f9131633951616be835fc018b9

Observation 18c8f537-ed32-4557-a495-86e6bfb803a1 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Activation Reward Models for Few-Shot Model Alignment Steering Llama 2 via Contrastive Activation Addition

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.443531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.443531Z digest=sha256:74a4e64021891e0c19ad391a424c87f712353bcaee6cd9202e2c693a433e0252

Observation 6ef9ecfa-3019-442a-90d9-88d37689425d · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Activation Reward Models for Few-Shot Model Alignment Pytorch: An imperative style, high-performance deep learning library

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.446511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.446511Z digest=sha256:f108eb0d96c297a99e0ef172225eb3e092770da11f771130155cfd3b89c6a93d

Observation bdc8313f-0b58-48f9-b918-8a89e203f6ba · outbound

This paper cites Red Teaming Language Models with Language Models.

Activation Reward Models for Few-Shot Model Alignment Red Teaming Language Models with Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.449385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.449385Z digest=sha256:5d28f77eb5bcb011e70c61e10e60dba7a2a3b5a5ecf791a5f1ec21d96fd255e6

Observation 89618b88-420f-45bd-ae5d-8471c85695b0 · outbound

This paper cites Improving language understanding by generative pre-training.

Activation Reward Models for Few-Shot Model Alignment Improving language understanding by generative pre-training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.452328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.452328Z digest=sha256:4d6405f95bde16e76a9b230003079d6ac9b65de715bb0b84c3a83c2c896eea3c

Observation f0234c10-4552-46d9-a723-ca951d5ca19e · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Activation Reward Models for Few-Shot Model Alignment Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.455001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.455001Z digest=sha256:bcd558707661fde0a82b8a8745bcc6fd7d60a3f8bff0c2fe038263ded4ab2271

Observation 9b59a616-ab3d-40e1-a961-41b0ddcf8921 · outbound

This paper cites GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation.

Activation Reward Models for Few-Shot Model Alignment GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.457877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.457877Z digest=sha256:66ac5bd50cafbcf71f1bd50dce72e204edaef74b8e0884a3a5356bce66015ef7

Observation 08a76f31-3f9e-4e7a-923b-1454c5381435 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Activation Reward Models for Few-Shot Model Alignment Proximal Policy Optimization Algorithms

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.461296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.461296Z digest=sha256:59ff85e8ebd02ed2c7485bd3bd038315f58ff1f5b08059031e2967be60820b02

Observation 68bbd057-2773-4fa5-a82f-ded588030d21 · outbound

This paper cites Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations.

Activation Reward Models for Few-Shot Model Alignment Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.463921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.463921Z digest=sha256:a1eb4e92493a07783ff47d748371dc674886481930c8c834d394a86e9d7fb44d

Observation a6b977d1-c958-4b0e-8354-29a257170a05 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Activation Reward Models for Few-Shot Model Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.466647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.466647Z digest=sha256:f36b5daa275fd4dff0c80d42b720e8a5643dd12ae6adaf693be2c4d560a2bd98

Observation 416fede6-7ea0-435b-b28f-6c9561e43e39 · outbound

This paper cites FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users.

Activation Reward Models for Few-Shot Model Alignment FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.469412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.469412Z digest=sha256:ca5b0ca1c6900df2e7dd6c75a88e1e4680c107df1272209e819a7e64469f58cd

Observation 01895fb7-6d99-4260-ad20-7e6d22c9fcd4 · outbound

This paper cites Alpaca: A strong, replicable instruction- following model.

Activation Reward Models for Few-Shot Model Alignment Alpaca: A strong, replicable instruction- following model

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.199194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.472472Z digest=sha256:33068d6ea3c4e0aa95e5cc20edd95b624056fbc589b76cbd031e5fb28f023ab3

Observation 6a231fd1-79dc-4ccd-9ef3-66cb51d24458 · outbound

This paper cites Learning to summarize with human feedback.

Activation Reward Models for Few-Shot Model Alignment Learning to summarize with human feedback

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.475288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.475288Z digest=sha256:d2d88d80dd95df6551c388a78a4a75b7b547553daca3528f7807772976329df5

Observation 01348201-5f79-4333-a21a-8e69d5d86b58 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, Paul Christiano, Jan Leike, and Others.

Activation Reward Models for Few-Shot Model Alignment Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, Paul Christiano, Jan Leike, and Others

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.183491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.477978Z digest=sha256:0aa4b55c622767fba7c3dc807a6adbe1591d4742348b1a5141d926c89adb9a16

Observation 7597d84e-84c7-4f0c-ae61-22ea44d0e9ff · outbound

This paper cites Extracting latent steering vectors from pretrained language models.

Activation Reward Models for Few-Shot Model Alignment Extracting latent steering vectors from pretrained language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.173972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.480561Z digest=sha256:7c3588a4474ac3e23e4a520f9a5d4b7135079fa5d0f2db0274e9020f02fd5431

Observation 6c308a2a-af65-4032-94a0-27c9dc16d7ee · outbound

This paper cites Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence.

Activation Reward Models for Few-Shot Model Alignment Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.483528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.483528Z digest=sha256:47d2b30269321069fc83c7b69813e26c3b7d1270fe7cb0951fb3234b985fba09

Observation d87fff3e-fc59-4c55-8b05-953b8182f488 · outbound

This paper cites Li, Arnab Sen Sharma, Aaron Mueller, Byron C.

Activation Reward Models for Few-Shot Model Alignment Li, Arnab Sen Sharma, Aaron Mueller, Byron C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.163979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.486151Z digest=sha256:2364adef32072d6ac9fb1fd1ee3f7410f0552caa7dc5e8469da647036257ef69

Observation 5b0d810c-d934-4b46-a0ab-e1a0b4d32077 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Activation Reward Models for Few-Shot Model Alignment LLaMA: Open and Efficient Foundation Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.488755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.488755Z digest=sha256:830dc5daf2ebc4889bcc52abe9cc901b1790493a377313287efd2f9774d4b549

Observation d292f8cf-a3bb-4424-bba5-1c9b73ac1628 · outbound

This paper cites Steering Language Models With Activation Engineering.

Activation Reward Models for Few-Shot Model Alignment Steering Language Models With Activation Engineering

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.491453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.491453Z digest=sha256:bc8c3bd4cbb748fb99ea20a25d48079726ed3db95db5f978f973729ae25ea297

Observation c0166f73-c608-4948-9c7d-e922a1945039 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Activation Reward Models for Few-Shot Model Alignment Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.494467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.494467Z digest=sha256:1fe923610977ba575491607fcdad4a668f41fb08c8573c68c5b705998225d939

Observation 885d840c-9f1f-4e78-bc23-e9c322835eef · outbound

This paper cites Large language models are not fair evaluators.

Activation Reward Models for Few-Shot Model Alignment Large language models are not fair evaluators

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.154796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.497393Z digest=sha256:c93cf693147a7c95d48e975a910b6f63337fc18216bfb20d46c5da1aa93e3ea3

Observation e8db6c18-6cf8-4e95-9d2c-fd2e7aa05d7e · outbound

This paper cites Williams.

Activation Reward Models for Few-Shot Model Alignment Williams

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.144570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.500110Z digest=sha256:6911330291fea574ea87ab3952e410e9adc6057ac7f00c9e49e0f0c5a46c577d

Observation 5c653327-9475-4a00-8979-ce6e5ce22713 · outbound

This paper cites rewordbench: Benchmarking and improving the robustness of reward models with transformed inputs.

Activation Reward Models for Few-Shot Model Alignment rewordbench: Benchmarking and improving the robustness of reward models with transformed inputs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.502647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.502647Z digest=sha256:b71ea126b002270347af7f41846b1158bccd04ccbf0c8fea69230cb0d5f835ba

Observation e2908560-f838-4916-a340-625aab08f338 · outbound

This paper cites Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models.

Activation Reward Models for Few-Shot Model Alignment Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.505501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.505501Z digest=sha256:19eea363104da1adb8ade3198c327f4926817540c49a78b562fdf0d33dfde957

Observation b9ac0d03-a65d-4f0b-ad1a-586be8f418bb · outbound

This paper cites Zettlemoyer, and Marjan Ghazvininejad.

Activation Reward Models for Few-Shot Model Alignment Zettlemoyer, and Marjan Ghazvininejad

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.135993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.508936Z digest=sha256:8319626e2057d75c5fe005a700e954d2b12f6e2193aea164671e8c78381461b9

Observation a898471a-c5b5-4a9c-bb5c-6c89f20c46d1 · outbound

This paper cites Which attention heads matter for in-context learning? In Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025.

Activation Reward Models for Few-Shot Model Alignment Which attention heads matter for in-context learning? In Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.127194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.511545Z digest=sha256:3896218cab7f591f6288364a277c94a360194815553faacffb7e068ecfb3b1eb

Observation fdfc61a0-60bb-41b1-8192-1c5893abc6df · outbound

This paper cites ICPL: Few-shot In-context Preference Learning via LLMs.

Activation Reward Models for Few-Shot Model Alignment ICPL: Few-shot In-context Preference Learning via LLMs

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.514668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.514668Z digest=sha256:c68648e80bf502210c534e89a1713aaba89cee0ab6ed25f518748b4e13733990

Observation 4d7b7262-f799-465f-a7e2-0578d1c3d2a1 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

Activation Reward Models for Few-Shot Model Alignment Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.517736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.517736Z digest=sha256:0ec8ee2788a057546db1d693b318a8b24d50fa761bd4ad3c4eb5d5a4af03d017

Observation e8f5b9bb-35ec-4f6f-84c6-d5d8950363da · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

Activation Reward Models for Few-Shot Model Alignment RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.520388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.520388Z digest=sha256:0250729a2de29452244171e326fee208077e176e57b807a848a764b477607f5e

Observation eaef4565-ec43-4aa0-8dd9-7c359a83de49 · outbound

This paper cites Rag-reward: Optimizing rag with reward modeling and rlhf.

Activation Reward Models for Few-Shot Model Alignment Rag-reward: Optimizing rag with reward modeling and rlhf

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.523148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.523148Z digest=sha256:41d88caf59b38458d4382e297dc98beaf8aadeb39cf326607f1acb7dfdfa2b9b

Observation 206120a9-5041-4a56-a862-e9e6169bdc5a · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

Activation Reward Models for Few-Shot Model Alignment Generative verifiers: Reward modeling as next-token prediction

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.110641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.526351Z digest=sha256:048ca9eab9d4de2613c84d3eceeb15c6f72770069b71fde1e83e2bc8af092a86

Observation 0ac44e47-5c44-4ac3-9f80-4bbcdfea8319 · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

Activation Reward Models for Few-Shot Model Alignment Generative verifiers: Reward modeling as next-token prediction

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.100892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.529059Z digest=sha256:59a8fbb6b465d0bd117f4fdba527db75e43d80e292e7343c39881715bf269cb4

Observation 7508d138-a168-4880-8c5b-b1f4a844fd42 · outbound

This paper cites MM-RLHF: The Next Step Forward in Multimodal LLM Alignment.

Activation Reward Models for Few-Shot Model Alignment MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.531667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.531667Z digest=sha256:dd4bc6699ceb12162e76891ee67a36597b7e7158509c3539a4bfb3fd5d4a8bbc

Observation 39407f16-bc69-4dcc-afc7-8c670e2c90e6 · outbound

This paper cites Interpreting deep visual representations via network dissection.

Activation Reward Models for Few-Shot Model Alignment Interpreting deep visual representations via network dissection

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.090900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.534698Z digest=sha256:66fa82cfdd60cce9d0620483aa2e8919bd791e86389acc64c69bc5a64d6115dd

Observation a0b6a405-a8af-46ec-8d6b-3882c7fdc397 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Activation Reward Models for Few-Shot Model Alignment Fine-Tuning Language Models from Human Preferences

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.537188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.537188Z digest=sha256:1096d148c91d68b31799e0d841d5768840ff728c8097899e89f69dab1039e636

Observation 97ed4056-3f3b-4b9c-8b63-918eb0850aa2 · outbound

This paper cites an unresolved cited work.

Activation Reward Models for Few-Shot Model Alignment Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:43.352767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.354176Z digest=sha256:df581b556a5e7d23aa069fbe57bb927e514d65a9013c7eaeba04eabdec7cb4d7

Pith citing papers

Observation df15135a-6ad7-4bf5-9d50-2dfb086d01af · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Activation Reward Models for Few-Shot Model Alignment

Reference 222

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.437494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:24dff51563f48bcde4fdb8ba4089730ac28a8d25f76ed6ada725d4dbb8efb9b1

Observation c697a69a-b83f-45aa-8222-9158dc93e9e7 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Activation Reward Models for Few-Shot Model Alignment

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.573192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:50493ad5f77e1f9628bcedd8a9d93bd9f9fc8c1266e8d3fd1402c82011869de2

Observation e10fc225-4be8-474f-9947-e9332a7ffc41 · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning Activation Reward Models for Few-Shot Model Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:2d12b829946fc76b1179da4aa941d4b2bdd529f0dc94237cf2c0bf8f53af326a