Pith. sign in

Paper Citation Record · LEDGER

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment

As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2506.12446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12446 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:57:56.794234Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:15:37.083838Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:17:30.651801Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4912a7b3-e65e-41ef-8303-80c8726670f7 · outbound

This paper cites GPT-4 Technical Report.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.681605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.681605Z digest=sha256:edec23f8f3e03dde497c8d36a5c15045692eb5aebed8063ad18fb2283d2f278c

Observation 36b19547-6eef-4627-ae27-4dcc131ed85a · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.685536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.685536Z digest=sha256:f27cceb346460ad313d71d8a0ed75d43ea7f423dae8d291ffa4f7d661118509b

Observation 5a301271-6912-400a-b58d-7c8d8420db2b · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.689287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.689287Z digest=sha256:8c316abaecc0d65beccc03f354fa9a6cae1b1f521b53162f2d3a85003f683720

Observation 64b8a358-7d5d-46f4-b1bd-763f83fc57b1 · outbound

This paper cites Transfer Q Star: Principled Decoding for LLM Alignment.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Transfer Q Star: Principled Decoding for LLM Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.692630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.692630Z digest=sha256:4fc191475863e359a7fe367d5c10fa8220ccb4fd1bc16d30325156bf0b8a1324

Observation bd759edf-e7f4-41a2-b8c4-a9a8ed339231 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.695734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.695734Z digest=sha256:713d7abead4b3cb54361d4f1d255a6ff224be38f4e52a904ac7a3e958570fc6e

Observation be50a20e-9127-4da4-90a6-8a96f6782934 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.228881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.698732Z digest=sha256:528a7c2f6b2a0877e24e971690c64e57cfe27476665fc0d4cdf34c1735acc372

Observation 5ba7954f-67b5-4ade-a03c-f8e5177afea9 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.218508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.701902Z digest=sha256:8bf36dfc8f86a08f4f0d48bda7344381e1e4231ab2c036a3f1544c964194b59a

Observation f4b0db48-f16d-4119-9208-e0a95dcea152 · outbound

This paper cites The Llama 3 Herd of Models.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.704410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.704410Z digest=sha256:8ee466cdfd6304ef1ce74d69cf7b9e63e7c09b4d4764405c582e2d39e5154833

Observation 7c3de4f8-d9f8-4c1e-8feb-4f2346d6cdee · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment KTO: Model Alignment as Prospect Theoretic Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.707328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.707328Z digest=sha256:f373ccd2859525dde4dd8c9acd24aeaddbac684cdfd9061db99ebf7e5eebf142

Observation f374665c-7c97-4133-b1d5-1da0c4d3329d · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.207916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.710065Z digest=sha256:239842849ce762ce24904cc4f5de400f11a81ce2f0d24cf940b38c9a988e01a2

Observation 750e2506-feb8-4427-b82d-5e28a9a52517 · outbound

This paper cites Value Augmented Sampling for Language Model Alignment and Personalization.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Value Augmented Sampling for Language Model Alignment and Personalization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.712904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.712904Z digest=sha256:5dbe792b69cb1305fd6086316b0f92ef196ff10646a95ad208b7fba2d11acaeb

Observation ff29bdb1-8785-49a8-a405-a1467e673158 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.716070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.716070Z digest=sha256:84c7bb1a7a083197e4f5f8e9751d63db809697b97e4194631d229e64e472252e

Observation 30413e44-3eea-4f9f-99c3-37fa6e7a001b · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.198063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.718781Z digest=sha256:a2b58d339cddda822a4a3ab8292af5d8a8800ad970b219612c37c64587a46fc5

Observation 2d75fb37-7379-41ef-9b21-9d795adc87fe · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.188134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.721421Z digest=sha256:710b43d4f313e2f947581a50a8decec9f1c41c854738b47364b3e81d37c7cb6b

Observation e2c5cddc-97b0-4acb-9b40-bb6750515811 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.178228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.724342Z digest=sha256:baf3f40dd8602b35ff0e83dd67f0c1a6e1d934a09dba03340e9325448b6504e3

Observation d2fea9c8-b210-42d0-8d78-732189a2b6d3 · outbound

This paper cites DeepSeek-V3 Technical Report.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment DeepSeek-V3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.727160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.727160Z digest=sha256:22dc98b3d78db25fe623f412fec3e92e0a0282775d07afe99bbd4076ea9e9999

Observation 63beea3b-8356-4f0c-96e8-868abdd8ef81 · outbound

This paper cites Chain of Hindsight Aligns Language Models with Feedback.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Chain of Hindsight Aligns Language Models with Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.730028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.730028Z digest=sha256:c3a45b6c07d8bc9da70def6ab101b2bccb3301dc77ee431abed778195ee0d126

Observation 90b1afa9-5f82-4ca0-a092-fe93568fe61a · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.168071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.733069Z digest=sha256:f2ce40607f911c5ac7d2f1e8427fbbf1da2e5e83bc4c60203ab33159b3983ba3

Observation b034acf2-58e1-40be-b0af-8d389838d0e4 · outbound

This paper cites Controlled Decoding from Language Models.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Controlled Decoding from Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.735807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.735807Z digest=sha256:7aeff3cbbc1d3f1baaa15279cbde536e31a853626204a582e3b9a2f70c8927ea

Observation 1367ce34-8eba-4b7f-a077-1e9959661349 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment WebGPT: Browser-assisted question-answering with human feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.738980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.738980Z digest=sha256:efb24719751d99a355fc6334aa71389435c05fc3bfe8badfbb2a9237270df38e

Observation cd673f1c-19fe-4958-a75d-b6523f875956 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.742029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.742029Z digest=sha256:e7f06a794ec0eff116aa24c2efcfae0c4b51474ef9690a5c6152c8b3d7a048fc

Observation 922294fd-5611-495f-9001-32f2266d9f50 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.744661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.744661Z digest=sha256:f68088a15f0068e12cfc1fbc18f80bff0e17acc7f172cde6eb471b52dbcf6f0e

Observation 143a9fd4-240f-44aa-bc2f-f615c4e8db85 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.747399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.747399Z digest=sha256:39fa69ad0341b6863d696ba036abf7b8d0f6dea726b4b650107dd9bc739c6820

Observation 3199fd39-8f9e-45ef-ae6a-b68988faee9d · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.750234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.750234Z digest=sha256:a32bd0cc868a78eb6d21dd4faa1730efdde17b31570d6af7f0db29d095001b48

Observation a7787224-2806-4057-ac3b-82b765031810 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.131045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.753004Z digest=sha256:f24743a54dd312e9ec75f2077d647d2e8c72afd2c788a680edbd96489a0c9aaa

Observation 78ed9d1c-1e6c-47a6-8d80-cb1291d8bbef · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment LLaMA: Open and Efficient Foundation Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.755879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.755879Z digest=sha256:b0d4833eb5b6de8fba72eaaedd8415c9ee2e51ac318ec2c5ecebff6decca3be8

Observation 62b97dd5-3282-43c5-8392-2a43316d5d4c · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.119669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.758899Z digest=sha256:4087da6cdb6583e2bf015e33b55615291646dd9afe52359a569cf62228907cf4

Observation b835bba0-5dc3-49a4-a69b-fdd4e7ec0ca3 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Aligning Large Language Models with Human: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.761990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.761990Z digest=sha256:7afb81ee6e174dbd6f08faaa316c96babb10e8993a380d40ccd2097e97c7a79a

Observation 4e89d054-047c-48db-be8a-bd8991fc20d9 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.109086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.764821Z digest=sha256:3c74edae6b1e41d0ad613c56250398e41e8b16568a02666fa7ce40749aae30c1

Observation ac75c50f-f1d4-4a42-9063-586384f48b1e · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.097961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.767542Z digest=sha256:8a280f32f4ba475bdf8a1f76cb63be1844469bb4dabe73c272d8d969559a9ca2

Observation affda21e-2a61-4a2e-a931-ae8f8241e175 · outbound

This paper cites GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.770058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.770058Z digest=sha256:2ea44e314e6a61907da67fc47247b21e4d081e1a579b5ba78699cd6ac5c548d4

Observation aca8fcc6-d181-4273-9fc6-653d0d315f8c · outbound

This paper cites Inference-time alignment in continuous space.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Inference-time alignment in continuous space

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:57.087942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.772720Z digest=sha256:b889b1388027808f9d3eed8359d37c6fd11764c80d5235dad7f30743c316c1e5

Observation 71993794-c053-464e-aee0-5b8a05fd2004 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Secrets of RLHF in Large Language Models Part I: PPO

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.775426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.775426Z digest=sha256:f72ef9943d571efb435d96b04e3b2bfbcfc0c78acc5d2465246ae8338fb1409e

Observation 1b4a77da-8cf0-4db5-ad51-48d485e998c6 · outbound

This paper cites Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.778437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.778437Z digest=sha256:f8ebb0f4c79b843421689f4e49e9aafbd5a45f7421855017433252104a1ab710

Observation f423252d-98fe-4eb1-b8ba-ddcb0c443197 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.075895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.781278Z digest=sha256:2aed495f1554925db7c947f75cbc7df30f011eba7aa35a6cca5c0d0a42b6db77

Observation 5cdc6326-be2b-460e-bb3f-bacb141ca3a5 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Fine-Tuning Language Models from Human Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.785108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.785108Z digest=sha256:fe8647860ab6457bedc3649b0cca49734739cfd827c575ce2c7c7fd284e3dbba

Observation a3988930-adf1-455e-8a44-bdacb97b37f9 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.788097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.788097Z digest=sha256:69cf28a97bb619197365835653fe41c7f0f8a47cb55f1770e78649d18a10a9b1

Observation 89243b78-9e17-42f3-8ef2-3cb6b72e0e5b · outbound

This paper cites online" 'onlinestring :=.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment online" 'onlinestring :=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.790928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.790928Z digest=sha256:645b172dd5fb1937148da920125f916e81d3e3786b3d64dacd318395981bd5e9

Observation abab279f-b4a8-4765-8e2c-21f35eb5758f · outbound

This paper cites write newline.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment write newline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.794234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.794234Z digest=sha256:1146e587733cb89a1332aae6b44e1283e3ad7f3304b5a5032406cb352abcf659

Pith citing papers

Observation 3b792169-b2d2-4d3e-9265-9f0db8bf8c72 · inbound

Gradient-Guided Reward Optimization for Inference-time Alignment cites this paper.

Gradient-Guided Reward Optimization for Inference-time Alignment From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:30.653219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:43:04.984866Z digest=sha256:0a4699cd29a8ee10814c8df2700792da51cf8095190e83d4d0fffe2db62fb87f

Observation 033c006d-117b-49b8-afbe-e84dce3a2697 · inbound

Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation cites this paper.

Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:37.083838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:37.083838Z digest=sha256:2ca81dd47258a44cff41c5ac63d292f24bb3d47b57453147140f7a67bb004cfa