Pith. sign in

Paper Citation Record · LEDGER

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment

As of 4 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2605.13537.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.13537 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T20:27:26.552596Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact11
  • verified fuzzy5
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38c7f9fa-91b3-4485-b188-07783ea51d97 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:39:29.265050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:62722830fb0cca47f3b31cda4847bade299573c3e6d6681f1bd1ad3c68c38413

Observation 8584fc7a-97be-49fe-a67a-187801539ab3 · outbound

This paper cites Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:39:29.247752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:135dfebec38195e267993037d5568616792e3d1c53538ce1d6616d9f5a0a0f94

Observation b485feae-cc7e-449b-8996-f25be0ac63d2 · outbound

This paper cites Qwen3-VL Technical Report.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Qwen3-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:39:29.235963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:a01523440ddd36a221fd7eb600a29ecd3b1cd5724e72263a3a8a11f6f08fa425

Observation 40f68bf7-bc59-4722-8ec2-116e64ff3637 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:39:29.253089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:6185d8120fa648da7c40791d29151ead3dfabd65426a4a5f4138d796f4d9ece4

Observation 4e673fbb-612c-4c36-8366-790b2d7043fe · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:39:29.240816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:fdc5ea4f79fd0acdcd8f171b501ad116e9d269da68c574b3c6540ef27517f465

Observation ddaab207-5768-41a3-a33a-2335c5c399d0 · outbound

This paper cites Reward-augmented decoding: Efficient controlled text generation with a unidirectional reward model.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Reward-augmented decoding: Efficient controlled text generation with a unidirectional reward model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:21:11.804102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:7d08a997bf562d24be5fcae2d88dae3a6d3aba9a88f3aa4c21ea6ee5281395a7

Observation bcd1b07e-c8b5-4cb1-95f0-4b379e0754c0 · outbound

This paper cites Gemma 3 Technical Report.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Gemma 3 Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:39:29.270766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:5354b395e0bd0bfed2302c3ada99045622f03c7094c64c3638342aaa5b25b0aa

Observation 2f0faf65-83d7-4d47-9a1e-0d49c7716eab · outbound

This paper cites Regularized best-of-n sampling with minimum bayes risk objective for language model alignment.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Regularized best-of-n sampling with minimum bayes risk objective for language model alignment

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:21:11.799864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:f999764b0b594b5ce0643582c31741ac2a5860f5920c4362bdf98a1bc2531826

Observation d7bd3608-fc68-4a01-b1d8-a444a69a3c70 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Adam: A Method for Stochastic Optimization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:39:29.258460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:c0b3af4e24792fa01c11ba84ddaf9b9b2848d2a856951e3f780b01cf75faa3ea

Observation f21ccb7c-2cf7-4590-9315-40e9ce3e666a · outbound

This paper cites On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:21:11.808393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:342da20980a94facd7786e191278c3a552758544663c215b3c3d38114b306abc

Observation aca132c2-62de-46a9-b25c-9b675de70f39 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment WebGPT: Browser-assisted question-answering with human feedback

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:39:29.225175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T01:08:09.995583+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:dcf293a06f20caf1c2ce288ea30291cd078525e7df5343e7a452f5e856fa7553

Observation 61a03144-eb7c-4d9f-904a-c758f6753342 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:39:29.231231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:1f923f26203b313f8315eb6ae25813e0c336f660842441222d3827270b9ead10

Observation 7ed7e962-1724-40b9-9122-22b5361ae2c3 · outbound

This paper cites Qwen2 Technical Report.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Qwen2 Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:39:29.219280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:220f4f950b191d6b904d74057be4861c1612d03f2f4327d7bb8e6c57a73688e1

Observation a2fa2806-1291-4929-beb4-9026148a945e · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Fine-Tuning Language Models from Human Preferences

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:39:29.213700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:cffe5ab5b3f933a6e4c8cb16440d7eaf0a8dbd7a5da5ec3cd8dba55a878f1afc

Observation 82087f67-e075-469a-9176-62ae87afbc4f · outbound

This paper cites For this example, the expected reward is given by Ex∼D,y∼π ∗ω(y|x)[g(x, y)] =ρ ω1 log p1(1|0) p1(0|0) +ρ ω1 log p1(0|1) p1(1|1) ,(22) as a function of two log-likelihood ratios.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment For this example, the expected reward is given by Ex∼D,y∼π ∗ω(y|x)[g(x, y)] =ρ ω1 log p1(1|0) p1(0|0) +ρ ω1 log p1(0|1) p1(1|1) ,(22) as a function of two log-likelihood ratios

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:21:11.792023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:5b93e23c05f273a6c52498befff8fe9c5af4c4ac19823bd3b676b0feaf57ddab

Observation 28c1645e-3ffd-47a0-b347-52cdd45a4a5b · outbound

This paper cites Final answer.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Final answer

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:21:11.795673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:8abcee6d38ff88e7496e576639528c339c2624a1658fa8f7cd142b9d950a9f1f

Observation 60facd7e-a407-4129-aeaa-bf03cf493f98 · outbound

This paper cites an unresolved cited work.

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:21:11.787163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:27:26.552596Z digest=sha256:3d3d801ea09830dac22acf0e615935b3859991f1fd88a433ba3cf1f84c6d3a89

Pith citing papers

No inbound Pith citation observations are available.