Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-14T20:27:26.552596Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2605.13537.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-14T20:27:26.552596Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
17 of 17 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 38c7f9fa-91b3-4485-b188-07783ea51d97 · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8584fc7a-97be-49fe-a67a-187801539ab3 · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b485feae-cc7e-449b-8996-f25be0ac63d2 · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 40f68bf7-bc59-4722-8ec2-116e64ff3637 · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4e673fbb-612c-4c36-8366-790b2d7043fe · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ddaab207-5768-41a3-a33a-2335c5c399d0 · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Reward-augmented decoding: Efficient controlled text generation with a unidirectional reward model
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bcd1b07e-c8b5-4cb1-95f0-4b379e0754c0 · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Gemma 3 Technical Report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2f0faf65-83d7-4d47-9a1e-0d49c7716eab · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Regularized best-of-n sampling with minimum bayes risk objective for language model alignment
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d7bd3608-fc68-4a01-b1d8-a444a69a3c70 · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Adam: A Method for Stochastic Optimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f21ccb7c-2cf7-4590-9315-40e9ce3e666a · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation aca132c2-62de-46a9-b25c-9b675de70f39 · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment WebGPT: Browser-assisted question-answering with human feedback
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T01:08:09.995583+00:00.
Observation 61a03144-eb7c-4d9f-904a-c758f6753342 · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7ed7e962-1724-40b9-9122-22b5361ae2c3 · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Qwen2 Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a2fa2806-1291-4929-beb4-9026148a945e · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Fine-Tuning Language Models from Human Preferences
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 82087f67-e075-469a-9176-62ae87afbc4f · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment For this example, the expected reward is given by Ex∼D,y∼π ∗ω(y|x)[g(x, y)] =ρ ω1 log p1(1|0) p1(0|0) +ρ ω1 log p1(0|1) p1(1|1) ,(22) as a function of two log-likelihood ratios
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 28c1645e-3ffd-47a0-b347-52cdd45a4a5b · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Final answer
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 60facd7e-a407-4129-aeaa-bf03cf493f98 · outbound
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
No inbound Pith citation observations are available.