Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T14:05:37.120262Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 2 inbound Pith citation observations for arXiv:2601.21484.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T14:05:37.120262Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T07:26:46.105050Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-29T07:33:14.104597Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ef8e1f2d-2d33-4bf8-9154-c4357321ca13 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7e031410-7850-43bb-a61a-57e5111c4ca8 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 114285a7-d4c2-4d75-b1a5-181c265bfa08 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3729bb81-31cf-4524-b97c-e89ea17c170d · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Step-level verifier- guided hybrid test-time scaling for large language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9faa33c2-721a-448b-aa35-0800432d54c1 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Evaluating Large Language Models Trained on Code
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 50e3e2fd-8e02-47ce-9143-e9da9eba0189 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Inference-aware fine-tuning for best-of-n sampling in large language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 387f31dc-d891-4a83-8207-eed2c09e3c6f · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Training Verifiers to Solve Math Word Problems
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a45c89ce-0c7c-4762-a974-7bb75bbb874c · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4ea25cbc-3cc5-4549-bef9-331d6a9e9558 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment On grpo collapse in search-r1: The lazy likelihood- displacement death spiral
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4c092ba1-561c-46a3-a115-ae396192aaa4 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ca713a13-d20d-4186-a2a2-7bc2a3e542b1 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6f5bf541-14e5-4ffd-8b3e-1a5725aef206 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 05dca34f-b046-4843-bb3b-4f77cd29b5e4 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment S., Seo, J.-s., Zhang, Z., and Gupta, U
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 45178250-d860-4872-b3a2-fa959dc34ecf · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 38315df3-a3e1-41cc-8d71-5e74fb1602f8 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Language Models (Mostly) Know What They Know
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 58f156b8-f32d-4ad5-8501-24ad5c9d98dd · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Scalable best-of-n selection for large language models via self-certainty.arXiv preprint arXiv:2502.18581
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6359aaa2-1066-43cc-b494-f3c027ab0c01 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Reasoning with Sampling: Your Base Model is Smarter Than You Think
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6b2e9ac0-d626-4e05-bbdc-582c1c2cc4d7 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6aaa1066-340d-4cc3-9fdd-a6052b819ec2 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment LLM Post-Training: A Deep Dive into Reasoning Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5b05f9b4-3f8a-4611-bd4c-4c92da5f2043 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3e1d58b8-e6ce-4fcb-bdc9-a3c5ca538821 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Towards a theoretical understanding to the generalization of rlhf
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f0c8d3af-fff1-4936-9936-13fb9537a5c6 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fc4de321-91a3-40a4-b09d-f67609b6292e · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment WebGPT: Browser-assisted question-answering with human feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation afe8b470-0139-40f1-8199-7a358ef12e1d · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Large Language Diffusion Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c8c0a958-b130-4195-8d7c-afeb86cbdf46 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bd461575-c5ef-4c14-b844-3b19f37026e0 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Proximal Policy Optimization Algorithms
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b24ce457-a6d1-4ff7-8ec1-c86db7239192 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Can large reasoning models self-train?
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 56619bfe-1c99-44fa-8045-5d20715d702d · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f9f81cf2-6ed0-41d9-bb4c-9a639d9be577 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment HybridFlow: A Flexible and Efficient RLHF Framework
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d3b769b0-f82e-4df6-ad2b-bd496be1b478 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment A General Framework for Inference-time Scaling and Steering of Diffusion Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9e3224b3-69c0-4c35-b52f-3aa3bb9f6bac · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2f636823-3802-4b33-bad4-320f77a7a7d2 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment LongCat-Image Technical Report
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6173065f-af2e-4d9c-9c8f-e4c4a21d425b · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ea60f26c-e0a8-4a16-8ee2-505577de2822 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a977a383-cd55-4062-b03f-c10eebb8a2c8 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Transformers: State-of-the-art natural language processing
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 976f0292-ec29-485f-8f4c-9e73d25be15e · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5cefdcf1-46e6-4933-96f5-79e65ee335e7 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Qwen3 Technical Report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6d2fbad8-8e70-43d4-b9e2-5c2df885c689 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 70a2a09d-4379-4f5a-ae18-7e03d495d561 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ae350ee7-47dc-4ae1-abca-813d6df0cf58 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment 1 M MX m=1 E(y,x ti−1 (m)) +ϵ C−ϵ−h(ϵ, M, λ, D) f(x ti−1 (m)) # +E q(xti−1 |y,xti )[1Ac ti f(x ti−1 )] ≤ C C−ϵ−h(ϵ, M, λ, D) Exti−1
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 15d562f2-aab2-48b4-90da-e969ef298187 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment The number of tokens generated overIsteps is Ntokens = IX i=1 M(B+K(d x −iB)) =M d x +IM Kd x − 1 2 (I+ 1)M Kd x =M 1 + I−1 2 K dx
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dd9e1dd6-e54d-4660-afb1-4445cf78e362 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Best-of-N is naturally integrated into our ETS framework as a special case, with detailed hyperparameters provided in Appendix C.2
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation af774b80-8076-4032-9792-75d4d8ccc42d · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment For DLMs, we implement beam search ourselves; however, due to their iterative generation nature, DLMs cannot be accelerated via batching in the same way as ARMs
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dcadba1b-9e35-4970-84a1-d11a96d7bd56 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment We ablate the temperature on Qwen3-8B and plot GPQA accuracies (left) with corresponding latencies (right)
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 34c34621-212a-45da-b0e8-f1d034f42558 · outbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Based on this efficiency trade-off, we fix dx = 512for all main experiments on ARMs
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9361fab5-b076-4d63-8fa7-e8edc2a4ba2e · inbound
Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b0475f41-f8d4-4fcb-9dbb-8a32d296b39a · inbound
HTAM: Hierarchical Transition-Attended Memory for Operator Optimization ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.