Pith. sign in

Paper Citation Record · LEDGER

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment

As of 14 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 2 inbound Pith citation observations for arXiv:2601.21484.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.21484 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T14:05:37.120262Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T07:26:46.105050Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-29T07:33:14.104597Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact33
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef8e1f2d-2d33-4bf8-9154-c4357321ca13 · outbound

This paper cites GPT-4 Technical Report.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment GPT-4 Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T14:10:13.494087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:0e11e7659a19f44aff7479712bb50966c245399ee849c235dfa5d6b5c20e0527

Observation 7e031410-7850-43bb-a61a-57e5111c4ca8 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.537044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:e920b34d5daf1854485dd2a496fededed056d2b1a83d532585e68d935f448ed2

Observation 114285a7-d4c2-4d75-b1a5-181c265bfa08 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T14:10:13.449970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:c84d0d89427610966ff20f5d806e110a4d40bcedd8879f8b715662caf29045aa

Observation 3729bb81-31cf-4524-b97c-e89ea17c170d · outbound

This paper cites Step-level verifier- guided hybrid test-time scaling for large language models.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Step-level verifier- guided hybrid test-time scaling for large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T14:10:13.943440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:aefad36da3ab211061bb694096d9911d11d800cb39c0401ed4e05b6bf0ea48bb

Observation 9faa33c2-721a-448b-aa35-0800432d54c1 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Evaluating Large Language Models Trained on Code

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.470562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:6e2c4e636fb426cadbad7c384cf2fbf9d353cb50828bbb420394493a8e6964aa

Observation 50e3e2fd-8e02-47ce-9143-e9da9eba0189 · outbound

This paper cites Inference-aware fine-tuning for best-of-n sampling in large language models.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Inference-aware fine-tuning for best-of-n sampling in large language models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.481343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:7ab00cafd37376ff89ff8a735d1ab915a9bdeeeb394043d7552dd71e1c27570c

Observation 387f31dc-d891-4a83-8207-eed2c09e3c6f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.446956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:c55ac089f3411b582aa8bb6f70cb0a2e7b4a9f95ee83a9cf6c48182143c58afe

Observation a45c89ce-0c7c-4762-a974-7bb75bbb874c · outbound

This paper cites Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.437124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:a2fa2ba8916b1443a33866930251a90fcd91301f5c6adc5772c7138c2a32d8b4

Observation 4ea25cbc-3cc5-4549-bef9-331d6a9e9558 · outbound

This paper cites On grpo collapse in search-r1: The lazy likelihood- displacement death spiral.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment On grpo collapse in search-r1: The lazy likelihood- displacement death spiral

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.527496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:d33fd78f674722ff3091270fb93fe100bda747fef4784d3580522b5f1708203e

Observation 4c092ba1-561c-46a3-a115-ae396192aaa4 · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.456057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:9307a4b2a72f15075d343605beed54311c5c095a08f32bd0e6b5da7c9b0d807d

Observation ca713a13-d20d-4186-a2a2-7bc2a3e542b1 · outbound

This paper cites He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T14:10:13.485186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:0becba2b0f0fc71702029240112a6ce415aac718d132696e0099f8e48bf7a681

Observation 6f5bf541-14e5-4ffd-8b3e-1a5725aef206 · outbound

This paper cites ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.533374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:e21c5abf6a1ed158d1880ccada71942dc2ed7e4a21cea5a5f0df9a00d8c72672

Observation 05dca34f-b046-4843-bb3b-4f77cd29b5e4 · outbound

This paper cites S., Seo, J.-s., Zhang, Z., and Gupta, U.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment S., Seo, J.-s., Zhang, Z., and Gupta, U

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.460621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:711c2d8eb54fd8e64ef04177320c070b4b635128e2721188e086e8348f9adb26

Observation 45178250-d860-4872-b3a2-fa959dc34ecf · outbound

This paper cites Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.542812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:81d4cb08f22b0200893675e2c99a1013bb8856415c5f5238cc920a0c52ab09ce

Observation 38315df3-a3e1-41cc-8d71-5e74fb1602f8 · outbound

This paper cites Language Models (Mostly) Know What They Know.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Language Models (Mostly) Know What They Know

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.575828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:d3daf211f48de08cce77cebd1f9e1db0636aeb253225a67b749962d072b639ae

Observation 58f156b8-f32d-4ad5-8501-24ad5c9d98dd · outbound

This paper cites Scalable best-of-n selection for large language models via self-certainty.arXiv preprint arXiv:2502.18581.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Scalable best-of-n selection for large language models via self-certainty.arXiv preprint arXiv:2502.18581

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.552143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:09b85378524ad38fff84f4bb13c866f3435d4398c9f951eed9cca82946fb0868

Observation 6359aaa2-1066-43cc-b494-f3c027ab0c01 · outbound

This paper cites Reasoning with Sampling: Your Base Model is Smarter Than You Think.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Reasoning with Sampling: Your Base Model is Smarter Than You Think

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.477944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:01c06f2b77b27f6aae72d464b605eda377af527ca6fba21700b27a6e173ab5a6

Observation 6b2e9ac0-d626-4e05-bbdc-582c1c2cc4d7 · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.474054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:066d54c4cc5ce92d9095af86c17d3941300c3baf72fdb1ee27ebc2e6a7635d06

Observation 6aaa1066-340d-4cc3-9fdd-a6052b819ec2 · outbound

This paper cites LLM Post-Training: A Deep Dive into Reasoning Large Language Models.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T14:10:13.580164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:3f3fd334ea2d81d848ae7c2b46cc3a9b1001668cb4c58ea784b19f917c08510a

Observation 5b05f9b4-3f8a-4611-bd4c-4c92da5f2043 · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.467310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:94dfe7d45e046fd7b7ea3444954c8a8cc7665d51324574a8cae75d04c2945bcd

Observation 3e1d58b8-e6ce-4fcb-bdc9-a3c5ca538821 · outbound

This paper cites Towards a theoretical understanding to the generalization of rlhf.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Towards a theoretical understanding to the generalization of rlhf

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.522043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:6bdf4eeb13ae7c8f6bb079584db6fef89c9e448c8580a491a2d19dc1d0a9d87d

Observation f0c8d3af-fff1-4936-9936-13fb9537a5c6 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.433825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:96348a3f4f009763bf7b3e1d66bda99ca83a545b615668af5565f92947a8dcfc

Observation fc4de321-91a3-40a4-b09d-f67609b6292e · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment WebGPT: Browser-assisted question-answering with human feedback

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.464048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:e2a52895feaa3d64e5db7543f25bd82d72005afb41ea38285be2b702f66bd562

Observation afe8b470-0139-40f1-8199-7a358ef12e1d · outbound

This paper cites Large Language Diffusion Models.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Large Language Diffusion Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.566324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:a13a4f008b07d74a934d5ed85a3d4e60e61bd3bcdd7e1c6a28942c46007e735a

Observation c8c0a958-b130-4195-8d7c-afeb86cbdf46 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.443897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:c6d4a6824e8f083be2f6c1b224b901e958eb5bda4fecd538007edbd37b3f384b

Observation bd461575-c5ef-4c14-b844-3b19f37026e0 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Proximal Policy Optimization Algorithms

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.488106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:72ebf238fc740381fa16878983b575b765f58480281f3c548a62f7010e11cb02

Observation b24ce457-a6d1-4ff7-8ec1-c86db7239192 · outbound

This paper cites Can large reasoning models self-train?.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Can large reasoning models self-train?

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.561469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:66ab2876404caffa549022b8ab616d108a3eac5d775234b7255a7d6ca108fcdd

Observation 56619bfe-1c99-44fa-8045-5d20715d702d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.516942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:d0692d0e5acbae5575bfa7b582e7ddd3a71eabb48341030f42225a3f70843d25

Observation f9f81cf2-6ed0-41d9-bb4c-9a639d9be577 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment HybridFlow: A Flexible and Efficient RLHF Framework

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.569876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:84798d74a6bbce02d5b874e062cce0f3a88590a8a33ac923844b8df098e25aa2

Observation d3b769b0-f82e-4df6-ad2b-bd496be1b478 · outbound

This paper cites A General Framework for Inference-time Scaling and Steering of Diffusion Models.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment A General Framework for Inference-time Scaling and Steering of Diffusion Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.499261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:64ddc5e485c6f5b7cc7ebcfa4e8877de0f2e28c45d937325049e4dd09e3d0138

Observation 9e3224b3-69c0-4c35-b52f-3aa3bb9f6bac · outbound

This paper cites Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.548018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:9b9ee02b88345d089801704f87f12e4b409971319dc1d17192243d7a15e8b71a

Observation 2f636823-3802-4b33-bad4-320f77a7a7d2 · outbound

This paper cites LongCat-Image Technical Report.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment LongCat-Image Technical Report

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T14:10:13.506735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:8c037f2c6e783fe8d6ab6a688acfb3b984ed6a38209b1010b3bb4e5bfa12eeaa

Observation 6173065f-af2e-4d9c-9c8f-e4c4a21d425b · outbound

This paper cites Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.556700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:13b6653066d90de3aecba47fc02b012ee193b8ec145bc5c5f1bec7fb31907825

Observation ea60f26c-e0a8-4a16-8ee2-505577de2822 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.427068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:47c2fc1f7fce0dd160344ec91e5d4532a51a9773bc527012a7a519493db5ac15

Observation a977a383-cd55-4062-b03f-c10eebb8a2c8 · outbound

This paper cites Transformers: State-of-the-art natural language processing.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Transformers: State-of-the-art natural language processing

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T14:10:13.936639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:6354536ee7d56ce61b09e525a9a1d073224453922548c4d8ef175bbebfaffd64

Observation 976f0292-ec29-485f-8f4c-9e73d25be15e · outbound

This paper cites Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.430330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:6dda1c675d639042c09190d9c07eb9a4c4e8228245f1b2f5a19e644443edcca2

Observation 5cefdcf1-46e6-4933-96f5-79e65ee335e7 · outbound

This paper cites Qwen3 Technical Report.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Qwen3 Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.453187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:1e8be6aa4e83bad5b79260a0f16dae1c83f34fd0024c72b341d7edc2e2b20e44

Observation 6d2fbad8-8e70-43d4-b9e2-5c2df885c689 · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.491193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:b98fad0f2594782fc3e5ed21303120f25d9509114798fc1a2479ba5698778bc1

Observation 70a2a09d-4379-4f5a-ae18-7e03d495d561 · outbound

This paper cites LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.511937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:d93dc8d398f7206cfc9f6a45bcb1cc7e24902436030746ecb3c562d3286ce9ed

Observation ae350ee7-47dc-4ae1-abca-813d6df0cf58 · outbound

This paper cites 1 M MX m=1 E(y,x ti−1 (m)) +ϵ C−ϵ−h(ϵ, M, λ, D) f(x ti−1 (m)) # +E q(xti−1 |y,xti )[1Ac ti f(x ti−1 )] ≤ C C−ϵ−h(ϵ, M, λ, D) Exti−1.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment 1 M MX m=1 E(y,x ti−1 (m)) +ϵ C−ϵ−h(ϵ, M, λ, D) f(x ti−1 (m)) # +E q(xti−1 |y,xti )[1Ac ti f(x ti−1 )] ≤ C C−ϵ−h(ϵ, M, λ, D) Exti−1

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T14:10:13.931796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:c1b38bafefc5c536e2533571953fb6e0de61849a17038a9a508438ca9f0c8572

Observation 15d562f2-aab2-48b4-90da-e969ef298187 · outbound

This paper cites The number of tokens generated overIsteps is Ntokens = IX i=1 M(B+K(d x −iB)) =M d x +IM Kd x − 1 2 (I+ 1)M Kd x =M 1 + I−1 2 K dx.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment The number of tokens generated overIsteps is Ntokens = IX i=1 M(B+K(d x −iB)) =M d x +IM Kd x − 1 2 (I+ 1)M Kd x =M 1 + I−1 2 K dx

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T14:10:13.929579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:064d483d95d8d8df0130610f0647b2c7c5645a42a95dd4dccc396e20ebd740d7

Observation dd9e1dd6-e54d-4660-afb1-4445cf78e362 · outbound

This paper cites Best-of-N is naturally integrated into our ETS framework as a special case, with detailed hyperparameters provided in Appendix C.2.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Best-of-N is naturally integrated into our ETS framework as a special case, with detailed hyperparameters provided in Appendix C.2

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T14:10:13.934233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:af3d40f205f4be1cfa2537edfc32d5bc86956b0ad6427acfb15185491f8ab2b4

Observation af774b80-8076-4032-9792-75d4d8ccc42d · outbound

This paper cites For DLMs, we implement beam search ourselves; however, due to their iterative generation nature, DLMs cannot be accelerated via batching in the same way as ARMs.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment For DLMs, we implement beam search ourselves; however, due to their iterative generation nature, DLMs cannot be accelerated via batching in the same way as ARMs

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T14:10:13.939017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:cb695dddc7a03c3a93cb9438a8f0e949be37dd66d3e7762744080d04b2c82618

Observation dcadba1b-9e35-4970-84a1-d11a96d7bd56 · outbound

This paper cites We ablate the temperature on Qwen3-8B and plot GPQA accuracies (left) with corresponding latencies (right).

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment We ablate the temperature on Qwen3-8B and plot GPQA accuracies (left) with corresponding latencies (right)

Reference 44

Resolution
malformed identifier
raw_fallback, observed 2026-05-21T14:10:13.941243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:c9bacbecf340d7a7bbb2ca202d2a2a41cd550d41b2b8480eaf139e90f28ae933

Observation 34c34621-212a-45da-b0e8-f1d034f42558 · outbound

This paper cites Based on this efficiency trade-off, we fix dx = 512for all main experiments on ARMs.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Based on this efficiency trade-off, we fix dx = 512for all main experiments on ARMs

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.440830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:b84dd6d3bb4ff6d325b69719e88062f0f6b0dc5e259db32670d7937091f19076

Pith citing papers

Observation 9361fab5-b076-4d63-8fa7-e8edc2a4ba2e · inbound

Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting cites this paper.

Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T07:23:06.921587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-20T07:20:50.835826Z digest=sha256:d2aa32828987b31f433b7b585ca6324f8130884cbb593fa2624ae6fbf51da7c4

Observation b0475f41-f8d4-4fcb-9dbb-8a32d296b39a · inbound

HTAM: Hierarchical Transition-Attended Memory for Operator Optimization cites this paper.

HTAM: Hierarchical Transition-Attended Memory for Operator Optimization ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:14.105842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T07:26:46.105050Z digest=sha256:483dd8a8a2a5a22ea539d1956d490b641084a00b551e6ca604d4114c88f95bf8