Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2605.29782.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T04:26:24.074391Z
A source-named dated measurement, never combined with another source.
Source: cited_works
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c6e7e5f7-7236-48cd-a1f1-03b73300bfd1 · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29c8779f-fdda-409e-80f1-508e656c1a1f · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a303672-891d-4b0b-9604-177825d56b39 · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f312830-4cf0-4aa6-affd-cbadde38aa03 · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Assumption A.1(Answer Parsing).The final answer for both states is parsed from one action a, which corresponds to one token
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d212c651-8b0a-4eaa-b0c5-a1c230eb079b · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fab8912-a53d-4dbf-8ea5-5f512bcc4d39 · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Thus, the corresponding eigenvalue isλ ∥ = dϵ ∥x∥2 2+dϵ
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f473ed68-7dda-4f91-8346-1d113bd5bf45 · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Suppose we continue the generation of s1 and s2 until (T−1) -th state located at the token index ℓ
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c78888f1-c705-46c0-b21e-f0723b0e069c · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning A General Case In previous section, we only prove the correlation between MinDistanceof hidden states and final reward in a simple case, where
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be727993-136a-4e6c-8cbd-2ead0138ad68 · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7956b08b-ea97-42dc-be55-b03b2198d1ae · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5831e3c3-ec11-4cef-a11d-a72dbd323a95 · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning To expand Theorem A.11 to general case, we need to re- view these three assumptions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb323b79-50ff-480c-8f3f-c00ae34c329a · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Now, the problem becomes, the relation between the final reward of ˆs2 and s2
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d50afa9f-a2d0-42ad-a1e7-4faa3360f6f2 · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning , η 2}, the dot product score p1,i is replicated ˆri times in the first block of p2, where ˆri are nonnegative integers satisfyingPη2 i=1 ˆri =η 1
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9977850-c895-4943-8fe6-87ce659b43e2 · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning , ℓ}, the entry p2,η1+i−η2 in p2 equalsp 1,i +δ i for someδ i ∈R
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc1ee54-db30-4a84-bab3-7340a6b767aa · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Without loss of generality, we assume s1’s hidden states Xl 1 ∈R η1,d and s2’s hidden states Xl 2 ∈R η2,d, where η1 > η2
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ec788a2-686d-4ecf-8d56-389ead59fee9 · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Now, we need to build correlation between ˆXl 2 and Xl
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df0f422-faef-42d5-9afb-28ed16c62a59 · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning important
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 710fa810-ac94-409c-bb64-0de3069c356f · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da1feee1-300c-4718-86d9-f081ce785f5b · outbound
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning The result is presented in Table 10
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f36e8105-ebcc-453b-9b13-0eabe99137a5 · inbound
APeB: Benchmarking Personalization Ability of Large Language Model Agents Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.