Pith. sign in

Paper Citation Record · LEDGER

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2605.29782.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.29782 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T04:26:24.074391Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved14
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6e7e5f7-7236-48cd-a1f1-03b73300bfd1 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:43:15.445367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:d457e08fa039d6af3f4843f851d98da7b2b8d07765039f7937ec8fd89d03242c

Observation 29c8779f-fdda-409e-80f1-508e656c1a1f · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:43:14.669708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:2284e3518d7841ad12895ed6a019c7e4ca004a1e317b297a9edc24c10ef86e99

Observation 0a303672-891d-4b0b-9604-177825d56b39 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:43:14.666563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:219793c8c010eaa8e0204012d59e1c1452f57a8383b2fcdff9869ea6204a2035

Observation 5f312830-4cf0-4aa6-affd-cbadde38aa03 · outbound

This paper cites Assumption A.1(Answer Parsing).The final answer for both states is parsed from one action a, which corresponds to one token.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Assumption A.1(Answer Parsing).The final answer for both states is parsed from one action a, which corresponds to one token

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:0565b79a079a4d45af6aa229e8653e12517ad1baebc7758f15e332ce2801efe4

Observation d212c651-8b0a-4eaa-b0c5-a1c230eb079b · outbound

This paper cites an unresolved cited work.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:f873f05c40ceb852c50a47e165db24c3523252f71b84a1fa1614adc070f19686

Observation 4fab8912-a53d-4dbf-8ea5-5f512bcc4d39 · outbound

This paper cites Thus, the corresponding eigenvalue isλ ∥ = dϵ ∥x∥2 2+dϵ.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Thus, the corresponding eigenvalue isλ ∥ = dϵ ∥x∥2 2+dϵ

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:950be3f2ab423cd61110be81ced7a099e738adb40edf23fde9f443fbb724e55f

Observation f473ed68-7dda-4f91-8346-1d113bd5bf45 · outbound

This paper cites Suppose we continue the generation of s1 and s2 until (T−1) -th state located at the token index ℓ.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Suppose we continue the generation of s1 and s2 until (T−1) -th state located at the token index ℓ

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:a3b368c0e1689fcaf5f184e50ea85889252725e7db1379c4a95b0731ec2bfd07

Observation c78888f1-c705-46c0-b21e-f0723b0e069c · outbound

This paper cites A General Case In previous section, we only prove the correlation between MinDistanceof hidden states and final reward in a simple case, where.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning A General Case In previous section, we only prove the correlation between MinDistanceof hidden states and final reward in a simple case, where

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:f24b304ac638077fa35bfa8000847996be6fe8660f5914949d41a958ca6525b5

Observation be727993-136a-4e6c-8cbd-2ead0138ad68 · outbound

This paper cites an unresolved cited work.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:8a9a192fc87c8d7e077c842574504e608e29510bbc3dca82c76076951ff235f9

Observation 7956b08b-ea97-42dc-be55-b03b2198d1ae · outbound

This paper cites an unresolved cited work.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:eff0ea7853a6f70a9147a8a2f4c233a1f1c00c6c0fd5004de5adf514f054d2be

Observation 5831e3c3-ec11-4cef-a11d-a72dbd323a95 · outbound

This paper cites To expand Theorem A.11 to general case, we need to re- view these three assumptions.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning To expand Theorem A.11 to general case, we need to re- view these three assumptions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:f06384521883349bac06c7d2af90cbc88c10572c5e2a14f0ad5dab6c8f171c6b

Observation cb323b79-50ff-480c-8f3f-c00ae34c329a · outbound

This paper cites Now, the problem becomes, the relation between the final reward of ˆs2 and s2.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Now, the problem becomes, the relation between the final reward of ˆs2 and s2

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:24d9890cc560fa31e1ecc7c403d0f28393bf5d253ecc7684378eea7630d5e60e

Observation d50afa9f-a2d0-42ad-a1e7-4faa3360f6f2 · outbound

This paper cites , η 2}, the dot product score p1,i is replicated ˆri times in the first block of p2, where ˆri are nonnegative integers satisfyingPη2 i=1 ˆri =η 1.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning , η 2}, the dot product score p1,i is replicated ˆri times in the first block of p2, where ˆri are nonnegative integers satisfyingPη2 i=1 ˆri =η 1

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:a4ff520c67152167ac475d760b06a1c7ddbf5151332276ad1742acbfd77ee0c4

Observation a9977850-c895-4943-8fe6-87ce659b43e2 · outbound

This paper cites , ℓ}, the entry p2,η1+i−η2 in p2 equalsp 1,i +δ i for someδ i ∈R.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning , ℓ}, the entry p2,η1+i−η2 in p2 equalsp 1,i +δ i for someδ i ∈R

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:47356175d7ec4b23e6b7cd26eae07d991268eacff1b3854fd19f47878b18625a

Observation 9bc1ee54-db30-4a84-bab3-7340a6b767aa · outbound

This paper cites Without loss of generality, we assume s1’s hidden states Xl 1 ∈R η1,d and s2’s hidden states Xl 2 ∈R η2,d, where η1 > η2.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Without loss of generality, we assume s1’s hidden states Xl 1 ∈R η1,d and s2’s hidden states Xl 2 ∈R η2,d, where η1 > η2

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:25f9d8cc1fb5085124fe678150b4d3d337ee05fbb88f368a53f942d25c65af0d

Observation 8ec788a2-686d-4ecf-8d56-389ead59fee9 · outbound

This paper cites Now, we need to build correlation between ˆXl 2 and Xl.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Now, we need to build correlation between ˆXl 2 and Xl

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:9fc73288547a68b7120b9762c577811f43c3c5ac1cc06c12a10e0c618fe97b6c

Observation 0df0f422-faef-42d5-9afb-28ed16c62a59 · outbound

This paper cites important.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning important

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:1884dcf28036068e2c62f1c46515ea12c4b925b350f7a7efd9cffc1ef40daf6c

Observation 710fa810-ac94-409c-bb64-0de3069c356f · outbound

This paper cites an unresolved cited work.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:ae06917bb411ccd65420293f7a6dc8f9d2d0569a15ac15be64603b4edad7715a

Observation da1feee1-300c-4718-86d9-f081ce785f5b · outbound

This paper cites The result is presented in Table 10.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning The result is presented in Table 10

Reference 19

Resolution
malformed identifier
arxiv_id, observed 2026-06-29T08:43:15.447991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:c043595ca5bea27d4fbd5b09aadfc4d887693f22704a2271e14527ea9ecd184a

Pith citing papers

Observation f36e8105-ebcc-453b-9b13-0eabe99137a5 · inbound

APeB: Benchmarking Personalization Ability of Large Language Model Agents cites this paper.

APeB: Benchmarking Personalization Ability of Large Language Model Agents Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T04:26:24.074391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:26:24.074391Z digest=sha256:52145b896b009a0c52a30454ed46b919302a650a18eaac6f3075fae2df522a23