Pith. sign in

Paper Citation Record · LEDGER

Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2402.06700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.06700 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:29.215085Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T10:18:11.872878Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f0f69048-4940-4a7c-8c9b-2503ce729b36 · inbound

Advancing Decoding Strategies: Enhancements in Locally Typical Sampling for LLMs cites this paper.

Advancing Decoding Strategies: Enhancements in Locally Typical Sampling for LLMs Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:29.215085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:29.215085Z digest=sha256:4ad53e6d4a5ba3b578a6f570e389073c35bdcfb85f720ad26c3d233153b51b7c

Observation bcf07401-23ed-448d-a23a-661e2eea50c4 · inbound

Multi-Amateur Contrastive Decoding for Text Generation cites this paper.

Multi-Amateur Contrastive Decoding for Text Generation Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:15.174905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:15.174905Z digest=sha256:7bd9517ebafb8dc503cbf9ef287783e725643930c46086f5da4a7ce99a4def5c

Observation 31b0c102-880e-4755-bbaa-8d84bca1bb8a · inbound

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning cites this paper.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.614727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.614727Z digest=sha256:b0c0f63d81afa5346717c5f63682813539c8bf513990366ad0be816969577b4f

Observation 0e5d7203-d1d9-4838-8800-dbbb6a3e972e · inbound

Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation cites this paper.

Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.875767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T10:15:22.633882Z digest=sha256:90f0426605490f3ee7ba1ce5187294abd560553604d663a212c3a6ad377b1608