Pith. sign in

Paper Citation Record · LEDGER

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator

As of 8 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2511.16886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.16886 v5

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T21:06:30.259347Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T14:03:01.974171Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ff9e43c3-9687-454e-ae99-7924ad2edec2 · outbound

This paper cites Universal Transformers.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Universal Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.296424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.296424Z digest=sha256:17e14e71d2025ad762fab8b4500a18a48838d35ab3a5242a4d827adc9605a4ee

Observation faa93361-0335-4d8e-aeb3-ccb0b08bc471 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Classifier-Free Diffusion Guidance

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.717017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.717017Z digest=sha256:ba1b8e0bf5477a0edd9365b182eafc0bd246cbe0a087d3c2b03605963a7df4df

Observation d9965a2c-7799-4554-9e43-271b2abce97e · outbound

This paper cites The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:30.009675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:30.009675Z digest=sha256:3e0a9ac7d4b6d3b0ec0db65be303c92a7cd7970bba18e151f9e2cc96ab894705

Observation f05a28e1-9b46-4a09-90ea-5f15886d9229 · outbound

This paper cites Looped Transformers are Better at Learning Learning Algorithms.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Looped Transformers are Better at Learning Learning Algorithms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:30.161963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:30.161963Z digest=sha256:0df0404b9e59c8aba520b2624aa9d0f42c53a72c9b22ecea385583cfc061530f

Observation fd0c9e37-1de4-447d-8634-215537f7b208 · outbound

This paper cites Alternative Improvement Generators.As described in §4.2, there are several viable methods to generate intermediate steps.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Alternative Improvement Generators.As described in §4.2, there are several viable methods to generate intermediate steps

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:30.259347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:30.259347Z digest=sha256:3609daeea213531e9ed6d76f1767e914562f602a0a3290205ace454020d17116

Observation 21bcbd0d-a95d-42de-9a24-1c41b8f413f4 · outbound

This paper cites Hierarchical Reasoning Model.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Hierarchical Reasoning Model

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:30.093930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:30.093930Z digest=sha256:c1b606e6f9faffc13322751909ab44eee3a2b184e757d5740d3c0be0e837f528

Observation 22b9ddee-1aa1-49cd-a886-ee384d2d7b4a · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Dream to Control: Learning Behaviors by Latent Imagination

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.597655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.597655Z digest=sha256:34bd10d7468e0782d8c608835bbe7f51da0305d18e4eaacdfcfc24b4f96ea582

Observation 21705f32-f9ce-49ff-a85b-0490e878683c · outbound

This paper cites Diffusion Guidance Is a Controllable Policy Improvement Operator.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.374560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.374560Z digest=sha256:8d976ae43f335ac56b819785b81bc4d7b04a45ff1b119b74f3077dde4b02ca2b

Observation 4f18f9f3-50c1-426c-954e-c5937e77cede · outbound

This paper cites On the Measure of Intelligence.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator On the Measure of Intelligence

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.227914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.227914Z digest=sha256:aaf1819b3f2f572107791729c5909472c7984470febcfe711db9d70cd1e32155

Observation 883b64f5-5f91-4300-9ae1-1a2d68106f17 · outbound

This paper cites Less is More: Recursive Reasoning with Tiny Networks.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Less is More: Recursive Reasoning with Tiny Networks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.803009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.803009Z digest=sha256:ab5a05354c6d2be98b03bc62d88a317ce418c4c3c25c9a32627f66e151ed536e

Observation 53302236-baab-4592-86ae-5d7e6e2ac11a · outbound

This paper cites Adaptive Computation Time for Recurrent Neural Networks.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Adaptive Computation Time for Recurrent Neural Networks

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.486681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.486681Z digest=sha256:d3e3a037eb2156224f8d2a333183a9afbae0ac817e795c48bb6368113257fa97

Observation f10b11f9-6161-40e2-b920-75d1698c4166 · outbound

This paper cites Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.881012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.881012Z digest=sha256:67fddb4b9fd5c8993c99115baa15a12ddd280d3f3f25bbbb2d5d3e2997494ce5

Pith citing papers

Observation fa662205-7a44-4221-88c9-4e704fca57ba · inbound

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook cites this paper.

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Latent Reasoning in TRMs is Secretly a Policy Improvement Operator

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-13T14:09:38.535527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-13T14:03:01.974171Z digest=sha256:cd22a988d61ef64dab0995cdae2803a62255141a627031f570756a835b0f4c24