Pith. sign in

Paper Citation Record · LEDGER

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection

As of 10 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 3 inbound Pith citation observations for arXiv:2505.19475.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19475 v2

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:17:44.447421Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T11:21:30.867480Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T11:24:08.689134Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 807d5c9c-325e-4275-b43c-c5fb24b057f5 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Measuring Mathematical Problem Solving With the MATH Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:43.662673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:43.662673Z digest=sha256:1743fd440fe8552378e18378b75d4aea2559794850739b4dedca0f35bcfc52cd

Observation f4a35f45-02bc-4213-9e61-4aff232acf98 · outbound

This paper cites SuperHF: Supervised Iterative Learning from Human Feedback.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection SuperHF: Supervised Iterative Learning from Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:43.978557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:43.978557Z digest=sha256:cf5a96693f195b03bd766949fb5e80ae129b6b44800c759d3c4f83a3b6ae069e

Observation c0b18e1d-6719-47ea-b3fb-c519270c0560 · outbound

This paper cites The Entropy Enigma: Success and Failure of Entropy Minimization.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection The Entropy Enigma: Success and Failure of Entropy Minimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:44.046534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:44.046534Z digest=sha256:f1892d4e83697c2ba67dc7f2808aaf9db1413b75c607993f23930a903a3625e5

Observation 7f1808f0-f761-4bde-b22c-8a0385d9eb8c · outbound

This paper cites The effect of sampling temperature on problem solving in large language models.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection The effect of sampling temperature on problem solving in large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:44.962617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:17:44.125520Z digest=sha256:a9c84cd9cf2d41d6dc967519973124f11a3e3603eba463c60dfc016c14654610

Observation 693ddb9e-c624-41dc-aaee-bdf452aec13a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Proximal Policy Optimization Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:44.208326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:44.208326Z digest=sha256:a12b20d9cdbd9770df3b4b2cf85cabe249add7c04e7d1a8ef9c9ab366a607fd3

Observation 93af7555-85bd-4448-88db-01c371e75911 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:44.288121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:44.288121Z digest=sha256:14ddf186e8ee5788ed460896aea2af505ca6095ec73bf4c60e75039393669cc2

Observation f2549d4f-86c9-4f5e-bba8-49617eedade3 · outbound

This paper cites Continual Learning for Large Language Models: A Survey.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Continual Learning for Large Language Models: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:44.344757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:44.344757Z digest=sha256:8d1efc71116b54fc5a7473b73f7ab91f7d20aa0c8136d3edbe48b06fc58cc4a4

Observation 69a951f2-45b4-4564-9711-2c0316699745 · outbound

This paper cites Beyond Model Adaptation at Test Time: A Survey.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Beyond Model Adaptation at Test Time: A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:44.395312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:44.395312Z digest=sha256:84ce387585e8f2964f3a50f1121610a41b60353157544c0f2824ab615ae843ea

Observation 06c5c7e8-624a-4d61-877a-f1af05548987 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection TTRL: Test-Time Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:44.447421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:44.447421Z digest=sha256:482b1573151e88c68b7acefa692f83f0a5a20521113559aaf4bbbc907fe71a20

Observation f36fc033-e6de-4467-b241-d481a7ef774f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Training Verifiers to Solve Math Word Problems

Reference 1988

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:43.253745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:43.253745Z digest=sha256:f3ea035f2674eee14a3420bc0fb6c93c6330c1265e1b7d3e8737ea8fc78f1220

Observation 159f8971-b300-4993-9824-c833f327f7fd · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:43.075394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:43.075394Z digest=sha256:737bfdd7a6f757b81039fdd66065e1083467826256c1a9c4106017b1ca5a22de

Observation 38e18a3b-642c-459f-8caa-ccc9876c8e26 · outbound

This paper cites Scaling Test-Time Compute Without Verification or RL is Suboptimal.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:44.275296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:44.275296Z digest=sha256:2dc180dfa6b9010649866720b22d7451bea1864e75cd179172f4f998b96fcc12

Observation 08e01256-2de9-41f1-9149-40516eca05f1 · outbound

This paper cites Tent: Fully Test-time Adaptation by Entropy Minimization.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Tent: Fully Test-time Adaptation by Entropy Minimization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:44.295987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:44.295987Z digest=sha256:ea0b5e2a8013221c8f40f8ded10728b2af2aadbd7f882b6b459fdb9a6bdff3be

Observation 6b85c2f8-e47f-4d9d-bd4a-e8f099e93c21 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Reinforced Self-Training (ReST) for Language Modeling

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:43.424702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:43.424702Z digest=sha256:633935fc00e9ef2edeea0422431acd8d7b3b734ebd2175b72e2683ea0adbb82f

Observation 4a4e6df7-7e9a-421e-964f-fbba1689d469 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:43.753618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:43.753618Z digest=sha256:c9cf8ee1296559dcf5cbafcbbbc4e4788fdee4e96497b632c540227e5b913a47

Observation d435c54c-67a7-453a-8e56-f683ca89d209 · outbound

This paper cites Test-Time Training on Nearest Neighbors for Large Language Models.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Test-Time Training on Nearest Neighbors for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:43.596078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:43.596078Z digest=sha256:dbab0c5c2657a454288b8e4e996e9171a38aa9a3178c25246bc06d4147a6a1d4

Observation 58da377b-4834-4a5c-bc68-60066dce9c07 · outbound

This paper cites A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:43.887609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:43.887609Z digest=sha256:3f9c5cb9885dd1e9693b53efd94559232b8a448d4dda715fcf011ec81d67e641

Observation d1ea5909-4513-4eaf-87c1-a98325cd8196 · outbound

This paper cites Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:43.818843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:43.818843Z digest=sha256:dcf930d00019e0d30eb0102d0f3100cecc4a4502cc82f1899722b4a16b8b03f4

Pith citing papers

Observation 8c8461f9-393c-4608-9713-a5662f24200b · inbound

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration cites this paper.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.467718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:1ecfab6bc55cd4da468da365c271f7b88a207a32dba94056e01268bf569956e9

Observation 9490c672-7ce5-4cd1-a42a-9df04c413960 · inbound

Epistemic Uncertainty for Test-Time Discovery cites this paper.

Epistemic Uncertainty for Test-Time Discovery Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.136696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:52:41.192353Z digest=sha256:5c484c22350d6dfb4fe60a7680de5043a0bc93a4baa9532dfb93455027d52f98

Observation 1a97d93e-2d8f-4ce8-9198-cdd718d374a2 · inbound

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation cites this paper.

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:24:08.690998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T11:21:30.867480Z digest=sha256:39ba416fb8a57da0d1da989727290ac7e9f1979c587742e4700979185548fefe